Diego Moussallem, Thiago Ferreira, Marcos Zampieri, Maria Claudia Cavalcanti, Geraldo Xexéo, Mariana Neves, Axel-Cyrille Ngonga Ngomo
| Challenge: | Existing approaches to generate natural language from RDF data have been proposed to generate texts in Brazilian Portuguese. |
| Approach: | They propose a rule-based approach to verbalize RDF data to Brazilian Portuguese language. |
| Outcome: | The proposed approach generates text similar to that generated by humans and can hence be easily understood. |
Similar Papers
Natural Language Generation: Recently Learned Lessons, Directions for Semantic Representation-based Approaches, and the Case of Brazilian Portuguese Language (P19-2)
Copied to clipboard
| Challenge: | Natural Language Generation (NLG) is a promising area in Natural Language Processing (NLP) . |
| Approach: | They present a review of the literature on Natural Language Generation in Brazilian Portuguese. |
| Outcome: | The proposed approaches are based on the Abstract Meaning Representation formalism and have potential future directions. |
InferBR: A Natural Language Inference Dataset in Portuguese (2024.lrec-main)
Copied to clipboard
| Challenge: | Portuguese has few NLI-annotated datasets created through automatic translation followed by manual checking. |
| Approach: | They propose to generate premises and hypotheses using a semiautomatic process to generate sentences and manually check the annotations. |
| Outcome: | The proposed dataset is better at recognizing entailment classes in other Portuguese datasets than the reverse. |
Back-Translation as Strategy to Tackle the Lack of Corpus in Natural Language Generation from Semantic Representations (D19-63)
Copied to clipboard
| Challenge: | Abstract Meaning Representation and Brazilian Portuguese (BP) are selected as semantic representation and language, respectively. |
| Approach: | They propose to use Brazilian Portuguese and Abstract Meaning Representation as semantic representations for NLG. |
| Outcome: | The proposed methods were evaluated on two datasets (one automatically generated and another human-generated) to compare the performance in a real context. |
BlogSet-BR: A Brazilian Portuguese Blog Corpus (L18-1)
Copied to clipboard
| Challenge: | Several efforts have been made to build a corpus based on user-generated content . however, there is still a lack of a large semi-structured corpus that also contains author profiles in Brazilian Portuguese. |
| Approach: | They propose to build a Brazilian Portuguese corpus with 2.1 billion words extracted from 7.4 million posts over 808 thousand different Brazilian blogs. |
| Outcome: | The proposed corpus contains 2.1 billion words extracted from 7.4 million posts over 808 thousand different Brazilian blogs. |
DORE: A Dataset for Portuguese Definition Generation (2024.lrec-main)
Copied to clipboard
| Challenge: | Definition modelling (DM) is the task of automatically generating a dictionary definition of a specific word. |
| Approach: | They propose to create a dataset for definition modelling for Portuguese with 100,000 definitions and evaluate several deep learning based DM models on the dataset. |
| Outcome: | The proposed dataset will facilitate research and study of Portuguese in wider contexts. |
The brWaC Corpus: A New Open Resource for Brazilian Portuguese (L18-1)
Copied to clipboard
| Challenge: | a large corpus for Brazilian Portuguese is needed for NLP applications . the corpus is 2.7 billion tokens, and domain diversity is maximized . |
| Approach: | They propose to build a large Web corpus for Brazilian Portuguese with 2.7 billion tokens . they also propose an updated sentence-level approach for the strict removal of duplicated content . |
| Outcome: | The proposed corpus is based on a pipeline methodology and is available for querying and downloading. |
Multimodality for NLP-Centered Applications: Resources, Advances and Frontiers (2022.lrec-1)
Copied to clipboard
| Challenge: | resurgence of multimodal datasets has attracted significant research interest, but there is no comprehensive survey for this task. |
| Approach: | They present a survey of a multimodal dataset with different modalities according to the applications. |
| Outcome: | The proposed datasets are available online and discuss the new frontier and motivate future researches. |
Finely Tuned, 2 Billion Token Based Word Embeddings for Portuguese (L18-1)
Copied to clipboard
| Challenge: | A distributional semantics model is instrumental to improve the performance of many applications and processing tasks for any language. |
| Approach: | They propose to develop an advanced distributional model for Portuguese with the largest vocabulary and best evaluation scores published so far. |
| Outcome: | The proposed model has the largest vocabulary and the best evaluation scores published so far. |
Towards AMR-BR: A SemBank for Brazilian Portuguese Language (L18-1)
Copied to clipboard
| Challenge: | Abstract Meaning Representation (AMR) is a recent and prominent meaning representation with good acceptance and several applications in the Natural Language Processing area. |
| Approach: | They propose to build an AMR annotated corpus for Brazilian Portuguese using an alignment-based approach. |
| Outcome: | The proposed corpus is based on the Little Prince book, which went into the public domain and explored some language-specific annotation issues. |
Synthetic Data in the Era of Large Language Models (2025.acl-tutorials)
Copied to clipboard
| Challenge: | 'synthetic data' is a data generated with the assistance of large language models to make dataset construction faster and cheaper. |
| Approach: | This tutorial seeks to build a shared understanding of recent progress in synthetic data generation from NLP and related fields by grouping and describing major methods, applications, and open problems. |
| Outcome: | This tutorial will describe methods, applications, and open problems that have been developed and are being used to improve the quality and efficiency of synthetic data generation. |