Papers by Thiago Pardo
Natural Language Generation: Recently Learned Lessons, Directions for Semantic Representation-based Approaches, and the Case of Brazilian Portuguese Language (P19-2)
Copied to clipboard
| Challenge: | Natural Language Generation (NLG) is a promising area in Natural Language Processing (NLP) . |
| Approach: | They present a review of the literature on Natural Language Generation in Brazilian Portuguese. |
| Outcome: | The proposed approaches are based on the Abstract Meaning Representation formalism and have potential future directions. |
PortiLexicon-UD: a Portuguese Lexical Resource according to Universal Dependencies Model (2022.lrec-1)
Copied to clipboard
| Challenge: | lexical resource for Brazilian Portuguese with 1,221,218 entries, according to the Universal Dependencies model and guidelines. |
| Approach: | They propose to build a large and freely available lexicon for Portuguese that delivers morphosyntactic information according to the Universal Dependencies model. |
| Outcome: | The proposed lexical resource has high language coverage and good quality data. |
Semantically Inspired AMR Alignment for the Portuguese Language (2020.emnlp-main)
Copied to clipboard
| Challenge: | Abstract Meaning Representation (AMR) parsers require alignment between nodes and words of the sentence. |
| Approach: | They propose to use a more semantically matched word-concept pair to align graphs with words in Portuguese . they performed intrinsic and extrinsic evaluations and found it outperforms the English alignment strategies. |
| Outcome: | The proposed method outperforms the existing methods for English and achieves competitive results with a parser designed for the Portuguese language. |
Towards AMR-BR: A SemBank for Brazilian Portuguese Language (L18-1)
Copied to clipboard
| Challenge: | Abstract Meaning Representation (AMR) is a recent and prominent meaning representation with good acceptance and several applications in the Natural Language Processing area. |
| Approach: | They propose to build an AMR annotated corpus for Brazilian Portuguese using an alignment-based approach. |
| Outcome: | The proposed corpus is based on the Little Prince book, which went into the public domain and explored some language-specific annotation issues. |
HateBR: A Large Expert Annotated Corpus of Brazilian Instagram Comments for Offensive Language and Hate Speech Detection (2022.lrec-1)
Copied to clipboard
| Challenge: | In Brazil, hate speech is prohibited, however the regulation is not effective due to the difficulty of identifying, quantifying and classifying this kind of online content. |
| Approach: | They propose to annotate a large corpus of Brazilian Instagram comments manually and to use it to detect hate speech and offensive language. |
| Outcome: | The HateBR corpus was collected from the comment section of Brazilian politicians’ accounts on Instagram and manually annotated by specialists, reaching a high inter-annotator agreement. |
Measuring the Impact of Readability Features in Fake News Detection (2020.lrec-1)
Copied to clipboard
Roney Santos, Gabriela Pedro, Sidney Leal, Oto Vale, Thiago Pardo, Kalina Bontcheva, Carolina Scarton
| Challenge: | Recent efforts to detect fake news use language-based approaches to detect news articles . authors show that readability features can improve classification accuracy . |
| Approach: | They propose to use readability features to detect fake news in the Brazilian Portuguese language . they show that such features can achieve up to 92% classification accuracy . |
| Outcome: | The proposed features achieve up to 92% accuracy and may improve previous classification results. |
Rhetorical Structure Approach for Online Deception Detection: A Survey (2022.lrec-1)
Copied to clipboard
| Challenge: | Existing studies on how people use language to inform and misinform are relevant. |
| Approach: | They analyze how discourse structure is applied to fake news detection on the web and social media. |
| Outcome: | The proposed framework is applied to fake news and fake reviews detection on the web and social media. |
Back-Translation as Strategy to Tackle the Lack of Corpus in Natural Language Generation from Semantic Representations (D19-63)
Copied to clipboard
| Challenge: | Abstract Meaning Representation and Brazilian Portuguese (BP) are selected as semantic representation and language, respectively. |
| Approach: | They propose to use Brazilian Portuguese and Abstract Meaning Representation as semantic representations for NLG. |
| Outcome: | The proposed methods were evaluated on two datasets (one automatically generated and another human-generated) to compare the performance in a real context. |