Papers by Francielle Vargas
MFTCXplain: A Multilingual Benchmark Dataset for Evaluating the Moral Reasoning of LLMs through Multi-hop Hate Speech Explanation (2025.findings-emnlp)
Copied to clipboard
Jackson Trager, Francielle Vargas, Diego Alves, Matteo Guida, Mikel K. Ngueajio, Ameeta Agrawal, Yalda Daryani, Farzan Karimi Malekabadi, Flor Miriam Plaza-del-Arco
| Challenge: | Existing evaluation benchmarks for large language models lack annotations that justify moral classifications and focus on English constrain moral reasoning across diverse cultural settings. |
| Approach: | They propose a multilingual benchmark dataset for evaluating moral reasoning of large language models . it includes 3,000 tweets annotated with binary hate speech labels, moral categories and rationales . |
| Outcome: | The proposed dataset shows a misalignment between LLM outputs and human annotations in moral reasoning tasks. |
Self-Explaining Hate Speech Detection with Moral Rationales (2026.findings-acl)
Copied to clipboard
Francielle Vargas, Jackson Trager, Diego Alves, Matteo Guida, Surendrabikram Thapa, Berk Atıl, Daryna Dementieva, Andrew J Smart, Ameeta Agrawal
| Challenge: | Existing models for hate speech detection are opaque and rely on surface-level cues. Existing approaches often encode biases originating from training data and annotation processes. |
| Approach: | They propose a framework that integrates moral rationale supervision into training . they propose SMRA for self-explaining hate speech detection . |
| Outcome: | The proposed framework improves performance across binary hate speech detection and multi-label moral sentiment classification. |
HateBR: A Large Expert Annotated Corpus of Brazilian Instagram Comments for Offensive Language and Hate Speech Detection (2022.lrec-1)
Copied to clipboard
| Challenge: | In Brazil, hate speech is prohibited, however the regulation is not effective due to the difficulty of identifying, quantifying and classifying this kind of online content. |
| Approach: | They propose to annotate a large corpus of Brazilian Instagram comments manually and to use it to detect hate speech and offensive language. |
| Outcome: | The HateBR corpus was collected from the comment section of Brazilian politicians’ accounts on Instagram and manually annotated by specialists, reaching a high inter-annotator agreement. |
Rhetorical Structure Approach for Online Deception Detection: A Survey (2022.lrec-1)
Copied to clipboard
| Challenge: | Existing studies on how people use language to inform and misinform are relevant. |
| Approach: | They analyze how discourse structure is applied to fake news detection on the web and social media. |
| Outcome: | The proposed framework is applied to fake news and fake reviews detection on the web and social media. |
HateBRXplain: A Benchmark Dataset with Human-Annotated Rationales for Explainable Hate Speech Detection in Brazilian Portuguese (2025.coling-main)
Copied to clipboard
| Challenge: | Hate speech detection systems have been developed to inhibit offensive and hateful language from being published or spread on the Web and social media. |
| Approach: | They propose to use a Portuguese dataset to provide rationales for hate speech detection with text span annotations. |
| Outcome: | The proposed models outperform the baselines in Portuguese and showed that they provide plausible explanations when compared to human annotations. |