Papers by Mattia Samory
Counterfactually Augmented Data and Unintended Bias: The Case of Sexism and Hate Speech Detection (2022.naacl-main)
Copied to clipboard
| Challenge: | sexism and hate speech detection models may be over-relying on core features . construct-driven CAD may induce models to ignore context in which core features are used . |
| Approach: | They propose to use construct-driven and construct-agnostic CAD to reduce model bias . sexism and hate speech detection models are trained on counterfactually augmented data . |
| Outcome: | Using a diverse set of CAD—construct-driven and construct-agnostic—reduces unintended bias. |
Robustness and Confounders in the Demographic Alignment of LLMs with Human Perceptions of Offensiveness (2025.findings-acl)
Copied to clipboard
| Challenge: | despite evidence of demographic bias, reports with whom they align best are hard to generalize or contradictory . confounders introduced in the annotation process account for more variation in alignment patterns than demographic traits . |
| Approach: | They examine the alignment of large language models with human annotations in offensive language datasets. |
| Outcome: | The results show that LLMs align better with human annotations than other models. |
People Make Better Edits: Measuring the Efficacy of LLM-Generated Counterfactually Augmented Data for Harmful Language Detection (2023.emnlp-main)
Copied to clipboard
| Challenge: | Past work has shown that counterfactually augmented data (CADs) can improve models' performance on out-of-domain tests. |
| Approach: | They use Polyjuice, ChatGPT, and Flan-T5 to automatically generate CADs . they find that CAD generates a model that flips the original label with minimal changes . |
| Outcome: | The proposed model improves model robustness on out-of-domain test sets and individual data points. |
How Does Counterfactually Augmented Data Impact Models for Social Computing Constructs? (2021.emnlp-main)
Copied to clipboard
| Challenge: | Recent studies have shown that models trained on CAD can learn cues in the dataset which are spuriously correlated with the construct. |
| Approach: | They focus on sentiment, sexism, and hate speech as social constructs to investigate their effects on model performance. |
| Outcome: | The proposed model generalizes better on out-of-domain datasets while relying less on spurious features. |