Papers by Mattia Samory

4 papers
Counterfactually Augmented Data and Unintended Bias: The Case of Sexism and Hate Speech Detection (2022.naacl-main)

Copied to clipboard

Challenge: sexism and hate speech detection models may be over-relying on core features . construct-driven CAD may induce models to ignore context in which core features are used .
Approach: They propose to use construct-driven and construct-agnostic CAD to reduce model bias . sexism and hate speech detection models are trained on counterfactually augmented data .
Outcome: Using a diverse set of CAD—construct-driven and construct-agnostic—reduces unintended bias.
Robustness and Confounders in the Demographic Alignment of LLMs with Human Perceptions of Offensiveness (2025.findings-acl)

Copied to clipboard

Challenge: despite evidence of demographic bias, reports with whom they align best are hard to generalize or contradictory . confounders introduced in the annotation process account for more variation in alignment patterns than demographic traits .
Approach: They examine the alignment of large language models with human annotations in offensive language datasets.
Outcome: The results show that LLMs align better with human annotations than other models.
People Make Better Edits: Measuring the Efficacy of LLM-Generated Counterfactually Augmented Data for Harmful Language Detection (2023.emnlp-main)

Copied to clipboard

Challenge: Past work has shown that counterfactually augmented data (CADs) can improve models' performance on out-of-domain tests.
Approach: They use Polyjuice, ChatGPT, and Flan-T5 to automatically generate CADs . they find that CAD generates a model that flips the original label with minimal changes .
Outcome: The proposed model improves model robustness on out-of-domain test sets and individual data points.
How Does Counterfactually Augmented Data Impact Models for Social Computing Constructs? (2021.emnlp-main)

Copied to clipboard

Challenge: Recent studies have shown that models trained on CAD can learn cues in the dataset which are spuriously correlated with the construct.
Approach: They focus on sentiment, sexism, and hate speech as social constructs to investigate their effects on model performance.
Outcome: The proposed model generalizes better on out-of-domain datasets while relying less on spurious features.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations