Challenge: toxicity annotations are often ignored because of its subjective nature and lack of nuance.
Approach: They examine the effect of annotator identities and beliefs on toxic language annotations by considering posts with three characteristics: anti-Black language, African American English (AAE) dialect, and vulgarity.
Outcome: The findings show strong associations between annotator identity and beliefs and ratings of toxicity.

Similar Papers

The Risk of Racial Bias in Hate Speech Detection (P19-1)

Copied to clipboard

Challenge: Annotators’ insensitivity to differences in dialect can lead to racial bias in automatic hate speech detection models, potentially amplifying harm against minority populations.
Approach: They propose *dialect* and *race priming* as ways to reduce the racial bias in hate speech detection models by detecting differences in dialects in annotated tweets.
Outcome: The proposed models acquire and propagate these biases, such that AAE tweets and tweets by self-identified African Americans are up to two times more likely to be labelled as offensive compared to others.
Challenges in Automated Debiasing for Toxic Language Detection (2021.eacl-main)

Copied to clipboard

Challenge: Existing methods for debiasing toxic language data are limited in their ability to prevent biased behavior in toxic language detection systems.
Approach: They propose to debiase toxic language detection models using lexical and dialectal markers using synthetic labels instead of traditional methods.
Outcome: The proposed method reduces dialectal associations with toxicity despite the use of synthetic labels .
On the Role of Speech Data in Reducing Toxicity Detection Bias (2025.naacl-long)

Copied to clipboard

Challenge: Text toxicity detection systems produce disproportionate rates of false positives on demographic groups . toxicity classification systems often misinterpret benign group mentions as toxic .
Approach: They use group annotations to compare text-based and speech-based toxicity detection systems.
Outcome: The results show that access to speech data supports reduced bias against group mentions . the authors recommend improving classifiers, rather than transcription pipelines if possible .
Unveiling Identity Biases in Toxicity Detection : A Game-Focused Dataset and Reactivity Analysis Approach (2023.emnlp-industry)

Copied to clipboard

Challenge: Existing datasets focused on gender or racial biases are not designed for the gaming industry, a concern for models built for toxicity detection in videogames’ written chat.
Approach: They propose to use reactivity analysis to highlight oversensitive terms using a language model developed by Ubisoft for toxicity detection on videogame’s written chat and Perspective API to generate a list of terms that trigger the models to varying degrees.
Outcome: The proposed model can detect and amplify identity biases in annotated language models and is compared with a language model developed by Ubisoft for toxicity detection on videogames’ written chat and Perspective API.
Accurate and Data-Efficient Toxicity Prediction when Annotators Disagree (2024.emnlp-main)

Copied to clipboard

Challenge: Disagreement among annotators can reveal nuances in subjective tasks that lack a simple ground truth .
Approach: They propose three approaches to predict annotator ratings on the toxicity of text . they integrate annotators' history, demographics, survey information into their models .
Outcome: The proposed approach outperforms other methods in toxicity rating prediction.
Toxic, Hateful, Offensive or Abusive? What Are We Really Classifying? An Empirical Analysis of Hate Speech Datasets (2020.lrec-1)

Copied to clipboard

Challenge: a recent study shows that many definitions are being used for equivalent concepts, making most datasets incompatible.
Approach: They analyze six publicly available datasets to determine their similarity and compatibility . they propose to use Fast Text word vectors to analyze similarity between different datasets .
Outcome: The proposed model performs better on similar datasets and worse on more non-offensive samples.
Re-examining Sexism and Misogyny Classification with Annotator Attitudes (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing datasets for content moderation fail to capture plurality of possible annotator perspectives or ensure representation of affected groups.
Approach: They examine the relationship between annotator identities and attitudes and the responses they give to two GBV labelling tasks.
Outcome: The results show that higher Right Wing Authoritarianism scores are associated with a higher propensity to label text as sexist . higher scores are also associated with negative attitudes towards sexism and neosexist attitudes .
Robustness and Confounders in the Demographic Alignment of LLMs with Human Perceptions of Offensiveness (2025.findings-acl)

Copied to clipboard

Challenge: despite evidence of demographic bias, reports with whom they align best are hard to generalize or contradictory . confounders introduced in the annotation process account for more variation in alignment patterns than demographic traits .
Approach: They examine the alignment of large language models with human annotations in offensive language datasets.
Outcome: The results show that LLMs align better with human annotations than other models.
Explaining Toxic Text via Knowledge Enhanced Text Generation (2022.naacl-main)

Copied to clipboard

Challenge: Existing work on toxic speech classification relies on generic and repetitive explanations . elucidating toxic speech can help with downstream tasks such as debiasing .
Approach: They propose a knowledge-informed encoder-decoder framework to generate toxic text explanations . they use multiple knowledge sources to generate detailed explanations of toxic text .
Outcome: The proposed model outperforms state-of-the-art models significantly in generating toxic explanations . the proposed model can generate detailed explanations of toxic speech compared to baselines compared with baseline models .
ModelCitizens: Representing Community Voices in Online Safety (2025.emnlp-main)

Copied to clipboard

Challenge: Existing toxic language detection models are trained on annotations that collapse diverse perspectives into a single ground truth.
Approach: They propose to augment social media posts with conversational scenarios to reflect the impact of conversational context on toxicity.
Outcome: The proposed model outperforms existing models on social media with conversational scenarios.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations