Challenge: Existing datasets focused on gender or racial biases are not designed for the gaming industry, a concern for models built for toxicity detection in videogames’ written chat.
Approach: They propose to use reactivity analysis to highlight oversensitive terms using a language model developed by Ubisoft for toxicity detection on videogame’s written chat and Perspective API to generate a list of terms that trigger the models to varying degrees.
Outcome: The proposed model can detect and amplify identity biases in annotated language models and is compared with a language model developed by Ubisoft for toxicity detection on videogames’ written chat and Perspective API.

Similar Papers

Towards Detecting Contextual Real-Time Toxicity for In-Game Chat (2023.findings-emnlp)

Copied to clipboard

Challenge: ToxBuster is a simple and scalable model that reliably detects toxic content in real-time for a line of chat by including chat history and metadata.
Approach: They propose a model that detects toxic content in real-time for a line of chat by including chat history and metadata.
Outcome: The proposed model outperforms conventional toxicity models across popular multiplayer games including Rainbow Six Siege, For Honor, and DOTA 2 and 6% of unreported toxic players can be proactively moderated.
Annotators with Attitudes: How Annotator Beliefs And Identities Bias Toxic Language Detection (2022.naacl-main)

Copied to clipboard

Challenge: toxicity annotations are often ignored because of its subjective nature and lack of nuance.
Approach: They examine the effect of annotator identities and beliefs on toxic language annotations by considering posts with three characteristics: anti-Black language, African American English (AAE) dialect, and vulgarity.
Outcome: The findings show strong associations between annotator identity and beliefs and ratings of toxicity.
ModelCitizens: Representing Community Voices in Online Safety (2025.emnlp-main)

Copied to clipboard

Challenge: Existing toxic language detection models are trained on annotations that collapse diverse perspectives into a single ground truth.
Approach: They propose to augment social media posts with conversational scenarios to reflect the impact of conversational context on toxicity.
Outcome: The proposed model outperforms existing models on social media with conversational scenarios.
GameTox: A Comprehensive Dataset and Analysis for Enhanced Toxicity Detection in Online Gaming Communities (2025.naacl-short)

Copied to clipboard

Challenge: Existing methods to detect toxic behavior in online gaming environments are limited by utterance-level annotation.
Approach: They propose to annotate game chat utterances for toxicity detection through intent classification and slot filling.
Outcome: The proposed model improves the detection of toxic speech in online gaming environments and reveals limitations of current models.
Hollywood Identity Bias Dataset: A Context Oriented Bias Analysis of Movie Dialogues (2022.lrec-1)

Copied to clipboard

Challenge: Movies reflect society and also hold power to transform opinions.
Approach: They propose to annotate movie scripts for identity bias using a dataset that is annotated for gender, race/ethnicity, religion, age, occupation, LGBTQ, and other .
Outcome: The proposed dataset contains dialogue turns annotated for gender, race/ethnicity, religion, age, occupation, LGBTQ, and other, which contains biases like body shaming, personality bias, etc.
ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection (2022.acl-long)

Copied to clipboard

Challenge: Toxic language detection systems often falsely flag text that contains minority group mentions as toxic . this over-reliance on spurious correlations also causes systems to struggle with detecting implicitly toxic language.
Approach: They develop a machine-generated dataset of toxic and benign statements about 13 minority groups that generates subtly toxic and harmless text with a massive pretrained language model.
Outcome: The proposed method can detect toxic and benign statements on a large scale . it can also detect hate speech on 94.5% of the toxic examples .
Centering the Margins: Outlier-Based Identification of Harmed Populations in Toxicity Detection (2023.emnlp-main)

Copied to clipboard

Challenge: toxicity detection models focus on marginalized groups, but they obscure harms faced by intersectional subgroups.
Approach: They use outlier detection to identify text about people with demographic attributes distant from the "norm" they find model performance is worse for demographic outliers than non-outliers .
Outcome: The proposed model performance is worse for outliers than non-outliers, the authors say . their analysis also shows that outlier analysis can identify harms faced by intersectional groups .
CONDA: a CONtextual Dual-Annotated dataset for in-game toxicity understanding and detection (2021.findings-acl)

Copied to clipboard

Challenge: Existing toxic language detection models focus on the single utterance level without deeper understanding of context.
Approach: They propose a dataset for in-game toxic language detection enabling joint intent classification and slot filling analysis, which is the core task of Natural Language Understanding (NLU).
Outcome: The proposed framework handles utterance and token-level patterns, and rich contextual chatting history.
On the Role of Speech Data in Reducing Toxicity Detection Bias (2025.naacl-long)

Copied to clipboard

Challenge: Text toxicity detection systems produce disproportionate rates of false positives on demographic groups . toxicity classification systems often misinterpret benign group mentions as toxic .
Approach: They use group annotations to compare text-based and speech-based toxicity detection systems.
Outcome: The results show that access to speech data supports reduced bias against group mentions . the authors recommend improving classifiers, rather than transcription pipelines if possible .
Challenges in Automated Debiasing for Toxic Language Detection (2021.eacl-main)

Copied to clipboard

Challenge: Existing methods for debiasing toxic language data are limited in their ability to prevent biased behavior in toxic language detection systems.
Approach: They propose to debiase toxic language detection models using lexical and dialectal markers using synthetic labels instead of traditional methods.
Outcome: The proposed method reduces dialectal associations with toxicity despite the use of synthetic labels .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations