Challenge: Large Language Models (LLMs) generate ethically inappropriate texts even for seemingly innocuous contexts.
Approach: They propose to use large language models to detect and filter toxic content in text prediction tasks by evaluating their toxicity detection approaches against a manually crafted CheckList of harms.
Outcome: The proposed methods are compared against a checklist of harms targeted at different groups and different levels of severity in English.

Similar Papers

Challenges in Detoxifying Language Models (2021.findings-emnlp)

Copied to clipboard

Challenge: Prior work often relies on automatic evaluation of LM toxicity.
Approach: They evaluate toxicity mitigation strategies for automated and human evaluations . they find human raters disagree with high automatic toxicity scores after strong toxicity reduction interventions .
Outcome: The proposed methods reduce LM toxicity but lower coverage for marginalized texts . human raters disagree with high toxicity scores after strong toxicity reduction interventions .
A Survey of Toxicity Mitigation Strategies for Multilingual Language Models (2026.findings-acl)

Copied to clipboard

Challenge: Large language models can reproduce and amplify toxic content, including hate speech, harassment, and bias.
Approach: They propose a comprehensive survey of the many detoxification methods tailored to multilingual LLMs.
Outcome: The proposed methods are based on data filtering, style transfer, expert-based logit steering, retrieval augmentation, and human feedback.
Realistic Evaluation of Toxicity in Large Language Models (2024.findings-acl)

Copied to clipboard

Challenge: a large amount of data exposes large language models to toxicity and bias . prompt engineering can be easily bypassed with minimal prompt engineering.
Approach: They propose a dataset that uses manually crafted prompts to nullify protective layers of large language models.
Outcome: The proposed dataset shows that prompts can nullify protective layers of large language models.
Towards Building a Robust Toxicity Predictor (2023.acl-industry)

Copied to clipboard

Challenge: Recent studies have focused on robustness of toxicity language predictors, but this is problematic for real-world toxicity detection.
Approach: They propose a novel adversarial attack that exploits greedy search strategies to fool toxic text classifiers.
Outcome: The proposed attack can detect weaker toxicity language detectors even against unseen attacks.
On the Role of Speech Data in Reducing Toxicity Detection Bias (2025.naacl-long)

Copied to clipboard

Challenge: Text toxicity detection systems produce disproportionate rates of false positives on demographic groups . toxicity classification systems often misinterpret benign group mentions as toxic .
Approach: They use group annotations to compare text-based and speech-based toxicity detection systems.
Outcome: The results show that access to speech data supports reduced bias against group mentions . the authors recommend improving classifiers, rather than transcription pipelines if possible .
Probing Toxic Content in Large Pre-Trained Language Models (2021.acl-long)

Copied to clipboard

Challenge: Existing studies on pre-trained language models have shown that they carry harmful biases towards different social groups.
Approach: They propose a method to probe English, French, and Arabic PTLMs and quantify the potentially harmful content they convey with respect to a set of templates.
Outcome: The proposed method analyzes PTLMs to predict masked tokens at the end of sentences to assess their toxicity.
From One to Many: Expanding the Scope of Toxicity Mitigation in Language Models (2024.findings-acl)

Copied to clipboard

Challenge: toxicity mitigation in language models has been focused on single-language settings . however, widespread adoption of LLMs has introduced a range of unknown -harms .
Approach: They employ translated data to evaluate and enhance mitigation techniques in the absence of sufficient annotated datasets across languages.
Outcome: The proposed approach compares translation quality and retrieval-augmented mitigation techniques under static and continual toxicity mitigation scenarios.
Mitigating Societal Harms in Large Language Models (2023.emnlp-tutorial)

Copied to clipboard

Challenge: Recent studies have highlighted societal harms that can be caused by language generation models deployed in the wild.
Approach: They propose to use a typology of technical approaches to mitigating harms of language generation models to provide an overview of potential social issues in language generation including toxicity, social biases, misinformation, factual inconsistency, and privacy violations.
Outcome: The proposed typology addresses toxicity, biases, misinformation, factual inconsistency, and privacy violations in language generation models.
♪ Something Just Like TRuST ♪ *: Toxicity Recognition of Span and Target (2026.findings-acl)

Copied to clipboard

Challenge: Toxic language is pervasive online, and because LLMs are trained on web data, it generates such content.
Approach: They propose a large-scale dataset that synthesizes toxicity definitions and an annotation scheme . they use a rigorous human annotation process to evaluate the diversity of the annotations .
Outcome: The proposed model outperforms existing models on three tasks and is not reliable.
Transfer Learning and Prediction Consistency for Detecting Offensive Spans of Text (2022.findings-acl)

Copied to clipboard

Challenge: Existing models for toxic span detection only classify text snippets as offensive or not . a novel model seeks to simultaneously predict offensive words and opinion phrases .
Approach: They propose a novel model that seeks to predict offensive words and opinion phrases simultaneously . they also introduce a regularization mechanism to encourage consistency of the model predictions .
Outcome: The proposed model performs well compared to baselines on toxic span detection tasks . it predicts offensive words and opinion phrases to leverage inter-dependencies .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations