Challenge: Neologisms can foster new linguistic consensus by stabilizing shared meanings and usage in common communicative norms.
Approach: They propose a taxonomy that captures the origins and consensus-verification criteria of toxic neologisms . they propose 'SeTox' framework that integrates real-time web context for naeologim detection .
Outcome: The proposed framework outperforms large-scale models in detecting neologism toxicity.

Similar Papers

Enhancing LLM-based Hatred and Toxicity Detection with Meta-Toxic Knowledge Graph (2025.findings-acl)

Copied to clipboard

Challenge: Existing methods to address toxicity issues with large language models are inadequate . lack of domain-specific knowledge leads to false negatives and excessive sensitivity to toxic speech limits freedom of speech.
Approach: They propose a method that leverages graph search on a meta-toxic knowledge graph to enhance hatred and toxicity detection.
Outcome: The proposed method lowers false positive rate and improves toxicity detection performance in out-of-domain scenarios.
Exploring Multimodal Challenges in Toxic Chinese Detection: Taxonomy, Benchmark, and Findings (2025.findings-acl)

Copied to clipboard

Challenge: Recent studies show that character substitutions in toxic Chinese text can confuse state-of-the-art LLMs.
Approach: They propose a taxonomy of 3 perturbation strategies and 8 specific approaches in Chinese text to assess if they can detect perturbed Chinese toxic contents.
Outcome: The proposed model can detect perturbed Chinese text with 8 different approaches . the proposed model is compared with 9 other LLMs from the US and China .
♪ Something Just Like TRuST ♪ *: Toxicity Recognition of Span and Target (2026.findings-acl)

Copied to clipboard

Challenge: Toxic language is pervasive online, and because LLMs are trained on web data, it generates such content.
Approach: They propose a large-scale dataset that synthesizes toxicity definitions and an annotation scheme . they use a rigorous human annotation process to evaluate the diversity of the annotations .
Outcome: The proposed model outperforms existing models on three tasks and is not reliable.
Facilitating Fine-grained Detection of Chinese Toxic Language: Hierarchical Taxonomy, Resources, and Benchmarks (2023.acl-long)

Copied to clipboard

Challenge: Existing datasets suffer from a lack of fine-grained annotations, such as the toxic type and expressions with indirect toxicity.
Approach: They propose a benchmark model to detect toxic language by incorporating lexical features into a Chinese dataset to facilitate fine-grained annotations.
Outcome: The proposed model is based on insulting vocabulary containing implicit profanity and is able to detect toxic language with lexical features.
Chinese Toxic Language Mitigation via Sentiment Polarity Consistent Rewrites (2025.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) can be effective at rewriting toxic content, but they often default to overly polite rewrites, distorting the emotional tone and communicative intent.
Approach: They evaluate 17 large language models with variant architectures to evaluate their ability to rewrite toxic content while preserving the speaker's original intent.
Outcome: The first Chinese detoxification dataset explicitly designed to preserve sentiment polarity is evaluated across five real-world scenarios.
Can LLMs Recognize Toxicity? A Structured Investigation Framework and Toxicity Metric (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing toxicity metrics rely on encoder models trained on specific toxicity datasets, which are susceptible to out-of-distribution (OOD) problems and depend on the dataset’s definition of toxicity.
Approach: They propose a robust metric grounded on LLMs to flexibly measure toxicity according to the given definition by analysing toxicity factors and intrinsic toxic attributes.
Outcome: The proposed metric improves on conventional metrics by 12 points in the F1 score and shows that upstream toxicity significantly influences downstream metrics, suggesting that LLMs are unsuitable for toxicity evaluations within unverified factors.
Nuanced Toxicity Detection in Spanish: A New Corpus and Benchmark Study (2026.findings-eacl)

Copied to clipboard

Challenge: Existing corpora for Spanish are under-resourced for toxic content detection . sarcasm, indirect aggression, irony, and other toxicity are not detected in English .
Approach: They propose to extend the NECOS-TOX corpus to include 4,011 Spanish comments . each comment is annotated across three levels of toxicity, with substantial inter-annotator agreement .
Outcome: The proposed model performs on par with larger models and is released publicly . the proposed model is based on a human-in-the-loop active learning strategy .
A Survey of Toxicity Mitigation Strategies for Multilingual Language Models (2026.findings-acl)

Copied to clipboard

Challenge: Large language models can reproduce and amplify toxic content, including hate speech, harassment, and bias.
Approach: They propose a comprehensive survey of the many detoxification methods tailored to multilingual LLMs.
Outcome: The proposed methods are based on data filtering, style transfer, expert-based logit steering, retrieval augmentation, and human feedback.
ModelCitizens: Representing Community Voices in Online Safety (2025.emnlp-main)

Copied to clipboard

Challenge: Existing toxic language detection models are trained on annotations that collapse diverse perspectives into a single ground truth.
Approach: They propose to augment social media posts with conversational scenarios to reflect the impact of conversational context on toxicity.
Outcome: The proposed model outperforms existing models on social media with conversational scenarios.
MICo: Preventative Detoxification of Large Language Models through Inhibition Control (2024.findings-naacl)

Copied to clipboard

Challenge: Large Language Models (LLMs) have a tendency to devolve into toxic degeneration . model may classify prompts as toxic or non-toxic and categorically refuse to respond to those deemed toxic.
Approach: They propose a mechanism for LLM detoxification by labeling acceptable and unacceptable examples and including a corresponding acceptable rewrite with every unacceptable example.
Outcome: The proposed model improves on the baseline model and shows that it detects and rewrites toxic and harmful examples.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations