Challenge: Existing toxic content detection methods focus on sentence-level classification but fail to provide readable and contiguous toxic evidence spans.
Approach: They propose an explainability-oriented method for Chinese toxic content detection methods . they refine saliency cues into fine-grained toxic spans with lightweight LLM guidance .
Outcome: The proposed method improves classification accuracy and toxic span extraction while preserving efficient encoder-based inference and producing more coherent explanations.

Similar Papers

TOXIFRENCH: Benchmarking and Enhancing Language Models via CoT Fine-Tuning for French Toxicity Detection (2026.findings-acl)

Copied to clipboard

Challenge: toxicity detection in French remains underdeveloped due to the lack of culturally relevant, human-annotated, large-scale datasets.
Approach: They propose a method that generalizes French online comments using a semi-automated annotation pipeline that reduces manual labeling to only 10% through high-confidence LLM-based pre-annotation and human verification.
Outcome: The proposed model outperforms GPT-4o and DeepSeek-R1 on the benchmark while maintaining cross-lingual capabilities.
ToxiGAN: Toxic Data Augmentation via LLM-Guided Directional Adversarial Generation (2026.eacl-long)

Copied to clipboard

Challenge: Generating toxic text for augmentation is a sensitive and challenging task . toxic language remains pervasive, often escaping automated moderation systems .
Approach: They propose a class-aware text augmentation framework that combines adversarial generation with semantic guidance from large language models to provide balanced guidance.
Outcome: Experiments on four hate speech benchmarks show that ToxiGAN outperforms traditional and LLM-based augmentation methods.
ToXCL: A Unified Framework for Toxic Speech Detection and Explanation (2024.naacl-long)

Copied to clipboard

Challenge: Existing models that focus on explicit toxic speech detection and explanation are prone to error propagation problems . et al., 2018) show that toxic speech models can be prone for generating errors .
Approach: They propose a framework that can detect and explain toxic speech using a target group generator and an encoder-decoder model.
Outcome: The proposed model outperforms baseline models and achieves state-of-the-art effectiveness . the proposed model generates a toxic explanation that matches the ground truth explanation .
ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection (2022.acl-long)

Copied to clipboard

Challenge: Toxic language detection systems often falsely flag text that contains minority group mentions as toxic . this over-reliance on spurious correlations also causes systems to struggle with detecting implicitly toxic language.
Approach: They develop a machine-generated dataset of toxic and benign statements about 13 minority groups that generates subtly toxic and harmless text with a massive pretrained language model.
Outcome: The proposed method can detect toxic and benign statements on a large scale . it can also detect hate speech on 94.5% of the toxic examples .
Facilitating Fine-grained Detection of Chinese Toxic Language: Hierarchical Taxonomy, Resources, and Benchmarks (2023.acl-long)

Copied to clipboard

Challenge: Existing datasets suffer from a lack of fine-grained annotations, such as the toxic type and expressions with indirect toxicity.
Approach: They propose a benchmark model to detect toxic language by incorporating lexical features into a Chinese dataset to facilitate fine-grained annotations.
Outcome: The proposed model is based on insulting vocabulary containing implicit profanity and is able to detect toxic language with lexical features.
Penetrating Linguistic Disguises: A Slang-aware Label-Aligned Framework for Fine-Grained Toxicity Extraction in Chinese Hate Speech Detection (2026.findings-acl)

Copied to clipboard

Challenge: Flexible word boundaries and linguistic obfuscation, particularly slang, challenge precise span-level hate speech detection in Chinese.
Approach: They propose a Slang-aware Label-Aligned Framework that maps slang to explicit hate semantics and uses task-specific branches to mitigate feature interference.
Outcome: The proposed framework reduces ambiguity by mapping obscure slang to explicit hate semantics.
ToVo: Toxicity Taxonomy via Voting (2025.findings-naacl)

Copied to clipboard

Challenge: Existing toxic content detection models face limitations due to the closed-source nature of training data and the paucity of explanations for their evaluation mechanism.
Approach: They propose a mechanism that integrates voting and chain-of-thought processes to produce a high-quality open-source dataset for toxic content detection.
Outcome: The proposed model improves transparency and customizability while facilitating better fine-tuning for specific use cases.
Adaptive Detoxification: Safeguarding General Capabilities of LLMs through Toxicity-Aware Knowledge Editing (2025.findings-acl)

Copied to clipboard

Challenge: Existing knowledge editing methods for large language models (LLMs) suffer from over-editing, where detoxified models reject legitimate queries, compromising overall performance.
Approach: They propose a toxicity-aware knowledge editing approach that dynamically detects toxic activation patterns during forward propagation and then routes computations through adaptive inter-layer pathways to mitigate toxicity effectively.
Outcome: The proposed method outperforms existing methods on large language models and enhances the SafeEdit benchmark.
DetoxLLM: A Framework for Detoxification with Explanations (2024.emnlp-main)

Copied to clipboard

Challenge: DetoxLLM is a comprehensive end-to-end detoxification framework for toxic language.
Approach: They propose a comprehensive end-to-end detoxification framework that tackles toxic language across platforms.
Outcome: The proposed detoxification framework outperforms the SoTA model on human-annotated parallel corpus and offers explanation to promote transparency and trustworthiness.
ToxVidLM: A Multimodal Framework for Toxicity Detection in Code-Mixed Videos (2024.findings-acl)

Copied to clipboard

Challenge: Using a dataset of 931 videos with 4021 code-mixed Hindi-English utterances, we find that video content with multiple modalities is more accurate and more accurate than textual content.
Approach: They propose to use a dataset to analyze toxic content in video content in non-English languages by leveraging language models.
Outcome: The proposed framework achieves an Accuracy and Weighted F1 score of 94.29% and 94.35% for the first time in its class.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations