Fortifying Toxic Speech Detectors Against Veiled Toxicity (2020.emnlp-main)

Copied to clipboard

Challenge: Modern toxic speech detectors are incompetent in recognizing disguised offensive language, such as adversarial attacks that deliberately avoid known toxic lexicons.
Approach: They propose a framework that fortifies existing toxic speech detectors without a large labeled corpus of veiled toxicity.
Outcome: The proposed framework is aimed at fortifying existing toxic speech detectors without a large labeled corpus of disguised offensive language.

Similar Papers

On the Role of Speech Data in Reducing Toxicity Detection Bias (2025.naacl-long)

Copied to clipboard

Challenge: Text toxicity detection systems produce disproportionate rates of false positives on demographic groups . toxicity classification systems often misinterpret benign group mentions as toxic .
Approach: They use group annotations to compare text-based and speech-based toxicity detection systems.
Outcome: The results show that access to speech data supports reduced bias against group mentions . the authors recommend improving classifiers, rather than transcription pipelines if possible .
Towards Building a Robust Toxicity Predictor (2023.acl-industry)

Copied to clipboard

Challenge: Recent studies have focused on robustness of toxicity language predictors, but this is problematic for real-world toxicity detection.
Approach: They propose a novel adversarial attack that exploits greedy search strategies to fool toxic text classifiers.
Outcome: The proposed attack can detect weaker toxicity language detectors even against unseen attacks.
ToxiCloakCN: Evaluating Robustness of Offensive Language Detection in Chinese with Cloaking Perturbations (2024.emnlp-main)

Copied to clipboard

Challenge: Existing large language models struggle with systematically perturbed data designed to evade detection mechanisms.
Approach: They propose a large language model with homophonic substitutions and emoji transformations to test their models' robustness against cloaking perturbations.
Outcome: The proposed model underperforms in detecting offensive content when perturbations are applied to Chinese language datasets.
Challenges in Automated Debiasing for Toxic Language Detection (2021.eacl-main)

Copied to clipboard

Challenge: Existing methods for debiasing toxic language data are limited in their ability to prevent biased behavior in toxic language detection systems.
Approach: They propose to debiase toxic language detection models using lexical and dialectal markers using synthetic labels instead of traditional methods.
Outcome: The proposed method reduces dialectal associations with toxicity despite the use of synthetic labels .
♪ Something Just Like TRuST ♪ *: Toxicity Recognition of Span and Target (2026.findings-acl)

Copied to clipboard

Challenge: Toxic language is pervasive online, and because LLMs are trained on web data, it generates such content.
Approach: They propose a large-scale dataset that synthesizes toxicity definitions and an annotation scheme . they use a rigorous human annotation process to evaluate the diversity of the annotations .
Outcome: The proposed model outperforms existing models on three tasks and is not reliable.
Adding Instructions during Pretraining: Effective way of Controlling Toxicity in Language Models (2023.eacl-main)

Copied to clipboard

Challenge: Pretrained large language models generate harmful language encompassing hate speech, abusive language, social biases, and threats.
Approach: They propose two strategies that augment pretraining data to reduce model toxicity . MEDA adds raw toxicity score as meta-data and INST adds instructions indicating toxicity to pretraining samples.
Outcome: The proposed strategies reduce toxicity probability up to 61% while preserving accuracy on five benchmark NLP tasks and improving AUC scores on bias detection tasks by 1.3%.
A little goes a long way: Improving toxic language classification despite data scarcity (2020.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for toxic language classification have not been thoroughly explored.
Approach: They propose to use data augmentation to generate new synthetic data from labeled seed datasets to improve toxic language classification.
Outcome: The proposed techniques perform well on very scarce toxic language datasets while performing worse on shallower models.
Unveiling the Implicit Toxicity in Large Language Models (2023.emnlp-main)

Copied to clipboard

Challenge: Recent studies focus on probing toxic outputs that can be easily detected with existing toxicity classifiers, but LLMs can generate diverse implicit toxic output that are difficult to detect via simply zero-shot prompting.
Approach: They propose a reinforcement learning based attacking method to induce the implicit toxic outputs in large language models by fine-tuning toxicity classifiers.
Outcome: The proposed method generates implicit toxic outputs that are difficult to detect via zero-shot prompting on five widely-adopted toxicity classifiers.
Penetrating Linguistic Disguises: A Slang-aware Label-Aligned Framework for Fine-Grained Toxicity Extraction in Chinese Hate Speech Detection (2026.findings-acl)

Copied to clipboard

Challenge: Flexible word boundaries and linguistic obfuscation, particularly slang, challenge precise span-level hate speech detection in Chinese.
Approach: They propose a Slang-aware Label-Aligned Framework that maps slang to explicit hate semantics and uses task-specific branches to mitigate feature interference.
Outcome: The proposed framework reduces ambiguity by mapping obscure slang to explicit hate semantics.
ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection (2022.acl-long)

Copied to clipboard

Challenge: Toxic language detection systems often falsely flag text that contains minority group mentions as toxic . this over-reliance on spurious correlations also causes systems to struggle with detecting implicitly toxic language.
Approach: They develop a machine-generated dataset of toxic and benign statements about 13 minority groups that generates subtly toxic and harmless text with a massive pretrained language model.
Outcome: The proposed method can detect toxic and benign statements on a large scale . it can also detect hate speech on 94.5% of the toxic examples .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations