Challenge: Recent studies focus on probing toxic outputs that can be easily detected with existing toxicity classifiers, but LLMs can generate diverse implicit toxic output that are difficult to detect via simply zero-shot prompting.
Approach: They propose a reinforcement learning based attacking method to induce the implicit toxic outputs in large language models by fine-tuning toxicity classifiers.
Outcome: The proposed method generates implicit toxic outputs that are difficult to detect via zero-shot prompting on five widely-adopted toxicity classifiers.

Similar Papers

Realistic Evaluation of Toxicity in Large Language Models (2024.findings-acl)

Copied to clipboard

Challenge: a large amount of data exposes large language models to toxicity and bias . prompt engineering can be easily bypassed with minimal prompt engineering.
Approach: They propose a dataset that uses manually crafted prompts to nullify protective layers of large language models.
Outcome: The proposed dataset shows that prompts can nullify protective layers of large language models.
Revisiting Implicitly Abusive Language Detection: Evaluating LLMs in Zero-Shot and Few-Shot Settings (2025.coling-main)

Copied to clipboard

Challenge: Current research focuses on explicit abusive language, but subtler forms of IAL remain insufficiently studied.
Approach: They evaluate the models' capabilities in classifying sentences directly as either IAL or benign, and in extracting linguistic features associated with IAL.
Outcome: The proposed models outperform the best previously reported methods in classifying sentences directly as IAL or benign and extracting linguistic features associated with IAL.
LLM Sensitivity Challenges in Abusive Language Detection: Instruction-Tuned vs. Human Feedback (2025.coling-main)

Copied to clipboard

Challenge: Existing studies show that instruction-tuned LLMs under-predict positive classes . however, they are overly sensitive and can be applied for abuse detection without fine-tuning .
Approach: They show that instruction-tuned LLMs tend to under-predict positive classes . they also show that label frequency in the prompt helps with the significant over-prediction .
Outcome: The proposed models under-predict positive classes in social media, whereas they are overly sensitive.
Intrinsic Model Weaknesses: How Priming Attacks Unveil Vulnerabilities in Large Language Models (2025.findings-naacl)

Copied to clipboard

Challenge: Large language models (LLMs) have significant impact on various industries and societal functions due to advanced instruction-following capabilities.
Approach: They developed and tested novel attack strategies on popular LLMs to expose their vulnerabilities in generating harmful content.
Outcome: The proposed attacks achieved an ASR of 100% on open-source models, including Meta’s Llama-3.2, Google’s Gemma-2, Mistral’s Mistral-NeMo, Falcon’s Falcon-mamba, Apple’s DCLM, Microsoft’s Phi3, and Qwen’s Qwend2.5, among others.
Pragmatic Inference Chain (PIC) Improving LLMs’ Reasoning of Authentic Implicit Toxic Language (2025.emnlp-main)

Copied to clipboard

Challenge: Existing studies have shown that LLMs can detect toxicity by using a variety of inference-intensive tasks, such as understanding humour and metaphors.
Approach: They propose a new method to prompt LLMs to identify toxic language using a set of online data that are verified by human annotators.
Outcome: The proposed method significantly improves the success rate of GPT-4o, Llama-3.1-70B-Instruct, DeepSeek-v2.5, and DeepSeq-v3 in identifying implicit toxic language compared to five baseline prompts, such as CoT and rule-based baselines.
Challenges in Detoxifying Language Models (2021.findings-emnlp)

Copied to clipboard

Challenge: Prior work often relies on automatic evaluation of LM toxicity.
Approach: They evaluate toxicity mitigation strategies for automated and human evaluations . they find human raters disagree with high automatic toxicity scores after strong toxicity reduction interventions .
Outcome: The proposed methods reduce LM toxicity but lower coverage for marginalized texts . human raters disagree with high toxicity scores after strong toxicity reduction interventions .
Can LLMs Recognize Toxicity? A Structured Investigation Framework and Toxicity Metric (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing toxicity metrics rely on encoder models trained on specific toxicity datasets, which are susceptible to out-of-distribution (OOD) problems and depend on the dataset’s definition of toxicity.
Approach: They propose a robust metric grounded on LLMs to flexibly measure toxicity according to the given definition by analysing toxicity factors and intrinsic toxic attributes.
Outcome: The proposed metric improves on conventional metrics by 12 points in the F1 score and shows that upstream toxicity significantly influences downstream metrics, suggesting that LLMs are unsuitable for toxicity evaluations within unverified factors.
Probing LLMs for hate speech detection: strengths and vulnerabilities (2023.findings-emnlp)

Copied to clipboard

Challenge: Recent efforts to detect hateful or toxic language using large language models have not used explanation, additional context and victim community information in the detection process.
Approach: They use different prompt variations, input information and victim community information to evaluate large language models in zero shot setting without adding any in-context examples.
Outcome: The proposed models perform significantly better when included in the pipeline than baseline models.
How Susceptible are Large Language Models to Ideological Manipulation? (2024.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) have the potential to exert substantial influence on public perceptions and interactions with information.
Approach: They examine how LLMs can learn and generalize ideological biases from their instruction-tuning data.
Outcome: The LLMs show a startling ability to absorb ideology from one topic and generalize it to even unrelated ones.
A Survey of Toxicity Mitigation Strategies for Multilingual Language Models (2026.findings-acl)

Copied to clipboard

Challenge: Large language models can reproduce and amplify toxic content, including hate speech, harassment, and bias.
Approach: They propose a comprehensive survey of the many detoxification methods tailored to multilingual LLMs.
Outcome: The proposed methods are based on data filtering, style transfer, expert-based logit steering, retrieval augmentation, and human feedback.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations