Challenge: Existing work to generate adversarial attacks is costly and not scalable . despite the abundance of research in this area, little attention has been given to adversarials .
Approach: They propose an adversarial attack mechanism that mitigates toxic language generation . they propose a defense mechanism that is scalable and can be generalized .
Outcome: The proposed defense is effective at avoiding toxic language generation even against imperceptible toxicity triggers while preserving conversational flow.

Similar Papers

Towards Building a Robust Toxicity Predictor (2023.acl-industry)

Copied to clipboard

Challenge: Recent studies have focused on robustness of toxicity language predictors, but this is problematic for real-world toxicity detection.
Approach: They propose a novel adversarial attack that exploits greedy search strategies to fool toxic text classifiers.
Outcome: The proposed attack can detect weaker toxicity language detectors even against unseen attacks.
Bot-Adversarial Dialogue for Safe Conversational Agents (2021.naacl-main)

Copied to clipboard

Challenge: a new method for evaluating chatbot safety is proposed to mimic human-generated data . a bot-adversarial dialogue model learns undesirable features from this data, a study finds .
Approach: They propose a human-and-model-in-the-loop framework for evaluating toxicity of chatbots . they propose two methods for safe conversational agents by either training on data or ”baking-in” safety to the generative model itself.
Outcome: The proposed methods are safer than existing models while maintaining usability metrics, the authors say . they show that the proposed methods can be used to make safer models with human-model interactions .
Unveiling the Implicit Toxicity in Large Language Models (2023.emnlp-main)

Copied to clipboard

Challenge: Recent studies focus on probing toxic outputs that can be easily detected with existing toxicity classifiers, but LLMs can generate diverse implicit toxic output that are difficult to detect via simply zero-shot prompting.
Approach: They propose a reinforcement learning based attacking method to induce the implicit toxic outputs in large language models by fine-tuning toxicity classifiers.
Outcome: The proposed method generates implicit toxic outputs that are difficult to detect via zero-shot prompting on five widely-adopted toxicity classifiers.
Contrastive Perplexity for Controlled Generation: An Application in Detoxifying Large Language Models (2025.acl-long)

Copied to clipboard

Challenge: Existing approaches to generate toxic content by large language models are based on pipelines . current approaches focus on preserving performance while effectively mitigating toxicity .
Approach: They propose a framework for implicit knowledge editing and controlled text generation by using hard negatives.
Outcome: The proposed framework significantly reduces toxic generation while maintaining strong performance on downstream tasks.
Attacking Misinformation Detection Using Adversarial Examples Generated by Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: Large language models can be used to attack content filtering algorithms in social media platforms.
Approach: They propose to generate adversarial examples to test the robustness of social media content filtering algorithms.
Outcome: The proposed model outperforms existing models in the case of propaganda, false claims, rumours and hyperpartisan news.
Unlocking Anticipatory Text Generation: A Constrained Approach for Large Language Models Decoding (2024.emnlp-main)

Copied to clipboard

Challenge: Large language models have shown a powerful ability for text generation, but undesired behaviors such as toxicity and hallucinations can manifest.
Approach: They propose to formalize text generation as a future-constrained generation problem to minimize undesirable behaviors and enforce faithfulness to instructions.
Outcome: The proposed approach is effective across three tasks, including keyword-constrained generation, toxicity reduction, and factual correctness in question-answering.
Challenges in Detoxifying Language Models (2021.findings-emnlp)

Copied to clipboard

Challenge: Prior work often relies on automatic evaluation of LM toxicity.
Approach: They evaluate toxicity mitigation strategies for automated and human evaluations . they find human raters disagree with high automatic toxicity scores after strong toxicity reduction interventions .
Outcome: The proposed methods reduce LM toxicity but lower coverage for marginalized texts . human raters disagree with high toxicity scores after strong toxicity reduction interventions .
ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection (2022.acl-long)

Copied to clipboard

Challenge: Toxic language detection systems often falsely flag text that contains minority group mentions as toxic . this over-reliance on spurious correlations also causes systems to struggle with detecting implicitly toxic language.
Approach: They develop a machine-generated dataset of toxic and benign statements about 13 minority groups that generates subtly toxic and harmless text with a massive pretrained language model.
Outcome: The proposed method can detect toxic and benign statements on a large scale . it can also detect hate speech on 94.5% of the toxic examples .
Toxicity in chatgpt: Analyzing persona-assigned language models (2023.findings-emnlp)

Copied to clipboard

Challenge: Large language models (LLMs) have shown incredible capabilities and transcended the natural language processing community.
Approach: They evaluate toxicity in over half a million generations of ChatGPT by assigning it a persona . they find that outputs engage in incorrect stereotypes, harmful dialogue, hurtful opinions .
Outcome: a new study shows that assigning a persona to a chatbot can increase toxicity in half a million generations.
Language Generation Models Can Cause Harm: So What Can We Do About It? An Actionable Survey (2023.eacl-main)

Copied to clipboard

Challenge: Recent advances in the capacity of large language models to generate human-like text have prompted a heated discourse around the risks of societal harms they introduce.
Approach: They propose a taxonomy of interventions organized around the different phases where they can be adopted to mitigate harms.
Outcome: The proposed methods are based on several prior works’ taxonomies of language model risks and provide an overview of strategies for detecting and ameliorating different kinds of risks/harms.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations