Challenge: Existing datasets do not provide enough annotation to explain unsafe behavior . current chatbots generate toxic and offensive responses, which can be dangerous .
Approach: They construct a dataset called SafeConv that provides comprehensive annotations for chatbots . they compare safe alternatives to rewrite unsafe responses .
Outcome: The proposed model can explain unsafe behavior and detoxify chatbots, the authors show . the proposed model is able to detect unsafe utterances, extract unsafe spans, and convert unsafe responses to safe versions.

Similar Papers

On the Safety of Conversational Models: Taxonomy, Dataset, and Benchmark (2022.findings-acl)

Copied to clipboard

Challenge: Dialogue safety problems severely limit the real-world deployment of generative conversational models.
Approach: They propose a taxonomy for dialogue safety specifically designed to capture unsafe behaviors in human-bot dialogue settings.
Outcome: The proposed taxonomy captures unsafe behaviors in human-bot dialogue settings with rich context-sensitive unsafe examples.
SaFeRDialogues: Taking Feedback Gracefully after Conversational Safety Failures (2022.acl-long)

Copied to clipboard

Challenge: Existing open-domain conversational models can easily be made to talk in inadequate ways.
Approach: They propose a task and dataset of graceful responses to safety feedback . they collect 8k dialogues demonstrating safety failures, feedback signaling them, and a response acknowledging feedback.
Outcome: The proposed model improves on a dataset of 8k dialogues demonstrating safety failures, feedback signaling them, and a response acknowledging the feedback.
SafetyKit: First Aid for Measuring Safety in Open-domain Conversational Systems (2022.acl-long)

Copied to clipboard

Challenge: Several studies discuss the potential harms and benefits of large language models (LLMs) large neural models can replicate and even amplify negative, stereotypical, and derogatory associations in the data.
Approach: They propose to use a first aid kit to assess the safety of conversational AI in various settings . they propose several future directions and discuss ethical considerations .
Outcome: The proposed tools can provide estimates of the relative safety of systems in various settings, but they still have several shortcomings.
GrounDial: Human-norm Grounded Safe Dialog Response Generation (2024.findings-eacl)

Copied to clipboard

Challenge: Recent conversational AI systems generate unsafe responses agreeing to offensive user input or including toxic content.
Approach: They propose a method where response safety is achieved by grounding responses to commonsense social rules without fine-tuning.
Outcome: The proposed approach is quantitatively and qualitatively safer even without additional data or tuning.
ProsocialDialog: A Prosocial Backbone for Conversational Agents (2022.emnlp-main)

Copied to clipboard

Challenge: Existing dialogue systems fail to respond properly to potentially unsafe user utterances . existing systems either ignore or passively agree with unsafe content .
Approach: They introduce a dataset to teach conversational agents to respond to problematic content following social norms.
Outcome: The proposed dataset shows that ProsocialDialog generates more socially acceptable dialogues than existing models.
Bot-Adversarial Dialogue for Safe Conversational Agents (2021.naacl-main)

Copied to clipboard

Challenge: a new method for evaluating chatbot safety is proposed to mimic human-generated data . a bot-adversarial dialogue model learns undesirable features from this data, a study finds .
Approach: They propose a human-and-model-in-the-loop framework for evaluating toxicity of chatbots . they propose two methods for safe conversational agents by either training on data or ”baking-in” safety to the generative model itself.
Outcome: The proposed methods are safer than existing models while maintaining usability metrics, the authors say . they show that the proposed methods can be used to make safer models with human-model interactions .
Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety (2025.emnlp-main)

Copied to clipboard

Challenge: Existing surveys focus on interpretation or safety, but safety and understanding are core motivations for interpretation research.
Approach: They propose a framework that connects interpretation methods, enhancements they inform, and tools that operationalize them.
Outcome: The proposed framework summarizes nearly 70 studies at their intersections and concludes with open challenges and future directions.
Using In-Context Learning to Improve Dialogue Safety (2023.findings-emnlp)

Copied to clipboard

Challenge: Recent work has highlighted safety issues with large neural-based conversational models.
Approach: They propose a retrieval-based approach for reducing bias and toxicity in chatbot responses . they retrieve demonstrations of safe responses to similar dialogue contexts to generate a response .
Outcome: The proposed method reduces bias and toxicity in three chatbot models . it can be used in compliment to existing dialogue safety approaches, such as RLHF.
Attacks, Defenses and Evaluations for LLM Conversation Safety: A Survey (2024.naacl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) are now commonplace in conversation applications, but their misuse for generating harmful responses has raised serious societal concerns.
Approach: They provide a comprehensive overview of recent studies covering attacks, defenses, and evaluations of Large Language Models (LLMs) .
Outcome: The proposed review summarizes three aspects of LLM conversation safety: attacks, defenses, and evaluations.
XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models (2024.naacl-long)

Copied to clipboard

Challenge: Large language models (LLMs) are now being used by millions of people across the world.
Approach: They propose a test suite called XSTest to identify such eXaggerated Safety behaviours in a systematic way.
Outcome: The proposed test suite identifies eXaggerated Safety behaviours in a systematic way.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations