Challenge: Prior work focused on characterizing and detecting content exhibiting antisocial online behavior.
Approach: They propose a task of predicting from the very start of a conversation whether it will get out of hand.
Outcome: The proposed framework can detect early warning signs of antisocial behavior in online conversations.

Similar Papers

Trouble on the Horizon: Forecasting the Derailment of Online Conversations as they Develop (D19-1)

Copied to clipboard

Challenge: Recent efforts focused on detecting antisocial behavior after the fact . a forecasting model needs to capture the flow of the conversation, not individual comments . real conversations have an unknown horizon; therefore a practical forecasting system needs to assess the risk .
Approach: They propose a conversational forecasting model that learns conversational dynamics and exploits it to predict derailment as the conversation develops.
Outcome: The proposed model outperforms state-of-the-art models at forecasting derailment . it learns an unsupervised representation of conversational dynamics and exploits it to predict future derailments .
A Theoretically Grounded Approach to Summarizing Conversation Dynamics for Forecasting the Derailment of Online Conversations (2026.acl-long)

Copied to clipboard

Challenge: Recent work on conversation derailment prediction relies on linguistic features rooted in linguistic and social theories.
Approach: They propose a system that predicts from the start of a conversation whether it will derail into toxic exchanges.
Outcome: The proposed system achieves 10% performance increase over baseline and 6.47% increase on benchmark dataset.
How did we get here? Summarizing conversation dynamics (2024.naacl-long)

Copied to clipboard

Challenge: Throughout a conversation, the way participants interact with each other is in constant flux.
Approach: They propose to summarize conversations by constructing human-written summaries and exploring automated baselines.
Outcome: The summarizing tools help both humans and automated systems forecast toxic behavior in conversations.
Can LLMs Be Efficient Predictors of Conversational Derailment? (2025.findings-emnlp)

Copied to clipboard

Challenge: Conversational derailment is a common issue on online platforms due to toxic or inappropriate remarks.
Approach: They prompt pre-trained large language models to predict conversational derailment without fine-tuning . they compare chain-of-thought reasoning and few-shot exemplars to predict derailments .
Outcome: The proposed model predicts conversational derailment without task-specific fine-tuning without fine-cuning.
Early Rumour Detection (N19-1)

Copied to clipboard

Challenge: Existing studies on rumour detection are concerned with timing, but few are interested in how early we can detect them.
Approach: They propose a method that integrates reinforcement learning to learn the minimum number of posts required before classifying an event as a rumour.
Outcome: The proposed model detects rumours earlier than state-of-the-art systems while maintaining comparable accuracy.
Sketching a Linguistically-Driven Reasoning Dialog Model for Social Talk (2022.acl-srw)

Copied to clipboard

Challenge: a new study shows that dialog systems that can hold social talk and make sense of conversational content are not efficient for context-sensitive natural language understanding and reasoning.
Approach: They propose a linguistically-informed architecture to handle social talk in English . they propose linguistic models that fit the context-sensitive components into a Bayesian game-theoretic model .
Outcome: The proposed architecture is based on corpus-based methods but does not track what is happening in a conversation.
ProsocialDialog: A Prosocial Backbone for Conversational Agents (2022.emnlp-main)

Copied to clipboard

Challenge: Existing dialogue systems fail to respond properly to potentially unsafe user utterances . existing systems either ignore or passively agree with unsafe content .
Approach: They introduce a dataset to teach conversational agents to respond to problematic content following social norms.
Outcome: The proposed dataset shows that ProsocialDialog generates more socially acceptable dialogues than existing models.
SaFeRDialogues: Taking Feedback Gracefully after Conversational Safety Failures (2022.acl-long)

Copied to clipboard

Challenge: Existing open-domain conversational models can easily be made to talk in inadequate ways.
Approach: They propose a task and dataset of graceful responses to safety feedback . they collect 8k dialogues demonstrating safety failures, feedback signaling them, and a response acknowledging feedback.
Outcome: The proposed model improves on a dataset of 8k dialogues demonstrating safety failures, feedback signaling them, and a response acknowledging the feedback.
SafetyKit: First Aid for Measuring Safety in Open-domain Conversational Systems (2022.acl-long)

Copied to clipboard

Challenge: Several studies discuss the potential harms and benefits of large language models (LLMs) large neural models can replicate and even amplify negative, stereotypical, and derogatory associations in the data.
Approach: They propose to use a first aid kit to assess the safety of conversational AI in various settings . they propose several future directions and discuss ethical considerations .
Outcome: The proposed tools can provide estimates of the relative safety of systems in various settings, but they still have several shortcomings.
Wait! There’s a Way Out: A Decision Mechanism for Forecasting Conversational Derailment (2026.acl-long)

Copied to clipboard

Challenge: Existing approaches make decision to "trigger" based on the estimated likelihood of derailment given the preceding utterances, implicitly assuming that the conversation’s future trajectory is fixed.
Approach: They propose a method for decoupling the decision to trigger from derailment likelihood estimation.
Outcome: The proposed method is inspired by the first human baseline on this task, which shows that humans achieve dramatically lower false positive rates by selectively deferring their decision to trigger when they anticipate that tension is likely to subside.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations