Challenge: Existing work on analysing textual dialogues that derailed into toxic content ignores visual information, such as images and videos.
Approach: They propose a new multimodal conversational architecture that utilises visual and conversational contexts to classify comments for derailment.
Outcome: The proposed approach outperforms existing methods and is more robust to textual noise.

Similar Papers

A Theoretically Grounded Approach to Summarizing Conversation Dynamics for Forecasting the Derailment of Online Conversations (2026.acl-long)

Copied to clipboard

Challenge: Recent work on conversation derailment prediction relies on linguistic features rooted in linguistic and social theories.
Approach: They propose a system that predicts from the start of a conversation whether it will derail into toxic exchanges.
Outcome: The proposed system achieves 10% performance increase over baseline and 6.47% increase on benchmark dataset.
Trouble on the Horizon: Forecasting the Derailment of Online Conversations as they Develop (D19-1)

Copied to clipboard

Challenge: Recent efforts focused on detecting antisocial behavior after the fact . a forecasting model needs to capture the flow of the conversation, not individual comments . real conversations have an unknown horizon; therefore a practical forecasting system needs to assess the risk .
Approach: They propose a conversational forecasting model that learns conversational dynamics and exploits it to predict derailment as the conversation develops.
Outcome: The proposed model outperforms state-of-the-art models at forecasting derailment . it learns an unsupervised representation of conversational dynamics and exploits it to predict future derailments .
Can LLMs Be Efficient Predictors of Conversational Derailment? (2025.findings-emnlp)

Copied to clipboard

Challenge: Conversational derailment is a common issue on online platforms due to toxic or inappropriate remarks.
Approach: They prompt pre-trained large language models to predict conversational derailment without fine-tuning . they compare chain-of-thought reasoning and few-shot exemplars to predict derailments .
Outcome: The proposed model predicts conversational derailment without task-specific fine-tuning without fine-cuning.
MMCoQA: Conversational Question Answering over Text, Tables, and Images (2022.acl-long)

Copied to clipboard

Challenge: Existing conversational QA systems only use a single knowledge source, e.g., paragraphs or knowledge graph, and assume it contains enough evidence to extract answers to users' questions.
Approach: They propose a task to answer users' questions with multimodal knowledge sources via multi-turn conversations using a multimodal dataset.
Outcome: The proposed task brings a series of research challenges, including but not limited to priority, consistency, and complementarity of multimodal knowledge.
“Mm, Wat?” Detecting Other-initiated Repair Requests in Dialogue (2025.emnlp-main)

Copied to clipboard

Challenge: Current conversational agents (CAs) do not recognize repair initiation, leading to breakdowns or disengagement.
Approach: They propose a multimodal model to automatically detect repair initiation in Dutch dialogues by integrating linguistic and prosodic features grounded in Conversation Analysis.
Outcome: The proposed model integrates linguistic and prosodic features grounded in Conversation Analysis to detect repair initiation in Dutch dialogues.
Multimodal Conversation Structure Understanding (2026.eacl-long)

Copied to clipboard

Challenge: a new set of tasks is being developed to parse the structure of conversation . female characters are 1.2 times more likely to be cast as an addressee or side-participant .
Approach: They propose a set of tasks and release an annotated dataset for multimodal conversation structure understanding.
Outcome: The proposed model outperforms the baseline model, but performance drops when character identities are anonymized.
MMChat: Multi-Modal Chat Dataset on Social Media (2022.lrec-1)

Copied to clipboard

Challenge: Incorporating multi-modal contexts in conversation is important for developing engaging dialogue systems.
Approach: They propose a large scale Chinese multi-modal dialogue corpus that contains image-grounded dialogues from real conversations on social media.
Outcome: The proposed model can handle sparsity issues in dialogue generation tasks by incorporating image features.
MM-Claims: A Dataset for Multimodal Claim Detection in Social Media (2022.findings-naacl)

Copied to clipboard

Challenge: Using image and text, we investigate the role of image and texts in fake news detection . claim detection is a step in fighting misinformation and as a precursor to prioritize potentially false information for fact-checking.
Approach: They propose a dataset that consists of tweets and corresponding images for claim detection . they evaluate strong unimodal and multimodal baselines and analyze drawbacks of current models .
Outcome: The proposed dataset evaluates strong unimodal and multimodal baselines and examines drawbacks of existing models.
Evaluating Cognitive-Behavioral Fixation via Multimodal User Viewing Patterns on Social Media (2025.emnlp-main)

Copied to clipboard

Challenge: Digital media platforms often contribute to cognitive-behavioral fixation, a phenomenon in which users exhibit sustained and repetitive engagement with narrow content domains.
Approach: They propose a multimodal topic extraction module and a cognitive-behavioral fixation quantification module that collaboratively enable adaptive, hierarchical, and interpretable assessment of user behavior.
Outcome: The proposed framework lays the groundwork for scalable computational analysis of cognitive fixation.
Dynamic Forecasting of Conversation Derailment (2021.emnlp-main)

Copied to clipboard

Challenge: a pretrained language encoder can predict derailment in online conversations . this is a useful task for detecting and preventing abusive language .
Approach: They extend a task to predict derailment in online conversations by using a pretrained language encoder.
Outcome: The proposed task outperforms previous approaches in terms of performance and quality.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations