Challenge: Existing methods for multi-turn, multi-speaker multimodal affect understanding are difficult to maintain conversation-level consistency under within-speaks' emotion shifts.
Approach: They propose a framework that combines appraisal-guided structured generation with graph-structured reinforcement learning to extract triplets from multi-turn multimodal conversations.
Outcome: The proposed framework outperforms baselines on public MECTEC benchmarks and improves structure-aware metrics on emotion shift coherence and core events.

Similar Papers

Generative Emotion Cause Triplet Extraction in Conversations with Commonsense Knowledge (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing studies on ECTEC focus on Causal Emotion Entailment and Emotion-Cause Pair Extraction in Conversations.
Approach: They propose to decompose the ECTEC task into multiple subtasks and solve them in a pipeline manner.
Outcome: The proposed model outperforms competing systems on two benchmark datasets.
M3HG: Multimodal, Multi-scale, and Multi-type Node Heterogeneous Graph for Emotion Cause Triplet Extraction in Conversations (2025.findings-acl)

Copied to clipboard

Challenge: Existing methods for ECAC focus on textual contexts, overlooking other modalities.
Approach: They propose a multimodal, multi-scenario MECTEC dataset that captures emotional and causal contexts and effectively fuses contextual information at different levels.
Outcome: The proposed model captures emotional and causal contexts and effectively fuses contextual information at both inter- and intra-utterance levels.
Locate and Explain: Joint Multimodal Emotion Cause Extraction and Summarization in Conversation (2026.acl-long)

Copied to clipboard

Challenge: Existing studies focus on utterance-level emotion cause extraction and multimodal emotion cause generation, resulting in subjective and inconsistent annotations.
Approach: They propose a task that extracts emotion cause utterances and generates cause summaries . they propose utterrance-level emotion cause extraction and multimodal emotion cause generation tasks .
Outcome: The proposed task extracts emotion cause utterances and generates cause summaries . the proposed task establishes strong benchmark results for the proposed project .
One Unified Model for Diverse Tasks: Emotion Cause Analysis via Self-Promote Cognitive Structure Modeling (2025.naacl-long)

Copied to clipboard

Challenge: Existing models for emotion cause analysis overlook common ground rooted in cognitive emotion theories, in particular, the cognitive structure of emotions.
Approach: They propose a unified model capable of tackling diverse emotion cause analysis tasks . they propose 'self-promote mechanism' that constructs the emotion cognitive structure through LLM .
Outcome: The proposed model outperforms existing models and baselines on multiple emotion cause analysis tasks.
UniMEEC: Towards Unified Multimodal Emotion Recognition and Emotion Cause (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing studies treat emotion recognition and emotion cause extraction as two individual problems, ignoring their natural causality.
Approach: They propose a Unified Multimodal Emotion recognition and Emotion-Cause analysis framework to explore the causality between emotion and emotion cause.
Outcome: The proposed framework reformulates MERC and MECPE tasks as mask prediction problems and unifies them with a causal prompt template.
Emotion–Cause Pair Extraction in Conversations via Semantic Decoupling and Alignment (2026.findings-acl)

Copied to clipboard

Challenge: Emotion-Cause Pair Extraction in Conversations (ECPEC) aims to identify the set of causal relations between emotion utterances and their triggering causes within a dialogue.
Approach: They propose a framework for Emotion-Cause Pair Extraction in Conversations that decouples emotion-oriented semantics from cause-oriented ones and employs optimal transport to enable many-to-many and globally consistent emotion-cause matching.
Outcome: The proposed framework achieves state-of-the-art on several benchmark datasets.
Enhancing Emotion-Cause Pair Extraction in Conversations via Center Event Detection and Reasoning (2024.findings-emnlp)

Copied to clipboard

Challenge: Emotion-Cause Pair Extraction in Conversations (ECPEC) aims to identify emotion utterances and their corresponding cause utterrances in unannotated conversations.
Approach: They propose a new method to identify emotion utterances and their corresponding cause utterrances in unannotated conversations by using a center event-aware graph.
Outcome: The proposed model outperforms existing methods and achieves state-of-the-art performance across three benchmark datasets.
ECC: An Emotion-Cause Conversation Dataset for Empathy Response (2025.emnlp-main)

Copied to clipboard

Challenge: Existing empathy dialogue datasets focus on emotion labels while cause annotations are added post hoc.
Approach: They propose an emotion-cause conversation dataset with 2.4K dialogues that can be scalable . they use a framework that utilizes knowledge and large language models to automatically generate dialogues .
Outcome: The proposed dataset can achieve comparable or even superior performance to existing empathy dialogue datasets.
A Multi-turn Machine Reading Comprehension Framework with Rethink Mechanism for Emotion-Cause Pair Extraction (2022.coling-1)

Copied to clipboard

Challenge: Emotion-cause pair extraction (ECPE) is an emerging task in emotion cause analysis, which extracts potential emotion-caused pairs from an emotional document.
Approach: They propose a document-level machine reading comprehension task to model complex relations between emotions and causes while avoiding generating the pairing matrix.
Outcome: The proposed framework outperforms existing state-of-the-art methods on the emotion cause corpus and can model complex relations between emotions and causes while avoiding pairing matrix.
A Facial Expression-Aware Multimodal Multi-task Learning Framework for Emotion Recognition in Multi-party Conversations (2023.acl-long)

Copied to clipboard

Challenge: Recent studies have shown the importance of visual information in multi-party conversations due to the complexity of visual scenes.
Approach: They propose a framework to extract face sequences as visual features from a real speaker's utterance and a pipeline method to extract the face sequence.
Outcome: The proposed framework extracts face sequences of the real speaker of each utterance and improves emotion prediction on the MELD dataset.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations