Challenge: Multimodal emotion recognition in conversation (MERC) aims to identify speakers’ emotional states by utilizing text, audio, and visual modalities.
Approach: They propose an adaptive modality selection framework for multimodal emotion recognition in conversation that integrates all available modalities into one .
Outcome: The proposed framework outperforms existing methods on multimodal dialogue datasets and is available at https://github.com/youflyaway/Modality-Selection-Enhanced-LoRA-Tuned-LLMs.

Similar Papers

A Dual Contrastive Learning Framework for Enhanced Multimodal Conversational Emotion Recognition (2025.coling-main)

Copied to clipboard

Challenge: Existing methods struggle to capture emotion shifts due to label replication and fail to preserve positive independent modality contributions during fusion.
Approach: They propose a Dual Contrastive Learning Framework that enhances existing MERC models without additional data.
Outcome: The proposed framework outperforms existing models on two MERC benchmark datasets and shows that it reduces label dependence and enhances emotion-sensitive independent modality features.
Self-supervised Cross-modal Pretraining for Speech Emotion Recognition and Sentiment Analysis (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches to multimodal speech emotion recognition and sentiment analysis have not improved results due to their relatively simple fusion mechanisms and lack of proper cross-modal pretraining.
Approach: They propose a deep-fused audio-text bi-modal transformer with carefully designed cross-modal fusion mechanism and stage-wise cross-mod pretraining scheme to facilitate cross-modulation.
Outcome: The proposed method exceeds benchmarks on public IEMOCAP emotion and CMU-MOSEI sentiment datasets by a large margin.
Multimodal Emotion Recognition in Conversations: A Survey of Methods, Trends, Challenges and Prospects (2025.findings-emnlp)

Copied to clipboard

Challenge: Multimodal Emotion Recognition in Conversations (MERC) is a new way to enhance human-computer interaction.
Approach: This survey offers a systematic overview of Multimodal Emotion Recognition in Conversations . it examines motivations, core tasks, representative methods, and evaluation strategies .
Outcome: The survey examines the effectiveness of MERC and its evaluation strategies.
Self-adaptive Context and Modal-interaction Modeling For Multimodal Emotion Recognition (2023.findings-acl)

Copied to clipboard

Challenge: Existing methods to predict emotion label for a given utterance lack modeling of diverse dependency ranges and inconsistent treatment of contribution for various modalities.
Approach: They propose a multimodal emotion recognition in conversation task that uses context and multiple modalities to predict emotion label for a given utterance.
Outcome: The proposed method outperforms the state-of-the-art methods on three multimodal datasets.
Learning Language-guided Adaptive Hyper-modality Representation for Multimodal Sentiment Analysis (2023.emnlp-main)

Copied to clipboard

Challenge: Multimodal Sentiment Analysis (MSA) is effective when using rich information from multiple sources, but the potential sentiment-irrelevant information across modalities may hinder the performance from being further improved.
Approach: They propose an Adaptive Language-guided Multimodal Transformer (ALMT) that learns an irrelevance/conflict-suppressing representation from visual and audio features under guidance of language features at different scales.
Outcome: The proposed model achieves state-of-the-art on several popular datasets and an abundance of ablation shows the effectiveness of the proposed model.
Topic and Style-aware Transformer for Multimodal Emotion Recognition (2023.findings-acl)

Copied to clipboard

Challenge: Existing studies show that visual modality makes minimal contribution to multimodal emotion recognition due to its high dimensionality.
Approach: They propose to leverage the strong multimodality backbone VATT to project the visual signal to the common space with language and acoustic signals.
Outcome: The proposed model outperforms SOTA results and integrates visual signals and handles subjectivity issues by serving as content "normalization" previous studies show that visual modality makes minimal contribution to the performance of multimodal emotion recognition tasks due to high dimensionality.
Emotion-Wheel-Guided Audio-Referred Text Representation for Multimodal Emotion Recognition in Conversation (2026.acl-long)

Copied to clipboard

Challenge: Existing methods for Emotion Recognition in Conversation ignore their distinct communicative roles and information capacities and apply uniform penalties regardless of affective proximity.
Approach: They propose a modality-aware fusion strategy capturing linguistic features from text as the primary source and audio as a complementary component.
Outcome: The proposed method captures linguistic features from text as the primary source and audio as a complementary component and supervised contrastive loss to encode emotional proximity based on Russell’s circumplex model.
Chain of Strategy Optimization Makes Large Language Models Better Emotional Supporter (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing supervised fine-tuning (SFT) fails to address these issues, as it trains models on single gold-standard responses without modeling nuanced strategy trade-offs.
Approach: They propose a two-stage framework that optimizes strategy selection preferences at each dialogue turn.
Outcome: The proposed framework improves strategy selection preferences at each dialogue turn.
Multi-Condition Guided Diffusion Network for Multimodal Emotion Recognition in Conversation (2025.findings-naacl)

Copied to clipboard

Challenge: Current research emphasizes contextual factors, the speaker’s influence, and extracting complementary information across different modalities.
Approach: They propose a diffusion-based approach to address the challenges posed by redundant information and redundant information at the semantic level while robustly capturing shared semantics.
Outcome: The proposed model outperforms existing state-of-the-art models on two multimodal datasets and is generalizable and effective.
MultiEMO: An Attention-Based Correlation-Aware Multimodal Fusion Framework for Emotion Recognition in Conversations (2023.acl-long)

Copied to clipboard

Challenge: Emotion Recognition in Conversations (ERC) is an increasingly popular task in the field of Natural Language Processing.
Approach: They propose a framework that captures cross-modal mapping relationships across modalities . they propose 'multiemotion-aware' framework that integrates multimodal cues into the model .
Outcome: The proposed framework outperforms state-of-the-art models in all emotion categories on two benchmark datasets.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations