Multimodal and Multi-view Models for Emotion Recognition (P19-1)

Copied to clipboard

Challenge: combining lexical and acoustic information results in more robust and accurate models . combining both modalities may be a bottleneck in a deployment pipeline due to computational complexity or privacy constraints .
Approach: They propose to combine acoustic and lexical information to provide a deployable acustic model . they use multimodal models and two attention mechanisms to assess the benefits of lexicals .
Outcome: The proposed model outperforms the state-of-the-art on the USC-IEMOCAP dataset . it significantly surpasses models that have been exclusively trained with acoustic features .

Similar Papers

MultiEMO: An Attention-Based Correlation-Aware Multimodal Fusion Framework for Emotion Recognition in Conversations (2023.acl-long)

Copied to clipboard

Challenge: Emotion Recognition in Conversations (ERC) is an increasingly popular task in the field of Natural Language Processing.
Approach: They propose a framework that captures cross-modal mapping relationships across modalities . they propose 'multiemotion-aware' framework that integrates multimodal cues into the model .
Outcome: The proposed framework outperforms state-of-the-art models in all emotion categories on two benchmark datasets.
Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation (2025.findings-emnlp)

Copied to clipboard

Challenge: AVLM integrates full-face visual cues into a pre-trained expressive speech model.
Approach: They propose an Audio-Visual Language Model (AVLM) for expressive speech generation by integrating full-face visual cues into a pre-trained expressive speech model.
Outcome: The proposed model incorporates full-face visual cues into a pre-trained expressive speech model.
Modality-Transferable Emotion Embeddings for Low-Resource Multimodal Emotion Recognition (2020.aacl-main)

Copied to clipboard

Challenge: despite recent advances in multimodal emotion recognition, two problems still exist: sub-optimal performance and low-resource emotions.
Approach: They propose a modality-transferable model with emotion embeddings to solve these problems . they use pre-trained word embedders to represent emotion categories for textual data .
Outcome: The proposed model outperforms baselines in zero-shot and few-shot scenarios for unseen emotions.
Multimodal Affective Analysis Using Hierarchical Attention Strategy with Word-Level Alignment (P18-1)

Copied to clipboard

Challenge: Existing approaches to classify human affect and subjective information from multiple data sources are limited by the lack of high-level feature associations.
Approach: They propose a hierarchical multimodal architecture with attention and word-level fusion to classify utterance-level sentiment and emotion from text and audio data.
Outcome: The proposed model outperforms state-of-the-art approaches on published datasets and visualizes and interprets synchronized attention over modalities.
Leveraging Unimodal Self-Supervised Learning for Multimodal Audio-Visual Speech Recognition (2022.acl-long)

Copied to clipboard

Challenge: Existing methods for audio-visual speech recognition use extra data to increase performance . a recent study shows that the use of unimodal self-supervised learning improves performance on multimodal tasks.
Approach: They propose to use unimodal self-supervised learning to train AVSR models on unlabelled unilateral data.
Outcome: The proposed model improves on lip reading sentences 2 by 30% even without an external language model.
Self-adaptive Context and Modal-interaction Modeling For Multimodal Emotion Recognition (2023.findings-acl)

Copied to clipboard

Challenge: Existing methods to predict emotion label for a given utterance lack modeling of diverse dependency ranges and inconsistent treatment of contribution for various modalities.
Approach: They propose a multimodal emotion recognition in conversation task that uses context and multiple modalities to predict emotion label for a given utterance.
Outcome: The proposed method outperforms the state-of-the-art methods on three multimodal datasets.
A Dual Contrastive Learning Framework for Enhanced Multimodal Conversational Emotion Recognition (2025.coling-main)

Copied to clipboard

Challenge: Existing methods struggle to capture emotion shifts due to label replication and fail to preserve positive independent modality contributions during fusion.
Approach: They propose a Dual Contrastive Learning Framework that enhances existing MERC models without additional data.
Outcome: The proposed framework outperforms existing models on two MERC benchmark datasets and shows that it reduces label dependence and enhances emotion-sensitive independent modality features.
Multi-Condition Guided Diffusion Network for Multimodal Emotion Recognition in Conversation (2025.findings-naacl)

Copied to clipboard

Challenge: Current research emphasizes contextual factors, the speaker’s influence, and extracting complementary information across different modalities.
Approach: They propose a diffusion-based approach to address the challenges posed by redundant information and redundant information at the semantic level while robustly capturing shared semantics.
Outcome: The proposed model outperforms existing state-of-the-art models on two multimodal datasets and is generalizable and effective.
UniMSE: Towards Unified Multimodal Sentiment Analysis and Emotion Recognition (2022.emnlp-main)

Copied to clipboard

Challenge: Existing studies study sentiment and emotion separately and do not fully exploit the complementary knowledge behind the two.
Approach: They propose a multimodal sentiment knowledge-sharing framework that unifies MSA and ERC tasks from features, labels, and models.
Outcome: The proposed framework achieves consistent improvements on four public benchmark datasets on MOSI, MOSEI, MELD, and IEMOCAP.
Do Audio LLMs Really LISTEN, or Just Transcribe? Measuring Lexical vs. Acoustic Emotion Cues Reliance (2026.eacl-long)

Copied to clipboard

Challenge: LISTEN is a controlled benchmark to disentangle lexical reliance from acoustic sensitivity in emotion understanding.
Approach: They propose a benchmark to disentangle lexical reliance from acoustic sensitivity in emotion understanding.
Outcome: LISTEN shows that current LALMs largely "transcribe" rather than "listen" authors note that models underutilize acoustic cues while relying on lexical semantics .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations