| Challenge: | combining lexical and acoustic information results in more robust and accurate models . combining both modalities may be a bottleneck in a deployment pipeline due to computational complexity or privacy constraints . |
| Approach: | They propose to combine acoustic and lexical information to provide a deployable acustic model . they use multimodal models and two attention mechanisms to assess the benefits of lexicals . |
| Outcome: | The proposed model outperforms the state-of-the-art on the USC-IEMOCAP dataset . it significantly surpasses models that have been exclusively trained with acoustic features . |
Similar Papers
MultiEMO: An Attention-Based Correlation-Aware Multimodal Fusion Framework for Emotion Recognition in Conversations (2023.acl-long)
Copied to clipboard
| Challenge: | Emotion Recognition in Conversations (ERC) is an increasingly popular task in the field of Natural Language Processing. |
| Approach: | They propose a framework that captures cross-modal mapping relationships across modalities . they propose 'multiemotion-aware' framework that integrates multimodal cues into the model . |
| Outcome: | The proposed framework outperforms state-of-the-art models in all emotion categories on two benchmark datasets. |
Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation (2025.findings-emnlp)
Copied to clipboard
| Challenge: | AVLM integrates full-face visual cues into a pre-trained expressive speech model. |
| Approach: | They propose an Audio-Visual Language Model (AVLM) for expressive speech generation by integrating full-face visual cues into a pre-trained expressive speech model. |
| Outcome: | The proposed model incorporates full-face visual cues into a pre-trained expressive speech model. |
Modality-Transferable Emotion Embeddings for Low-Resource Multimodal Emotion Recognition (2020.aacl-main)
Copied to clipboard
| Challenge: | despite recent advances in multimodal emotion recognition, two problems still exist: sub-optimal performance and low-resource emotions. |
| Approach: | They propose a modality-transferable model with emotion embeddings to solve these problems . they use pre-trained word embedders to represent emotion categories for textual data . |
| Outcome: | The proposed model outperforms baselines in zero-shot and few-shot scenarios for unseen emotions. |
Multimodal Affective Analysis Using Hierarchical Attention Strategy with Word-Level Alignment (P18-1)
Copied to clipboard
| Challenge: | Existing approaches to classify human affect and subjective information from multiple data sources are limited by the lack of high-level feature associations. |
| Approach: | They propose a hierarchical multimodal architecture with attention and word-level fusion to classify utterance-level sentiment and emotion from text and audio data. |
| Outcome: | The proposed model outperforms state-of-the-art approaches on published datasets and visualizes and interprets synchronized attention over modalities. |
Leveraging Unimodal Self-Supervised Learning for Multimodal Audio-Visual Speech Recognition (2022.acl-long)
Copied to clipboard
| Challenge: | Existing methods for audio-visual speech recognition use extra data to increase performance . a recent study shows that the use of unimodal self-supervised learning improves performance on multimodal tasks. |
| Approach: | They propose to use unimodal self-supervised learning to train AVSR models on unlabelled unilateral data. |
| Outcome: | The proposed model improves on lip reading sentences 2 by 30% even without an external language model. |
Self-adaptive Context and Modal-interaction Modeling For Multimodal Emotion Recognition (2023.findings-acl)
Copied to clipboard
| Challenge: | Existing methods to predict emotion label for a given utterance lack modeling of diverse dependency ranges and inconsistent treatment of contribution for various modalities. |
| Approach: | They propose a multimodal emotion recognition in conversation task that uses context and multiple modalities to predict emotion label for a given utterance. |
| Outcome: | The proposed method outperforms the state-of-the-art methods on three multimodal datasets. |
A Dual Contrastive Learning Framework for Enhanced Multimodal Conversational Emotion Recognition (2025.coling-main)
Copied to clipboard
| Challenge: | Existing methods struggle to capture emotion shifts due to label replication and fail to preserve positive independent modality contributions during fusion. |
| Approach: | They propose a Dual Contrastive Learning Framework that enhances existing MERC models without additional data. |
| Outcome: | The proposed framework outperforms existing models on two MERC benchmark datasets and shows that it reduces label dependence and enhances emotion-sensitive independent modality features. |
Multi-Condition Guided Diffusion Network for Multimodal Emotion Recognition in Conversation (2025.findings-naacl)
Copied to clipboard
| Challenge: | Current research emphasizes contextual factors, the speaker’s influence, and extracting complementary information across different modalities. |
| Approach: | They propose a diffusion-based approach to address the challenges posed by redundant information and redundant information at the semantic level while robustly capturing shared semantics. |
| Outcome: | The proposed model outperforms existing state-of-the-art models on two multimodal datasets and is generalizable and effective. |
UniMSE: Towards Unified Multimodal Sentiment Analysis and Emotion Recognition (2022.emnlp-main)
Copied to clipboard
| Challenge: | Existing studies study sentiment and emotion separately and do not fully exploit the complementary knowledge behind the two. |
| Approach: | They propose a multimodal sentiment knowledge-sharing framework that unifies MSA and ERC tasks from features, labels, and models. |
| Outcome: | The proposed framework achieves consistent improvements on four public benchmark datasets on MOSI, MOSEI, MELD, and IEMOCAP. |
Do Audio LLMs Really LISTEN, or Just Transcribe? Measuring Lexical vs. Acoustic Emotion Cues Reliance (2026.eacl-long)
Copied to clipboard
| Challenge: | LISTEN is a controlled benchmark to disentangle lexical reliance from acoustic sensitivity in emotion understanding. |
| Approach: | They propose a benchmark to disentangle lexical reliance from acoustic sensitivity in emotion understanding. |
| Outcome: | LISTEN shows that current LALMs largely "transcribe" rather than "listen" authors note that models underutilize acoustic cues while relying on lexical semantics . |