Challenge: Existing studies focus on developing models that exploit the unification of multiple modalities.
Approach: They propose to maintain modality independence by using a multi-modal transformer model that fuses all modalities.
Outcome: The proposed model outperforms state-of-the-art models in multi-modal emotion recognition.

Similar Papers

Multi-modal Multi-label Emotion Detection with Modality and Label Dependence (2020.emnlp-main)

Copied to clipboard

Challenge: Existing studies on multi-label emotion detection focus on one modality . current studies focus on label dependence, but there is no consensus on the model .
Approach: They propose a multi-modal sequence-to-set approach to model label dependence and modality dependence in a multiple-modal scenario.
Outcome: The proposed approach is able to model the label dependence and the modality dependence in a multi-modal scenario.
Self-supervised Cross-modal Pretraining for Speech Emotion Recognition and Sentiment Analysis (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches to multimodal speech emotion recognition and sentiment analysis have not improved results due to their relatively simple fusion mechanisms and lack of proper cross-modal pretraining.
Approach: They propose a deep-fused audio-text bi-modal transformer with carefully designed cross-modal fusion mechanism and stage-wise cross-mod pretraining scheme to facilitate cross-modulation.
Outcome: The proposed method exceeds benchmarks on public IEMOCAP emotion and CMU-MOSEI sentiment datasets by a large margin.
Modality-Transferable Emotion Embeddings for Low-Resource Multimodal Emotion Recognition (2020.aacl-main)

Copied to clipboard

Challenge: despite recent advances in multimodal emotion recognition, two problems still exist: sub-optimal performance and low-resource emotions.
Approach: They propose a modality-transferable model with emotion embeddings to solve these problems . they use pre-trained word embedders to represent emotion categories for textual data .
Outcome: The proposed model outperforms baselines in zero-shot and few-shot scenarios for unseen emotions.
Joyful: Joint Modality Fusion and Graph Contrastive Learning for Multimoda Emotion Recognition (2023.emnlp-main)

Copied to clipboard

Challenge: Existing graph-based methods fail to depict global contextual features and local diverse unimodal features in a dialogue.
Approach: They propose a method for joint modality fusion and graph contrastive learning for multimodal emotion recognition using a multimodal fusion mechanism and a graph contrastative learning framework.
Outcome: The proposed method improves multimodal emotion recognition on unbalanced and small-scale emotional datasets.
Self-adaptive Context and Modal-interaction Modeling For Multimodal Emotion Recognition (2023.findings-acl)

Copied to clipboard

Challenge: Existing methods to predict emotion label for a given utterance lack modeling of diverse dependency ranges and inconsistent treatment of contribution for various modalities.
Approach: They propose a multimodal emotion recognition in conversation task that uses context and multiple modalities to predict emotion label for a given utterance.
Outcome: The proposed method outperforms the state-of-the-art methods on three multimodal datasets.
Word-Aware Modality Stimulation for Multimodal Fusion (2024.lrec-main)

Copied to clipboard

Challenge: Multimodal learning is expected to make more accurate predictions than text-only analysis.
Approach: They propose a method for fusing multimodal inputs with text-based fusion methods . they propose fusion that integrates non-verbal modalities with text .
Outcome: The proposed method improves sentiment prediction by using non-verbal modalities with text . the proposed method is unsuitable for applying attention to text modality in the fusion phase .
Integrating Representation Subspace Mapping with Unimodal Auxiliary Loss for Attention-based Multimodal Emotion Recognition (2024.lrec-main)

Copied to clipboard

Challenge: Existing methods to identify emotions rely on a large modality gap in their representations .
Approach: They propose a representation subspace mapping module that maps each modality into two distinct subspaces and a cross-modality attention module that leverages auxiliary loss to remove the noise unrelated to emotion classification.
Outcome: The proposed approach achieves superior performance to state-of-the-art MER methods on the IEMOCAP and MSP-Improv datasets.
Capturing Latent Modal Association For Multimodal Entity Alignment (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for multimodal entity alignment overlook the quality of input modality embeddings during modality interaction, amplifying noise propagation while suppressing discriminative feature representations.
Approach: They propose a model for capturing latent modal association for multimodal entity alignment using a self-attention mechanism to enhance salient information while attenuating noise within individual modality embeddings.
Outcome: The proposed model achieves an absolute 3.1% higher Hits@1 score than the sota method.
Multimodal Multi-loss Fusion Network for Sentiment Analysis (2024.naacl-long)

Copied to clipboard

Challenge: This paper examines the optimal selection and fusion of feature encoders across multiple modalities and combines them in one neural network to improve sentiment detection.
Approach: They propose to combine feature encoders across multiple modalities into one neural network to improve sentiment detection.
Outcome: The proposed model achieves state-of-the-art performance for three datasets . it also shows that integrating context significantly improves model performance.
Context-aware Interactive Attention for Multi-modal Sentiment and Emotion Analysis (D19-1)

Copied to clipboard

Challenge: Multi-modal analysis is a field emerging in the fields of natural language processing, computer vision and speech processing . multimodal analysis uses a variety of information from multiple sources to build efficient systems . acoustic and visual information can provide better information for classification decisions .
Approach: They propose a recurrent neural network based approach for multi-modal sentiment and emotion analysis . they employ a context-aware attention module to exploit the correspondence among neighboring utterances .
Outcome: The proposed model learns inter-modal interaction among participating modalities through auto-encoder mechanism . it is compared with existing state-of-the-art models on five standard multi-modal affect analysis datasets .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations