Challenge: Existing approaches to classify human affect and subjective information from multiple data sources are limited by the lack of high-level feature associations.
Approach: They propose a hierarchical multimodal architecture with attention and word-level fusion to classify utterance-level sentiment and emotion from text and audio data.
Outcome: The proposed model outperforms state-of-the-art approaches on published datasets and visualizes and interprets synchronized attention over modalities.

Similar Papers

Modeling Intra- and Inter-Modal Relations: Hierarchical Graph Contrastive Learning for Multimodal Sentiment Analysis (2022.coling-1)

Copied to clipboard

Challenge: Existing studies in Multimodal Sentiment Analysis lack a mechanism to understand complex relations between different modalities.
Approach: They propose a hierarchical graph contrastive learning framework for multimodal sentiment analysis that explores the relationships between modality representations.
Outcome: The proposed framework outperforms the state-of-the-art in multimodal sentiment analysis on two benchmark datasets.
Joyful: Joint Modality Fusion and Graph Contrastive Learning for Multimoda Emotion Recognition (2023.emnlp-main)

Copied to clipboard

Challenge: Existing graph-based methods fail to depict global contextual features and local diverse unimodal features in a dialogue.
Approach: They propose a method for joint modality fusion and graph contrastive learning for multimodal emotion recognition using a multimodal fusion mechanism and a graph contrastative learning framework.
Outcome: The proposed method improves multimodal emotion recognition on unbalanced and small-scale emotional datasets.
CLMLF:A Contrastive Learning and Multi-Layer Fusion Method for Multimodal Sentiment Detection (2022.findings-naacl)

Copied to clipboard

Challenge: Existing methods for multimodal sentiment detection do not consider token-level feature fusion.
Approach: They propose a method for multimodal sentiment detection using a combination of text and image to encode and fuse token-level features.
Outcome: The proposed method can fuse multimodal features with token-level features on three publicly available multimodal datasets.
Multimodal Contrastive Learning via Uni-Modal Coding and Cross-Modal Prediction for Multimodal Sentiment Analysis (2022.findings-emnlp)

Copied to clipboard

Challenge: Recent work on multimodal representation learning has focused on uni-modality pre-training or cross-modalities integration.
Approach: They propose a framework for multimodal representation learning that uses uni-modal contrastive coding and an efficient unimodal feature augmentation strategy to capture intermodal dynamics.
Outcome: The proposed framework surpasses state-of-the-art methods on two public datasets.
Context-aware Interactive Attention for Multi-modal Sentiment and Emotion Analysis (D19-1)

Copied to clipboard

Challenge: Multi-modal analysis is a field emerging in the fields of natural language processing, computer vision and speech processing . multimodal analysis uses a variety of information from multiple sources to build efficient systems . acoustic and visual information can provide better information for classification decisions .
Approach: They propose a recurrent neural network based approach for multi-modal sentiment and emotion analysis . they employ a context-aware attention module to exploit the correspondence among neighboring utterances .
Outcome: The proposed model learns inter-modal interaction among participating modalities through auto-encoder mechanism . it is compared with existing state-of-the-art models on five standard multi-modal affect analysis datasets .
Improving Multimodal Fusion with Hierarchical Mutual Information Maximization for Multimodal Sentiment Analysis (2021.emnlp-main)

Copied to clipboard

Challenge: Existing work on multimodal sentiment analysis relies on back-propagated task loss or geometric property of feature spaces to produce favorable fusion results.
Approach: They propose a framework which hierarchically maximizes the Mutual Information (MI) in unimodal input pairs and between multimodal fusion result and unimod input to maintain task-related information through multimodal integration.
Outcome: The proposed framework maximizes the Mutual Information (MI) in unimodal input pairs and between multimodal fusion result and unimodulated input to maintain task-related information through multimodal integration.
MultiEMO: An Attention-Based Correlation-Aware Multimodal Fusion Framework for Emotion Recognition in Conversations (2023.acl-long)

Copied to clipboard

Challenge: Emotion Recognition in Conversations (ERC) is an increasingly popular task in the field of Natural Language Processing.
Approach: They propose a framework that captures cross-modal mapping relationships across modalities . they propose 'multiemotion-aware' framework that integrates multimodal cues into the model .
Outcome: The proposed framework outperforms state-of-the-art models in all emotion categories on two benchmark datasets.
Divide, Conquer and Combine: Hierarchical Feature Fusion Network with Local and Global Perspectives for Multimodal Affective Computing (P19-1)

Copied to clipboard

Challenge: Existing approaches to multimodal fusion are based on fusing features at holistic level instead of focusing on local and global interactions.
Approach: They propose a general strategy called ‘divide, conquer and combine’ for multimodal fusion that combines local and global interactions in a hierarchy.
Outcome: The proposed strategy achieves state-of-the-art performance on multimodal affective computing with higher efficiency.
Hybrid Attention based Multimodal Network for Spoken Language Classification (C18-1)

Copied to clipboard

Challenge: Using linguistic content and vocal characteristics for multimodal deep learning is difficult for computers to interpret human meaning .
Approach: They propose a deep multimodal network with feature attention and modality attention to classify utterance-level speech data.
Outcome: The proposed system achieves state-of-the-art or competitive results on three published multimodal datasets.
CTFN: Hierarchical Learning for Multimodal Sentiment Analysis Using Coupled-Translation Fusion Network (2021.acl-long)

Copied to clipboard

Challenge: Existing methods for multimodal sentiment analysis require all modalities as input, thus are sensitive to missing modality at predicting time.
Approach: They propose to model bi-direction interplay via couple learning and exploit multiple bi-directional translations to exploit multimodal fusion embeddings.
Outcome: The proposed framework achieves state-of-the-art or often competitive performance on two multimodal benchmarks with extensive ablation studies.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations