Challenge: Existing methods to analyze images focus on superficial features or descriptions, omitting subtle contextual information.
Approach: They propose a Visual Connotation and Aesthetic Attributes Understanding Network (Vanessa) for Multimodal Aspect-based Sentiment Analysis.
Outcome: The proposed network captures both implicit and explicit sentimental cues and can be used to enrich textual sentiment analysis.

Similar Papers

AoM: Detecting Aspect-oriented Information for Multimodal Aspect-Based Sentiment Analysis (2023.findings-acl)

Copied to clipboard

Challenge: Existing methods to extract aspects from text-image pairs and recognize their sentiments are noisy and coarsely establishing image-aspect alignment will interfere with aspect-relevant semantic and sentiment information.
Approach: They propose an Aspect-oriented method to detect aspect-relevant semantic and sentiment information by selecting textual tokens and image blocks that are semantically related to the aspects.
Outcome: The proposed method is superior to existing methods in the field of sentiment analysis.
Learning from Adjective-Noun Pairs: A Knowledge-enhanced Framework for Target-Oriented Multimodal Sentiment Classification (2022.coling-1)

Copied to clipboard

Challenge: Existing methods to determine sentiment polarity of opinion target are inconsistent and lack visual attention.
Approach: They propose a framework which can exploit adjective-noun pairs extracted from images to improve visual attention and sentiment prediction capability of the TMSC task.
Outcome: The proposed framework outperforms state-of-the-art on two public datasets.
Face-Sensitive Image-to-Emotional-Text Cross-modal Translation for Multimodal Aspect-based Sentiment Analysis (2022.emnlp-main)

Copied to clipboard

Challenge: Existing models focus on utilizing semantic information in the image but ignore using visual emotional cues.
Approach: They propose a face-sensitive image-to-emotional-text translation method that captures visual emotional cues through facial expressions and selectively matches and fuses with the textual content.
Outcome: The proposed method achieves state-of-the-art results on the Twitter-2015 and Twitter-2017 datasets.
Affective Knowledge Enhanced Multiple-Graph Fusion Networks for Aspect-based Sentiment Analysis (2022.emnlp-main)

Copied to clipboard

Challenge: Existing methods for sentiment analysis ignore the roles of syntax dependency relation labels and affective semantic information in determining the sentiment polarity of social media users.
Approach: They propose a new multi-graph fusion network to leverage the richer syntax dependency relation labels and affective semantic information of words.
Outcome: The proposed model outperforms state-of-the-art methods on three datasets.
MM-GATBT: Enriching Multimodal Representation Using Graph Attention Network (2022.naacl-srw)

Copied to clipboard

Challenge: Existing models that use a self-attention mechanism to create graphs with multiple modes ignore interaction between entities, multimodalities, or both.
Approach: They propose a multimodal graph representation learning model that captures relational semantics within one modality and interactions between different modalities.
Outcome: The proposed model outperforms existing models on the MM-IMDb dataset in all aspects of multimodal representation.
TMFN: A Target-oriented Multi-grained Fusion Network for End-to-end Aspect-based Multimodal Sentiment Analysis (2024.lrec-main)

Copied to clipboard

Challenge: Existing methods for multimodal aspect-based sentiment analysis focus on fusing image regional information and textual words.
Approach: They propose a multimodal aspect-based sentiment analysis method that integrates regional and global image information with global image data.
Outcome: Experiments show that the proposed method outperforms state-of-the-art methods on two benchmark datasets.
Aspect-Based Emotion Analysis and Multimodal Coreference: A Case Study of Customer Comments on Adidas Instagram Posts (2022.lrec-1)

Copied to clipboard

Challenge: Aspect-based sentiment analysis of user-generated content has been relatively unexplored in recent years.
Approach: They present a multimodal dataset for Aspect-Based Emotion Analysis (ABEA) they take the first steps in investigating the utility of multimodal coreference resolution in an ABEA framework.
Outcome: The proposed dataset consists of 4,900 comments on 175 images and is annotated with aspect and emotion categories and the emotional dimensions of valence and arousal.
Ameli: Enhancing Multimodal Entity Linking with Fine-Grained Attributes (2024.eacl-long)

Copied to clipboard

Challenge: Experimental results show that understanding attributes of mentions from text descriptions and visual images plays a vital role in multimodal entity linking.
Approach: They propose to integrate attributes into multimodal entity linking using a text-image-based knowledge base.
Outcome: The proposed approach integrates attributes into disambiguation.
Context-aware Interactive Attention for Multi-modal Sentiment and Emotion Analysis (D19-1)

Copied to clipboard

Challenge: Multi-modal analysis is a field emerging in the fields of natural language processing, computer vision and speech processing . multimodal analysis uses a variety of information from multiple sources to build efficient systems . acoustic and visual information can provide better information for classification decisions .
Approach: They propose a recurrent neural network based approach for multi-modal sentiment and emotion analysis . they employ a context-aware attention module to exploit the correspondence among neighboring utterances .
Outcome: The proposed model learns inter-modal interaction among participating modalities through auto-encoder mechanism . it is compared with existing state-of-the-art models on five standard multi-modal affect analysis datasets .
SWAFN: Sentimental Words Aware Fusion Network for Multimodal Sentiment Analysis (2020.coling-main)

Copied to clipboard

Challenge: Existing studies focus on learning the joint representation of multiple modalities, ignoring useful knowledge contained in language modal.
Approach: They propose to incorporate sentimental words knowledge into the fusion network to guide the learning of joint representation of multimodal features.
Outcome: The proposed method improves the fusion representation of multimodal features on a YouTube and video dataset.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations