From the Token to the Review: A Hierarchical Multimodal approach to Opinion Mining (D19-1)
Copied to clipboard
| Challenge: | Existing work on fine grained opinion annotations rely only on coarsely labeled opinions. |
| Approach: | They propose to use hierarchical structure of opinions to build a fine and coarse grained opinion model that exploits different views of the opinion expression. |
| Outcome: | The proposed model outperforms existing models on a recently released multimodal fine grained annotated corpus on IMDB and social networks. |
Similar Papers
Aspect-Based Sentiment Analysis as Fine-Grained Opinion Mining (2020.lrec-1)
Copied to clipboard
| Challenge: | a large body of research has been done on aspect-based sentiment analysis (ABSA) for almost two decades . aspect-Based sentiment analysis is a task that extracts sentiment/opinions from text in terms of targets . |
| Approach: | They propose a meaning-preserving annotation scheme for aspect-based sentiment analysis . they then apply it to two popular ABSA datasets to examine their results . |
| Outcome: | The proposed approach improves the state of aspect-based sentiment analysis (ABSA) by preserving the meaning of the sentiment. |
From Coarse to Fine: A Multi-Granularity Multimodal Framework for Teacher Sentiment Analysis (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing approaches to teacher sentiment analysis treat it as a static label . current approaches fail to capture structured heterogeneity of classroom expressions . |
| Approach: | They propose a coarse-to-fine multimodal framework that decomposes teacher sentiment into three granularities and employ CLS-guided cross-modal attention to recover effective signals from regulated displays. |
| Outcome: | The proposed framework outperforms state-of-the-art models on T-MED and CMU-MOSEI. |
Token Pruning in Multimodal Large Language Models: Are We Solving the Right Problem? (2025.findings-acl)
Copied to clipboard
| Challenge: | Multimodal large language models have shown remarkable performance for cross-modal understanding and generation, yet suffer from severe inference costs. |
| Approach: | They propose to prune redundant tokens in MLLMs to reduce computation and storage costs. |
| Outcome: | The proposed method reduces the computational and storage costs of MLLMs by identifying redundant tokens and pruning them. |
Self-Supervised Multimodal Opinion Summarization (2021.acl-long)
Copied to clipboard
| Challenge: | Existing methods for opinion summarization use text data, but non-text data are less abundant. |
| Approach: | They propose a self-supervised opinion summarization framework that uses non-text data to generate a summary from multiple reviews. |
| Outcome: | The proposed framework is superior to existing methods on Yelp and Amazon datasets. |
An Empirical Examination of Online Restaurant Reviews (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing methods for opinion mining and sentiment analysis focus on extracting either positive or negative opinions from texts and determining the targets of these opinions. |
| Approach: | They propose a corpus-based scheme that detects evaluative language at a finer-grained level. |
| Outcome: | The proposed scheme classifies each sentence into one of four evaluation types based on the proposed scheme. |
Multimodality for NLP-Centered Applications: Resources, Advances and Frontiers (2022.lrec-1)
Copied to clipboard
| Challenge: | resurgence of multimodal datasets has attracted significant research interest, but there is no comprehensive survey for this task. |
| Approach: | They present a survey of a multimodal dataset with different modalities according to the applications. |
| Outcome: | The proposed datasets are available online and discuss the new frontier and motivate future researches. |
Can Large Language Models be Effective Online Opinion Miners? (2025.emnlp-main)
Copied to clipboard
| Challenge: | OOMB is a novel benchmark designed to assess the ability of large language models (LLMs) to extract and analyze opinions from diverse and complex online environments. |
| Approach: | They propose an online opinion mining benchmark to assess the ability of large language models to extract and analyze opinions from diverse online environments. |
| Outcome: | The proposed benchmark assesses the ability of large language models to mine opinions effectively from diverse and complex online environments. |
TMFN: A Target-oriented Multi-grained Fusion Network for End-to-end Aspect-based Multimodal Sentiment Analysis (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing methods for multimodal aspect-based sentiment analysis focus on fusing image regional information and textual words. |
| Approach: | They propose a multimodal aspect-based sentiment analysis method that integrates regional and global image information with global image data. |
| Outcome: | Experiments show that the proposed method outperforms state-of-the-art methods on two benchmark datasets. |
Late Fusion with Triplet Margin Objective for Multimodal Ideology Prediction and Analysis (2022.emnlp-main)
Copied to clipboard
| Challenge: | Prior work on ideology prediction has focused on single modalities, i.e., text or images. |
| Approach: | They propose a task where a model predicts binary or five-point scale ideological leanings given a text-image pair with political content. |
| Outcome: | The proposed model outperforms the state-of-the-art model by almost 4% and a strong multimodal baseline with no pretraining by over 3%. |
Multimodal Affective Analysis Using Hierarchical Attention Strategy with Word-Level Alignment (P18-1)
Copied to clipboard
| Challenge: | Existing approaches to classify human affect and subjective information from multiple data sources are limited by the lack of high-level feature associations. |
| Approach: | They propose a hierarchical multimodal architecture with attention and word-level fusion to classify utterance-level sentiment and emotion from text and audio data. |
| Outcome: | The proposed model outperforms state-of-the-art approaches on published datasets and visualizes and interprets synchronized attention over modalities. |