Challenge: despite its importance, little attention has been paid to improving the robustness of multimodal models.
Approach: They propose simple diagnostic checks for modality robustness in a trained multimodal model . they find MSA models highly sensitive to a single modality, which creates issues .
Outcome: The proposed checks show that models are highly sensitive to a single modality, which creates issues in their robustness.

Similar Papers

Which is Making the Contribution: Modulating Unimodal and Cross-modal Dynamics for Multimodal Sentiment Analysis (2021.findings-emnlp)

Copied to clipboard

Challenge: Recent studies focus on learning cross-modal dynamics, but neglect to explore optimal solution for unimodal networks.
Approach: They propose a new MSA framework to identify contribution of modalities and reduce impact of noisy information.
Outcome: The proposed model outperforms state-of-the-art methods on publicly available datasets.
Proxy-Driven Robust Multimodal Sentiment Analysis with Incomplete Data (2025.acl-long)

Copied to clipboard

Challenge: Existing studies focus on optimizing model structures to handle uncertain missingness, but models still face challenges when dealing with uncertain missing data.
Approach: They propose a data-centric robust multimodal sentiment analysis method, Proxy-Driven Robust Multimodal Fusion, which maps unimodal data to the latent space of Gaussian distributions to capture core features and structure.
Outcome: The proposed method outperforms existing models in noise resistance and achieves state-of-the-art performance on multiple benchmark datasets.
Two Challenges, One Solution: Robust Multimodal Learning through Dynamic Modality Recognition and Enhancement (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing methods require full-modality data during training phase or require explicit annotations to detect missing modalities.
Approach: They propose a Dynamic modality Recognition and Enhancement for Adaptive Multimodal fusion framework that directs selective reconstruction of missing or underperforming modalities.
Outcome: The proposed framework outperforms several baseline and state-of-the-art models on three benchmark datasets.
DEAR: Distributional Error-Aware Reliability for Robust Multimodal Sentiment Analysis with Missing Modalities (2026.findings-acl)

Copied to clipboard

Challenge: Existing methods focus on feature completion but neglect semantic shifts caused by distribution gaps and decision risks under high uncertainty.
Approach: They propose a distributional error-aware reliability estimation framework for robust MSA . they propose reconstructed features to be explicitly aligned with original distributional manifold .
Outcome: The proposed framework mitigates semantic shifts by aligning reconstructed features with original distributional manifold . Extensive experiments on MOSI, MOSEI, and SIMS validate the framework .
QA-MoE: Towards a Continuous Reliability Spectrum with Quality-Aware Mixture of Experts for Robust Multimodal Sentiment Analysis (2026.acl-long)

Copied to clipboard

Challenge: Existing models that use multimodal inputs are often noisy or incomplete.
Approach: They propose a Quality-Aware Mixture-of-Experts framework that quantifies modality reliability via aleatoric uncertainty.
Outcome: The proposed framework is competitive or state-of-the-art across diverse degradation scenarios and exhibits a promising One-Checkpoint-for-all property in practice.
Modal Feature Optimization Network with Prompt for Multimodal Sentiment Analysis (2025.coling-main)

Copied to clipboard

Challenge: Multimodal sentiment analysis(MSA) is used to understand human emotional states through multimodal.
Approach: They propose a Modal Feature Optimization Network with a modal prompt attention mechanism to optimize the under-optimized modal representation by determining which modalities are under- optimized .
Outcome: The proposed method outperforms existing state-of-the-art models on public benchmark datasets.
Knowledge-Guided Dynamic Modality Attention Fusion Framework for Multimodal Sentiment Analysis (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing methods focus on equally treating the contribution of each modality or statically using text as the dominant modality to conduct interaction, which neglects the situation where each modal may become dominant.
Approach: They propose a Knowledge-Guided Dynamic Modality Attention Fusion Framework (KuDA) that uses sentiment knowledge to guide the model dynamically selecting the dominant modality and adjusting the contributions of each modality.
Outcome: The proposed model can be used to highlight the contribution of dominant modality through the correlation evaluation loss.
Learning Language-guided Adaptive Hyper-modality Representation for Multimodal Sentiment Analysis (2023.emnlp-main)

Copied to clipboard

Challenge: Multimodal Sentiment Analysis (MSA) is effective when using rich information from multiple sources, but the potential sentiment-irrelevant information across modalities may hinder the performance from being further improved.
Approach: They propose an Adaptive Language-guided Multimodal Transformer (ALMT) that learns an irrelevance/conflict-suppressing representation from visual and audio features under guidance of language features at different scales.
Outcome: The proposed model achieves state-of-the-art on several popular datasets and an abundance of ablation shows the effectiveness of the proposed model.
Self-Supervised Unimodal Label Generation Strategy Using Recalibrated Modality Representations for Multimodal Sentiment Analysis (2023.findings-eacl)

Copied to clipboard

Challenge: Multimodal sentiment analysis (MSA) has gained much attention over the last few years due to a lack of unimodal annotations in benchmark datasets.
Approach: They propose a framework which integrates multimodal and unimodal tasks to optimize learning representations from multimodal data.
Outcome: The proposed model learns to weight features differently based on features of other modalities and auto-generates unimodal annotations via a unimodule.
Multimodal Multi-loss Fusion Network for Sentiment Analysis (2024.naacl-long)

Copied to clipboard

Challenge: This paper examines the optimal selection and fusion of feature encoders across multiple modalities and combines them in one neural network to improve sentiment detection.
Approach: They propose to combine feature encoders across multiple modalities into one neural network to improve sentiment detection.
Outcome: The proposed model achieves state-of-the-art performance for three datasets . it also shows that integrating context significantly improves model performance.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations