Multimodal Routing: Improving Local and Global Interpretability of Multimodal Language Analysis (2020.emnlp-main)
Copied to clipboard
| Challenge: | Recent multimodal learning models with strong performances on human-centric tasks are often black-box with very limited interpretability. |
| Approach: | They propose a multimodal routing algorithm which dynamically adjusts weights between input and output modalities for each input sample. |
| Outcome: | The proposed model can interpret modality-prediction relationships globally and locally for each input sample while keeping competitive performance compared to state-of-the-art methods. |
Similar Papers
Multimodality for NLP-Centered Applications: Resources, Advances and Frontiers (2022.lrec-1)
Copied to clipboard
| Challenge: | resurgence of multimodal datasets has attracted significant research interest, but there is no comprehensive survey for this task. |
| Approach: | They present a survey of a multimodal dataset with different modalities according to the applications. |
| Outcome: | The proposed datasets are available online and discuss the new frontier and motivate future researches. |
Multimodal Language Analysis in the Wild: CMU-MOSEI Dataset and Interpretable Dynamic Fusion Graph (P18-1)
Copied to clipboard
| Challenge: | Analyzing human multimodal language is emerging area of research in NLP. |
| Approach: | They propose a multimodal fusion technique to exploit how modalities interact in multimodal language. |
| Outcome: | The proposed technique exploits how modalities interact with each other in human multimodal language. |
Analyzing Modality Robustness in Multimodal Sentiment Analysis (2022.naacl-main)
Copied to clipboard
| Challenge: | despite its importance, little attention has been paid to improving the robustness of multimodal models. |
| Approach: | They propose simple diagnostic checks for modality robustness in a trained multimodal model . they find MSA models highly sensitive to a single modality, which creates issues . |
| Outcome: | The proposed checks show that models are highly sensitive to a single modality, which creates issues in their robustness. |
Explainability and Interpretability of Multilingual Large Language Models: A Survey (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing literature on multilingual large language models lacks transparency in their internal processes. |
| Approach: | They propose to use multilingual large language models to examine their explainability and interpretability methods. |
| Outcome: | The present study examines the explainability and interpretability of multilingual large language models. |
Multimodal Grounding for Language Processing (C18-1)
Copied to clipboard
| Challenge: | Recent developments in multimodal processing facilitate conceptual grounding of language. |
| Approach: | They analyze multimodal processing to examine the benefits and challenges of multimodal grounding . they focus on multimodal linguistic grounding of verbs which play a crucial role in compositional power of language. |
| Outcome: | The proposed methods improve the cognitive models of human information processing and address the challenges that arise. |
Beyond Cross-Modal Alignment: Measuring and Leveraging Modality Gap in Vision-Language Models (2026.findings-acl)
Copied to clipboard
| Challenge: | a recent study shows that vision-language models have modality gaps that persist even in well-aligned models. |
| Approach: | They propose a modality-dominance score to measure and leverage modality gaps . they propose automatic interpretability metrics to evaluate these features in a scalable manner . |
| Outcome: | The proposed framework allows for training-free probing and editing methods for understanding model perception across genders and generating adversarial examples. |
Cross-Modal Attribute Insertions for Assessing the Robustness of Vision-and-Language Learning (2023.acl-long)
Copied to clipboard
| Challenge: | Existing approaches to model multimodal data do not leverage cross-modal information . augmenting input text using cross-module attribute insertions results in poor performance . |
| Approach: | They propose a multimodal deep learning approach that adds visual attributes to inputs to enhance model robustness. |
| Outcome: | The proposed approach is modular, controllable, and task-agnostic. |
Learning Language-guided Adaptive Hyper-modality Representation for Multimodal Sentiment Analysis (2023.emnlp-main)
Copied to clipboard
| Challenge: | Multimodal Sentiment Analysis (MSA) is effective when using rich information from multiple sources, but the potential sentiment-irrelevant information across modalities may hinder the performance from being further improved. |
| Approach: | They propose an Adaptive Language-guided Multimodal Transformer (ALMT) that learns an irrelevance/conflict-suppressing representation from visual and audio features under guidance of language features at different scales. |
| Outcome: | The proposed model achieves state-of-the-art on several popular datasets and an abundance of ablation shows the effectiveness of the proposed model. |
UniMSE: Towards Unified Multimodal Sentiment Analysis and Emotion Recognition (2022.emnlp-main)
Copied to clipboard
| Challenge: | Existing studies study sentiment and emotion separately and do not fully exploit the complementary knowledge behind the two. |
| Approach: | They propose a multimodal sentiment knowledge-sharing framework that unifies MSA and ERC tasks from features, labels, and models. |
| Outcome: | The proposed framework achieves consistent improvements on four public benchmark datasets on MOSI, MOSEI, MELD, and IEMOCAP. |
Sentiment Word Aware Multimodal Refinement for Multimodal Sentiment Analysis with ASR Errors (2022.findings-acl)
Copied to clipboard
| Challenge: | Existing models for multimodal sentiment analysis are limited in their capacity to be deployed in the real world. |
| Approach: | They propose a model that can dynamically refine erroneous sentiment words by leveraging multimodal sentiment clues. |
| Outcome: | The proposed model surpasses the state-of-the-art models on three datasets. |