Yao-Hung Hubert Tsai, Shaojie Bai, Paul Pu Liang, J. Zico Kolter, Louis-Philippe Morency, Ruslan Salakhutdinov
| Challenge: | Human language is often multimodal, which comprehends a mixture of natural language, facial gestures, and acoustic behaviors. |
| Approach: | They propose a multimodal model that extends the standard Transformer network to learn representations directly from unaligned multimodal streams. |
| Outcome: | The proposed model outperforms state-of-the-art methods on aligned and non-aligned data. |
Similar Papers
MTAG: Modal-Temporal Attention Graph for Unaligned Human Multimodal Language Sequences (2021.naacl-main)
Copied to clipboard
Jianing Yang, Yongxin Wang, Ruitao Yi, Yuying Zhu, Azaan Rehman, Amir Zadeh, Soujanya Poria, Louis-Philippe Morency
| Challenge: | a novel graph-based neural model for multimodal sequential data is proposed . fusion is the process of blending information from multiple modalities, usually preceded by alignment . |
| Approach: | They propose a graph-based neural model that converts unaligned data into a modal-temporal graph . they use a dynamic pruning and read-out technique to efficiently process the graph fusion operation . |
| Outcome: | The proposed model performs state-of-the-art on multimodal sentiment analysis and emotion recognition benchmarks while utilizing significantly fewer model parameters. |
Aligning Text/Speech Representations from Multimodal Models with MEG Brain Activity During Listening (2025.emnlp-main)
Copied to clipboard
Padakanti Srijith, Khushbu Pahwa, Radhika Mamidi, Bapi Raju Surampudi, Manish Gupta, Subba Reddy Oota
| Challenge: | Recent studies have found that speech language models fail to capture brain-relevant semantics beyond low-level features. |
| Approach: | They analyze multimodal models to assess their alignment with MEG brain recordings . they find text embeddings from multimodal and unimodal models significantly outperform unilateral models . |
| Outcome: | a new study shows that text-based models outperform unimodal models in alignment with brain recordings during naturalistic story listening. |
Dual-Encoder Transformers with Cross-modal Alignment for Multimodal Aspect-based Sentiment Analysis (2022.aacl-main)
Copied to clipboard
| Challenge: | Multimodal aspect-based sentiment analysis (MABSA) aims to extract aspect terms from text and image pairs, and then analyze their corresponding sentiment. |
| Approach: | They propose a dual-encoder transformer with cross-modal alignment to extract aspect terms from text and image pairs and then analyze their corresponding sentiments. |
| Outcome: | The proposed approach outperforms existing methods on two benchmarks. |
Rethinking the Multimodal Correlation of Multimodal Sequential Learning via Generalizable Attentional Results Alignment (2024.acl-long)
Copied to clipboard
| Challenge: | Existing studies have focused on the alignment of multimodal sequential learning using transformers. |
| Approach: | They propose a constrained scheme to align the multiple attentional results from both local and global perspectives. |
| Outcome: | The proposed scheme could align the multiple attentional results from both local and global perspectives, making the information capture more efficient. |
CLASP: Cross-modal Alignment Using Pre-trained Unimodal Models (2024.findings-acl)
Copied to clipboard
| Challenge: | Recent advances in speech-text pretraining rely on parallel speech- text data . however, data accessibility is a challenge due to the limited data available. |
| Approach: | They propose a framework for jointly performing speech and text processing without parallel corpora during pre-training but only downstream. |
| Outcome: | The proposed framework extracts distinct representations for speech and text, aligning them effectively in a newly defined space using a multi-level contrastive learning mechanism. |
Cross-lingual AMR Aligner: Paying Attention to Cross-Attention (2023.findings-acl)
Copied to clipboard
| Challenge: | Abstract Meaning Representation (AMR) graphs embed the semantics of a sentence in a directed acyclic graph, where concepts are represented by nodes, semantic relations between concepts by edges, and the co-references by reentrant nodes. |
| Approach: | They propose a novel aligner for Abstract Meaning Representation graphs that scales cross-lingually and can align units and spans in sentences of different languages. |
| Outcome: | The proposed aligner achieves state-of-the-art in the benchmarks and can scale cross-lingually. |
Multimodal Unlearning Across Vision, Language, Video, and Audio: Survey of Methods, Datasets, and Benchmarks (2026.findings-acl)
Copied to clipboard
| Challenge: | Survey aims to identify challenges of multimodal unlearning for vision, language, audio and video . retraining after deletion requests or policy updates is often impractical, survey finds . |
| Approach: | They propose to enable selective removal across modalities while retaining overall utility. |
| Outcome: | This study compares models with existing models to identify weaknesses and improves performance. |
Adaptive Transformers for Learning Multimodal Representations (2020.acl-srw)
Copied to clipboard
| Challenge: | Existing approaches for learning visiolinguistic representations with transformers are over-parametrized and require extensive training. |
| Approach: | They propose to extend attention spans, sparse, and structured dropout methods to learn more about how the network perceives the complexity of input sequences. |
| Outcome: | The proposed approaches improve on language semantics and visiolinguistic representations, but are often over-parametrized and require large amounts of computation. |
Understanding Cross-Lingual Alignment—A Survey (2024.findings-acl)
Copied to clipboard
| Challenge: | Cross-lingual alignment is the meaningful similarity of representations across languages in multilingual language models. |
| Approach: | They propose a taxonomy of methods to improve cross-lingual alignment . they argue that an effective trade-off between language-neutral and language-specific information is key . |
| Outcome: | The proposed methods can be applied to encoder models and encoder-decoder-only models . they show that language-neutral and language-specific information is key . |
Towards Unified Multimodal Large Language Models: A survey (2026.findings-acl)
Copied to clipboard
| Challenge: | unified multimodal large language models (MLLMs) are emerging but lack a systematic framework to connect them and situate current trends within a broader landscape. |
| Approach: | They present a systematic review of unified Multimodal Large Language Models . they outline the foundational concepts and prerequisites for understanding them . |
| Outcome: | The present review provides a systematic and systematic overview of unified MLLMs . it discusses persistent challenges and identify promising directions for future research . |