Challenge: Human language is often multimodal, which comprehends a mixture of natural language, facial gestures, and acoustic behaviors.
Approach: They propose a multimodal model that extends the standard Transformer network to learn representations directly from unaligned multimodal streams.
Outcome: The proposed model outperforms state-of-the-art methods on aligned and non-aligned data.

Similar Papers

MTAG: Modal-Temporal Attention Graph for Unaligned Human Multimodal Language Sequences (2021.naacl-main)

Copied to clipboard

Challenge: a novel graph-based neural model for multimodal sequential data is proposed . fusion is the process of blending information from multiple modalities, usually preceded by alignment .
Approach: They propose a graph-based neural model that converts unaligned data into a modal-temporal graph . they use a dynamic pruning and read-out technique to efficiently process the graph fusion operation .
Outcome: The proposed model performs state-of-the-art on multimodal sentiment analysis and emotion recognition benchmarks while utilizing significantly fewer model parameters.
Aligning Text/Speech Representations from Multimodal Models with MEG Brain Activity During Listening (2025.emnlp-main)

Copied to clipboard

Challenge: Recent studies have found that speech language models fail to capture brain-relevant semantics beyond low-level features.
Approach: They analyze multimodal models to assess their alignment with MEG brain recordings . they find text embeddings from multimodal and unimodal models significantly outperform unilateral models .
Outcome: a new study shows that text-based models outperform unimodal models in alignment with brain recordings during naturalistic story listening.
Dual-Encoder Transformers with Cross-modal Alignment for Multimodal Aspect-based Sentiment Analysis (2022.aacl-main)

Copied to clipboard

Challenge: Multimodal aspect-based sentiment analysis (MABSA) aims to extract aspect terms from text and image pairs, and then analyze their corresponding sentiment.
Approach: They propose a dual-encoder transformer with cross-modal alignment to extract aspect terms from text and image pairs and then analyze their corresponding sentiments.
Outcome: The proposed approach outperforms existing methods on two benchmarks.
Rethinking the Multimodal Correlation of Multimodal Sequential Learning via Generalizable Attentional Results Alignment (2024.acl-long)

Copied to clipboard

Challenge: Existing studies have focused on the alignment of multimodal sequential learning using transformers.
Approach: They propose a constrained scheme to align the multiple attentional results from both local and global perspectives.
Outcome: The proposed scheme could align the multiple attentional results from both local and global perspectives, making the information capture more efficient.
CLASP: Cross-modal Alignment Using Pre-trained Unimodal Models (2024.findings-acl)

Copied to clipboard

Challenge: Recent advances in speech-text pretraining rely on parallel speech- text data . however, data accessibility is a challenge due to the limited data available.
Approach: They propose a framework for jointly performing speech and text processing without parallel corpora during pre-training but only downstream.
Outcome: The proposed framework extracts distinct representations for speech and text, aligning them effectively in a newly defined space using a multi-level contrastive learning mechanism.
Cross-lingual AMR Aligner: Paying Attention to Cross-Attention (2023.findings-acl)

Copied to clipboard

Challenge: Abstract Meaning Representation (AMR) graphs embed the semantics of a sentence in a directed acyclic graph, where concepts are represented by nodes, semantic relations between concepts by edges, and the co-references by reentrant nodes.
Approach: They propose a novel aligner for Abstract Meaning Representation graphs that scales cross-lingually and can align units and spans in sentences of different languages.
Outcome: The proposed aligner achieves state-of-the-art in the benchmarks and can scale cross-lingually.
Multimodal Unlearning Across Vision, Language, Video, and Audio: Survey of Methods, Datasets, and Benchmarks (2026.findings-acl)

Copied to clipboard

Challenge: Survey aims to identify challenges of multimodal unlearning for vision, language, audio and video . retraining after deletion requests or policy updates is often impractical, survey finds .
Approach: They propose to enable selective removal across modalities while retaining overall utility.
Outcome: This study compares models with existing models to identify weaknesses and improves performance.
Adaptive Transformers for Learning Multimodal Representations (2020.acl-srw)

Copied to clipboard

Challenge: Existing approaches for learning visiolinguistic representations with transformers are over-parametrized and require extensive training.
Approach: They propose to extend attention spans, sparse, and structured dropout methods to learn more about how the network perceives the complexity of input sequences.
Outcome: The proposed approaches improve on language semantics and visiolinguistic representations, but are often over-parametrized and require large amounts of computation.
Understanding Cross-Lingual Alignment—A Survey (2024.findings-acl)

Copied to clipboard

Challenge: Cross-lingual alignment is the meaningful similarity of representations across languages in multilingual language models.
Approach: They propose a taxonomy of methods to improve cross-lingual alignment . they argue that an effective trade-off between language-neutral and language-specific information is key .
Outcome: The proposed methods can be applied to encoder models and encoder-decoder-only models . they show that language-neutral and language-specific information is key .
Towards Unified Multimodal Large Language Models: A survey (2026.findings-acl)

Copied to clipboard

Challenge: unified multimodal large language models (MLLMs) are emerging but lack a systematic framework to connect them and situate current trends within a broader landscape.
Approach: They present a systematic review of unified Multimodal Large Language Models . they outline the foundational concepts and prerequisites for understanding them .
Outcome: The present review provides a systematic and systematic overview of unified MLLMs . it discusses persistent challenges and identify promising directions for future research .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations