Representation Potentials of Foundation Models for Multimodal Alignment: A Survey (2025.emnlp-main)
Copied to clipboard
| Challenge: | foundation models learn highly transferable representations through large-scale pretraining on diverse data. |
| Approach: | They examine the representation potentials of foundation models by examining their latent capacity to capture task-specific information within a single modality while providing a transferable basis for alignment and unification across modalities. |
| Outcome: | The foundation models exhibit remarkable similarities across architectures and modalities, the authors show . the models can capture task-specific information within a single modality while providing a transferable basis for alignment and unification across modality. |
Similar Papers
How do Multimodal Foundation Models Encode Text and Speech? An Analysis of Cross-Lingual and Cross-Modal Representations (2025.naacl-short)
Copied to clipboard
| Challenge: | Recent advances in foundation models have sparked growing interest in expanding their text processing capabilities to speech. |
| Approach: | They analyze the model activations from semantically equivalent sentences across languages in the text and speech modalities and examine how text and spoken are represented in recent multimodal foundation models. |
| Outcome: | The proposed models exhibit cross-lingual differences, but are not explicitly trained for modality-agnostic representations. |
Multimodal Unlearning Across Vision, Language, Video, and Audio: Survey of Methods, Datasets, and Benchmarks (2026.findings-acl)
Copied to clipboard
| Challenge: | Survey aims to identify challenges of multimodal unlearning for vision, language, audio and video . retraining after deletion requests or policy updates is often impractical, survey finds . |
| Approach: | They propose to enable selective removal across modalities while retaining overall utility. |
| Outcome: | This study compares models with existing models to identify weaknesses and improves performance. |
Filling in the Mechanisms: How do LMs Learn Filler-Gap Dependencies under Developmental Constraints? (2026.findings-acl)
Copied to clipboard
| Challenge: | Language models lack language-specific biases, yet still posit some important syntactic generalizations. |
| Approach: | They applied Distributed Alignment Search to checkpoints of a language model from the BabyLM challenge to evaluate whether representations of filler-gap dependencies transfer between wh-questions and topicalization. |
| Outcome: | The results suggest shared, yet item-sensitive mechanisms may develop with limited training data. |
Word Representation Learning in Multimodal Pre-Trained Transformers: An Intrinsic Evaluation (2021.tacl-1)
Copied to clipboard
| Challenge: | Existing models for linguistic representations of words are based on information extracted from large text corpora, and the sensory-motor experiences humans have with the world play an important role in determining word meaning. |
| Approach: | They propose to use contextualized word representations to learn semantic representations of words that align with human semantic intuitions. |
| Outcome: | The proposed models are shown to be more efficient on concrete word pairs than on abstract ones. |
A Survey on Training-free Alignment of Large Language Models (2025.findings-emnlp)
Copied to clipboard
Birong Pan, Yongqi Li, Weiyu Zhang, Wenpeng Lu, Mayi Xu, Shen Zhou, Yuanyuan Zhu, Ming Zhong, Tieyun Qian
| Challenge: | a survey of large language models (LLMs) aims to ensure outputs adhere to human values, ethical standards, and legal norms. |
| Approach: | They present the first systematic review of TF alignment methods . they categorize them by stages of pre-decoding, in-decoder and post-decoration . |
| Outcome: | The proposed methods are based on training-free (TF) alignment techniques . they are able to be used in open-source and closed-source environments without retraining . |
Multimodality for NLP-Centered Applications: Resources, Advances and Frontiers (2022.lrec-1)
Copied to clipboard
| Challenge: | resurgence of multimodal datasets has attracted significant research interest, but there is no comprehensive survey for this task. |
| Approach: | They present a survey of a multimodal dataset with different modalities according to the applications. |
| Outcome: | The proposed datasets are available online and discuss the new frontier and motivate future researches. |
Foundation Models Meet Embodied Agents (2025.naacl-tutorial)
Copied to clipboard
| Challenge: | This tutorial will present a systematic overview of recent advances in foundation models for embodied agents . |
| Approach: | This tutorial will present a systematic overview of recent advances in foundation models for embodied agents. |
| Outcome: | This tutorial covers three types of foundation models for embodied agents . |
Retrieving Multimodal Information for Augmented Generation: A Survey (2023.findings-emnlp)
Copied to clipboard
Ruochen Zhao, Hailin Chen, Weishi Wang, Fangkai Jiao, Xuan Long Do, Chengwei Qin, Bosheng Ding, Xiaobao Guo, Minzhi Li, Xingxuan Li, Shafiq Joty
| Challenge: | Large Language Models (LLMs) are increasingly using multimodality to augment their generation ability, but there is no unified perception of at which stage and how to incorporate different modalities. |
| Approach: | They propose to use multimodality to augment Large Language Models (LLMs) this will provide scholars with a deeper understanding of the methods' applications and encourage them to adapt existing techniques to the fast-growing field of LLMs. |
| Outcome: | The proposed methods improve factuality, reasoning, interpretability, and robustness of the generated content. |
Enhancing Multimodal Unified Representations for Cross Modal Generalization (2025.findings-acl)
Copied to clipboard
Hai Huang, Yan Xia, Shengpeng Ji, Shulei Wang, Hanting Wang, Minghui Fang, Jieming Zhu, Zhenhua Dong, Sashuai Zhou, Zhou Zhao
| Challenge: | Existing studies on discrete unified representations overlook important distinctions between different dimensions of features. |
| Approach: | They propose to use a codebook to optimize unified representations from pretraining and fine- and coarse-grained disentangling to optimize the representations. |
| Outcome: | The proposed methods improve the interpretability of multimodal unified representations . they use training-free optimization of codebook and fine and coarse cross-modal disentangling . |
VIP5: Towards Multimodal Foundation Models for Recommendation (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Recent advances in foundation models have impeded the ability for these fields to benefit from each other’s advancements. |
| Approach: | They propose to use a multimodal foundation model to unify various modalities and recommendation tasks under the P5 recommendation paradigm to implement personalized prompts. |
| Outcome: | The proposed model will unify visual, textual, and personalization modalities under the P5 recommendation paradigm and will improve recommendation performance and efficiency. |