Vision-Language Models Align with Human Neural Representations in Concept Processing (2026.eacl-long)
Copied to clipboard
| Challenge: | Recent studies suggest that transformer-based vision-language models capture the multimodality of concept processing in the human brain. |
| Approach: | They analysed multiple VLMs employing different strategies to integrate visual and textual modalities, along with language-only counterparts. |
| Outcome: | The transformer-based vision-language models outperform language-only models in two experimental conditions, while only some outperformed the language-based models. |
Similar Papers
Seeing Through Words, Speaking Through Pixels: Deep Representational Alignment Between Vision and Language Models (2025.emnlp-main)
Copied to clipboard
| Challenge: | Recent studies show that deep vision-only and language-only models project inputs into a partially aligned representational space. |
| Approach: | They investigate whether a model's representational code is semantically shared . they find that alignment peaks in mid-to-late layers of both model types . |
| Outcome: | a forced-choice "Pick-a-Pic" task shows human preferences for image-caption matches are mirrored in embedding spaces across vision-language model pairs. |
Visio-Linguistic Brain Encoding (2022.coling-1)
Copied to clipboard
| Challenge: | Existing studies have failed to explore co-attentive multi-modal modeling for visual and text reasoning. |
| Approach: | They propose to use image and multi-modal Transformers to reconstruct fMRI brain activity . they use two popular datasets to study visual and text reasoning . |
| Outcome: | The proposed model outperforms existing models on two popular datasets . the results raise the question whether visual processing is affected implicitly by linguistic processing . |
Aligning Text/Speech Representations from Multimodal Models with MEG Brain Activity During Listening (2025.emnlp-main)
Copied to clipboard
Padakanti Srijith, Khushbu Pahwa, Radhika Mamidi, Bapi Raju Surampudi, Manish Gupta, Subba Reddy Oota
| Challenge: | Recent studies have found that speech language models fail to capture brain-relevant semantics beyond low-level features. |
| Approach: | They analyze multimodal models to assess their alignment with MEG brain recordings . they find text embeddings from multimodal and unimodal models significantly outperform unilateral models . |
| Outcome: | a new study shows that text-based models outperform unimodal models in alignment with brain recordings during naturalistic story listening. |
Speech language models lack important brain-relevant semantics (2024.acl-long)
Copied to clipboard
| Challenge: | Recent work shows that text-based language models predict both text- and speech-evoked brain activity. |
| Approach: | They remove low-level stimulus features from language models to assess their impact on alignment with fMRI brain recordings during reading and listening. |
| Outcome: | The proposed model removes low-level features from fMRI brain recordings to assess their impact on alignment with fmr recordings. |
Encoding and Decoding Language in the Brain with Language Models (2026.eacl-tutorials)
Copied to clipboard
| Challenge: | This tutorial introduces brain-language model alignment and recent advances in brain-informed fine-tuning and brain-based fine-caching with language models. |
| Approach: | This tutorial introduces brain-language model alignment and recent advances in brain-informed fine-tuning and scaling with language models. |
| Outcome: | This tutorial introduces brain-language model alignment and recent advances in brain-informed fine-tuning and decoding with language models. |
Grounding Visual Illusions in Language: Do Vision-Language Models Perceive Illusions Like Humans? (2023.emnlp-main)
Copied to clipboard
| Challenge: | Visual illusions are a phenomenon that is often seen in human perception but are not always faithful to the physical world. |
| Approach: | They build a dataset containing five types of visual illusions and formulate four tasks to examine visual illusion in state-of-the-art VLMs. |
| Outcome: | The proposed dataset reveals that larger models are closer to human perception and more susceptible to visual illusions. |
Re-Align: Aligning Vision Language Models via Retrieval-Augmented Direct Preference Optimization (2025.emnlp-main)
Copied to clipboard
Shuo Xing, Peiran Li, Yuping Wang, Ruizheng Bai, Yueqi Wang, Chan-Wei Hu, Chengxuan Qian, Huaxiu Yao, Zhengzhong Tu
| Challenge: | emergence of large Vision Language Models (VLMs) has broadened the capabilities of single-modal Large Language Model (LLM) but VLMs are prone to significant hallucinations, especially in the form of cross-modal inconsistencies. |
| Approach: | They propose a new alignment framework that leverages image retrieval to integrate both textual and visual preference signals. |
| Outcome: | The proposed framework mitigates hallucinations more effectively than previous methods . it maintains robustness and scalability across a wide range of VLM sizes and architectures . |
Can VLMs Actually See and Read? A Survey on Modality Collapse in Vision-Language Models (2025.findings-acl)
Copied to clipboard
| Challenge: | Vision-language models integrate textual and visual information, enabling them to process visual inputs and generate predictions. |
| Approach: | They review work on modality collapse analysis to provide insights into the reason for this unintended behavior and review probing studies for fine-grained vision-language understanding. |
| Outcome: | The proposed models can achieve competitive performance in vision-language tasks despite relying heavily on textual information and ignoring visual information. |
Evaluating Model Alignment with Human Perception: A Study on Shitsukan in LLMs and LVLMs (2025.coling-main)
Copied to clipboard
| Challenge: | This work examines the alignment of large language models and large vision-language models with human perception. |
| Approach: | They use a dataset of *shitsukan* terms elicited from individuals in response to object images to evaluate their understanding of the Japanese concept of shitukan. |
| Outcome: | The proposed models demonstrated mixed accuracy across benchmark tasks, with limited overlap between model- and human-generated terms. |
From Language to Cognition: How LLMs Outgrow the Human Language Network (2025.emnlp-main)
Copied to clipboard
Badr AlKhamissi, Greta Tuckute, Yingtian Tang, Taha Osama A Binhuraib, Antoine Bosselut, Martin Schrimpf
| Challenge: | Large language models exhibit remarkable similarity to neural activity in the human language network, but their properties remain unclear. |
| Approach: | They benchmark 34 training checkpoints spanning 300B tokens across 8 different model sizes . they find that brain alignment tracks the development of formal linguistic competence more closely than functional linguistic competency. |
| Outcome: | The results show that large language models exhibit similarity to human language networks . they show that the correlation between next-word prediction and brain alignment fades once models surpass human language proficiency. |