Aligning Text/Speech Representations from Multimodal Models with MEG Brain Activity During Listening (2025.emnlp-main)
Copied to clipboard
Padakanti Srijith, Khushbu Pahwa, Radhika Mamidi, Bapi Raju Surampudi, Manish Gupta, Subba Reddy Oota
| Challenge: | Recent studies have found that speech language models fail to capture brain-relevant semantics beyond low-level features. |
| Approach: | They analyze multimodal models to assess their alignment with MEG brain recordings . they find text embeddings from multimodal and unimodal models significantly outperform unilateral models . |
| Outcome: | a new study shows that text-based models outperform unimodal models in alignment with brain recordings during naturalistic story listening. |
Similar Papers
Speech language models lack important brain-relevant semantics (2024.acl-long)
Copied to clipboard
| Challenge: | Recent work shows that text-based language models predict both text- and speech-evoked brain activity. |
| Approach: | They remove low-level stimulus features from language models to assess their impact on alignment with fMRI brain recordings during reading and listening. |
| Outcome: | The proposed model removes low-level features from fMRI brain recordings to assess their impact on alignment with fmr recordings. |
Decoding the Multimodal Mind: Generalizable Brain-to-Text Translation via Multimodal Alignment and Adaptive Routing (2026.findings-acl)
Copied to clipboard
| Challenge: | Current approaches to decoding language from the human brain rely on unimodal representations, neglecting the brain’s inherently multimodal processing. |
| Approach: | They propose a framework that leverages Multimodal Large Language Models to align brain signals with a shared semantic space encompassing text, images, and audio. |
| Outcome: | The proposed framework achieves an 8.48% improvement on the most commonly used benchmark on fMRI datasets with textual, visual, and auditory stimuli. |
Unveiling Multi-level and Multi-modal Semantic Representations in the Human Brain using Large Language Models (2024.emnlp-main)
Copied to clipboard
Yuko Nakagi, Takuya Matsuyama, Naoko Koide-Majima, Hiroto Yamaguchi, Rieko Kubo, Shinji Nishimoto, Yu Takagi
| Challenge: | Recent studies have assessed different levels of semantic content, such as speech, objects, and stories, separately. |
| Approach: | They used functional magnetic resonance imaging to record brain activity while watching 8.3 hours of dramas and movies. |
| Outcome: | The findings show that LLMs predict human brain activity more accurately than traditional language models, particularly for complex background stories. |
How do Multimodal Foundation Models Encode Text and Speech? An Analysis of Cross-Lingual and Cross-Modal Representations (2025.naacl-short)
Copied to clipboard
| Challenge: | Recent advances in foundation models have sparked growing interest in expanding their text processing capabilities to speech. |
| Approach: | They analyze the model activations from semantically equivalent sentences across languages in the text and speech modalities and examine how text and spoken are represented in recent multimodal foundation models. |
| Outcome: | The proposed models exhibit cross-lingual differences, but are not explicitly trained for modality-agnostic representations. |
Vision-Language Models Align with Human Neural Representations in Concept Processing (2026.eacl-long)
Copied to clipboard
| Challenge: | Recent studies suggest that transformer-based vision-language models capture the multimodality of concept processing in the human brain. |
| Approach: | They analysed multiple VLMs employing different strategies to integrate visual and textual modalities, along with language-only counterparts. |
| Outcome: | The transformer-based vision-language models outperform language-only models in two experimental conditions, while only some outperformed the language-based models. |
Seeing Through Words, Speaking Through Pixels: Deep Representational Alignment Between Vision and Language Models (2025.emnlp-main)
Copied to clipboard
| Challenge: | Recent studies show that deep vision-only and language-only models project inputs into a partially aligned representational space. |
| Approach: | They investigate whether a model's representational code is semantically shared . they find that alignment peaks in mid-to-late layers of both model types . |
| Outcome: | a forced-choice "Pick-a-Pic" task shows human preferences for image-caption matches are mirrored in embedding spaces across vision-language model pairs. |
From Neurons to Semantics: Evaluating Cross-Linguistic Alignment Capabilities of Large Language Models via Neurons Alignment (2025.acl-long)
Copied to clipboard
| Challenge: | Existing alignment benchmarks focus on sentence embeddings, but prior research has shown that neural models tend to induce a non-smooth representation space, which impact of semantic alignment evaluation on low-resource languages. |
| Approach: | They propose a novel cross-lingual alignment evaluation method based on the consistency of parallel sentences to assess model alignment. |
| Outcome: | The proposed method achieves a correlation of 0.9556 with downstream tasks performance and 0.8524 with transferability even with a small dataset. |
Improve Language Model and Brain Alignment via Associative Memory (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing studies have shown that associative memory is essential for language comprehension and comprehension. |
| Approach: | They propose to integrate associative memory into language models to improve alignment . they find alignment is improved in brain regions closely related to associativ memory processing . |
| Outcome: | The proposed model improves in brain regions closely related to associative memory processing. |
Multimodal Transformer for Unaligned Multimodal Language Sequences (P19-1)
Copied to clipboard
Yao-Hung Hubert Tsai, Shaojie Bai, Paul Pu Liang, J. Zico Kolter, Louis-Philippe Morency, Ruslan Salakhutdinov
| Challenge: | Human language is often multimodal, which comprehends a mixture of natural language, facial gestures, and acoustic behaviors. |
| Approach: | They propose a multimodal model that extends the standard Transformer network to learn representations directly from unaligned multimodal streams. |
| Outcome: | The proposed model outperforms state-of-the-art methods on aligned and non-aligned data. |
CLASP: Cross-modal Alignment Using Pre-trained Unimodal Models (2024.findings-acl)
Copied to clipboard
| Challenge: | Recent advances in speech-text pretraining rely on parallel speech- text data . however, data accessibility is a challenge due to the limited data available. |
| Approach: | They propose a framework for jointly performing speech and text processing without parallel corpora during pre-training but only downstream. |
| Outcome: | The proposed framework extracts distinct representations for speech and text, aligning them effectively in a newly defined space using a multi-level contrastive learning mechanism. |