Challenge: Recent studies have found that speech language models fail to capture brain-relevant semantics beyond low-level features.
Approach: They analyze multimodal models to assess their alignment with MEG brain recordings . they find text embeddings from multimodal and unimodal models significantly outperform unilateral models .
Outcome: a new study shows that text-based models outperform unimodal models in alignment with brain recordings during naturalistic story listening.

Similar Papers

Speech language models lack important brain-relevant semantics (2024.acl-long)

Copied to clipboard

Challenge: Recent work shows that text-based language models predict both text- and speech-evoked brain activity.
Approach: They remove low-level stimulus features from language models to assess their impact on alignment with fMRI brain recordings during reading and listening.
Outcome: The proposed model removes low-level features from fMRI brain recordings to assess their impact on alignment with fmr recordings.
Decoding the Multimodal Mind: Generalizable Brain-to-Text Translation via Multimodal Alignment and Adaptive Routing (2026.findings-acl)

Copied to clipboard

Challenge: Current approaches to decoding language from the human brain rely on unimodal representations, neglecting the brain’s inherently multimodal processing.
Approach: They propose a framework that leverages Multimodal Large Language Models to align brain signals with a shared semantic space encompassing text, images, and audio.
Outcome: The proposed framework achieves an 8.48% improvement on the most commonly used benchmark on fMRI datasets with textual, visual, and auditory stimuli.
Unveiling Multi-level and Multi-modal Semantic Representations in the Human Brain using Large Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: Recent studies have assessed different levels of semantic content, such as speech, objects, and stories, separately.
Approach: They used functional magnetic resonance imaging to record brain activity while watching 8.3 hours of dramas and movies.
Outcome: The findings show that LLMs predict human brain activity more accurately than traditional language models, particularly for complex background stories.
How do Multimodal Foundation Models Encode Text and Speech? An Analysis of Cross-Lingual and Cross-Modal Representations (2025.naacl-short)

Copied to clipboard

Challenge: Recent advances in foundation models have sparked growing interest in expanding their text processing capabilities to speech.
Approach: They analyze the model activations from semantically equivalent sentences across languages in the text and speech modalities and examine how text and spoken are represented in recent multimodal foundation models.
Outcome: The proposed models exhibit cross-lingual differences, but are not explicitly trained for modality-agnostic representations.
Vision-Language Models Align with Human Neural Representations in Concept Processing (2026.eacl-long)

Copied to clipboard

Challenge: Recent studies suggest that transformer-based vision-language models capture the multimodality of concept processing in the human brain.
Approach: They analysed multiple VLMs employing different strategies to integrate visual and textual modalities, along with language-only counterparts.
Outcome: The transformer-based vision-language models outperform language-only models in two experimental conditions, while only some outperformed the language-based models.
Seeing Through Words, Speaking Through Pixels: Deep Representational Alignment Between Vision and Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: Recent studies show that deep vision-only and language-only models project inputs into a partially aligned representational space.
Approach: They investigate whether a model's representational code is semantically shared . they find that alignment peaks in mid-to-late layers of both model types .
Outcome: a forced-choice "Pick-a-Pic" task shows human preferences for image-caption matches are mirrored in embedding spaces across vision-language model pairs.
From Neurons to Semantics: Evaluating Cross-Linguistic Alignment Capabilities of Large Language Models via Neurons Alignment (2025.acl-long)

Copied to clipboard

Challenge: Existing alignment benchmarks focus on sentence embeddings, but prior research has shown that neural models tend to induce a non-smooth representation space, which impact of semantic alignment evaluation on low-resource languages.
Approach: They propose a novel cross-lingual alignment evaluation method based on the consistency of parallel sentences to assess model alignment.
Outcome: The proposed method achieves a correlation of 0.9556 with downstream tasks performance and 0.8524 with transferability even with a small dataset.
Improve Language Model and Brain Alignment via Associative Memory (2025.findings-acl)

Copied to clipboard

Challenge: Existing studies have shown that associative memory is essential for language comprehension and comprehension.
Approach: They propose to integrate associative memory into language models to improve alignment . they find alignment is improved in brain regions closely related to associativ memory processing .
Outcome: The proposed model improves in brain regions closely related to associative memory processing.
Multimodal Transformer for Unaligned Multimodal Language Sequences (P19-1)

Copied to clipboard

Challenge: Human language is often multimodal, which comprehends a mixture of natural language, facial gestures, and acoustic behaviors.
Approach: They propose a multimodal model that extends the standard Transformer network to learn representations directly from unaligned multimodal streams.
Outcome: The proposed model outperforms state-of-the-art methods on aligned and non-aligned data.
CLASP: Cross-modal Alignment Using Pre-trained Unimodal Models (2024.findings-acl)

Copied to clipboard

Challenge: Recent advances in speech-text pretraining rely on parallel speech- text data . however, data accessibility is a challenge due to the limited data available.
Approach: They propose a framework for jointly performing speech and text processing without parallel corpora during pre-training but only downstream.
Outcome: The proposed framework extracts distinct representations for speech and text, aligning them effectively in a newly defined space using a multi-level contrastive learning mechanism.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations