Papers by Mariya Toneva

3 papers
Speech language models lack important brain-relevant semantics (2024.acl-long)

Copied to clipboard

Challenge: Recent work shows that text-based language models predict both text- and speech-evoked brain activity.
Approach: They remove low-level stimulus features from language models to assess their impact on alignment with fMRI brain recordings during reading and listening.
Outcome: The proposed model removes low-level features from fMRI brain recordings to assess their impact on alignment with fmr recordings.
Language models and brains align due to more than next-word prediction and word-level information (2024.emnlp-main)

Copied to clipboard

Challenge: Pretrained language models have been shown to significantly predict brain recordings of people comprehending language.
Approach: They propose to use two perturbations to design contrasts that control for different types of information.
Outcome: The proposed model is largely agnostic about the exact linguistic information contained in the conceptual quantities "word-level information" and "multi-word information".
Perturbed examples reveal invariances shared by language models (2024.findings-acl)

Copied to clipboard

Challenge: Rapid growth in natural language processing (NLP) research has led to numerous new models outpacing our understanding of how they compare to established ones.
Approach: They propose a framework to compare two NLP models by revealing their shared invariance to interpretable input perturbations targeting a specific linguistic capability.
Outcome: The proposed framework can shed light on the types of invariances retained or emerging in new models.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations