Papers by Mariya Toneva
Speech language models lack important brain-relevant semantics (2024.acl-long)
Copied to clipboard
| Challenge: | Recent work shows that text-based language models predict both text- and speech-evoked brain activity. |
| Approach: | They remove low-level stimulus features from language models to assess their impact on alignment with fMRI brain recordings during reading and listening. |
| Outcome: | The proposed model removes low-level features from fMRI brain recordings to assess their impact on alignment with fmr recordings. |
Language models and brains align due to more than next-word prediction and word-level information (2024.emnlp-main)
Copied to clipboard
| Challenge: | Pretrained language models have been shown to significantly predict brain recordings of people comprehending language. |
| Approach: | They propose to use two perturbations to design contrasts that control for different types of information. |
| Outcome: | The proposed model is largely agnostic about the exact linguistic information contained in the conceptual quantities "word-level information" and "multi-word information". |
Perturbed examples reveal invariances shared by language models (2024.findings-acl)
Copied to clipboard
| Challenge: | Rapid growth in natural language processing (NLP) research has led to numerous new models outpacing our understanding of how they compare to established ones. |
| Approach: | They propose a framework to compare two NLP models by revealing their shared invariance to interpretable input perturbations targeting a specific linguistic capability. |
| Outcome: | The proposed framework can shed light on the types of invariances retained or emerging in new models. |