Papers by Marco Marelli
Scaling in Cognitive Modelling: a Multilingual Approach to Human Reading Times (2023.acl-short)
Copied to clipboard
| Challenge: | Neural language models provide conditional probability distributions over the lexicon that are predictive of human processing times. |
| Approach: | They propose to use a transformer-based model to generate probabilistic estimates that are less predictive of early eye-tracking measurements reflecting lexical access and early semantic integration. |
| Outcome: | The proposed models show that larger models capture late eye-tracking measurements that reflect the full integration of a word into the current language context. |
SWEAT: Scoring Polarization of Topics across Different Corpora (2021.emnlp-main)
Copied to clipboard
| Challenge: | Using two additional wordsets, we compute the relative polarization of a topical wordsetting across two distributional representations. |
| Approach: | They propose a new measure to compute the relative polarization of a topical wordset across two distributional representations using two additional wordsetes deemed to have opposite valence to represent two different poles. |
| Outcome: | The proposed measure is validated by a case study and validated in a randomized controlled trial. |
The Effects of Surprisal across Languages: Results from Native and Non-native Reading (2022.findings-aacl)
Copied to clipboard
| Challenge: | Context-dependent predictive processes have been proposed as a core component of the human cognitive system. |
| Approach: | They extract surprisal estimates from mBERT and assess their predictive power on the MECO corpus, a cross-linguistic dataset of eye movement behavior in reading. |
| Outcome: | The proposed model is based on a cross-linguistic dataset of eye movement behavior in reading. |
The Emergence of Semantic Units in Massively Multilingual Models (2024.lrec-main)
Copied to clipboard
| Challenge: | Massively multilingual models can process text in several languages relying on a shared set of parameters, but little is known about the encoding of multilingual information in single network units. |
| Approach: | They propose to use a shared set of parameters to encode multilingual information in single network units. |
| Outcome: | The proposed model achieves higher scores in semantic encoding in languages with more cross-lingual alignment than those with more shared cross-linguistic substrate. |