Papers by Marco Marelli

4 papers
Scaling in Cognitive Modelling: a Multilingual Approach to Human Reading Times (2023.acl-short)

Copied to clipboard

Challenge: Neural language models provide conditional probability distributions over the lexicon that are predictive of human processing times.
Approach: They propose to use a transformer-based model to generate probabilistic estimates that are less predictive of early eye-tracking measurements reflecting lexical access and early semantic integration.
Outcome: The proposed models show that larger models capture late eye-tracking measurements that reflect the full integration of a word into the current language context.
SWEAT: Scoring Polarization of Topics across Different Corpora (2021.emnlp-main)

Copied to clipboard

Challenge: Using two additional wordsets, we compute the relative polarization of a topical wordsetting across two distributional representations.
Approach: They propose a new measure to compute the relative polarization of a topical wordset across two distributional representations using two additional wordsetes deemed to have opposite valence to represent two different poles.
Outcome: The proposed measure is validated by a case study and validated in a randomized controlled trial.
The Effects of Surprisal across Languages: Results from Native and Non-native Reading (2022.findings-aacl)

Copied to clipboard

Challenge: Context-dependent predictive processes have been proposed as a core component of the human cognitive system.
Approach: They extract surprisal estimates from mBERT and assess their predictive power on the MECO corpus, a cross-linguistic dataset of eye movement behavior in reading.
Outcome: The proposed model is based on a cross-linguistic dataset of eye movement behavior in reading.
The Emergence of Semantic Units in Massively Multilingual Models (2024.lrec-main)

Copied to clipboard

Challenge: Massively multilingual models can process text in several languages relying on a shared set of parameters, but little is known about the encoding of multilingual information in single network units.
Approach: They propose to use a shared set of parameters to encode multilingual information in single network units.
Outcome: The proposed model achieves higher scores in semantic encoding in languages with more cross-lingual alignment than those with more shared cross-linguistic substrate.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations