Papers by Francesca Padovani
Child-Directed Language Does Not Consistently Boost Syntax Learning in Language Models (2025.emnlp-main)
Copied to clipboard
| Challenge: | prevailing view in language acquisition research has long held that child-directed language is more effective than adultdirected language (ADL) |
| Approach: | They propose a frequency-controlled testing methodology to enable balanced comparisons across training corpora. |
| Outcome: | The proposed method outperforms models trained on English Child-Directed Language (CDL) but it does not yield stronger generalizations for acquiring syntax. |
TurBLiMP: A Turkish Benchmark of Linguistic Minimal Pairs (2025.emnlp-main)
Copied to clipboard
| Challenge: | TurBLiMP is the first benchmark of linguistic minimal pairs for monolingual and multilingual language models . it covers 16 linguistic phenomena with 1000 minimal pairs each . a foundational insight in linguistics research is that applying minimal changes to a sentence can render it entirely acceptable or unacceptable to native speakers. |
| Approach: | They propose to use morphologically rich agglutinative language with highly flexible word order to evaluate linguistic abilities of monolingual and multilingual language models. |
| Outcome: | The proposed benchmark covers 16 linguistic phenomena with 1000 minimal pairs each. |
BabyBabelLM: A Multilingual Benchmark of Developmentally Plausible Training Data (2026.eacl-long)
Copied to clipboard
Jaap Jumelet, Abdellah Fourtassi, Akari Haga, Bastian Bunzeck, Bhargav Shandilya, Diana Galvan-Sosa, Faiz Ghifari Haznitrama, Francesca Padovani, Francois Meyer, Hai Hu, Julen Etxaniz, Laurent Prevot, Linyang He, María Grandury, Mila Marcheva, Negar Foroutan, Nikitas Theodoropoulos, Pouya Sadeghi, Siyuan Song, Suchir Salhan, Susana Zhou, Yurii Paniv, Ziyin Zhang, Arianna Bisazza, Alex Warstadt, Leshem Choshen
| Challenge: | prevailing trend in language modeling research is to prioritize scaling, authors say . from infancy to maturity, English learners acquire language through exposure to less than 100M words . |
| Approach: | They propose a multilingual collection of datasets modeling the language a person observes from birth until they acquire a native language. |
| Outcome: | The proposed models outperform models trained on a fixed, developmentally plausible English corpus on various benchmarks. |