Papers by Jaap Jumelet
Feature Interactions Reveal Linguistic Structure in Language Models (2023.findings-acl)
Copied to clipboard
| Challenge: | Existing features attribution methods for post-hoc interpretability ignore the existence of interactions between the effects of features on the prediction. |
| Approach: | They propose a grey box method to train models to perfection on a formal language classification task using PCFGs. |
| Outcome: | The proposed methods are able to uncover the grammatical rules acquired by the model under specific configurations and provide novel insights into the linguistic structure of the target models. |
Transparency at the Source: Evaluating and Interpreting Language Models With Access to the True Distribution (2023.findings-emnlp)
Copied to clipboard
| Challenge: | a new approach to train, evaluate and interpret neural language models uses artificial, language-like data. |
| Approach: | They propose a setup for training, evaluating and interpreting neural language models that uses artificial, language-like data. |
| Outcome: | The proposed model is based on a massive probabilistic grammar and a large natural language corpus, and provides complete control over the generative process. |
Assessing the Impact of Typological Features on Multilingual Machine Translation in the Age of Large Language Models (2026.eacl-long)
Copied to clipboard
| Challenge: | Existing evidence on the intrinsic difficulty of multilingual modeling is limited to small monolingual models or bilingual models trained from scratch. |
| Approach: | They propose to use typological properties to determine the difficulty of modeling a language . they analyze two large pre-trained multilingual translation models . |
| Outcome: | The proposed models are based on two large pre-trained models of encoder-decoder and decoder-only machine translation. |
Child-Directed Language Does Not Consistently Boost Syntax Learning in Language Models (2025.emnlp-main)
Copied to clipboard
| Challenge: | prevailing view in language acquisition research has long held that child-directed language is more effective than adultdirected language (ADL) |
| Approach: | They propose a frequency-controlled testing methodology to enable balanced comparisons across training corpora. |
| Outcome: | The proposed method outperforms models trained on English Child-Directed Language (CDL) but it does not yield stronger generalizations for acquiring syntax. |
Vocabulary Shapes Cross-Lingual Variation of Word-Order Learnability in Language Models (2026.acl-long)
Copied to clipboard
| Challenge: | a coarse distinction between free- and fixed-word-order languages does not explain cross-lingual variation . aaron e. smith: human languages have emerged over millennia through communicative and cognitive constraints . but, within those universally shared bounds, languages exhibit striking typological diversity . |
| Approach: | They propose to train transformer language models on a spectrum of synthetic word-order variants of natural languages. |
| Outcome: | a new study shows that word order irregularities raise model surprisal, but only weakly affects learnability . the study also shows that vocabulary structure emerges as a key driver of word-order learnability across languages. |
Transformer-specific Interpretability (2024.eacl-tutorials)
Copied to clipboard
| Challenge: | Transformers are dominant play-ers in various scientific fields, but their inner workings remain opaque. |
| Approach: | This tutorial presents a trending approach to interpreting Transformers . it uses specific features of the Transformer architecture to quantify context- mixing interactions . |
| Outcome: | This tutorial aims to show how a new trending approach can be applied to Transformer-based models. |
TurBLiMP: A Turkish Benchmark of Linguistic Minimal Pairs (2025.emnlp-main)
Copied to clipboard
| Challenge: | TurBLiMP is the first benchmark of linguistic minimal pairs for monolingual and multilingual language models . it covers 16 linguistic phenomena with 1000 minimal pairs each . a foundational insight in linguistics research is that applying minimal changes to a sentence can render it entirely acceptable or unacceptable to native speakers. |
| Approach: | They propose to use morphologically rich agglutinative language with highly flexible word order to evaluate linguistic abilities of monolingual and multilingual language models. |
| Outcome: | The proposed benchmark covers 16 linguistic phenomena with 1000 minimal pairs each. |
MultiBLiMP 1.0: A Massively Multilingual Benchmark of Linguistic Minimal Pairs (2026.tacl-1)
Copied to clipboard
| Challenge: | MultiBLiMP 1.0 is a massively multilingual benchmark of linguistic minimal pairs covering 101 languages and 2 types of subject-verb agreement. |
| Approach: | They propose to use multilingual benchmarks to evaluate linguistic minimal pairs in 101 languages and 2 types of subject-verb agreement to create the minimal pairs. |
| Outcome: | The proposed benchmark covers 101 languages and 2 types of subject-verb agreement, and contains more than 128,000 minimal pairs. |
Do Language Models Exhibit Human-like Structural Priming Effects? (2024.findings-acl)
Copied to clipboard
| Challenge: | a recent exposure to a structure facilitates processing of the same structure, a study finds . structural priming is well attested in humans, for both language production and comprehension . |
| Approach: | They use the structural priming paradigm to investigate where priming effects manifest . they find that rarer elements within a prime increase priming effect . |
| Outcome: | The findings provide an important piece in the puzzle of understanding how properties within their context affect structural prediction in language models. |
Language Models Use Monotonicity to Assess NPI Licensing (2021.findings-acl)
Copied to clipboard
| Challenge: | Neural language models (LMs) have become powerful approximators of human language . fewer studies have been done on what kind of formal semantic features are encoded by LMs . |
| Approach: | They propose a series of experiments that investigate the semantic knowledge of language models . they use diagnostic classifiers, linguistic acceptability tasks and a ranking method to investigate the models' inner workings. |
| Outcome: | The proposed method can be applied to LMs trained on filtered corpora and gain stronger insights into their generalizations. |
Language Modelling as a Multi-Task Problem (2021.eacl-main)
Copied to clipboard
| Challenge: | Using multitask learning, humans are optimising their behaviour towards a multitude of objectives to reach their goals in dayto-day life. |
| Approach: | They propose to study language modelling as a multi-task problem by examining the generalisation behaviour of language models as they learn the linguistic concept of Negative Polarity Items. |
| Outcome: | The proposed model is able to learn the linguistic concept of Negative Polarity Items (NPIs) and is a multi-task learning model. |
Structural Persistence in Language Models: Priming as a Window into Abstract Language Representations (2022.tacl-1)
Copied to clipboard
| Challenge: | a rich literature has emerged in the last few years addressing these questions, including whether specific LMs have acquired specific linguistic constructions. |
| Approach: | They introduce a novel metric and release Prime-LM, a large corpus where they control for various linguistic factors that interact with priming strength. |
| Outcome: | The proposed model can learn abstract structural information independent of the structure of a sentence and is able to perform tasks that require natural language understanding skills. |
Interpretability of Language Models via Task Spaces (2024.acl-long)
Copied to clipboard
| Challenge: | linguistic interpretability is a method used to assess language models' ability to interpret outputs. |
| Approach: | They propose a method to assess LMs' language conceptualisations by 'similarity probing' and a technique to fine tune them via gradient differentials to disentangle the learning signals of linguistic phenomena. |
| Outcome: | The proposed method generalises larger models to overarching general concepts for linguistic tasks, and the generalisation patterns are stable throughout training and not marked by incisive stages. |
BabyBabelLM: A Multilingual Benchmark of Developmentally Plausible Training Data (2026.eacl-long)
Copied to clipboard
Jaap Jumelet, Abdellah Fourtassi, Akari Haga, Bastian Bunzeck, Bhargav Shandilya, Diana Galvan-Sosa, Faiz Ghifari Haznitrama, Francesca Padovani, Francois Meyer, Hai Hu, Julen Etxaniz, Laurent Prevot, Linyang He, María Grandury, Mila Marcheva, Negar Foroutan, Nikitas Theodoropoulos, Pouya Sadeghi, Siyuan Song, Suchir Salhan, Susana Zhou, Yurii Paniv, Ziyin Zhang, Arianna Bisazza, Alex Warstadt, Leshem Choshen
| Challenge: | prevailing trend in language modeling research is to prioritize scaling, authors say . from infancy to maturity, English learners acquire language through exposure to less than 100M words . |
| Approach: | They propose a multilingual collection of datasets modeling the language a person observes from birth until they acquire a native language. |
| Outcome: | The proposed models outperform models trained on a fixed, developmentally plausible English corpus on various benchmarks. |
DecoderLens: Layerwise Interpretation of Encoder-Decoder Transformers (2024.findings-naacl)
Copied to clipboard
| Challenge: | Existing interpretability methods have been proposed to interpret the inner workings of Transformer models at different levels of precision and complexity. |
| Approach: | They propose a method to analyze encoder-decoder Transformers by using the decoder module Model Output encoder to cross-attend representations of intermediate encoder activations instead of using the default output. |
| Outcome: | The proposed method maps uninterpretable representations to human-interpreted sequences of words or symbols, shedding new light on the information flow in this popular but understudied class of models. |