Challenge: MultiBLiMP 1.0 is a massively multilingual benchmark of linguistic minimal pairs covering 101 languages and 2 types of subject-verb agreement.
Approach: They propose to use multilingual benchmarks to evaluate linguistic minimal pairs in 101 languages and 2 types of subject-verb agreement to create the minimal pairs.
Outcome: The proposed benchmark covers 101 languages and 2 types of subject-verb agreement, and contains more than 128,000 minimal pairs.

Similar Papers

UrBLiMP: A Benchmark for Evaluating the Linguistic Competence of Large Language Models in Urdu (2026.findings-acl)

Copied to clipboard

Challenge: Evaluating how large language models capture grammatical structure of low-resource languages remains underexplored.
Approach: They evaluate a set of 5,696 minimal pairs that contrast grammatical acceptability across ten core syntactic and morpho-syntactical phenomena in Urdu.
Outcome: The proposed framework compares multilingual models with the proprietary model . the proposed framework achieves the highest average accuracy on regular phenomena .
JBLiMP: Japanese Benchmark of Linguistic Minimal Pairs (2023.findings-eacl)

Copied to clipboard

Challenge: In this paper, we compare syntactic knowledge of language models across different languages.
Approach: They introduce a dataset for targeted syntactic evaluations of language models in Japanese.
Outcome: The proposed dataset compares the syntactic knowledge of language models across languages.
BLiMP: The Benchmark of Linguistic Minimal Pairs for English (2020.tacl-1)

Copied to clipboard

Challenge: Recent studies have examined how linguistic knowledge of language models (LMs) varies across English phenomena.
Approach: They propose a benchmark to evaluate linguistic knowledge of language models on major grammatical phenomena in English.
Outcome: The proposed benchmark evaluates the linguistic knowledge of language models on major grammatical phenomena in English.
QFrBLiMP: a Quebec-French Benchmark of Linguistic Minimal Pairs (2026.findings-eacl)

Copied to clipboard

Challenge: Specifically, these minimal pairs are created by manually modifying sentences extracted from an official online resource maintained by a Québec government institution.
Approach: They propose to use the Quebec-French Benchmark of Linguistic Minimal Pairs to evaluate LLMs’ linguistic knowledge of prominent grammatical phenomena in Quebec-french.
Outcome: The proposed corpus evaluates LLMs’ linguistic knowledge of prominent grammatical phenomena in Quebec-French.
TurBLiMP: A Turkish Benchmark of Linguistic Minimal Pairs (2025.emnlp-main)

Copied to clipboard

Challenge: TurBLiMP is the first benchmark of linguistic minimal pairs for monolingual and multilingual language models . it covers 16 linguistic phenomena with 1000 minimal pairs each . a foundational insight in linguistics research is that applying minimal changes to a sentence can render it entirely acceptable or unacceptable to native speakers.
Approach: They propose to use morphologically rich agglutinative language with highly flexible word order to evaluate linguistic abilities of monolingual and multilingual language models.
Outcome: The proposed benchmark covers 16 linguistic phenomena with 1000 minimal pairs each.
BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing multilingual benchmarks focus primarily on language understanding tasks.
Approach: They develop a multi-way multilingual benchmark that measures critical capabilities of large language models across languages.
Outcome: Extensive experiments on BenchMAX reveal uneven utilization of core capabilities across languages, emphasizing the performance gaps that scaling model size alone does not resolve.
A Systematic Assessment of Language Models with Linguistic Minimal Pairs in Chinese (2026.tacl-1)

Copied to clipboard

Challenge: Using sub-linear length normalized log-probabilities (SLLN-LP), we find unequal lengths of sentences in minimal pairs difficult for LMs even up to 32B parameters.
Approach: They propose to use ZhoBLiMP as a linguistic minimal pair benchmark for Chinese language models to mitigate biases.
Outcome: The proposed metric mitigates biases in Chinese language models with over 100 paradigms . Anaphor, Quantifiers, and Ellipsis are difficult for LMs even up to 32B parameters .
Linguistic Minimal Pairs Elicit Linguistic Similarity in Large Language Models (2025.coling-main)

Copied to clipboard

Challenge: a new analysis leverages linguistic minimal pairs to probe the internal linguistic representations of Large Language Models (LLMs).
Approach: They propose to use linguistic minimal pairs to probe the internal linguistic representations of Large Language Models (LLMs).
Outcome: The proposed analysis reveals that linguistic similarity is significantly influenced by training data exposure, leading to higher cross-LLM agreement in higher-resource languages.
Wiki-40B: Multilingual Language Model Dataset (2020.lrec-1)

Copied to clipboard

Challenge: We propose a new multilingual language model benchmark that is composed of 40+ languages spanning several scripts and linguistic families.
Approach: They propose a multilingual language model benchmark composed of 40+ languages . they train monolingual causal language models using a state-of-the-art model .
Outcome: The proposed model is composed of 40+ languages spanning several scripts and linguistic families.
Are Multilingual Models the Best Choice for Moderately Under-resourced Languages? A Comprehensive Assessment for Catalan (2021.findings-acl)

Copied to clipboard

Challenge: Multilingual language models have been a crucial breakthrough for under-resourced languages . however, the superiority of language-specific models has already been proven for underresourced ones .
Approach: They propose to build a monolingual monolingual model that is comparable to state-of-the-art large multilingual models.
Outcome: The proposed model consistently outperforms state-of-the-art models across tasks and settings.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations