Challenge: TurBLiMP is the first benchmark of linguistic minimal pairs for monolingual and multilingual language models . it covers 16 linguistic phenomena with 1000 minimal pairs each . a foundational insight in linguistics research is that applying minimal changes to a sentence can render it entirely acceptable or unacceptable to native speakers.
Approach: They propose to use morphologically rich agglutinative language with highly flexible word order to evaluate linguistic abilities of monolingual and multilingual language models.
Outcome: The proposed benchmark covers 16 linguistic phenomena with 1000 minimal pairs each.

Similar Papers

BLiMP: The Benchmark of Linguistic Minimal Pairs for English (2020.tacl-1)

Copied to clipboard

Challenge: Recent studies have examined how linguistic knowledge of language models (LMs) varies across English phenomena.
Approach: They propose a benchmark to evaluate linguistic knowledge of language models on major grammatical phenomena in English.
Outcome: The proposed benchmark evaluates the linguistic knowledge of language models on major grammatical phenomena in English.
MultiBLiMP 1.0: A Massively Multilingual Benchmark of Linguistic Minimal Pairs (2026.tacl-1)

Copied to clipboard

Challenge: MultiBLiMP 1.0 is a massively multilingual benchmark of linguistic minimal pairs covering 101 languages and 2 types of subject-verb agreement.
Approach: They propose to use multilingual benchmarks to evaluate linguistic minimal pairs in 101 languages and 2 types of subject-verb agreement to create the minimal pairs.
Outcome: The proposed benchmark covers 101 languages and 2 types of subject-verb agreement, and contains more than 128,000 minimal pairs.
RuBLiMP: Russian Benchmark of Linguistic Minimal Pairs (2024.emnlp-main)

Copied to clipboard

Challenge: Existing resources for minimal pairs address a limited number of languages and lack diversity of language-specific grammatical phenomena.
Approach: They propose to use a Russian benchmark of linguistic minimal pairs to evaluate grammatical knowledge of language models.
Outcome: The proposed benchmark includes 45k pairs of sentences that differ in grammaticality and isolate a morphological, syntactic, or semantic phenomenon.
UrBLiMP: A Benchmark for Evaluating the Linguistic Competence of Large Language Models in Urdu (2026.findings-acl)

Copied to clipboard

Challenge: Evaluating how large language models capture grammatical structure of low-resource languages remains underexplored.
Approach: They evaluate a set of 5,696 minimal pairs that contrast grammatical acceptability across ten core syntactic and morpho-syntactical phenomena in Urdu.
Outcome: The proposed framework compares multilingual models with the proprietary model . the proposed framework achieves the highest average accuracy on regular phenomena .
JBLiMP: Japanese Benchmark of Linguistic Minimal Pairs (2023.findings-eacl)

Copied to clipboard

Challenge: In this paper, we compare syntactic knowledge of language models across different languages.
Approach: They introduce a dataset for targeted syntactic evaluations of language models in Japanese.
Outcome: The proposed dataset compares the syntactic knowledge of language models across languages.
A Systematic Assessment of Language Models with Linguistic Minimal Pairs in Chinese (2026.tacl-1)

Copied to clipboard

Challenge: Using sub-linear length normalized log-probabilities (SLLN-LP), we find unequal lengths of sentences in minimal pairs difficult for LMs even up to 32B parameters.
Approach: They propose to use ZhoBLiMP as a linguistic minimal pair benchmark for Chinese language models to mitigate biases.
Outcome: The proposed metric mitigates biases in Chinese language models with over 100 paradigms . Anaphor, Quantifiers, and Ellipsis are difficult for LMs even up to 32B parameters .
TUMLU: A Unified and Native Language Understanding Benchmark for Turkic Languages (2025.acl-long)

Copied to clipboard

Challenge: preparing native language MMLU benchmarks is costly and limits representativeness of evaluation datasets.
Approach: They propose to use a Turkic language MMLU benchmark to assess massive multitask language understanding capabilities.
Outcome: The proposed benchmarks are based on a Turkic language morphosyntactic and cultural benchmark . the benchmarks evaluate a diverse range of open and proprietary multilingual large language models .
QFrBLiMP: a Quebec-French Benchmark of Linguistic Minimal Pairs (2026.findings-eacl)

Copied to clipboard

Challenge: Specifically, these minimal pairs are created by manually modifying sentences extracted from an official online resource maintained by a Québec government institution.
Approach: They propose to use the Quebec-French Benchmark of Linguistic Minimal Pairs to evaluate LLMs’ linguistic knowledge of prominent grammatical phenomena in Quebec-french.
Outcome: The proposed corpus evaluates LLMs’ linguistic knowledge of prominent grammatical phenomena in Quebec-French.
Mukayese: Turkish NLP Strikes Back (2022.findings-acl)

Copied to clipboard

Challenge: Having sufficient resources for language X lifts it from the under-resourced languages class, but not necessarily from the researched class.
Approach: They propose a set of NLP benchmarks for the Turkish language that contains several NLP tasks.
Outcome: The proposed benchmarks outperform previous work significantly in the Turkish language.
TurkishMMLU: Measuring Massive Multitask Language Understanding in Turkish (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing multiple choice question answering benchmarks employ automatic translation for multilingual evaluation, but this approach is error-prone and potentially introduces culturally biased questions.
Approach: They introduce the first multitask, multiple-choice Turkish QA benchmark, TurkishMMLU . they evaluate over 20 LLMs including open-source, closed-source and Turkish-adapted models .
Outcome: The proposed benchmarks evaluate the reasoning, comprehension, and mathematical abilities of large language models.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations