| Challenge: | TurBLiMP is the first benchmark of linguistic minimal pairs for monolingual and multilingual language models . it covers 16 linguistic phenomena with 1000 minimal pairs each . a foundational insight in linguistics research is that applying minimal changes to a sentence can render it entirely acceptable or unacceptable to native speakers. |
| Approach: | They propose to use morphologically rich agglutinative language with highly flexible word order to evaluate linguistic abilities of monolingual and multilingual language models. |
| Outcome: | The proposed benchmark covers 16 linguistic phenomena with 1000 minimal pairs each. |
Similar Papers
BLiMP: The Benchmark of Linguistic Minimal Pairs for English (2020.tacl-1)
Copied to clipboard
Alex Warstadt, Alicia Parrish, Haokun Liu, Anhad Mohananey, Wei Peng, Sheng-Fu Wang, Samuel R. Bowman
| Challenge: | Recent studies have examined how linguistic knowledge of language models (LMs) varies across English phenomena. |
| Approach: | They propose a benchmark to evaluate linguistic knowledge of language models on major grammatical phenomena in English. |
| Outcome: | The proposed benchmark evaluates the linguistic knowledge of language models on major grammatical phenomena in English. |
MultiBLiMP 1.0: A Massively Multilingual Benchmark of Linguistic Minimal Pairs (2026.tacl-1)
Copied to clipboard
| Challenge: | MultiBLiMP 1.0 is a massively multilingual benchmark of linguistic minimal pairs covering 101 languages and 2 types of subject-verb agreement. |
| Approach: | They propose to use multilingual benchmarks to evaluate linguistic minimal pairs in 101 languages and 2 types of subject-verb agreement to create the minimal pairs. |
| Outcome: | The proposed benchmark covers 101 languages and 2 types of subject-verb agreement, and contains more than 128,000 minimal pairs. |
RuBLiMP: Russian Benchmark of Linguistic Minimal Pairs (2024.emnlp-main)
Copied to clipboard
Ekaterina Taktasheva, Maxim Bazhukov, Kirill Koncha, Alena Fenogenova, Ekaterina Artemova, Vladislav Mikhailov
| Challenge: | Existing resources for minimal pairs address a limited number of languages and lack diversity of language-specific grammatical phenomena. |
| Approach: | They propose to use a Russian benchmark of linguistic minimal pairs to evaluate grammatical knowledge of language models. |
| Outcome: | The proposed benchmark includes 45k pairs of sentences that differ in grammaticality and isolate a morphological, syntactic, or semantic phenomenon. |
UrBLiMP: A Benchmark for Evaluating the Linguistic Competence of Large Language Models in Urdu (2026.findings-acl)
Copied to clipboard
| Challenge: | Evaluating how large language models capture grammatical structure of low-resource languages remains underexplored. |
| Approach: | They evaluate a set of 5,696 minimal pairs that contrast grammatical acceptability across ten core syntactic and morpho-syntactical phenomena in Urdu. |
| Outcome: | The proposed framework compares multilingual models with the proprietary model . the proposed framework achieves the highest average accuracy on regular phenomena . |
JBLiMP: Japanese Benchmark of Linguistic Minimal Pairs (2023.findings-eacl)
Copied to clipboard
| Challenge: | In this paper, we compare syntactic knowledge of language models across different languages. |
| Approach: | They introduce a dataset for targeted syntactic evaluations of language models in Japanese. |
| Outcome: | The proposed dataset compares the syntactic knowledge of language models across languages. |
A Systematic Assessment of Language Models with Linguistic Minimal Pairs in Chinese (2026.tacl-1)
Copied to clipboard
Yikang Liu, Yeting Shen, Hongao Zhu, Lilong Xu, Zhiheng Qian, Siyuan Song, Kejia Zhang, Jialong Tang, Pei Zhang, Baosong Yang, Rui Wang, Hai Hu
| Challenge: | Using sub-linear length normalized log-probabilities (SLLN-LP), we find unequal lengths of sentences in minimal pairs difficult for LMs even up to 32B parameters. |
| Approach: | They propose to use ZhoBLiMP as a linguistic minimal pair benchmark for Chinese language models to mitigate biases. |
| Outcome: | The proposed metric mitigates biases in Chinese language models with over 100 paradigms . Anaphor, Quantifiers, and Ellipsis are difficult for LMs even up to 32B parameters . |
TUMLU: A Unified and Native Language Understanding Benchmark for Turkic Languages (2025.acl-long)
Copied to clipboard
Jafar Isbarov, Arofat Akhundjanova, Mammad Hajili, Kavsar Huseynova, Dmitry Gaynullin, Anar Rzayev, Osman Tursun, Aizirek Turdubaeva, Ilshat Saetov, Rinat Kharisov, Saule Belginova, Ariana Kenbayeva, Amina Alisheva, Abdullatif Köksal, Samir Rustamov, Duygu Ataman
| Challenge: | preparing native language MMLU benchmarks is costly and limits representativeness of evaluation datasets. |
| Approach: | They propose to use a Turkic language MMLU benchmark to assess massive multitask language understanding capabilities. |
| Outcome: | The proposed benchmarks are based on a Turkic language morphosyntactic and cultural benchmark . the benchmarks evaluate a diverse range of open and proprietary multilingual large language models . |
QFrBLiMP: a Quebec-French Benchmark of Linguistic Minimal Pairs (2026.findings-eacl)
Copied to clipboard
| Challenge: | Specifically, these minimal pairs are created by manually modifying sentences extracted from an official online resource maintained by a Québec government institution. |
| Approach: | They propose to use the Quebec-French Benchmark of Linguistic Minimal Pairs to evaluate LLMs’ linguistic knowledge of prominent grammatical phenomena in Quebec-french. |
| Outcome: | The proposed corpus evaluates LLMs’ linguistic knowledge of prominent grammatical phenomena in Quebec-French. |
Mukayese: Turkish NLP Strikes Back (2022.findings-acl)
Copied to clipboard
| Challenge: | Having sufficient resources for language X lifts it from the under-resourced languages class, but not necessarily from the researched class. |
| Approach: | They propose a set of NLP benchmarks for the Turkish language that contains several NLP tasks. |
| Outcome: | The proposed benchmarks outperform previous work significantly in the Turkish language. |
TurkishMMLU: Measuring Massive Multitask Language Understanding in Turkish (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Existing multiple choice question answering benchmarks employ automatic translation for multilingual evaluation, but this approach is error-prone and potentially introduces culturally biased questions. |
| Approach: | They introduce the first multitask, multiple-choice Turkish QA benchmark, TurkishMMLU . they evaluate over 20 LLMs including open-source, closed-source and Turkish-adapted models . |
| Outcome: | The proposed benchmarks evaluate the reasoning, comprehension, and mathematical abilities of large language models. |