| Challenge: | In this paper, we compare syntactic knowledge of language models across different languages. |
| Approach: | They introduce a dataset for targeted syntactic evaluations of language models in Japanese. |
| Outcome: | The proposed dataset compares the syntactic knowledge of language models across languages. |
Similar Papers
MultiBLiMP 1.0: A Massively Multilingual Benchmark of Linguistic Minimal Pairs (2026.tacl-1)
Copied to clipboard
| Challenge: | MultiBLiMP 1.0 is a massively multilingual benchmark of linguistic minimal pairs covering 101 languages and 2 types of subject-verb agreement. |
| Approach: | They propose to use multilingual benchmarks to evaluate linguistic minimal pairs in 101 languages and 2 types of subject-verb agreement to create the minimal pairs. |
| Outcome: | The proposed benchmark covers 101 languages and 2 types of subject-verb agreement, and contains more than 128,000 minimal pairs. |
A Systematic Assessment of Language Models with Linguistic Minimal Pairs in Chinese (2026.tacl-1)
Copied to clipboard
Yikang Liu, Yeting Shen, Hongao Zhu, Lilong Xu, Zhiheng Qian, Siyuan Song, Kejia Zhang, Jialong Tang, Pei Zhang, Baosong Yang, Rui Wang, Hai Hu
| Challenge: | Using sub-linear length normalized log-probabilities (SLLN-LP), we find unequal lengths of sentences in minimal pairs difficult for LMs even up to 32B parameters. |
| Approach: | They propose to use ZhoBLiMP as a linguistic minimal pair benchmark for Chinese language models to mitigate biases. |
| Outcome: | The proposed metric mitigates biases in Chinese language models with over 100 paradigms . Anaphor, Quantifiers, and Ellipsis are difficult for LMs even up to 32B parameters . |
BLiMP: The Benchmark of Linguistic Minimal Pairs for English (2020.tacl-1)
Copied to clipboard
Alex Warstadt, Alicia Parrish, Haokun Liu, Anhad Mohananey, Wei Peng, Sheng-Fu Wang, Samuel R. Bowman
| Challenge: | Recent studies have examined how linguistic knowledge of language models (LMs) varies across English phenomena. |
| Approach: | They propose a benchmark to evaluate linguistic knowledge of language models on major grammatical phenomena in English. |
| Outcome: | The proposed benchmark evaluates the linguistic knowledge of language models on major grammatical phenomena in English. |
CLiMP: A Benchmark for Chinese Language Model Evaluation (2021.eacl-main)
Copied to clipboard
| Challenge: | Linguistically informed analyses of language models (LMs) contribute to understanding and improvement of such models. |
| Approach: | They introduce a corpus of Chinese linguistic minimal pairs (CLiMP) to investigate what knowledge Chinese LMs acquire. |
| Outcome: | The proposed corpus of Chinese linguistic minimal pairs (CLiMP) covers 9 major Chinese linguist phenomena. |
RuBLiMP: Russian Benchmark of Linguistic Minimal Pairs (2024.emnlp-main)
Copied to clipboard
Ekaterina Taktasheva, Maxim Bazhukov, Kirill Koncha, Alena Fenogenova, Ekaterina Artemova, Vladislav Mikhailov
| Challenge: | Existing resources for minimal pairs address a limited number of languages and lack diversity of language-specific grammatical phenomena. |
| Approach: | They propose to use a Russian benchmark of linguistic minimal pairs to evaluate grammatical knowledge of language models. |
| Outcome: | The proposed benchmark includes 45k pairs of sentences that differ in grammaticality and isolate a morphological, syntactic, or semantic phenomenon. |
JCoLA: Japanese Corpus of Linguistic Acceptability (2024.lrec-main)
Copied to clipboard
| Challenge: | Neural language models have exhibited outstanding performance in downstream tasks, yet there is limited understanding regarding the extent of their internalization of syntactic knowledge. |
| Approach: | They introduce a dataset that analyzes sentences annotated with binary acceptability judgments from linguistic textbooks and handbooks and splits them into in-domain and out-of-domain data. |
| Outcome: | The proposed datasets show that models can surpass human performance for in-domain data while no models can exceed human performance on out-of-domain datasets. |
UrBLiMP: A Benchmark for Evaluating the Linguistic Competence of Large Language Models in Urdu (2026.findings-acl)
Copied to clipboard
| Challenge: | Evaluating how large language models capture grammatical structure of low-resource languages remains underexplored. |
| Approach: | They evaluate a set of 5,696 minimal pairs that contrast grammatical acceptability across ten core syntactic and morpho-syntactical phenomena in Urdu. |
| Outcome: | The proposed framework compares multilingual models with the proprietary model . the proposed framework achieves the highest average accuracy on regular phenomena . |
TurBLiMP: A Turkish Benchmark of Linguistic Minimal Pairs (2025.emnlp-main)
Copied to clipboard
| Challenge: | TurBLiMP is the first benchmark of linguistic minimal pairs for monolingual and multilingual language models . it covers 16 linguistic phenomena with 1000 minimal pairs each . a foundational insight in linguistics research is that applying minimal changes to a sentence can render it entirely acceptable or unacceptable to native speakers. |
| Approach: | They propose to use morphologically rich agglutinative language with highly flexible word order to evaluate linguistic abilities of monolingual and multilingual language models. |
| Outcome: | The proposed benchmark covers 16 linguistic phenomena with 1000 minimal pairs each. |
JGLUE: Japanese General Language Understanding Evaluation (2022.lrec-1)
Copied to clipboard
| Challenge: | There is no benchmark for Japanese to evaluate and analyze NLU ability from different perspectives. |
| Approach: | They build a Japanese NLU benchmark from scratch without translation to measure general NLU ability in Japanese. |
| Outcome: | a Japanese NLU benchmark is built from scratch without translation to measure general NLU ability in Japanese. |
JMedBench: A Benchmark for Evaluating Japanese Biomedical Large Language Models (2025.coling-main)
Copied to clipboard
| Challenge: | Existing large language models (LLMs) focus on general domains, with fewer advancements in Japanese biomedical LLMs. |
| Approach: | They propose a benchmark for Japanese large language models with eight LLMs across four categories and 20 Japanese biomedical datasets for comparison. |
| Outcome: | The proposed benchmark includes eight LLMs across four categories and 20 Japanese biomedical datasets across five tasks. |