Challenge: In this paper, we compare syntactic knowledge of language models across different languages.
Approach: They introduce a dataset for targeted syntactic evaluations of language models in Japanese.
Outcome: The proposed dataset compares the syntactic knowledge of language models across languages.

Similar Papers

MultiBLiMP 1.0: A Massively Multilingual Benchmark of Linguistic Minimal Pairs (2026.tacl-1)

Copied to clipboard

Challenge: MultiBLiMP 1.0 is a massively multilingual benchmark of linguistic minimal pairs covering 101 languages and 2 types of subject-verb agreement.
Approach: They propose to use multilingual benchmarks to evaluate linguistic minimal pairs in 101 languages and 2 types of subject-verb agreement to create the minimal pairs.
Outcome: The proposed benchmark covers 101 languages and 2 types of subject-verb agreement, and contains more than 128,000 minimal pairs.
A Systematic Assessment of Language Models with Linguistic Minimal Pairs in Chinese (2026.tacl-1)

Copied to clipboard

Challenge: Using sub-linear length normalized log-probabilities (SLLN-LP), we find unequal lengths of sentences in minimal pairs difficult for LMs even up to 32B parameters.
Approach: They propose to use ZhoBLiMP as a linguistic minimal pair benchmark for Chinese language models to mitigate biases.
Outcome: The proposed metric mitigates biases in Chinese language models with over 100 paradigms . Anaphor, Quantifiers, and Ellipsis are difficult for LMs even up to 32B parameters .
BLiMP: The Benchmark of Linguistic Minimal Pairs for English (2020.tacl-1)

Copied to clipboard

Challenge: Recent studies have examined how linguistic knowledge of language models (LMs) varies across English phenomena.
Approach: They propose a benchmark to evaluate linguistic knowledge of language models on major grammatical phenomena in English.
Outcome: The proposed benchmark evaluates the linguistic knowledge of language models on major grammatical phenomena in English.
CLiMP: A Benchmark for Chinese Language Model Evaluation (2021.eacl-main)

Copied to clipboard

Challenge: Linguistically informed analyses of language models (LMs) contribute to understanding and improvement of such models.
Approach: They introduce a corpus of Chinese linguistic minimal pairs (CLiMP) to investigate what knowledge Chinese LMs acquire.
Outcome: The proposed corpus of Chinese linguistic minimal pairs (CLiMP) covers 9 major Chinese linguist phenomena.
RuBLiMP: Russian Benchmark of Linguistic Minimal Pairs (2024.emnlp-main)

Copied to clipboard

Challenge: Existing resources for minimal pairs address a limited number of languages and lack diversity of language-specific grammatical phenomena.
Approach: They propose to use a Russian benchmark of linguistic minimal pairs to evaluate grammatical knowledge of language models.
Outcome: The proposed benchmark includes 45k pairs of sentences that differ in grammaticality and isolate a morphological, syntactic, or semantic phenomenon.
JCoLA: Japanese Corpus of Linguistic Acceptability (2024.lrec-main)

Copied to clipboard

Challenge: Neural language models have exhibited outstanding performance in downstream tasks, yet there is limited understanding regarding the extent of their internalization of syntactic knowledge.
Approach: They introduce a dataset that analyzes sentences annotated with binary acceptability judgments from linguistic textbooks and handbooks and splits them into in-domain and out-of-domain data.
Outcome: The proposed datasets show that models can surpass human performance for in-domain data while no models can exceed human performance on out-of-domain datasets.
UrBLiMP: A Benchmark for Evaluating the Linguistic Competence of Large Language Models in Urdu (2026.findings-acl)

Copied to clipboard

Challenge: Evaluating how large language models capture grammatical structure of low-resource languages remains underexplored.
Approach: They evaluate a set of 5,696 minimal pairs that contrast grammatical acceptability across ten core syntactic and morpho-syntactical phenomena in Urdu.
Outcome: The proposed framework compares multilingual models with the proprietary model . the proposed framework achieves the highest average accuracy on regular phenomena .
TurBLiMP: A Turkish Benchmark of Linguistic Minimal Pairs (2025.emnlp-main)

Copied to clipboard

Challenge: TurBLiMP is the first benchmark of linguistic minimal pairs for monolingual and multilingual language models . it covers 16 linguistic phenomena with 1000 minimal pairs each . a foundational insight in linguistics research is that applying minimal changes to a sentence can render it entirely acceptable or unacceptable to native speakers.
Approach: They propose to use morphologically rich agglutinative language with highly flexible word order to evaluate linguistic abilities of monolingual and multilingual language models.
Outcome: The proposed benchmark covers 16 linguistic phenomena with 1000 minimal pairs each.
JGLUE: Japanese General Language Understanding Evaluation (2022.lrec-1)

Copied to clipboard

Challenge: There is no benchmark for Japanese to evaluate and analyze NLU ability from different perspectives.
Approach: They build a Japanese NLU benchmark from scratch without translation to measure general NLU ability in Japanese.
Outcome: a Japanese NLU benchmark is built from scratch without translation to measure general NLU ability in Japanese.
JMedBench: A Benchmark for Evaluating Japanese Biomedical Large Language Models (2025.coling-main)

Copied to clipboard

Challenge: Existing large language models (LLMs) focus on general domains, with fewer advancements in Japanese biomedical LLMs.
Approach: They propose a benchmark for Japanese large language models with eight LLMs across four categories and 20 Japanese biomedical datasets for comparison.
Outcome: The proposed benchmark includes eight LLMs across four categories and 20 Japanese biomedical datasets across five tasks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations