Challenge: Understanding language acquisition in language models remains an open question, yet many benchmarks focus on grammatical acceptability, with far less attention to interpreting meanings conveyed by grammatological forms.
Approach: They propose a benchmark to evaluate constructional understanding in language models using a controlled minimal-pair.
Outcome: The proposed benchmarks show that understanding of constructions develops more slowly and remains limited even in large language models (LLMs).

Similar Papers

BLiMP: The Benchmark of Linguistic Minimal Pairs for English (2020.tacl-1)

Copied to clipboard

Challenge: Recent studies have examined how linguistic knowledge of language models (LMs) varies across English phenomena.
Approach: They propose a benchmark to evaluate linguistic knowledge of language models on major grammatical phenomena in English.
Outcome: The proposed benchmark evaluates the linguistic knowledge of language models on major grammatical phenomena in English.
A Construction Grammar Corpus of Varying Schematicity: A Dataset for the Evaluation of Abstractions in Language Models (2024.lrec-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) have been developed without a theoretical framework . evaluating and improving LLMs will benefit from theoretical frameworks that enable comparison of structures of human language and model of language built up by LLM.
Approach: They propose to use a construction grammar schema corpus to compare human grammar to LLMs' model of language.
Outcome: The proposed corpus shows that even the largest LLMs are limited to more substantive constructions and do not recognize similarity of purely schematic constructions.
Enhancing Language Representation with Constructional Information for Natural Language Understanding (2023.acl-long)

Copied to clipboard

Challenge: Recent advances in natural language processing focus on acquiring lexico-semantic information.
Approach: They propose a construction grammar which highlights the pairings of form and meaning to enrich language representation.
Outcome: The proposed model is superior to existing models on a variety of NLU tasks.
CLiMP: A Benchmark for Chinese Language Model Evaluation (2021.eacl-main)

Copied to clipboard

Challenge: Linguistically informed analyses of language models (LMs) contribute to understanding and improvement of such models.
Approach: They introduce a corpus of Chinese linguistic minimal pairs (CLiMP) to investigate what knowledge Chinese LMs acquire.
Outcome: The proposed corpus of Chinese linguistic minimal pairs (CLiMP) covers 9 major Chinese linguist phenomena.
A Systematic Assessment of Language Models with Linguistic Minimal Pairs in Chinese (2026.tacl-1)

Copied to clipboard

Challenge: Using sub-linear length normalized log-probabilities (SLLN-LP), we find unequal lengths of sentences in minimal pairs difficult for LMs even up to 32B parameters.
Approach: They propose to use ZhoBLiMP as a linguistic minimal pair benchmark for Chinese language models to mitigate biases.
Outcome: The proposed metric mitigates biases in Chinese language models with over 100 paradigms . Anaphor, Quantifiers, and Ellipsis are difficult for LMs even up to 32B parameters .
BabyLM’s First Constructions: Causal interventions provide a signal of learning (2025.emnlp-main)

Copied to clipboard

Challenge: Recent work shows sensitivity to constructions in pretrained language models, but their relevance to human language learning is doubted.
Approach: They use construction grammars to demonstrate sensitivity to constructions in pretrained language models.
Outcome: The proposed models learn diverse constructions even hard cases that are superficially indistinguishable.
CxGBERT: BERT meets Construction Grammar (2020.coling-main)

Copied to clipboard

Challenge: lexico-semantic elements capture a large amount of linguistic information, but they do not capture all information contained in text.
Approach: They propose to use BERT to train a model that uses a deep bidirectional transformer to capture a significant amount of lexico-semantic information.
Outcome: The proposed model captures lexico-semantic information, but it is redundantly encoded in lexical information.
CxLM: A Construction and Context-aware Language Model (2022.lrec-1)

Copied to clipboard

Challenge: Constructions are direct form-meaning pairs with possible schematic slots . however, these slots are constrained by the embedded construction and the context . we propose that a conditional probability distribution could be described but language models cannot capture this distribution.
Approach: They propose that a conditional probability distribution could describe constructions’ schematic slots.
Outcome: The proposed model predicts masked slots more accurately than baselines and generates structurally and semantically plausible word samples.
The better your Syntax, the better your Semantics? Probing Pretrained Language Models for the English Comparative Correlative (2022.emnlp-main)

Copied to clipboard

Challenge: Construction Grammar posits constructions as the central building blocks of language . human-like performance of pretrained language models on many NLP tasks has been alleged .
Approach: They propose to use construction grammar to posit constructions as the central building blocks of language . they conduct experiments with three pretrained language models to examine their ability to classify and understand English comparative correlative .
Outcome: The proposed models are able to recognise the English comparative correlative (CC) but fail to use its meaning.
Discriminating Form and Meaning in Multilingual Models with Minimal-Pair ABX Tasks (2025.emnlp-main)

Copied to clipboard

Challenge: Existing studies have shown that multilingual models encode languagespecific information and language-agnostic features, but the nature and interaction of these representations is not fully understood.
Approach: They propose a set of training-free ABX-style discrimination tasks to evaluate how multilingual language models represent language identity (form) and semantic content (meaning).
Outcome: The proposed tasks show that language discrimination declines over training and strengthens over time and stabilizes in deeper layers.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations