Challenge: Large-scale pre-trained language models are driving recent improvements in perfromance on the Winograd Schema Challenge . a diagnostic dataset shows that these models are sensitive to linguistic perturbations that minimally affect human understanding .
Approach: They propose to use a dataset to test pre-trained language models for the Winograd Schema Challenge . they show that these models are sensitive to linguistic perturbations that minimally affect human understanding .
Outcome: The proposed models are sensitive to linguistic perturbations that minimally affect human understanding.

Similar Papers

On the Effect of Hyperparameters in Language Modeling for Computational Linguistics (2026.acl-long)

Copied to clipboard

Challenge: Training language models and examining their linguistic behaviors is a common protocol in computational linguistics for studying linguistic phenomena and modeling human language processing.
Approach: They replicate three prior studies with hyperparameters varied within a practical range and show that modest hyperparametric changes can alter qualitative conclusions about models’ linguistic abilities.
Outcome: The results show that hyperparameter changes can alter qualitative conclusions and reverse the ranking of models.
Do language models accommodate their users? A study of linguistic convergence (2026.eacl-long)

Copied to clipboard

Challenge: In this paper, we examine how large language models adapt their language use to the linguistic patterns of their user.
Approach: They examine whether large language models exhibit linguistic convergence, a pragmatic element of human language communication, and compare their results to original human responses.
Outcome: The proposed model language use is significantly different from that of humans.
EvoGrad: A Dynamic Take on the Winograd Schema Challenge with Human Adversaries (2024.lrec-main)

Copied to clipboard

Challenge: Large Language Models excel at the Winograd Schema Challenge, but struggle with instances that feature minor alterations or rewording.
Approach: They propose an open-source platform that harnesses a human-in-the-loop approach to create a dynamic dataset tailored to such altered WSC instances.
Outcome: The proposed model outperforms existing models in the Winograd Schema Challenge (WSC) a human-in-the-loop approach allows for a dynamic dataset tailored to such altered instances.
A fine-grained comparison of pragmatic language understanding in humans and language models (2023.acl-long)

Copied to clipboard

Challenge: Pragmatics and non-literal language understanding are essential to human communication . a long-standing challenge for artificial language models is to capture pragmatics .
Approach: They compare language models and humans on seven pragmatic phenomena using curated English materials.
Outcome: The proposed model achieves high accuracy and matches human error patterns . the results suggest pragmatic behaviors can emerge in models without explicit representations of mental states .
A Surprisingly Robust Trick for the Winograd Schema Challenge (P19-1)

Copied to clipboard

Challenge: The Winograd Schema Challenge (WSC) dataset WSC273 and its inference counterpart WNLI are popular benchmarks for natural language understanding and commonsense reasoning.
Approach: They propose to fine-tune language models on the Winograd Schema Challenge dataset WSC273 and its inference counterpart WNLI to achieve accuracies of 72.5% and 74.7%, respectively.
Outcome: The proposed language models achieve 72.5% and 74.7% accuracy on the WSC273 and WNLI datasets, respectively.
An Analysis of Dataset Overlap on Winograd-Style Tasks (2020.coling-main)

Copied to clipboard

Challenge: a large number of test instances overlap considerably with pretraining corpora, a study finds . for a number of years, models struggled to exceed chance-level performance .
Approach: They analyze the effects of varying degrees of overlaps that occur between pretraining corpora and test instances in WSC-style tasks.
Outcome: The WSC-Web dataset is the largest to date and has lower overlaps with current pretraining corpora.
Are Language Models Agnostic to Linguistically Grounded Perturbations? A Case Study of Indic Languages (2025.findings-naacl)

Copied to clipboard

Challenge: Existing studies do not focus on linguistically grounded attacks, but pre-trained models are susceptible to these perturbations.
Approach: They propose to examine whether pre-trained language models are agnostic to linguistically grounded attacks . they find that PLMs are less susceptible to linguistic perturbations than non-linguistic ones .
Outcome: The proposed model is agnostic to linguistically grounded attacks, but is less susceptible to linguist attacks than non-linguistic models.
Large Language Models: The Need for Nuance in Current Debates and a Pragmatic Perspective on Understanding (2023.emnlp-main)

Copied to clipboard

Challenge: Current Large Language Models (LLMs) are unparalleled in their ability to generate grammatically correct, fluent text.
Approach: They argue that LLMs only parrot statistical patterns in training data and that language learning in LLM cannot inform human language learning.
Outcome: The proposed model can generate grammatically correct, fluent text without requiring human intervention.
Perceptions of Linguistic Uncertainty by Language Models and Humans (2024.emnlp-main)

Copied to clipboard

Challenge: Prior work has shown that humans are well-attuned to the use of uncertainty expressions, exhibiting population-level agreement in mapping these expressions to numerical responses.
Approach: They propose to map linguistic expressions of uncertainty to numerical responses by using a theory of mind approach to understand the uncertainty of another agent.
Outcome: The proposed model can map expressions to probabilistic responses in a human-like manner, but different behavior depending on whether a statement is actually true or false.
Back to Square One: Artifact Detection, Training and Commonsense Disentanglement in the Winograd Schema (2021.emnlp-main)

Copied to clipboard

Challenge: Pre-trained language models have boosted performance on some WS benchmarks, but the source of improvement is not clear.
Approach: They propose a method that uses twin sentences for evaluation and two new baselines that account for artifacts in WS benchmarks.
Outcome: The proposed evaluation method is suboptimal for the Winograd Schema . it uses twin sentences to account for commonsense reasoning abilities .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations