The Sensitivity of Language Models and Humans to Winograd Schema Perturbations (2020.acl-main)
Copied to clipboard
| Challenge: | Large-scale pre-trained language models are driving recent improvements in perfromance on the Winograd Schema Challenge . a diagnostic dataset shows that these models are sensitive to linguistic perturbations that minimally affect human understanding . |
| Approach: | They propose to use a dataset to test pre-trained language models for the Winograd Schema Challenge . they show that these models are sensitive to linguistic perturbations that minimally affect human understanding . |
| Outcome: | The proposed models are sensitive to linguistic perturbations that minimally affect human understanding. |
Similar Papers
On the Effect of Hyperparameters in Language Modeling for Computational Linguistics (2026.acl-long)
Copied to clipboard
| Challenge: | Training language models and examining their linguistic behaviors is a common protocol in computational linguistics for studying linguistic phenomena and modeling human language processing. |
| Approach: | They replicate three prior studies with hyperparameters varied within a practical range and show that modest hyperparametric changes can alter qualitative conclusions about models’ linguistic abilities. |
| Outcome: | The results show that hyperparameter changes can alter qualitative conclusions and reverse the ranking of models. |
Do language models accommodate their users? A study of linguistic convergence (2026.eacl-long)
Copied to clipboard
| Challenge: | In this paper, we examine how large language models adapt their language use to the linguistic patterns of their user. |
| Approach: | They examine whether large language models exhibit linguistic convergence, a pragmatic element of human language communication, and compare their results to original human responses. |
| Outcome: | The proposed model language use is significantly different from that of humans. |
EvoGrad: A Dynamic Take on the Winograd Schema Challenge with Human Adversaries (2024.lrec-main)
Copied to clipboard
| Challenge: | Large Language Models excel at the Winograd Schema Challenge, but struggle with instances that feature minor alterations or rewording. |
| Approach: | They propose an open-source platform that harnesses a human-in-the-loop approach to create a dynamic dataset tailored to such altered WSC instances. |
| Outcome: | The proposed model outperforms existing models in the Winograd Schema Challenge (WSC) a human-in-the-loop approach allows for a dynamic dataset tailored to such altered instances. |
A fine-grained comparison of pragmatic language understanding in humans and language models (2023.acl-long)
Copied to clipboard
| Challenge: | Pragmatics and non-literal language understanding are essential to human communication . a long-standing challenge for artificial language models is to capture pragmatics . |
| Approach: | They compare language models and humans on seven pragmatic phenomena using curated English materials. |
| Outcome: | The proposed model achieves high accuracy and matches human error patterns . the results suggest pragmatic behaviors can emerge in models without explicit representations of mental states . |
A Surprisingly Robust Trick for the Winograd Schema Challenge (P19-1)
Copied to clipboard
| Challenge: | The Winograd Schema Challenge (WSC) dataset WSC273 and its inference counterpart WNLI are popular benchmarks for natural language understanding and commonsense reasoning. |
| Approach: | They propose to fine-tune language models on the Winograd Schema Challenge dataset WSC273 and its inference counterpart WNLI to achieve accuracies of 72.5% and 74.7%, respectively. |
| Outcome: | The proposed language models achieve 72.5% and 74.7% accuracy on the WSC273 and WNLI datasets, respectively. |
An Analysis of Dataset Overlap on Winograd-Style Tasks (2020.coling-main)
Copied to clipboard
| Challenge: | a large number of test instances overlap considerably with pretraining corpora, a study finds . for a number of years, models struggled to exceed chance-level performance . |
| Approach: | They analyze the effects of varying degrees of overlaps that occur between pretraining corpora and test instances in WSC-style tasks. |
| Outcome: | The WSC-Web dataset is the largest to date and has lower overlaps with current pretraining corpora. |
Are Language Models Agnostic to Linguistically Grounded Perturbations? A Case Study of Indic Languages (2025.findings-naacl)
Copied to clipboard
| Challenge: | Existing studies do not focus on linguistically grounded attacks, but pre-trained models are susceptible to these perturbations. |
| Approach: | They propose to examine whether pre-trained language models are agnostic to linguistically grounded attacks . they find that PLMs are less susceptible to linguistic perturbations than non-linguistic ones . |
| Outcome: | The proposed model is agnostic to linguistically grounded attacks, but is less susceptible to linguist attacks than non-linguistic models. |
Large Language Models: The Need for Nuance in Current Debates and a Pragmatic Perspective on Understanding (2023.emnlp-main)
Copied to clipboard
| Challenge: | Current Large Language Models (LLMs) are unparalleled in their ability to generate grammatically correct, fluent text. |
| Approach: | They argue that LLMs only parrot statistical patterns in training data and that language learning in LLM cannot inform human language learning. |
| Outcome: | The proposed model can generate grammatically correct, fluent text without requiring human intervention. |
Perceptions of Linguistic Uncertainty by Language Models and Humans (2024.emnlp-main)
Copied to clipboard
| Challenge: | Prior work has shown that humans are well-attuned to the use of uncertainty expressions, exhibiting population-level agreement in mapping these expressions to numerical responses. |
| Approach: | They propose to map linguistic expressions of uncertainty to numerical responses by using a theory of mind approach to understand the uncertainty of another agent. |
| Outcome: | The proposed model can map expressions to probabilistic responses in a human-like manner, but different behavior depending on whether a statement is actually true or false. |
Back to Square One: Artifact Detection, Training and Commonsense Disentanglement in the Winograd Schema (2021.emnlp-main)
Copied to clipboard
| Challenge: | Pre-trained language models have boosted performance on some WS benchmarks, but the source of improvement is not clear. |
| Approach: | They propose a method that uses twin sentences for evaluation and two new baselines that account for artifacts in WS benchmarks. |
| Outcome: | The proposed evaluation method is suboptimal for the Winograd Schema . it uses twin sentences to account for commonsense reasoning abilities . |