Challenge: cloze deletion test is a test that requires the learner to understand the context and vocabulary in order to identify the correct word.
Approach: They collect data from human participants and test various models in a local and a global context condition to examine the role of linguistic context in predicting quantifiers.
Outcome: The proposed models outperform humans in a local and global context and are only slightly better in the latter.

Similar Papers

Rarely a problem? Language models exhibit inverse scaling in their predictions following few-type quantifiers (2023.findings-acl)

Copied to clipboard

Challenge: Current work suggests that language models deal poorly with quantifiers-they struggle to predict which quantifier is used in a given context and also perform poorly at generating appropriate continuations following logical quantifier.
Approach: They propose to use 960 English sentence stimuli to build 22 autoregressive transformer models of different sizes to test their performance on ‘few’-type quantifiers.
Outcome: The proposed models perform poorly on ‘few’-type quantifiers, and the larger the model, the worse its performance.
Not all quantifiers are equal: Probing Transformer-based language models’ understanding of generalised quantifiers (2023.emnlp-main)

Copied to clipboard

Challenge: Recent popularity of generalised quantifiers and role in linguistics and logic raises the question of how they affect transformer-based language models (TLMs)
Approach: They propose to use textual entailment to assess the ability of TLMs to learn the meanings of generalised quantifiers by using a textual model-checking problem defined in a purely logical sense.
Outcome: The proposed method allows the automatic construction of datasets with respect to which we can assess the ability of TLMs to learn the meanings of generalised quantifiers.
Can Large Language Models Understand Context? (2024.findings-eacl)

Copied to clipboard

Challenge: Existing evaluation methodologies for Large Language Models (LLMs) have been inadequate to evaluate their ability to understand contextual features.
Approach: They propose a benchmark to assess large language models' ability to understand context by adapting existing datasets to suit their evaluation.
Outcome: The proposed model performs better under the in-context learning pretraining scenario than state-of-the-art models.
How Does Quantization Affect Multilingual LLMs? (2024.findings-emnlp)

Copied to clipboard

Challenge: Quantization is widely used to improve inference speed and deployment of large language models.
Approach: They conduct a thorough analysis of quantized multilingual LLMs . they find language disparately affected by quantization, non-Latin script languages worst . authors urge consideration of multilingual performance as evaluation criterion for efficient models .
Outcome: The results show that quantization has harmful effects on human evaluation . language performance is disparately affected by quantization, the authors say .
Do Language Models Exhibit Human-like Structural Priming Effects? (2024.findings-acl)

Copied to clipboard

Challenge: a recent exposure to a structure facilitates processing of the same structure, a study finds . structural priming is well attested in humans, for both language production and comprehension .
Approach: They use the structural priming paradigm to investigate where priming effects manifest . they find that rarer elements within a prime increase priming effect .
Outcome: The findings provide an important piece in the puzzle of understanding how properties within their context affect structural prediction in language models.
Negation, Coordination, and Quantifiers in Contextualized Language Models (2022.coling-1)

Copied to clipboard

Challenge: Recent work has focused on specific tasks and on the learning outcome.
Approach: They propose to decouple the weaknesses from specific tasks and focus on the embeddings per se and their mode of learning.
Outcome: The proposed model can learn semantic constraints and how the context impacts their embeddings.
How Quantization Shapes Bias in Large Language Models (2026.eacl-long)

Copied to clipboard

Challenge: a systematic review of quantization's effects on model biases focuses on stereotypes, fairness, toxicity, and sentiment.
Approach: They focus on weight and activation quantization strategies and examine their effects across bias types including stereotypes, fairness, toxicity, and sentiment.
Outcome: The proposed method can reduce stereotypes and unfairness, but it tends to increase stereotypes in generative tasks.
What Context Features Can Transformer Language Models Use? (2021.acl-long)

Copied to clipboard

Challenge: Recent studies show that transformer-based language models benefit from conditioning on contexts of hundreds to thousands of previous tokens.
Approach: They propose to use lexical and structural information to ablate usable information in transformer language models.
Outcome: The proposed model improves when conditioning on contexts of thousands of previous tokens.
Predicting Reference: What do Language Models Learn about Discourse Models? (2020.emnlp-main)

Copied to clipboard

Challenge: a growing literature that probes neural language models to assess their latent acquisition of grammatical knowledge has not investigated their acquisition of discourse modeling ability.
Approach: They draw on a psycholinguistic literature that has established how different contexts affect referential biases concerning who is likely to be referred to next.
Outcome: The proposed models do not resemble human language users, the authors show . their models capture the linguistic knowledge required to perform discourse modeling .
Generics are puzzling. Can language models find the missing piece? (2025.coling-main)

Copied to clipboard

Challenge: Generic sentences express generalisations about the world without explicit quantification . human biases in stereotypes can be observed in language models, authors say .
Approach: They analyze generic sentences to determine their quantification and quantify their implicit quantifications using language models.
Outcome: The proposed model shows that generics are more context-sensitive than determiner quantifiers and express weak generalisations.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations