Generics are puzzling. Can language models find the missing piece? (2025.coling-main)
Copied to clipboard
| Challenge: | Generic sentences express generalisations about the world without explicit quantification . human biases in stereotypes can be observed in language models, authors say . |
| Approach: | They analyze generic sentences to determine their quantification and quantify their implicit quantifications using language models. |
| Outcome: | The proposed model shows that generics are more context-sensitive than determiner quantifiers and express weak generalisations. |
Similar Papers
Quantifying Generalizations: Exploring the Divide Between Human and LLMs’ Sensitivity to Quantification (2024.acl-long)
Copied to clipboard
| Challenge: | Generics are expressions used to communicate abstractions about categories . they allow for exceptions, and they are a powerful way to express knowledge about the world . |
| Approach: | They examine how large language models interpret generics to understand their meanings . they find that the presence of a generic sentence as context influences quantifiers based on the generalization . |
| Outcome: | The proposed models do not exhibit a strong sensitivity to quantification, the study finds . the results suggest that the presence of a generic sentence as context influences quantifiers . |
Generics are not quantificational: A new path from language models to semantic theory (2026.findings-acl)
Copied to clipboard
| Challenge: | Generic sentences express generalizations that tolerate exceptions without explicitly communicating information about quantities. |
| Approach: | They compare generics and quantificational sentences to find out what quantifiers are . they argue that generics are not quantificationals, contrary to dominant views . |
| Outcome: | The proposed model recovers many semantic facts about quantifiers and their "quantificational counterparts". |
Not all quantifiers are equal: Probing Transformer-based language models’ understanding of generalised quantifiers (2023.emnlp-main)
Copied to clipboard
| Challenge: | Recent popularity of generalised quantifiers and role in linguistics and logic raises the question of how they affect transformer-based language models (TLMs) |
| Approach: | They propose to use textual entailment to assess the ability of TLMs to learn the meanings of generalised quantifiers by using a textual model-checking problem defined in a purely logical sense. |
| Outcome: | The proposed method allows the automatic construction of datasets with respect to which we can assess the ability of TLMs to learn the meanings of generalised quantifiers. |
Rarely a problem? Language models exhibit inverse scaling in their predictions following few-type quantifiers (2023.findings-acl)
Copied to clipboard
| Challenge: | Current work suggests that language models deal poorly with quantifiers-they struggle to predict which quantifier is used in a given context and also perform poorly at generating appropriate continuations following logical quantifier. |
| Approach: | They propose to use 960 English sentence stimuli to build 22 autoregressive transformer models of different sizes to test their performance on ‘few’-type quantifiers. |
| Outcome: | The proposed models perform poorly on ‘few’-type quantifiers, and the larger the model, the worse its performance. |
Some of Them Can be Guessed! Exploring the Effect of Linguistic Context in Predicting Quantifiers (P18-2)
Copied to clipboard
| Challenge: | cloze deletion test is a test that requires the learner to understand the context and vocabulary in order to identify the correct word. |
| Approach: | They collect data from human participants and test various models in a local and a global context condition to examine the role of linguistic context in predicting quantifiers. |
| Outcome: | The proposed models outperform humans in a local and global context and are only slightly better in the latter. |
Probing for idiomaticity in vector space models (2021.eacl-main)
Copied to clipboard
| Challenge: | Contextualised word representation models are used to represent idiomaticity in language. |
| Approach: | They propose probing measures to assess if some of the expected linguistic properties of noun compounds are readily available in some standard and widely used representations. |
| Outcome: | The proposed models show that idiomaticity is not yet accurately represented by contextualised models. |
Generalized Quantifiers as a Source of Error in Multilingual NLU Benchmarks (2022.naacl-main)
Copied to clipboard
| Challenge: | Quantifiers are pervasive in NLU benchmarks and their occurrence at test time is associated with performance drops. |
| Approach: | They propose a generalized quantifier NLI task to quantify their contribution to the errors of NLU models. |
| Outcome: | The proposed model is based on a generalized quantifier theory and is compared with pre-trained models. |
Negation, Coordination, and Quantifiers in Contextualized Language Models (2022.coling-1)
Copied to clipboard
| Challenge: | Recent work has focused on specific tasks and on the learning outcome. |
| Approach: | They propose to decouple the weaknesses from specific tasks and focus on the embeddings per se and their mode of learning. |
| Outcome: | The proposed model can learn semantic constraints and how the context impacts their embeddings. |
Pragmatic Reasoning Unlocks Quantifier Semantics for Foundation Models (2023.emnlp-main)
Copied to clipboard
| Challenge: | Generalized quantifiers are used to indicate the proportions predicates satisfy (e.g., some apples are red). |
| Approach: | They propose a framework to model quantifier semantics for textbased foundation models by combining natural language inference and the Rational Speech Acts framework. |
| Outcome: | The proposed framework shows a 20% improvement over a literal listener baseline in predicting percentage scopes for quantifier comprehension even with no training. |
Discourse Realization of Generics in Human and LLM-generated Texts (2026.acl-long)
Copied to clipboard
| Challenge: | Large Language Models produce texts that appear coherent and credible, even when their factual reliability is uncertain. |
| Approach: | They propose a text-level genericity score derived from clause-level annotations and apply it to argumentative essays produced by humans and LLMs. |
| Outcome: | The proposed model is less generic than LLM-produced arguments, the study shows . higher genericity correlates with less structured, paratactic structures, the research shows a. |