| Challenge: | a dataset containing 20,300 human ratings on quantified statements is used to evaluate the appropriateness of vague quantifiers in visual contexts. |
| Approach: | They use a visual-language-models-based dataset to evaluate the appropriateness of vague quantifiers. |
| Outcome: | The proposed model is based on a visual-visual-language-model-based dataset . it shows that the model is compatible with humans when producing or judging vague quantifiers . |
Similar Papers
Naming, Describing, and Quantifying Visual Objects in Humans and LLMs (2024.acl-short)
Copied to clipboard
| Challenge: | Recent work has highlighted that speakers display a wide range of variability when asked to utter sentences, resulting in inter-speaker variability but also variability over time for the same speaker. |
| Approach: | They evaluate Vision & Language Large Language Models (VLLMs) on three categories where humans show great subjective variability concerning the distribution over plausible labels. |
| Outcome: | The proposed models can mimic human distributions over plausible labels, but fail to assign quantifiers, a task that requires more accurate, high-level reasoning. |
Some of Them Can be Guessed! Exploring the Effect of Linguistic Context in Predicting Quantifiers (P18-2)
Copied to clipboard
| Challenge: | cloze deletion test is a test that requires the learner to understand the context and vocabulary in order to identify the correct word. |
| Approach: | They collect data from human participants and test various models in a local and a global context condition to examine the role of linguistic context in predicting quantifiers. |
| Outcome: | The proposed models outperform humans in a local and global context and are only slightly better in the latter. |
Pragmatic Reasoning Unlocks Quantifier Semantics for Foundation Models (2023.emnlp-main)
Copied to clipboard
| Challenge: | Generalized quantifiers are used to indicate the proportions predicates satisfy (e.g., some apples are red). |
| Approach: | They propose a framework to model quantifier semantics for textbased foundation models by combining natural language inference and the Rational Speech Acts framework. |
| Outcome: | The proposed framework shows a 20% improvement over a literal listener baseline in predicting percentage scopes for quantifier comprehension even with no training. |
Where is this coming from? Making groundedness count in the evaluation of Document VQA models (2025.findings-naacl)
Copied to clipboard
| Challenge: | Document Visual Question Answering (VQA) models have come close to or matching human performance on some benchmarks. |
| Approach: | They propose a method that accounts for the semantic and multimodal groundedness of a model’s outputs and can be parameterized so that users can configure the score according to their preferences. |
| Outcome: | The proposed method produces scores that are a better indicator of a model’s robustness and tends to give higher rewards to better-calibrated answers. |
Rarely a problem? Language models exhibit inverse scaling in their predictions following few-type quantifiers (2023.findings-acl)
Copied to clipboard
| Challenge: | Current work suggests that language models deal poorly with quantifiers-they struggle to predict which quantifier is used in a given context and also perform poorly at generating appropriate continuations following logical quantifier. |
| Approach: | They propose to use 960 English sentence stimuli to build 22 autoregressive transformer models of different sizes to test their performance on ‘few’-type quantifiers. |
| Outcome: | The proposed models perform poorly on ‘few’-type quantifiers, and the larger the model, the worse its performance. |
Quantifying Generalizations: Exploring the Divide Between Human and LLMs’ Sensitivity to Quantification (2024.acl-long)
Copied to clipboard
| Challenge: | Generics are expressions used to communicate abstractions about categories . they allow for exceptions, and they are a powerful way to express knowledge about the world . |
| Approach: | They examine how large language models interpret generics to understand their meanings . they find that the presence of a generic sentence as context influences quantifiers based on the generalization . |
| Outcome: | The proposed models do not exhibit a strong sensitivity to quantification, the study finds . the results suggest that the presence of a generic sentence as context influences quantifiers . |
How Does Quantization Affect Multilingual LLMs? (2024.findings-emnlp)
Copied to clipboard
Kelly Marchisio, Saurabh Dash, Hongyu Chen, Dennis Aumiller, Ahmet Üstün, Sara Hooker, Sebastian Ruder
| Challenge: | Quantization is widely used to improve inference speed and deployment of large language models. |
| Approach: | They conduct a thorough analysis of quantized multilingual LLMs . they find language disparately affected by quantization, non-Latin script languages worst . authors urge consideration of multilingual performance as evaluation criterion for efficient models . |
| Outcome: | The results show that quantization has harmful effects on human evaluation . language performance is disparately affected by quantization, the authors say . |
The Art of Saying "Maybe": A Conformal Lens for Uncertainty Benchmarking in VLMs (2026.findings-eacl)
Copied to clipboard
| Challenge: | Recent advances in large vision-language models have led to remarkable progress in complex visual understanding across scientific and reasoning tasks. |
| Approach: | They evaluate 18 state-of-the-art vision-language models across 6 multimodal datasets with 3 distinct scoring functions and develop instruction-guided likelihood proxies for closed-source models lacking token-level logprob access. |
| Outcome: | The proposed model is able to achieve higher accuracy on multimodal benchmarks while performing poorer on reasoning tasks. |
Visual Commonsense in Pretrained Unimodal and Multimodal Models (2022.naacl-main)
Copied to clipboard
| Challenge: | Fig. 1 shows how text-only and image-only models can capture commonsense visual attributes, but reporting bias affects their performance. |
| Approach: | They use a Visual Commonsense Tests dataset to validate their findings . they find multimodal models better reconstruct attribute distributions, but are still subject to reporting bias . |
| Outcome: | The proposed model improves on the unimodal and multimodal models, but is still subject to reporting bias. |
FOCUS: Evaluating Pre-trained Vision-Language Models on Underspecification Reasoning (2025.acl-long)
Copied to clipboard
| Challenge: | a new dataset evaluates whether vision-language models have underspecification reasoning abilities . underspecifications are often left incomplete or vague, and are often ignored for mutual understanding . |
| Approach: | They propose a probing dataset to evaluate whether VLMs have underspecification reasoning . they find that pre-trained vision-language models lack this ability . |
| Outcome: | The proposed probing dataset shows that pre-trained vision-language models lack underspecification reasoning abilities. |