Challenge: a dataset containing 20,300 human ratings on quantified statements is used to evaluate the appropriateness of vague quantifiers in visual contexts.
Approach: They use a visual-language-models-based dataset to evaluate the appropriateness of vague quantifiers.
Outcome: The proposed model is based on a visual-visual-language-model-based dataset . it shows that the model is compatible with humans when producing or judging vague quantifiers .

Similar Papers

Naming, Describing, and Quantifying Visual Objects in Humans and LLMs (2024.acl-short)

Copied to clipboard

Challenge: Recent work has highlighted that speakers display a wide range of variability when asked to utter sentences, resulting in inter-speaker variability but also variability over time for the same speaker.
Approach: They evaluate Vision & Language Large Language Models (VLLMs) on three categories where humans show great subjective variability concerning the distribution over plausible labels.
Outcome: The proposed models can mimic human distributions over plausible labels, but fail to assign quantifiers, a task that requires more accurate, high-level reasoning.
Some of Them Can be Guessed! Exploring the Effect of Linguistic Context in Predicting Quantifiers (P18-2)

Copied to clipboard

Challenge: cloze deletion test is a test that requires the learner to understand the context and vocabulary in order to identify the correct word.
Approach: They collect data from human participants and test various models in a local and a global context condition to examine the role of linguistic context in predicting quantifiers.
Outcome: The proposed models outperform humans in a local and global context and are only slightly better in the latter.
Pragmatic Reasoning Unlocks Quantifier Semantics for Foundation Models (2023.emnlp-main)

Copied to clipboard

Challenge: Generalized quantifiers are used to indicate the proportions predicates satisfy (e.g., some apples are red).
Approach: They propose a framework to model quantifier semantics for textbased foundation models by combining natural language inference and the Rational Speech Acts framework.
Outcome: The proposed framework shows a 20% improvement over a literal listener baseline in predicting percentage scopes for quantifier comprehension even with no training.
Where is this coming from? Making groundedness count in the evaluation of Document VQA models (2025.findings-naacl)

Copied to clipboard

Challenge: Document Visual Question Answering (VQA) models have come close to or matching human performance on some benchmarks.
Approach: They propose a method that accounts for the semantic and multimodal groundedness of a model’s outputs and can be parameterized so that users can configure the score according to their preferences.
Outcome: The proposed method produces scores that are a better indicator of a model’s robustness and tends to give higher rewards to better-calibrated answers.
Rarely a problem? Language models exhibit inverse scaling in their predictions following few-type quantifiers (2023.findings-acl)

Copied to clipboard

Challenge: Current work suggests that language models deal poorly with quantifiers-they struggle to predict which quantifier is used in a given context and also perform poorly at generating appropriate continuations following logical quantifier.
Approach: They propose to use 960 English sentence stimuli to build 22 autoregressive transformer models of different sizes to test their performance on ‘few’-type quantifiers.
Outcome: The proposed models perform poorly on ‘few’-type quantifiers, and the larger the model, the worse its performance.
Quantifying Generalizations: Exploring the Divide Between Human and LLMs’ Sensitivity to Quantification (2024.acl-long)

Copied to clipboard

Challenge: Generics are expressions used to communicate abstractions about categories . they allow for exceptions, and they are a powerful way to express knowledge about the world .
Approach: They examine how large language models interpret generics to understand their meanings . they find that the presence of a generic sentence as context influences quantifiers based on the generalization .
Outcome: The proposed models do not exhibit a strong sensitivity to quantification, the study finds . the results suggest that the presence of a generic sentence as context influences quantifiers .
How Does Quantization Affect Multilingual LLMs? (2024.findings-emnlp)

Copied to clipboard

Challenge: Quantization is widely used to improve inference speed and deployment of large language models.
Approach: They conduct a thorough analysis of quantized multilingual LLMs . they find language disparately affected by quantization, non-Latin script languages worst . authors urge consideration of multilingual performance as evaluation criterion for efficient models .
Outcome: The results show that quantization has harmful effects on human evaluation . language performance is disparately affected by quantization, the authors say .
The Art of Saying "Maybe": A Conformal Lens for Uncertainty Benchmarking in VLMs (2026.findings-eacl)

Copied to clipboard

Challenge: Recent advances in large vision-language models have led to remarkable progress in complex visual understanding across scientific and reasoning tasks.
Approach: They evaluate 18 state-of-the-art vision-language models across 6 multimodal datasets with 3 distinct scoring functions and develop instruction-guided likelihood proxies for closed-source models lacking token-level logprob access.
Outcome: The proposed model is able to achieve higher accuracy on multimodal benchmarks while performing poorer on reasoning tasks.
Visual Commonsense in Pretrained Unimodal and Multimodal Models (2022.naacl-main)

Copied to clipboard

Challenge: Fig. 1 shows how text-only and image-only models can capture commonsense visual attributes, but reporting bias affects their performance.
Approach: They use a Visual Commonsense Tests dataset to validate their findings . they find multimodal models better reconstruct attribute distributions, but are still subject to reporting bias .
Outcome: The proposed model improves on the unimodal and multimodal models, but is still subject to reporting bias.
FOCUS: Evaluating Pre-trained Vision-Language Models on Underspecification Reasoning (2025.acl-long)

Copied to clipboard

Challenge: a new dataset evaluates whether vision-language models have underspecification reasoning abilities . underspecifications are often left incomplete or vague, and are often ignored for mutual understanding .
Approach: They propose a probing dataset to evaluate whether VLMs have underspecification reasoning . they find that pre-trained vision-language models lack this ability .
Outcome: The proposed probing dataset shows that pre-trained vision-language models lack underspecification reasoning abilities.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations