Challenge: Vision-Language Models (VLMs) perform on par with larger models in general domain visual grounding and question-answering benchmarks.
Approach: They propose a "Uncontextualized Uncommon Objects" benchmark to evaluate their performance on common datasets.
Outcome: The proposed benchmark focuses on systematically testing VLMs with both large and small parameter counts on rare and specialized objects.

Similar Papers

Response Wide Shut? Surprising Observations in Basic Vision Language Model Capabilities (2025.acl-long)

Copied to clipboard

Challenge: Vision-language Models have been shown to be highly capable but lacking basic visual understanding skills.
Approach: They propose to examine the limitations of vision-language models on visual tasks by constructing a series of tests that probe which components of design may be lacking.
Outcome: The proposed tests compare VLMs to other models on visual encoders, intermediate vision-language projection and LLM-decoder outputs.
Beyond Visual Understanding Introducing PARROT-360V for Vision Language Model Benchmarking (2025.coling-industry)

Copied to clipboard

Challenge: Current benchmarks for evaluating Vision Language Models (VLMs) often fail to thoroughly assess these models’ abilities to understand complex visual and textual content.
Approach: They propose a benchmark that features 2487 visual puzzles designed to test VLMs on complex visual reasoning tasks.
Outcome: The PARROT-360V Benchmark features 2487 visual puzzles designed to test VLMs on complex visual reasoning tasks.
EffiVLM-BENCH: A Comprehensive Benchmark for Evaluating Training-Free Acceleration in Large Vision-Language Models (2025.acl-long)

Copied to clipboard

Challenge: Existing methods for accelerating Large Vision-Language Models lack comprehensive evaluation across diverse backbones, benchmarks, and metrics.
Approach: They propose EffiVLM-BENCH framework for evaluating absolute performance and generalization and loyalty.
Outcome: The proposed framework offers insights into optimal strategies for accelerating LVLMs.
African or European Swallow? Benchmarking Large Vision-Language Models for Fine-Grained Object Classification (2024.emnlp-main)

Copied to clipboard

Challenge: Recent Large Vision Language Models demonstrate impressive abilities on image understanding and reasoning tasks.
Approach: They propose a benchmark for fine-grained object classification that is difficult to evaluate . they benchmark 12 public LVLMs on and show CLIP models exhibit better performance .
Outcome: The proposed model improves on 12 public LVLMs on image understanding and reasoning tasks.
MEGAVERSE: Benchmarking Large Language Models Across Languages, Modalities, Models and Tasks (2024.naacl-long)

Copied to clipboard

Challenge: Several new LLMs have been introduced necessitating their evaluation on non-English languages.
Approach: They perform a thorough evaluation of the non-English capabilities of SoTA LLMs by comparing them on the same set of multilingual datasets.
Outcome: The proposed model outperforms models on multilingual datasets on 22 languages including low-resource African languages.
Losing Visual Needles in Image Haystacks: Vision Language Models are Easily Distracted in Short and Long Contexts (2024.findings-emnlp)

Copied to clipboard

Challenge: evaluators of long-context vision language models (VLMs) have not kept up with the rapid development of open-weight long-constraint language models.
Approach: They propose a dynamic benchmark generator for evaluating long-context reasoning in vision language models.
Outcome: The proposed model can ignore irrelevant information when answering queries, showing that current models lack this capability.
VLURes: Benchmarking Long-Text Grounding and Cross-Lingual Robustness in Vision Language Models (2026.findings-acl)

Copied to clipboard

Challenge: ***VLURes** provides a practical testbed for long-text grounding and multilingual robustness in web-realistic agent settings.
Approach: They propose a multilingual benchmark for evaluating vision-language models under long-text grounding.
Outcome: ***VLURes** provides a testbed for long-text grounding and multilingual robustness in web-realistic agent settings.
Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward (2025.findings-acl)

Copied to clipboard

Challenge: Large Vision Language Models (LVLMs) have shown impressive performance on various vision-language tasks.
Approach: They propose a benchmark framework for evaluating Visual Variation Robustness of Large Vision Language Models that incorporates automated evaluation dataset generation and principled metrics for thorough robustness assessment.
Outcome: The proposed framework identifies a vulnerability to visual variations affecting even advanced models that excel at complex vision-language tasks but significantly underperform on simple tasks like object recognition.
The Art of Saying "Maybe": A Conformal Lens for Uncertainty Benchmarking in VLMs (2026.findings-eacl)

Copied to clipboard

Challenge: Recent advances in large vision-language models have led to remarkable progress in complex visual understanding across scientific and reasoning tasks.
Approach: They evaluate 18 state-of-the-art vision-language models across 6 multimodal datasets with 3 distinct scoring functions and develop instruction-guided likelihood proxies for closed-source models lacking token-level logprob access.
Outcome: The proposed model is able to achieve higher accuracy on multimodal benchmarks while performing poorer on reasoning tasks.
Finer: Investigating and Enhancing Fine-Grained Visual Concept Recognition in Large Vision Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: Recent advances in instruction-tuned Large Vision-Language Models (LVLMs) have imbued the models with the ability to generate high-level, image-grounded explanations with ease.
Approach: They propose to use a multiple granularity attribute-centric benchmark and training mixture to evaluate LVLMs’ fine-grained visual comprehension ability.
Outcome: The proposed model improves on LLaVa-1.5, InstructBLIP and GPT-4V and demonstrates that they struggle to generate descriptive visual attributes based on a concept that appears within an input image despite their prominent zero-shot image captioning ability.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations