Challenge: Vision-language models excel at tasks requiring joint understanding of visual information and natural language.
Approach: They propose to use choropleth maps to answer questions from three geographical regions in the United States, India, China as question templates.
Outcome: The proposed model outperforms other models in the area of visual language and visual question answering.

Similar Papers

SpaRE: Enhancing Spatial Reasoning in Vision-Language Models with Synthetic Data (2025.acl-long)

Copied to clipboard

Challenge: Vision-language models struggle with spatial reasoning, a skill that humans excel at.
Approach: They propose to use a spatial-reasoning Enhanced (SpaRE) VLM to improve spatial reasoning in visual question answering and robotics.
Outcome: The proposed model achieves a 49% performance gain on the What's Up benchmark while maintaining strong results on general tasks.
Can VLMs Actually See and Read? A Survey on Modality Collapse in Vision-Language Models (2025.findings-acl)

Copied to clipboard

Challenge: Vision-language models integrate textual and visual information, enabling them to process visual inputs and generate predictions.
Approach: They review work on modality collapse analysis to provide insights into the reason for this unintended behavior and review probing studies for fine-grained vision-language understanding.
Outcome: The proposed models can achieve competitive performance in vision-language tasks despite relying heavily on textual information and ignoring visual information.
Evaluation of Multilingual Ability to Use Spatial Deictic Expressions in Vision-Language Models (2026.acl-srw)

Copied to clipboard

Challenge: Existing studies have focused on the ability of vision-language models to utilize spatial deictic expressions, which depend on the situation of utterance.
Approach: They develop a benchmark to evaluate the multilingual ability of VLMs to use spatial deictic expressions in four languages.
Outcome: The proposed models use demonstratives in a different manner from humans, particularly in selecting demonstrative based on distance from the object.
Unraveling the Truth: Do VLMs really Understand Charts? A Deep Dive into Consistency and Robustness (2024.findings-emnlp)

Copied to clipboard

Challenge: Chart question answering (CQA) is a crucial area of Visual Language Understanding.
Approach: They evaluate the robustness and consistency of current Visual Language Models on a dataset encompassing diverse question categories and chart formats.
Outcome: The proposed models handle varying levels of chart and question complexity and are robust across different visual representations of the same underlying data.
Beyond Visual Understanding Introducing PARROT-360V for Vision Language Model Benchmarking (2025.coling-industry)

Copied to clipboard

Challenge: Current benchmarks for evaluating Vision Language Models (VLMs) often fail to thoroughly assess these models’ abilities to understand complex visual and textual content.
Approach: They propose a benchmark that features 2487 visual puzzles designed to test VLMs on complex visual reasoning tasks.
Outcome: The PARROT-360V Benchmark features 2487 visual puzzles designed to test VLMs on complex visual reasoning tasks.
NTSEBENCH: Cognitive Reasoning Benchmark for Vision Language Models (2025.findings-naacl)

Copied to clipboard

Challenge: Recent advances in large language models have demonstrated their strong performance on IQ test questions, achieving high scores across many languages.
Approach: They propose a dataset to evaluate cognitive multimodal reasoning and problem-solving skills of large models.
Outcome: The proposed dataset contains 2,728 multiple-choice questions and 4,642 images spanning 26 categories.
Hospitality-VQA: Decision-Oriented Informativeness Evaluation for Vision–Language Models (2026.eacl-srw)

Copied to clipboard

Challenge: Existing VQA benchmarks focus on factual correctness but rarely capture what information users actually find useful.
Approach: They propose a framework to quantify how much information an image–question pair provides . they conduct experiments with several state-of-the-art VLMs to determine their reliability .
Outcome: The proposed framework quantifies how much information an image–question pair provides in hospitality contexts.
Response Wide Shut? Surprising Observations in Basic Vision Language Model Capabilities (2025.acl-long)

Copied to clipboard

Challenge: Vision-language Models have been shown to be highly capable but lacking basic visual understanding skills.
Approach: They propose to examine the limitations of vision-language models on visual tasks by constructing a series of tests that probe which components of design may be lacking.
Outcome: The proposed tests compare VLMs to other models on visual encoders, intermediate vision-language projection and LLM-decoder outputs.
Are Large Vision Language Models up to the Challenge of Chart Comprehension and Reasoning (2024.findings-emnlp)

Copied to clipboard

Challenge: Recent studies have demonstrated that large vision language models (LVLMs) are not multi-modal and lack multi-tasking capabilities.
Approach: They evaluate the performance of large vision language models (LVLMs) for chart understanding and reasoning tasks and compare them to open-source models.
Outcome: The proposed models demonstrate impressive abilities in generating fluent texts covering high-level data insights, but they also encounter common problems like hallucinations, factual errors, and data bias.
Benchmarking Vision Language Models for Cultural Understanding (2024.emnlp-main)

Copied to clipboard

Challenge: Recent multimodal vision-language models have shown impressive performance in tasks such as image-to-text generation, visual question answering, and image captioning.
Approach: They propose a visual question-answering benchmark to assess VLMs' cultural understanding of various facets of culture from 11 countries across 5 continents.
Outcome: The visual question-answering benchmark aims to assess VLMs' cultural understanding across regions.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations