MAPWise: Evaluating Vision-Language Models for Advanced Map Queries (2025.naacl-long)
Copied to clipboard
| Challenge: | Vision-language models excel at tasks requiring joint understanding of visual information and natural language. |
| Approach: | They propose to use choropleth maps to answer questions from three geographical regions in the United States, India, China as question templates. |
| Outcome: | The proposed model outperforms other models in the area of visual language and visual question answering. |
Similar Papers
SpaRE: Enhancing Spatial Reasoning in Vision-Language Models with Synthetic Data (2025.acl-long)
Copied to clipboard
| Challenge: | Vision-language models struggle with spatial reasoning, a skill that humans excel at. |
| Approach: | They propose to use a spatial-reasoning Enhanced (SpaRE) VLM to improve spatial reasoning in visual question answering and robotics. |
| Outcome: | The proposed model achieves a 49% performance gain on the What's Up benchmark while maintaining strong results on general tasks. |
Can VLMs Actually See and Read? A Survey on Modality Collapse in Vision-Language Models (2025.findings-acl)
Copied to clipboard
| Challenge: | Vision-language models integrate textual and visual information, enabling them to process visual inputs and generate predictions. |
| Approach: | They review work on modality collapse analysis to provide insights into the reason for this unintended behavior and review probing studies for fine-grained vision-language understanding. |
| Outcome: | The proposed models can achieve competitive performance in vision-language tasks despite relying heavily on textual information and ignoring visual information. |
Evaluation of Multilingual Ability to Use Spatial Deictic Expressions in Vision-Language Models (2026.acl-srw)
Copied to clipboard
| Challenge: | Existing studies have focused on the ability of vision-language models to utilize spatial deictic expressions, which depend on the situation of utterance. |
| Approach: | They develop a benchmark to evaluate the multilingual ability of VLMs to use spatial deictic expressions in four languages. |
| Outcome: | The proposed models use demonstratives in a different manner from humans, particularly in selecting demonstrative based on distance from the object. |
Unraveling the Truth: Do VLMs really Understand Charts? A Deep Dive into Consistency and Robustness (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Chart question answering (CQA) is a crucial area of Visual Language Understanding. |
| Approach: | They evaluate the robustness and consistency of current Visual Language Models on a dataset encompassing diverse question categories and chart formats. |
| Outcome: | The proposed models handle varying levels of chart and question complexity and are robust across different visual representations of the same underlying data. |
Beyond Visual Understanding Introducing PARROT-360V for Vision Language Model Benchmarking (2025.coling-industry)
Copied to clipboard
| Challenge: | Current benchmarks for evaluating Vision Language Models (VLMs) often fail to thoroughly assess these models’ abilities to understand complex visual and textual content. |
| Approach: | They propose a benchmark that features 2487 visual puzzles designed to test VLMs on complex visual reasoning tasks. |
| Outcome: | The PARROT-360V Benchmark features 2487 visual puzzles designed to test VLMs on complex visual reasoning tasks. |
NTSEBENCH: Cognitive Reasoning Benchmark for Vision Language Models (2025.findings-naacl)
Copied to clipboard
| Challenge: | Recent advances in large language models have demonstrated their strong performance on IQ test questions, achieving high scores across many languages. |
| Approach: | They propose a dataset to evaluate cognitive multimodal reasoning and problem-solving skills of large models. |
| Outcome: | The proposed dataset contains 2,728 multiple-choice questions and 4,642 images spanning 26 categories. |
Hospitality-VQA: Decision-Oriented Informativeness Evaluation for Vision–Language Models (2026.eacl-srw)
Copied to clipboard
Jeongwoo Lee, Baek Duhyeong, Eungyeol Han, Soyeon Shin, Gukin Han, Seungduk Kim, Jaehyun Jeon, Taewoo Jeong
| Challenge: | Existing VQA benchmarks focus on factual correctness but rarely capture what information users actually find useful. |
| Approach: | They propose a framework to quantify how much information an image–question pair provides . they conduct experiments with several state-of-the-art VLMs to determine their reliability . |
| Outcome: | The proposed framework quantifies how much information an image–question pair provides in hospitality contexts. |
Response Wide Shut? Surprising Observations in Basic Vision Language Model Capabilities (2025.acl-long)
Copied to clipboard
| Challenge: | Vision-language Models have been shown to be highly capable but lacking basic visual understanding skills. |
| Approach: | They propose to examine the limitations of vision-language models on visual tasks by constructing a series of tests that probe which components of design may be lacking. |
| Outcome: | The proposed tests compare VLMs to other models on visual encoders, intermediate vision-language projection and LLM-decoder outputs. |
Are Large Vision Language Models up to the Challenge of Chart Comprehension and Reasoning (2024.findings-emnlp)
Copied to clipboard
Mohammed Saidul Islam, Raian Rahman, Ahmed Masry, Md Tahmid Rahman Laskar, Mir Tafseer Nayeem, Enamul Hoque
| Challenge: | Recent studies have demonstrated that large vision language models (LVLMs) are not multi-modal and lack multi-tasking capabilities. |
| Approach: | They evaluate the performance of large vision language models (LVLMs) for chart understanding and reasoning tasks and compare them to open-source models. |
| Outcome: | The proposed models demonstrate impressive abilities in generating fluent texts covering high-level data insights, but they also encounter common problems like hallucinations, factual errors, and data bias. |
Benchmarking Vision Language Models for Cultural Understanding (2024.emnlp-main)
Copied to clipboard
Shravan Nayak, Kanishk Jain, Rabiul Awal, Siva Reddy, Sjoerd Steenkiste, Lisa Hendricks, Karolina Stanczak, Aishwarya Agrawal
| Challenge: | Recent multimodal vision-language models have shown impressive performance in tasks such as image-to-text generation, visual question answering, and image captioning. |
| Approach: | They propose a visual question-answering benchmark to assess VLMs' cultural understanding of various facets of culture from 11 countries across 5 continents. |
| Outcome: | The visual question-answering benchmark aims to assess VLMs' cultural understanding across regions. |