Challenge: Existing benchmarks for visual question answering lack in visual grounding and complexity, particularly in evaluating spatial reasoning skills.
Approach: They propose to use flowcharts as visual contexts to assess the capabilities of visual question-answering multimodal language models in reasoning.
Outcome: The proposed benchmarks evaluate models' ability to follow visual information without pre-existing knowledge on a suite of open-source and proprietary multimodal language models using various strategies, followed by an analysis of directional bias.

Similar Papers

MultiChartQA: Benchmarking Vision-Language Models on Multi-Chart Problems (2025.naacl-long)

Copied to clipboard

Challenge: Existing chart understanding benchmarks focus on single-chart tasks, neglecting multi-hop reasoning required to extract and integrate information from multiple charts.
Approach: They propose a benchmark that evaluates MLLMs’ capabilities in four key areas: direct question answering, parallel question answering and comparative reasoning.
Outcome: The proposed benchmark evaluates MLLMs’ capabilities in four key areas: direct question answering, parallel question answering and comparative reasoning.
ComicVQA: A Benchmark for Visual Reasoning in Multimodal LLMs (2026.findings-acl)

Copied to clipboard

Challenge: ComicVQA is a visual reasoning benchmark for comics.
Approach: They propose a comics-based benchmark for evaluating MLLMs on visual reasoning.
Outcome: The proposed model achieves 62.6% accuracy on Missing Panel Prediction and 46.4% on Panel Sorting, compared to open-source models.
DashboardQA: Benchmarking Multimodal Agents for Question Answering on Interactive Dashboards (2026.findings-eacl)

Copied to clipboard

Challenge: Existing question-answering benchmarks for data visualizations focus on static charts instead of interactive dashboards.
Approach: They propose a benchmark to assess how vision-language GUI agents comprehend and interact with real-world dashboards.
Outcome: The first benchmark explicitly designed to assess how vision-language GUI agents comprehend and interact with real-world dashboards.
MTabVQA: Evaluating Multi-Tabular Reasoning of Language Models in Visual Space (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing benchmarks address single tables or non-visual data, leaving a critical gap . MTabVQA comprises 3,745 complex question-answer pairs .
Approach: They propose a benchmark specifically designed for multi-tabular visual question answering that measures the ability to parse diverse table images and correlate information across them.
Outcome: The proposed benchmarks show that fine-tuning VLMs with MTabVQA-Instruct significantly improves their reasoning abilities.
Language Guided Visual Question Answering: Elevate Your Multimodal Language Model Using Knowledge-Enriched Prompts (2023.findings-emnlp)

Copied to clipboard

Challenge: Visual question answering (VQA) is a task that requires an understanding of both the image and the question to provide a natural language answer.
Approach: They propose a multimodal framework that leverages language guidance to answer questions more accurately.
Outcome: The proposed framework improves on the multi-choice question-answering task using CLIP and BLIP models.
ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning (2022.findings-acl)

Copied to clipboard

Challenge: Existing datasets that focus on complex reasoning questions do not address such questions as they are template-based and answers come from a fixed-vocabulary.
Approach: They propose a large-scale benchmark that uses visual and logical reasoning to answer questions using a transformer-based model.
Outcome: The proposed models achieve state-of-the-art on the previous datasets and on the current one, but also show that they have several challenges in answering complex reasoning questions.
xGQA: Cross-Lingual Visual Question Answering (2022.findings-acl)

Copied to clipboard

Challenge: a lack of multilingual multimodal datasets has hindered multimodal vision and language modeling efforts.
Approach: They propose a multilingual evaluation benchmark for the visual question answering task . they extend the established English GQA dataset to 7 typologically diverse languages .
Outcome: The proposed methods outperform current state-of-the-art models in zero-shot cross-lingual settings, but the accuracy remains low across languages.
Beyond Single Plots: A Benchmark for Question Answering on Multi-Charts (2026.findings-acl)

Copied to clipboard

Challenge: Existing research on chart understanding has been limited to single chart images.
Approach: They propose a dataset specifically designed for question answering over multi-chart images.
Outcome: The proposed method shows a 27.4% LLM-based accuracy drop on human-authored questions and a 5.39% gain in the human-generated questions.
POLYCHARTQA: Benchmarking Large Vision-Language Models with Multilingual Chart Question Answering (2026.acl-long)

Copied to clipboard

Challenge: Existing chart understanding benchmarks are overwhelmingly English-centric, limiting their accessibility and relevance to global audiences.
Approach: They propose a multilingual chart question answering benchmark that enables efficient multilingual generation via data translation and code reuse.
Outcome: The proposed benchmark systematically evaluates multilingual chart understanding on state-of-the-art LVLMs and shows a significant performance gap between English and other languages.
All-in-one: Understanding and Generation in Multimodal Reasoning with the MAIA Benchmark (2025.findings-emnlp)

Copied to clipboard

Challenge: MAIA evaluates visual language models on video-related tasks using reasoning categories that aim to disentangle language and vision relations.
Approach: a native-italian benchmark is designed for fine-grained investigation of the reasoning abilities of visual language models on videos.
Outcome: The benchmark evaluates visual language models on two aligned tasks and a visual question-answering task.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations