FlowVQA: Mapping Multimodal Logic in Visual Question Answering with Flowcharts (2024.findings-acl)
Copied to clipboard
Shubhankar Singh, Purvi Chaurasia, Yerram Varun, Pranshu Pandya, Vatsal Gupta, Vivek Gupta, Dan Roth
| Challenge: | Existing benchmarks for visual question answering lack in visual grounding and complexity, particularly in evaluating spatial reasoning skills. |
| Approach: | They propose to use flowcharts as visual contexts to assess the capabilities of visual question-answering multimodal language models in reasoning. |
| Outcome: | The proposed benchmarks evaluate models' ability to follow visual information without pre-existing knowledge on a suite of open-source and proprietary multimodal language models using various strategies, followed by an analysis of directional bias. |
Similar Papers
MultiChartQA: Benchmarking Vision-Language Models on Multi-Chart Problems (2025.naacl-long)
Copied to clipboard
| Challenge: | Existing chart understanding benchmarks focus on single-chart tasks, neglecting multi-hop reasoning required to extract and integrate information from multiple charts. |
| Approach: | They propose a benchmark that evaluates MLLMs’ capabilities in four key areas: direct question answering, parallel question answering and comparative reasoning. |
| Outcome: | The proposed benchmark evaluates MLLMs’ capabilities in four key areas: direct question answering, parallel question answering and comparative reasoning. |
ComicVQA: A Benchmark for Visual Reasoning in Multimodal LLMs (2026.findings-acl)
Copied to clipboard
| Challenge: | ComicVQA is a visual reasoning benchmark for comics. |
| Approach: | They propose a comics-based benchmark for evaluating MLLMs on visual reasoning. |
| Outcome: | The proposed model achieves 62.6% accuracy on Missing Panel Prediction and 46.4% on Panel Sorting, compared to open-source models. |
DashboardQA: Benchmarking Multimodal Agents for Question Answering on Interactive Dashboards (2026.findings-eacl)
Copied to clipboard
Aaryaman Kartha, Ahmed Masry, Mohammed Saidul Islam, Thinh Lang, Shadikur Rahman, Ridwan Mahbub, Mizanur Rahman, Mahir Ahmed, Md Rizwan Parvez, Enamul Hoque, Shafiq Joty
| Challenge: | Existing question-answering benchmarks for data visualizations focus on static charts instead of interactive dashboards. |
| Approach: | They propose a benchmark to assess how vision-language GUI agents comprehend and interact with real-world dashboards. |
| Outcome: | The first benchmark explicitly designed to assess how vision-language GUI agents comprehend and interact with real-world dashboards. |
MTabVQA: Evaluating Multi-Tabular Reasoning of Language Models in Visual Space (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Existing benchmarks address single tables or non-visual data, leaving a critical gap . MTabVQA comprises 3,745 complex question-answer pairs . |
| Approach: | They propose a benchmark specifically designed for multi-tabular visual question answering that measures the ability to parse diverse table images and correlate information across them. |
| Outcome: | The proposed benchmarks show that fine-tuning VLMs with MTabVQA-Instruct significantly improves their reasoning abilities. |
Language Guided Visual Question Answering: Elevate Your Multimodal Language Model Using Knowledge-Enriched Prompts (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Visual question answering (VQA) is a task that requires an understanding of both the image and the question to provide a natural language answer. |
| Approach: | They propose a multimodal framework that leverages language guidance to answer questions more accurately. |
| Outcome: | The proposed framework improves on the multi-choice question-answering task using CLIP and BLIP models. |
ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning (2022.findings-acl)
Copied to clipboard
| Challenge: | Existing datasets that focus on complex reasoning questions do not address such questions as they are template-based and answers come from a fixed-vocabulary. |
| Approach: | They propose a large-scale benchmark that uses visual and logical reasoning to answer questions using a transformer-based model. |
| Outcome: | The proposed models achieve state-of-the-art on the previous datasets and on the current one, but also show that they have several challenges in answering complex reasoning questions. |
xGQA: Cross-Lingual Visual Question Answering (2022.findings-acl)
Copied to clipboard
Jonas Pfeiffer, Gregor Geigle, Aishwarya Kamath, Jan-Martin Steitz, Stefan Roth, Ivan Vulić, Iryna Gurevych
| Challenge: | a lack of multilingual multimodal datasets has hindered multimodal vision and language modeling efforts. |
| Approach: | They propose a multilingual evaluation benchmark for the visual question answering task . they extend the established English GQA dataset to 7 typologically diverse languages . |
| Outcome: | The proposed methods outperform current state-of-the-art models in zero-shot cross-lingual settings, but the accuracy remains low across languages. |
Beyond Single Plots: A Benchmark for Question Answering on Multi-Charts (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing research on chart understanding has been limited to single chart images. |
| Approach: | They propose a dataset specifically designed for question answering over multi-chart images. |
| Outcome: | The proposed method shows a 27.4% LLM-based accuracy drop on human-authored questions and a 5.39% gain in the human-generated questions. |
POLYCHARTQA: Benchmarking Large Vision-Language Models with Multilingual Chart Question Answering (2026.acl-long)
Copied to clipboard
| Challenge: | Existing chart understanding benchmarks are overwhelmingly English-centric, limiting their accessibility and relevance to global audiences. |
| Approach: | They propose a multilingual chart question answering benchmark that enables efficient multilingual generation via data translation and code reuse. |
| Outcome: | The proposed benchmark systematically evaluates multilingual chart understanding on state-of-the-art LVLMs and shows a significant performance gap between English and other languages. |
All-in-one: Understanding and Generation in Multimodal Reasoning with the MAIA Benchmark (2025.findings-emnlp)
Copied to clipboard
Davide Testa, Giovanni Bonetta, Raffaella Bernardi, Alessandro Bondielli, Alessandro Lenci, Alessio Miaschi, Lucia Passaro, Bernardo Magnini
| Challenge: | MAIA evaluates visual language models on video-related tasks using reasoning categories that aim to disentangle language and vision relations. |
| Approach: | a native-italian benchmark is designed for fine-grained investigation of the reasoning abilities of visual language models on videos. |
| Outcome: | The benchmark evaluates visual language models on two aligned tasks and a visual question-answering task. |