| Challenge: | Existing models of geometric reasoning are based on visual representations of objects and objects, but they are not based in symbols or words. |
| Approach: | They propose a new deep network architecture that specializes in answering questions that admit latent visual representations and learns to generate and reason over such representations. |
| Outcome: | The proposed model can generate and reason over latent visual representations and is validated by two synthetic benchmarks. |
Similar Papers
Multimodal Neural Graph Memory Networks for Visual Question Answering (2020.acl-main)
Copied to clipboard
| Challenge: | Visual question answering (VQA) is a new challenge for AI. |
| Approach: | They propose a graph neural network architecture based on the recently proposed Graph Network (GN) . they generate visual features and encoded captions for an image to generate two GNs . |
| Outcome: | The proposed model rivals the state-of-the-art models on Visual7W, VQA-v2.0, and CLEVR datasets. |
Answering Cross-Dimensional Geometric Visual Questions by Multi-constraint Spatial Reasoning (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing methods for solving complex visual questions are limited in their ability to represent in a cross-dimensional space. |
| Approach: | They propose a method that can answer complex visual questions using cross-dimensional reasoning. |
| Outcome: | The proposed method can answer complex visual questions in 2D to 3D space with great application value. |
ProtoVQA: An Adaptable Prototypical Framework for Explainable Fine-Grained Visual Question Answering (2025.emnlp-main)
Copied to clipboard
| Challenge: | Visual Question Answering (VQA) is increasingly used in diverse applications where models must provide accurate answers and explanations that humans can easily understand and verify. |
| Approach: | They propose a unified prototypical framework that learns question-aware prototypes that serve as reasoning anchors and applies spatially constrained matching to ensure that the selected evidence is coherent and semantically relevant. |
| Outcome: | The proposed framework yields faithful, fine-grained explanations while maintaining competitive accuracy. |
CommVQA: Situating Visual Question Answering in Communicative Contexts (2024.emnlp-main)
Copied to clipboard
| Challenge: | Current visual question answering models are trained on image-question pairs in isolation, but the questions people ask are dependent on their informational needs and prior knowledge about the image content. |
| Approach: | They propose a visual question-answer-as-question dataset that contains 1000 images and 8,949 question-announcer pairs to evaluate how situating images within naturalistic contexts shapes visual questions. |
| Outcome: | The proposed dataset contains 1000 images and 8,949 question-answer pairs. |
Answering Complex Geographic Questions by Adaptive Reasoning with Visual Context and External Commonsense Knowledge (2025.acl-long)
Copied to clipboard
| Challenge: | a new task of answering geographic reasoning questions based on the given image is proposed . the task requires identifying the objects in the image and understanding the background context . |
| Approach: | They propose a task of answering geographic reasoning questions based on the given image . they analyze the image and describe its fine-grained content by text and keywords . |
| Outcome: | The proposed method can be used to answer geographic reasoning questions based on an image . it can be applied to a large-scale dataset with 41,329 samples . |
Can We Learn Question, Answer, and Distractors All from an Image? A New Task for Multiple-choice Visual Question Answering (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing studies focus on generating QADs from image and question, but a novel task is needed to generate meaningful questions, correct answers, and challenging distractors. |
| Approach: | They propose a task to generate QADs from images and encode images together . they use contrastive learning to ensure consistency of QAD generated and tested . |
| Outcome: | Empirical evaluations on the benchmark dataset validate the performance of the proposed task. |
ConceptBert: Concept-Aware Representation for Visual Question Answering (2020.findings-emnlp)
Copied to clipboard
| Challenge: | Visual Question Answering (VQA) is a challenging task that has received increasing attention from both the computer vision and the natural language processing communities. |
| Approach: | They propose an algorithm which learns a joint Concept-Vision-Language embedding for questions which require common sense knowledge from external structured content. |
| Outcome: | The proposed model is based on the Outer Knowledge-VQA and VQA datasets. |
Question Modifiers in Visual Question Answering (2022.lrec-1)
Copied to clipboard
| Challenge: | Visual Question Answering (VQA) is a multi-disciplinary task that requires integration of several key disciplines. |
| Approach: | They develop a model that adds modifiers to questions based on object properties and spatial relationships using Amazon Mechanical Turk data. |
| Outcome: | The proposed model can improve when questions are modified to include more details. |
Beyond Embeddings: The Promise of Visual Table in Visual Reasoning (2024.emnlp-main)
Copied to clipboard
| Challenge: | Visual representation learning has been a cornerstone in computer vision for decades. |
| Approach: | They propose a visual representation tailored for visual reasoning that provides instance-level world knowledge and detailed attributes that are essential for visual reason. |
| Outcome: | The proposed visual tables outperform existing models on 11 visual reasoning benchmarks. |
VISREAS: Complex Visual Reasoning with Unanswerable Questions (2024.findings-acl)
Copied to clipboard
| Challenge: | Logic2Vision is a visual question-answering dataset that validates question authenticity with the corresponding image and then reasoning over it. |
| Approach: | They propose a compositional visual question-answering dataset, VisReas, that consists of answerable and unanswerable visual queries . they use visual genome scene graphs to generate the query and the reasoning steps to generate it. |
| Outcome: | The proposed model outperforms generative models and the existing classification models and outperformed existing models. |