Challenge: Existing models of geometric reasoning are based on visual representations of objects and objects, but they are not based in symbols or words.
Approach: They propose a new deep network architecture that specializes in answering questions that admit latent visual representations and learns to generate and reason over such representations.
Outcome: The proposed model can generate and reason over latent visual representations and is validated by two synthetic benchmarks.

Similar Papers

Multimodal Neural Graph Memory Networks for Visual Question Answering (2020.acl-main)

Copied to clipboard

Challenge: Visual question answering (VQA) is a new challenge for AI.
Approach: They propose a graph neural network architecture based on the recently proposed Graph Network (GN) . they generate visual features and encoded captions for an image to generate two GNs .
Outcome: The proposed model rivals the state-of-the-art models on Visual7W, VQA-v2.0, and CLEVR datasets.
Answering Cross-Dimensional Geometric Visual Questions by Multi-constraint Spatial Reasoning (2026.findings-acl)

Copied to clipboard

Challenge: Existing methods for solving complex visual questions are limited in their ability to represent in a cross-dimensional space.
Approach: They propose a method that can answer complex visual questions using cross-dimensional reasoning.
Outcome: The proposed method can answer complex visual questions in 2D to 3D space with great application value.
ProtoVQA: An Adaptable Prototypical Framework for Explainable Fine-Grained Visual Question Answering (2025.emnlp-main)

Copied to clipboard

Challenge: Visual Question Answering (VQA) is increasingly used in diverse applications where models must provide accurate answers and explanations that humans can easily understand and verify.
Approach: They propose a unified prototypical framework that learns question-aware prototypes that serve as reasoning anchors and applies spatially constrained matching to ensure that the selected evidence is coherent and semantically relevant.
Outcome: The proposed framework yields faithful, fine-grained explanations while maintaining competitive accuracy.
CommVQA: Situating Visual Question Answering in Communicative Contexts (2024.emnlp-main)

Copied to clipboard

Challenge: Current visual question answering models are trained on image-question pairs in isolation, but the questions people ask are dependent on their informational needs and prior knowledge about the image content.
Approach: They propose a visual question-answer-as-question dataset that contains 1000 images and 8,949 question-announcer pairs to evaluate how situating images within naturalistic contexts shapes visual questions.
Outcome: The proposed dataset contains 1000 images and 8,949 question-answer pairs.
Answering Complex Geographic Questions by Adaptive Reasoning with Visual Context and External Commonsense Knowledge (2025.acl-long)

Copied to clipboard

Challenge: a new task of answering geographic reasoning questions based on the given image is proposed . the task requires identifying the objects in the image and understanding the background context .
Approach: They propose a task of answering geographic reasoning questions based on the given image . they analyze the image and describe its fine-grained content by text and keywords .
Outcome: The proposed method can be used to answer geographic reasoning questions based on an image . it can be applied to a large-scale dataset with 41,329 samples .
Can We Learn Question, Answer, and Distractors All from an Image? A New Task for Multiple-choice Visual Question Answering (2024.lrec-main)

Copied to clipboard

Challenge: Existing studies focus on generating QADs from image and question, but a novel task is needed to generate meaningful questions, correct answers, and challenging distractors.
Approach: They propose a task to generate QADs from images and encode images together . they use contrastive learning to ensure consistency of QAD generated and tested .
Outcome: Empirical evaluations on the benchmark dataset validate the performance of the proposed task.
ConceptBert: Concept-Aware Representation for Visual Question Answering (2020.findings-emnlp)

Copied to clipboard

Challenge: Visual Question Answering (VQA) is a challenging task that has received increasing attention from both the computer vision and the natural language processing communities.
Approach: They propose an algorithm which learns a joint Concept-Vision-Language embedding for questions which require common sense knowledge from external structured content.
Outcome: The proposed model is based on the Outer Knowledge-VQA and VQA datasets.
Question Modifiers in Visual Question Answering (2022.lrec-1)

Copied to clipboard

Challenge: Visual Question Answering (VQA) is a multi-disciplinary task that requires integration of several key disciplines.
Approach: They develop a model that adds modifiers to questions based on object properties and spatial relationships using Amazon Mechanical Turk data.
Outcome: The proposed model can improve when questions are modified to include more details.
Beyond Embeddings: The Promise of Visual Table in Visual Reasoning (2024.emnlp-main)

Copied to clipboard

Challenge: Visual representation learning has been a cornerstone in computer vision for decades.
Approach: They propose a visual representation tailored for visual reasoning that provides instance-level world knowledge and detailed attributes that are essential for visual reason.
Outcome: The proposed visual tables outperform existing models on 11 visual reasoning benchmarks.
VISREAS: Complex Visual Reasoning with Unanswerable Questions (2024.findings-acl)

Copied to clipboard

Challenge: Logic2Vision is a visual question-answering dataset that validates question authenticity with the corresponding image and then reasoning over it.
Approach: They propose a compositional visual question-answering dataset, VisReas, that consists of answerable and unanswerable visual queries . they use visual genome scene graphs to generate the query and the reasoning steps to generate it.
Outcome: The proposed model outperforms generative models and the existing classification models and outperformed existing models.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations