Challenge: Existing Large Vision-Language Models (LVLMs) lack integrated commonsense knowledge . lack of integrated common knowledge limits their robustness and accuracy in VQA .
Approach: They propose a framework to enhance multimodal inference by integrating commonsense reasoning.
Outcome: MAGIC-VQA improves comprehensive benchmark datasets, surpassing existing models in tasks requiring advanced commonsense reasoning.

Similar Papers

KG-ViP: Bridging Knowledge Grounding and Visual Perception in Multi-modal LLMs for Visual Question Answering (2026.acl-long)

Copied to clipboard

Challenge: Existing approaches to Visual Question Answering lack synergistic potential of scene graphs and scene graph.
Approach: They propose a retrieval-and-fusion pipeline that fuses scene graphs and commonsense graphs to enable multi-modal reasoning.
Outcome: Experiments on FVQA 2.0+ and MVQA benchmarks show that KG-ViP outperforms existing methods.
NLKI: A Lightweight Natural Language Knowledge Integration Framework for Improving Small VLMs in Commonsense VQA Tasks (2025.findings-emnlp)

Copied to clipboard

Challenge: Small vision-language models lag behind their larger generative counterparts due to lack of knowledge.
Approach: They propose a framework that integrates commonsense knowledge into small vision-language models . the framework retrieves natural language facts and prompts an LLM to craft natural language explanations .
Outcome: The proposed framework retrieves natural language facts and prompts an LLM to craft natural language explanations.
Language Guided Visual Question Answering: Elevate Your Multimodal Language Model Using Knowledge-Enriched Prompts (2023.findings-emnlp)

Copied to clipboard

Challenge: Visual question answering (VQA) is a task that requires an understanding of both the image and the question to provide a natural language answer.
Approach: They propose a multimodal framework that leverages language guidance to answer questions more accurately.
Outcome: The proposed framework improves on the multi-choice question-answering task using CLIP and BLIP models.
ConceptBert: Concept-Aware Representation for Visual Question Answering (2020.findings-emnlp)

Copied to clipboard

Challenge: Visual Question Answering (VQA) is a challenging task that has received increasing attention from both the computer vision and the natural language processing communities.
Approach: They propose an algorithm which learns a joint Concept-Vision-Language embedding for questions which require common sense knowledge from external structured content.
Outcome: The proposed model is based on the Outer Knowledge-VQA and VQA datasets.
Modality-Aware Integration with Large Language Models for Knowledge-Based Visual Question Answering (2024.acl-long)

Copied to clipboard

Challenge: Existing methods to integrate multimodal knowledge in a modality-agnostic manner can be sub-optimal.
Approach: They propose a modality-aware integration with large language models (LLMs) that leverages multimodal knowledge for both image understanding and knowledge reasoning.
Outcome: The proposed model is able to bridge a tight inter-modal exchange while preserving insightful intra-modal learning.
A Simple Baseline for Knowledge-Based Visual Question Answering (2023.emnlp-main)

Copied to clipboard

Challenge: Recent studies emphasize the importance of incorporating both explicit and implicit knowledge to answer questions requiring external knowledge.
Approach: They propose a pipeline that incorporates both explicit and implicit knowledge . their method is training-free and does not require access to external databases or APIs .
Outcome: The proposed method achieves state-of-the-art accuracy on OK-VQA and A-OK-VQ datasets.
Progressive Multimodal Search and Reasoning for Knowledge-Intensive Visual Question Answering (2026.acl-long)

Copied to clipboard

Challenge: Existing approaches to knowledge-intensive visual question answering lack mechanisms to revise misdirected reasoning.
Approach: They propose a framework that progressively constructs a structured reasoning trajectory . they use dual-scope queries to retrieve diverse knowledge from heterogeneous knowledge bases .
Outcome: The proposed framework improves retrieval recall and end-to-end answer accuracy.
ProtoVQA: An Adaptable Prototypical Framework for Explainable Fine-Grained Visual Question Answering (2025.emnlp-main)

Copied to clipboard

Challenge: Visual Question Answering (VQA) is increasingly used in diverse applications where models must provide accurate answers and explanations that humans can easily understand and verify.
Approach: They propose a unified prototypical framework that learns question-aware prototypes that serve as reasoning anchors and applies spatially constrained matching to ensure that the selected evidence is coherent and semantically relevant.
Outcome: The proposed framework yields faithful, fine-grained explanations while maintaining competitive accuracy.
In Factuality: Efficient Integration of Relevant Facts for Visual Question Answering (2021.acl-short)

Copied to clipboard

Challenge: Current Visual Question Answering (VQA) models are trained on labelled data that may be insufficient to learn complex knowledge representations.
Approach: They propose a method to integrate external knowledge into a visual pre-trained model by integrating facts extracted from a knowledge base.
Outcome: The proposed method outperforms baseline models on the KVQA dataset benchmark by 19% and shows that it is weaker than previous models.
EchoSight: Advancing Visual-Language Models with Wiki Knowledge (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing knowledge-based visual question answering systems struggle with these tasks due to limited integration of external knowledge.
Approach: They propose a framework that enables large language models to answer visual questions requiring encyclopedic knowledge.
Outcome: The proposed framework improves retrieval outcomes and accuracy of knowledge-based visual question answering tasks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations