Challenge: Existing knowledge-based visual question answering systems struggle with these tasks due to limited integration of external knowledge.
Approach: They propose a framework that enables large language models to answer visual questions requiring encyclopedic knowledge.
Outcome: The proposed framework improves retrieval outcomes and accuracy of knowledge-based visual question answering tasks.

Similar Papers

WikiSeeker: Rethinking the Role of Vision-Language Models in Knowledge-Based Visual Question Answering (2026.findings-acl)

Copied to clipboard

Challenge: Existing methods for Knowledge-Based Visual Question Answering rely on images as the retrieval key, and often overlook or misplace the role of Vision-Language Models (VLMs)
Approach: They propose a multi-modal RAG framework that assigns VLMs two specialized agents: a Refiner and an Inspector.
Outcome: Experiments on EVQA, InfoSeek, and M2KR show that the proposed framework achieves state-of-the-art performance with significant improvements in both retrieval accuracy and answer quality.
OMGM: Orchestrate Multiple Granularities and Modalities for Efficient Multimodal Retrieval (2025.acl-long)

Copied to clipboard

Challenge: Existing methods for Knowledge-Based Visual Question Answering lack multimodal retrieval . large language models (LLMs) have demonstrated remarkable generalization and reasoning capabilities in text-based systems.
Approach: They propose a multimodal vision-language retrieval-augmented generation system that harmonizes multiple modalities and modality to enhance retrieval.
Outcome: The proposed system achieves state-of-the-art retrieval performance and competitive answers on InfoSeek and Encyclopedic-VQA benchmarks.
Seeing Beyond: Enhancing Visual Question Answering with Multi-Modal Retrieval (2025.coling-industry)

Copied to clipboard

Challenge: Multi-modal Large language models still suffer from model hallucination and lack of specific knowledge when answering challenging questions.
Approach: They propose to use a multi-modal retrieval augmented generation method to integrate knowledge from all modalities into a model to enable alignment between query and knowledge.
Outcome: The proposed method achieves significant performance improvement on the VQA dataset.
Visual Text Matters: Improving Text-KVQA with Visual Text Entity Knowledge-aware Large Multimodal Assistant (2024.emnlp-main)

Copied to clipboard

Challenge: Existing knowledge-aware text-based visual question answering methods are based on textual entities in images.
Approach: They propose a visual text entity linking module that harnesses a state-of-the-art visual text recognition engine and the power of a large multimodal model to perform visual text-entity linking.
Outcome: The proposed approach surpasses the previous best approach by 23.3% on an absolute scale and establishes a new state of the art.
Multi-Level Information Retrieval Augmented Generation for Knowledge-based Visual Question Answering (2024.emnlp-main)

Copied to clipboard

Challenge: Knowledge-Aware Visual Question Answering about Entity tasks require two separate steps to generate accurate answers.
Approach: They propose a multi-level information RAG approach that enhances answer generation through entity retrieval and query expansion.
Outcome: The proposed approach improves answer generation through entity retrieval and query expansion.
Modality-Aware Integration with Large Language Models for Knowledge-Based Visual Question Answering (2024.acl-long)

Copied to clipboard

Challenge: Existing methods to integrate multimodal knowledge in a modality-agnostic manner can be sub-optimal.
Approach: They propose a modality-aware integration with large language models (LLMs) that leverages multimodal knowledge for both image understanding and knowledge reasoning.
Outcome: The proposed model is able to bridge a tight inter-modal exchange while preserving insightful intra-modal learning.
MAGIC-VQA: Multimodal And Grounded Inference with Commonsense Knowledge for Visual Question Answering (2025.findings-acl)

Copied to clipboard

Challenge: Existing Large Vision-Language Models (LVLMs) lack integrated commonsense knowledge . lack of integrated common knowledge limits their robustness and accuracy in VQA .
Approach: They propose a framework to enhance multimodal inference by integrating commonsense reasoning.
Outcome: MAGIC-VQA improves comprehensive benchmark datasets, surpassing existing models in tasks requiring advanced commonsense reasoning.
MegaRAG: Multimodal Knowledge Graph-Based Retrieval Augmented Generation (2026.acl-long)

Copied to clipboard

Challenge: Existing RAG solutions for large language models are limited by context windows limiting their ability to process long-form, domain-specific content.
Approach: They propose a multimodal knowledge graph-based RAG that enables cross-modal reasoning . their method incorporates visual cues into the construction of knowledge graphs, retrieval phase, and answer generation process .
Outcome: Experimental results show that the proposed approach outperforms existing approaches on textual and multimodal benchmarks.
MuKA: Multimodal Knowledge Augmented Visual Information-Seeking (2025.coling-main)

Copied to clipboard

Challenge: Existing methods for visual information-seeking tasks rely on textual knowledge . existing methods can impair information retrieval and confuse MLLMs .
Approach: They propose a framework which leverages a multimodal knowledge base to address these limitations.
Outcome: The proposed framework outperforms state-of-the-art methods on the InfoSeek and E-VQA benchmarks.
Large Language Models Know What is Key Visual Entity: An LLM-assisted Multimodal Retrieval for VQA (2024.emnlp-main)

Copied to clipboard

Challenge: Existing visual language models struggle to capture longtail knowledge in the real world due to redundant visual information.
Approach: They propose a method leveraging the reasoning capability of a large language model to identify key visual entities.
Outcome: The proposed method outperforms other strong visual language model-based systems in two knowledge-intensive VQA benchmarks and performs comparably to models with 1-2 orders larger parameters.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations