Stacking with Auxiliary Features for Visual Question Answering (N18-1)

Copied to clipboard

Challenge: Visual Question Answering (VQA) is a challenging task that requires systems to reason about natural language and vision.
Approach: They propose four categories of auxiliary features for ensembling for VQA . three out of the four categories can be inferred from an image-question pair . fourth category uses model-specific explanations .
Outcome: The proposed techniques improve performance for visual question answering (VQA) given an image and a natural language question, the task is to provide an accurate natural language answer.

Similar Papers

Tell-and-Answer: Towards Explainable Visual Question Answering using Attributes and Captions (D18-1)

Copied to clipboard

Challenge: Existing approaches to visual question answering represent images using pre-trained CNNs . but they rarely provide any insight, apart from the answer, into the VQA process .
Approach: They propose to break up the end-to-end VQA into two steps: explaining and reasoning . they first extract attributes and generate descriptions as explanations for an image . a reasoning module utilizes these explanations in place of the image to infer an answer .
Outcome: The proposed system achieves comparable performance with baselines, but with added benefits of explanability and the ability to improve with higher quality explanations.
ProtoVQA: An Adaptable Prototypical Framework for Explainable Fine-Grained Visual Question Answering (2025.emnlp-main)

Copied to clipboard

Challenge: Visual Question Answering (VQA) is increasingly used in diverse applications where models must provide accurate answers and explanations that humans can easily understand and verify.
Approach: They propose a unified prototypical framework that learns question-aware prototypes that serve as reasoning anchors and applies spatially constrained matching to ensure that the selected evidence is coherent and semantically relevant.
Outcome: The proposed framework yields faithful, fine-grained explanations while maintaining competitive accuracy.
Co-VQA : Answering by Interactive Sub Question Sequence (2022.findings-acl)

Copied to clipboard

Challenge: Existing approaches to Visual Question Answering (VQA) answer questions directly, but people usually decompose a complex question into a sequence of simple sub questions.
Approach: They propose a conversation-based VQA framework that decomposes questions into sub questions and answers them one-by-one.
Outcome: The proposed framework achieves state-of-the-art on VQA 2.0 and VQA-CP v2 datasets.
Open-Ended Visual Question Answering by Multi-Modal Domain Adaptation (2020.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches to visual question answering (VQA) are not suitable for real-world applications.
Approach: They propose a supervised multi-modal domain adaptation method for visual question answering in images that exploits supervised domain adaptation.
Outcome: The proposed method outperforms state-of-the-art methods on the benchmark VQA 2.0 and VizWiz datasets.
Cross-Modal Retrieval Augmentation for Multi-Modal Classification (2021.findings-emnlp)

Copied to clipboard

Challenge: Recent advances in using retrieval components over external knowledge sources have shown impressive results for a variety of downstream tasks in natural language processing.
Approach: They propose a retrieval-augmented multi-modal transformer architecture for embedding images and captions in the same space.
Outcome: The proposed approach improves visual question answering over strong baselines and hot-swapping indices.
Towards One-to-Many Visual Question Answering (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing Visual Question Answering systems are constrained to support domain-specific questions . a model trained on a single specific domain may not be competent for real-world application.
Approach: They propose a task to enable a single model to answer as many different domains of questions as possible . they break the task down into the integration of three key abilities .
Outcome: The proposed model can answer as many domains of questions as possible, the authors argue . the proposed model generalizes well to three extra zero-shot datasets, and the results are published.
Modular Visual Question Answering via Code Generation (2023.acl-short)

Copied to clipboard

Challenge: a framework for visual question answering is based on modular code generation . the scope of reasoning needed for visual questions is vast, and requires many skills .
Approach: They propose a framework that formulates visual question answering as modular code generation.
Outcome: The proposed framework improves accuracy on COVR and GQA datasets by 3% and 2% compared to the few-shot baseline that does not employ code generation.
All You May Need for VQA are Image Captions (2022.naacl-main)

Copied to clipboard

Challenge: Visual Question Answering (VQA) has benefited from increasingly sophisticated models, but has not enjoyed the same level of engagement in terms of data creation.
Approach: They propose a method that automatically derives VQA examples at volume by leveraging existing image-caption annotations combined with neural models for textual question generation.
Outcome: The proposed method improves state-of-the-art zero-shot accuracy by double digits and achieves robustness that lacks in the same model trained on human-annotated VQA data.
Intrinsic Subgraph Generation for Interpretable Graph Based Visual Question Answering (2024.lrec-main)

Copied to clipboard

Challenge: Visual Question Answering (VQA) is acknowledged as a challenging multi-modal task for Machine Learning (ML).
Approach: They propose an interpretable approach for graph-based Visual Question Answering . their model is designed to intrinsically produce a subgraph during the question-answering process as its explanation .
Outcome: The proposed model outperforms existing explainable methods on a graph-based VQA dataset.
In Factuality: Efficient Integration of Relevant Facts for Visual Question Answering (2021.acl-short)

Copied to clipboard

Challenge: Current Visual Question Answering (VQA) models are trained on labelled data that may be insufficient to learn complex knowledge representations.
Approach: They propose a method to integrate external knowledge into a visual pre-trained model by integrating facts extracted from a knowledge base.
Outcome: The proposed method outperforms baseline models on the KVQA dataset benchmark by 19% and shows that it is weaker than previous models.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations