Challenge: Visual question answering and image captioning require a shared body of general knowledge connecting language and vision.
Approach: They propose a method that exploits a shared body of general knowledge connecting language and vision by jointly generating captions.
Outcome: The proposed approach obtains state-of-the-art performance on the VQA v2 challenge . it uses human annotated captions to generate question-relevant captions .

Similar Papers

All You May Need for VQA are Image Captions (2022.naacl-main)

Copied to clipboard

Challenge: Visual Question Answering (VQA) has benefited from increasingly sophisticated models, but has not enjoyed the same level of engagement in terms of data creation.
Approach: They propose a method that automatically derives VQA examples at volume by leveraging existing image-caption annotations combined with neural models for textual question generation.
Outcome: The proposed method improves state-of-the-art zero-shot accuracy by double digits and achieves robustness that lacks in the same model trained on human-annotated VQA data.
Improving Visual Question Answering by Referring to Generated Paragraph Captions (P19-1)

Copied to clipboard

Challenge: Empirical results show that paragraph captions help answer more visual questions .
Approach: They propose a visual and textual question answering model which uses paragraph captions as input . they use cross-attention to extract related information, then consensus to fuse the inputs .
Outcome: Empirical results show that paragraph captions help answer more visual questions . the proposed model significantly improves the baseline model .
Can We Learn Question, Answer, and Distractors All from an Image? A New Task for Multiple-choice Visual Question Answering (2024.lrec-main)

Copied to clipboard

Challenge: Existing studies focus on generating QADs from image and question, but a novel task is needed to generate meaningful questions, correct answers, and challenging distractors.
Approach: They propose a task to generate QADs from images and encode images together . they use contrastive learning to ensure consistency of QAD generated and tested .
Outcome: Empirical evaluations on the benchmark dataset validate the performance of the proposed task.
Tell-and-Answer: Towards Explainable Visual Question Answering using Attributes and Captions (D18-1)

Copied to clipboard

Challenge: Existing approaches to visual question answering represent images using pre-trained CNNs . but they rarely provide any insight, apart from the answer, into the VQA process .
Approach: They propose to break up the end-to-end VQA into two steps: explaining and reasoning . they first extract attributes and generate descriptions as explanations for an image . a reasoning module utilizes these explanations in place of the image to infer an answer .
Outcome: The proposed system achieves comparable performance with baselines, but with added benefits of explanability and the ability to improve with higher quality explanations.
WeaQA: Weak Supervision via Captions for Visual Question Answering (2021.findings-acl)

Copied to clipboard

Challenge: Existing methods for training visual question answering models rely on datasets with human-annotated image-quest-answer triplets.
Approach: They propose a method to train models with synthetic Q-A pairs generated procedurally from captions.
Outcome: The proposed method trains models with synthetic Q-A pairs generated from captions on three VQA benchmarks.
Modular Visual Question Answering via Code Generation (2023.acl-short)

Copied to clipboard

Challenge: a framework for visual question answering is based on modular code generation . the scope of reasoning needed for visual questions is vast, and requires many skills .
Approach: They propose a framework that formulates visual question answering as modular code generation.
Outcome: The proposed framework improves accuracy on COVR and GQA datasets by 3% and 2% compared to the few-shot baseline that does not employ code generation.
Augmenting Image Question Answering Dataset by Exploiting Image Captions (L18-1)

Copied to clipboard

Challenge: Image question answering requires large amounts of human-annotated data to achieve optimal performance.
Approach: They propose a framework to augment training data by generating additional examples from unannotated pairs of an image and captions.
Outcome: The proposed framework augments training data by generating additional examples from unannotated pairs of an image and captions.
Open-Ended Visual Question Answering by Multi-Modal Domain Adaptation (2020.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches to visual question answering (VQA) are not suitable for real-world applications.
Approach: They propose a supervised multi-modal domain adaptation method for visual question answering in images that exploits supervised domain adaptation.
Outcome: The proposed method outperforms state-of-the-art methods on the benchmark VQA 2.0 and VizWiz datasets.
QACE: Asking Questions to Evaluate an Image Caption (2021.findings-emnlp)

Copied to clipboard

Challenge: Existing metric for image captioning evaluation is based on n-gram similarity metrics but these fail to capture semantic errors in captions.
Approach: They propose a new metric based on Question Answering for Caption Evaluation to evaluate image captioning based upon Question Generation and Question Answers systems.
Outcome: The proposed metric is multi-modal, reference-less and explainable.
Diversity and Consistency: Exploring Visual Question-Answer Pair Generation (2021.findings-emnlp)

Copied to clipboard

Challenge: Existing tasks to generate question-answer pairs from visual images are under-explored.
Approach: They propose a task that targets question-answer pair generation from visual images.
Outcome: The proposed model can generate diverse or consistent QAPs on two benchmarks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations