Papers by Robik Shrestha

2 papers
A negative case analysis of visual grounding methods for VQA (2020.acl-main)

Copied to clipboard

Challenge: Existing Visual Question Answering (VQA) methods exploit dataset biases and spurious statistical correlations instead of producing correct answers for the right reasons.
Approach: They propose to incorporate visual cues to better ground VQA models . they also propose a regularization effect which prevents over-fitting to linguistic priors .
Outcome: The proposed method outperforms existing methods on the Visual Question Answering (VQA) dataset.
BloomVQA: Assessing Hierarchical Multi-modal Comprehension (2024.findings-acl)

Copied to clipboard

Challenge: Recent advances of machine intelligence solutions have demonstrated tremendous success in a wide range of language and multi-modal tasks over diverse domains.
Approach: They propose a VQA dataset to facilitate comprehensive evaluation of large vision-language models on comprehension tasks.
Outcome: The proposed dataset shows improved accuracy over all comprehension levels and a tendency to bypass visual inputs especially for higher-level tasks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations