Papers by Debjyoti Mondal

2 papers
RG-VQA: Leveraging Retriever-Generator Pipelines for Knowledge Intensive Visual Question Answering (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing methods to improve the reasoning capabilities of VQA systems are limited due to complexity of graph neural networks and end-to-end training.
Approach: They propose a method to integrate Dense Passage Retrievers with Vision Language Models to boost the reasoning capabilities of VQA systems.
Outcome: The proposed method outperforms human accuracy and GPT-4 in the ScienceQA dataset.
From Perception to Reasoning: Enhancing Vision-Language Models for Mobile UI Understanding (2025.findings-acl)

Copied to clipboard

Challenge: Accurately grounding visual and textual elements within mobile user interfaces remains a challenge for Vision-Language Models (VLMs).
Approach: They propose a mobile UI understanding model trained on a dataset specifically tailored for mobile screen understanding and grounding.
Outcome: The proposed model achieves significant gains in accuracy across all perception tasks and on reasoning benchmarks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations