Papers by Anand Mishra
Visual Text Matters: Improving Text-KVQA with Visual Text Entity Knowledge-aware Large Multimodal Assistant (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing knowledge-aware text-based visual question answering methods are based on textual entities in images. |
| Approach: | They propose a visual text entity linking module that harnesses a state-of-the-art visual text recognition engine and the power of a large multimodal model to perform visual text-entity linking. |
| Outcome: | The proposed approach surpasses the previous best approach by 23.3% on an absolute scale and establishes a new state of the art. |
When Big Models Train Small Ones: Label-Free Model Parity Alignment for Efficient Visual Question Answering using Small VLMs (2025.emnlp-main)
Copied to clipboard
| Challenge: | Large vision and language models have demonstrated remarkable performance in visual question answering tasks. |
| Approach: | They introduce a framework to optimize L-VLMs by leveraging unlabeled images . they conduct extensive experiments on four diverse VQA benchmarks . |
| Outcome: | The proposed framework improves L-VLMs on four visual question answering benchmarks. |
COFAR: Commonsense and Factual Reasoning in Image Search (2022.aacl-main)
Copied to clipboard
Prajwal Gatti, Abhirama Subramanyam Penamakuri, Revant Teotia, Anand Mishra, Shubhashis Sengupta, Roshni Ramnani
| Challenge: | Existing approaches to retrieve relevant images for natural language searches are limited by visual recognition and lack of commonsense reasoning. |
| Approach: | They propose a framework that leverages visual content and natural language queries to enable commonsense reasoning and factual reasoning in the image search. |
| Outcome: | The proposed framework enables commonsense and factual reasoning in image search on a COFAR dataset. |
VisToT: Vision-Augmented Table-to-Text Generation (2022.emnlp-main)
Copied to clipboard
| Challenge: | Existing models for data-to-text generation are wrongly generating estate in the output text. |
| Approach: | They propose a task that incorporates visual cues from tables and associated images to generate relevant text. |
| Outcome: | The proposed task incorporates visual cues from tables and associated images to generate relevant text. |