Papers by Anil Batra
CAST: Cross-modal Alignment Similarity Test for Vision Language Models (2025.coling-main)
Copied to clipboard
| Challenge: | Vision Language Models (VLMs) are typically evaluated with Visual Question Answering tasks which assess a model’s understanding of scenes. |
| Approach: | They propose to use visual question answering (VQA) to assess a model's understanding of scenes to probe for self-consistency across modalities. |
| Outcome: | The proposed test does not focus on objective accuracy but rather on whether VLMs are internally consistent in their outputs. |
Predicting Implicit Arguments in Procedural Video Instructions (2025.acl-long)
Copied to clipboard
| Challenge: | Prior SRL benchmarks often miss implicit arguments, leading to incomplete understanding. |
| Approach: | They propose a dataset that necessitates inferring implicit and explicit arguments from contextual information in multimodal cooking procedures. |
| Outcome: | The proposed dataset achieves a 17% relative improvement in F1-score for what-implicit and a 14.7% improvement for where/with-implicative semantic roles over GPT-4o. |