Papers by Vibhav Vineet

6 papers
Exposing the Achilles’ Heel: Evaluating LLMs Ability to Handle Mistakes in Mathematical Reasoning (2025.acl-long)

Copied to clipboard

Challenge: Existing evaluations focus on final accuracy, neglecting the critical aspect of reasoning capabilities.
Approach: They propose to evaluate LLMs’ abilities to detect and correct reasoning mistakes by using rule-based methods and smaller language models.
Outcome: The proposed model outperforms existing models such as GPT-4o and GPT4 in both accuracy and accuracy, but lacks data contamination and memorization concerns.
Navigating Hallucinations for Reasoning of Unintentional Activities (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing models of intentionality recognition struggle to understand the reasoning behind unintentional actions.
Approach: They propose a novel prompting technique which allows the model to navigate through hallucinated thoughts to achieve better reasoning.
Outcome: The proposed prompting technique outperforms standard prompting while minimizing hallucinations.
RiTTA: Modeling Event Relations in Text-to-Audio Generation (2025.emnlp-main)

Copied to clipboard

Challenge: Existing text-to-audio (TTA) generation methods have not explored audio event relation modeling, nor proposed any new framework to enhance this capability.
Approach: They propose a comprehensive relation corpus covering all potential relations in real-world scenarios and a new audio event corpus encompassing commonly heard audios.
Outcome: The proposed framework improves existing models’ relation modeling capability with negligible extra parameters.
Out of Sight, Not Out of Context? Egocentric Spatial Reasoning in VLMs Across Disjoint Frames (2025.emnlp-main)

Copied to clipboard

Challenge: Disjoint-3DQA evaluates the spatial reasoning ability of embodied AI assistants based on egocentric video . it aims to catalyze future research at the intersection of vision, language, and embodie .
Approach: They propose a generative QA benchmark that evaluates the ability of embodied AI assistants to integrate spatial cues across time by asking object pairs that are not co-visible in the same frame.
Outcome: The proposed benchmark compares seven state-of-the-art VLMs and finds that they lag behind human performance by 28%, with steeper declines as the temporal gap widens.
Grounding Task Assistance with Multimodal Cues from a Single Demonstration (2025.findings-acl)

Copied to clipboard

Challenge: RGB video often fails to capture fine-grained contextual cues such as intent, safety-critical environmental factors, and subtle preferences embedded in human behavior.
Approach: They propose a framework that integrates eye gaze and speech cues to improve conversational agents for task assistance by integrating eye gaze with speech cuests.
Outcome: The proposed framework captures fine-grained intent and user-specific cues, enabling richer contextual grounding for visual question answering.
Image Retrieval from Contextual Descriptions (2022.acl-long)

Copied to clipboard

Challenge: a new multimodal challenge challenges vision-and-language models to integrate context into their representations.
Approach: They propose a multimodal challenge to integrate context into vision-and-language models . they benchmark several state-of-the-art models using cross-encoders and bi-encodings .
Outcome: The proposed model lags behind human models on imageCoDe, compared with human models.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations