Papers by Nishit Anand

4 papers
EGOILLUSION: Benchmarking Hallucinations in Egocentric Video Understanding (2025.emnlp-main)

Copied to clipboard

Challenge: Multimodal Large Language Models excel at visual perception and reasoning in third-person and egocentric videos, but are prone to hallucinations, generating coherent yet inaccurate responses.
Approach: They propose to use a benchmark to evaluate MLLM hallucinations in egocentric videos.
Outcome: EGOILLUSION comprises 1,400 videos paired with 8,000 human-annotated open and closed-ended questions designed to trigger hallucinations in both visual and auditory cues in egocentric videos.
MULTIVOX: A Benchmark for Evaluating Voice Assistants for Multimodal Interactions (2025.emnlp-main)

Copied to clipboard

Challenge: omni models lack spoken dialogues, which is essential for assessing conversational and auditory capabilities of voice assistants.
Approach: They propose a benchmark to evaluate the ability of voice assistants to integrate paralinguistic speech features into their models.
Outcome: The multivox voice assistant benchmark evaluates the ability of models to integrate spoken and visual cues including paralinguistic speech features for truly multimodal understanding.
FIGMA: Towards FIne-Grained Music retrievAl (2026.acl-long)

Copied to clipboard

Challenge: Existing music retrieval models fail to retrieve fine-grained musical attributes when using coarse semantic queries.
Approach: They propose a multi-view contrastive architecture that captures high-level semantic context and fine-grained musical attributes within a unified representation space.
Outcome: The proposed method outperforms existing CLAP-based music retrieval models on multiple benchmarks.
Do Audio-Language Models Understand Linguistic Variations? (2025.naacl-short)

Copied to clipboard

Challenge: Existing open-vocabulary audio language models struggle to generalize to linguistic variations in textual queries.
Approach: They propose a novel technique to learn audio-language representations agnostic to linguistic variations by reformulating contrastive loss used in CLAP architectures.
Outcome: The proposed approach improves the performance of the open-vocabulary audio language models by 0.8%-13% across benchmarks and enhances robustness to linguistic variation.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations