Papers by Ramani Duraiswami

5 papers
EGOILLUSION: Benchmarking Hallucinations in Egocentric Video Understanding (2025.emnlp-main)

Copied to clipboard

Challenge: Multimodal Large Language Models excel at visual perception and reasoning in third-person and egocentric videos, but are prone to hallucinations, generating coherent yet inaccurate responses.
Approach: They propose to use a benchmark to evaluate MLLM hallucinations in egocentric videos.
Outcome: EGOILLUSION comprises 1,400 videos paired with 8,000 human-annotated open and closed-ended questions designed to trigger hallucinations in both visual and auditory cues in egocentric videos.
GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities (2024.emnlp-main)

Copied to clipboard

Challenge: We propose a novel large-scale audio-language model with advanced audio understanding and reasoning abilities.
Approach: They propose a general-purpose large audio-language model with advanced audio understanding and reasoning abilities that integrates an LLM with multiple types of audio representations.
Outcome: The proposed model outperforms existing models on audio understanding tasks by 1%-84%.
ProSE: Diffusion Priors for Speech Enhancement (2025.naacl-long)

Copied to clipboard

Challenge: deterministic deep learning models have been used for speech enhancement, but generative models have shown promise.
Approach: They propose a method to apply diffusion probabilistic models to speech enhancement using priors in a latent space.
Outcome: The proposed method achieves state-of-the-art performance on synthetic and real-world datasets while consuming less computational costs.
FIGMA: Towards FIne-Grained Music retrievAl (2026.acl-long)

Copied to clipboard

Challenge: Existing music retrieval models fail to retrieve fine-grained musical attributes when using coarse semantic queries.
Approach: They propose a multi-view contrastive architecture that captures high-level semantic context and fine-grained musical attributes within a unified representation space.
Outcome: The proposed method outperforms existing CLAP-based music retrieval models on multiple benchmarks.
PolyAudio: Advancing Multi-Audio Reasoning in Large Audio Language Models with Interleaved Multi-Audio Contexts (2026.findings-acl)

Copied to clipboard

Challenge: Large Audio Language Models have shown impressive performance on single-clip tasks . however, their ability to reason over interleaved multi-audio contexts remains limited .
Approach: They propose a LALM that targets multi-audio understanding via instruction tuning rather than massive-scale pre-training.
Outcome: The proposed model outperforms baseline models on multi-audio tasks while maintaining robustness.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations