Papers by Ramani Duraiswami
EGOILLUSION: Benchmarking Hallucinations in Egocentric Video Understanding (2025.emnlp-main)
Copied to clipboard
Ashish Seth, Utkarsh Tyagi, Ramaneswaran Selvakumar, Nishit Anand, Sonal Kumar, Sreyan Ghosh, Ramani Duraiswami, Chirag Agarwal, Dinesh Manocha
| Challenge: | Multimodal Large Language Models excel at visual perception and reasoning in third-person and egocentric videos, but are prone to hallucinations, generating coherent yet inaccurate responses. |
| Approach: | They propose to use a benchmark to evaluate MLLM hallucinations in egocentric videos. |
| Outcome: | EGOILLUSION comprises 1,400 videos paired with 8,000 human-annotated open and closed-ended questions designed to trigger hallucinations in both visual and auditory cues in egocentric videos. |
GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities (2024.emnlp-main)
Copied to clipboard
Sreyan Ghosh, Sonal Kumar, Ashish Seth, Chandra Kiran Evuru, Utkarsh Tyagi, S Sakshi, Oriol Nieto, Ramani Duraiswami, Dinesh Manocha
| Challenge: | We propose a novel large-scale audio-language model with advanced audio understanding and reasoning abilities. |
| Approach: | They propose a general-purpose large audio-language model with advanced audio understanding and reasoning abilities that integrates an LLM with multiple types of audio representations. |
| Outcome: | The proposed model outperforms existing models on audio understanding tasks by 1%-84%. |
ProSE: Diffusion Priors for Speech Enhancement (2025.naacl-long)
Copied to clipboard
Sonal Kumar, Sreyan Ghosh, Utkarsh Tyagi, Anton Jeran Ratnarajah, Chandra Kiran Reddy Evuru, Ramani Duraiswami, Dinesh Manocha
| Challenge: | deterministic deep learning models have been used for speech enhancement, but generative models have shown promise. |
| Approach: | They propose a method to apply diffusion probabilistic models to speech enhancement using priors in a latent space. |
| Outcome: | The proposed method achieves state-of-the-art performance on synthetic and real-world datasets while consuming less computational costs. |
FIGMA: Towards FIne-Grained Music retrievAl (2026.acl-long)
Copied to clipboard
| Challenge: | Existing music retrieval models fail to retrieve fine-grained musical attributes when using coarse semantic queries. |
| Approach: | They propose a multi-view contrastive architecture that captures high-level semantic context and fine-grained musical attributes within a unified representation space. |
| Outcome: | The proposed method outperforms existing CLAP-based music retrieval models on multiple benchmarks. |
PolyAudio: Advancing Multi-Audio Reasoning in Large Audio Language Models with Interleaved Multi-Audio Contexts (2026.findings-acl)
Copied to clipboard
Sonal Kumar, Sreyan Ghosh, Yueqian Lin, S Sakshi, Ashish Seth, Yiran Chen, Ramani Duraiswami, Dinesh Manocha
| Challenge: | Large Audio Language Models have shown impressive performance on single-clip tasks . however, their ability to reason over interleaved multi-audio contexts remains limited . |
| Approach: | They propose a LALM that targets multi-audio understanding via instruction tuning rather than massive-scale pre-training. |
| Outcome: | The proposed model outperforms baseline models on multi-audio tasks while maintaining robustness. |