Papers by Sreyan Ghosh
Do Vision-Language Models Understand Compound Nouns? (2024.naacl-short)
Copied to clipboard
| Challenge: | Open-vocabulary vision-language models (CLIP) are emerging as a promising new paradigm for text-to-image retrieval. |
| Approach: | They propose a benchmark to evaluate the effectiveness of open-vocabulary vision-language models (CLIP) for text-to-image retrieval using contrastive loss. |
| Outcome: | The proposed framework improves CN understanding of CLIP by 8.25% on Compun. |
ASPIRE: Language-Guided Data Augmentation for Improving Robustness Against Spurious Correlations (2024.findings-acl)
Copied to clipboard
Sreyan Ghosh, Chandra Kiran Evuru, Sonal Kumar, Utkarsh Tyagi, S Sakshi, Sanjoy Chowdhury, Dinesh Manocha
| Challenge: | Neural image classifiers often rely on non-predictive features that are spuriously correlated with the class labels in training data. |
| Approach: | They propose a language-guided data augmented with images without spurious correlations that can be used to augment training datasets for robust learning. |
| Outcome: | The proposed model improves the worst-group classification accuracy of prior methods by 1% - 38%. |
EGOILLUSION: Benchmarking Hallucinations in Egocentric Video Understanding (2025.emnlp-main)
Copied to clipboard
Ashish Seth, Utkarsh Tyagi, Ramaneswaran Selvakumar, Nishit Anand, Sonal Kumar, Sreyan Ghosh, Ramani Duraiswami, Chirag Agarwal, Dinesh Manocha
| Challenge: | Multimodal Large Language Models excel at visual perception and reasoning in third-person and egocentric videos, but are prone to hallucinations, generating coherent yet inaccurate responses. |
| Approach: | They propose to use a benchmark to evaluate MLLM hallucinations in egocentric videos. |
| Outcome: | EGOILLUSION comprises 1,400 videos paired with 8,000 human-annotated open and closed-ended questions designed to trigger hallucinations in both visual and auditory cues in egocentric videos. |
CoDa: Constrained Generation based Data Augmentation for Low-Resource NLP (2024.findings-naacl)
Copied to clipboard
| Challenge: | a low-resource dataset is limited in training data, so generating task-specific data is challenging. |
| Approach: | They propose a data augmentation technique that prompts off-the-shelf instruction-following Large Language Models to generate augmentations. |
| Outcome: | The proposed technique outperforms baselines on 11 datasets spanning 3 tasks and 3 low-resource settings. |
ABEX: Data Augmentation for Low-Resource NLU via Expanding Abstract Descriptions (2024.acl-long)
Copied to clipboard
Sreyan Ghosh, Utkarsh Tyagi, Sonal Kumar, Chandra Kiran Evuru, Ramaneswaran S, S Sakshi, Dinesh Manocha
| Challenge: | ABEX is a novel and effective generative data augmentation methodology for low-resource Natural Language Understanding (NLU) tasks. |
| Approach: | They propose a novel generative data augmentation methodology for low-resource Natural Language Understanding (NLU) tasks based on a paradigm for generating diverse forms of an input document . |
| Outcome: | The proposed method outperforms all baselines qualitatively with improvements of 0.04% - 38.8%. |
GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities (2024.emnlp-main)
Copied to clipboard
Sreyan Ghosh, Sonal Kumar, Ashish Seth, Chandra Kiran Evuru, Utkarsh Tyagi, S Sakshi, Oriol Nieto, Ramani Duraiswami, Dinesh Manocha
| Challenge: | We propose a novel large-scale audio-language model with advanced audio understanding and reasoning abilities. |
| Approach: | They propose a general-purpose large audio-language model with advanced audio understanding and reasoning abilities that integrates an LLM with multiple types of audio representations. |
| Outcome: | The proposed model outperforms existing models on audio understanding tasks by 1%-84%. |
PAT: Parameter-Free Audio-Text Aligner to Boost Zero-Shot Audio Classification (2025.naacl-long)
Copied to clipboard
| Challenge: | Audio-Language Models (ALMs) have demonstrated remarkable performance in zero-shot audio classification. |
| Approach: | They propose a training-free method that enhances audio and language representations using mutual feedback. |
| Outcome: | The proposed method outperforms vanilla zero-shot evaluation with significant margins of 0.42%-27.0%. |
Speech-Hands: A Self-Reflection Voice Agentic Approach to Speech Recognition and Audio Reasoning with Omni Perception (2026.acl-long)
Copied to clipboard
Zhen Wan, Chao-Han Huck Yang, Jinchuan Tian, Hanrong Ye, Ankita Pasad, Szu-Wei Fu, Arushi Goel, Ryo Hachiuma, Shizhe Diao, Kunal Dhawan, Sreyan Ghosh, Yusuke Hirota, Zhehuai Chen, Rafael Valle, Chenhui Chu, Shinji Watanabe, Boris Ginsburg, Yu-Chiang Frank Wang
| Challenge: | naively fine-tuning an omni-model on speech recognition and external sound understanding tasks often degrades performance . Xie and Wu's framework, Speech-Hands, recasts the problem as an explicit self-reflection decision. |
| Approach: | They propose a voice-agentic framework that learns one critical omni-understanding skill: trusting itself versus external audio perception. |
| Outcome: | The proposed framework outperforms baseline models on the OpenASR leaderboard by 12.1% WER and high F1 on audio QA decisions. |
CoSyn: Detecting Implicit Hate Speech in Online Conversations Using a Context Synergized Hyperbolic Network (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing work on detecting explicit hate speech has focused on indirect or coded language. |
| Approach: | They propose a context synergized neural network that integrates user- and conversational-contexts for detecting implicit hate speech in online conversations. |
| Outcome: | The proposed framework outperforms baselines on 6 hate speech datasets and shows that it is highly efficient. |
MULTIVOX: A Benchmark for Evaluating Voice Assistants for Multimodal Interactions (2025.emnlp-main)
Copied to clipboard
Ramaneswaran Selvakumar, Ashish Seth, Nishit Anand, Utkarsh Tyagi, Sonal Kumar, Sreyan Ghosh, Dinesh Manocha
| Challenge: | omni models lack spoken dialogues, which is essential for assessing conversational and auditory capabilities of voice assistants. |
| Approach: | They propose a benchmark to evaluate the ability of voice assistants to integrate paralinguistic speech features into their models. |
| Outcome: | The multivox voice assistant benchmark evaluates the ability of models to integrate spoken and visual cues including paralinguistic speech features for truly multimodal understanding. |
EH-MAM: Easy-to-Hard Masked Acoustic Modeling for Self-Supervised Speech Representation Learning (2024.emnlp-main)
Copied to clipboard
| Challenge: | EH-MAM is a self-supervised learning approach for speech representation learning . prior methods used random masking schemes to learn speech representations . |
| Approach: | They propose a self-supervised approach that automatically selects hard regions during SSL training and introduces them to the model for reconstruction. |
| Outcome: | The proposed approach outperforms state-of-the-art models across low-resource speech recognition and SUPERB benchmarks by 5%-10%. |
Failing Forward: Improving Generative Error Correction for ASR with Synthetic Data and Retrieval Augmentation (2025.findings-acl)
Copied to clipboard
Sreyan Ghosh, Mohammad Sadegh Rasooli, Michael Levit, Peidong Wang, Jian Xue, Dinesh Manocha, Jinyu Li
| Challenge: | Generative Error Correction (GEC) is a powerful post-processing method to boost the performance of Automatic Speech Recognition systems. |
| Approach: | They propose a method to augment GEC models with retrieved entities to improve accuracy in out-of-domain and out-od scenarios. |
| Outcome: | The proposed method outperforms baseline models on multiple datasets and settings. |
ProSE: Diffusion Priors for Speech Enhancement (2025.naacl-long)
Copied to clipboard
Sonal Kumar, Sreyan Ghosh, Utkarsh Tyagi, Anton Jeran Ratnarajah, Chandra Kiran Reddy Evuru, Ramani Duraiswami, Dinesh Manocha
| Challenge: | deterministic deep learning models have been used for speech enhancement, but generative models have shown promise. |
| Approach: | They propose a method to apply diffusion probabilistic models to speech enhancement using priors in a latent space. |
| Outcome: | The proposed method achieves state-of-the-art performance on synthetic and real-world datasets while consuming less computational costs. |
FIGMA: Towards FIne-Grained Music retrievAl (2026.acl-long)
Copied to clipboard
| Challenge: | Existing music retrieval models fail to retrieve fine-grained musical attributes when using coarse semantic queries. |
| Approach: | They propose a multi-view contrastive architecture that captures high-level semantic context and fine-grained musical attributes within a unified representation space. |
| Outcome: | The proposed method outperforms existing CLAP-based music retrieval models on multiple benchmarks. |
ACLM: A Selective-Denoising based Generative Data Augmentation Approach for Low-Resource Complex NER (2023.acl-long)
Copied to clipboard
| Challenge: | Named Entity Recognition (NER) is a task of detecting linguistically complex named entities in low-context text. |
| Approach: | They propose a keyword-based augmentation approach to address the context-entity mismatch issue in complex name recognition (NER) they use selective masking to retain the named entities and certain keywords in the input sentence that provide contextually relevant additional knowledge or hints about the named entity. |
| Outcome: | The proposed approach outperforms baseline methods on monolingual, cross-lingual, and multilingual complex NER in various low-resource settings. |
Do Audio-Language Models Understand Linguistic Variations? (2025.naacl-short)
Copied to clipboard
Ramaneswaran Selvakumar, Sonal Kumar, Hemant Kumar Giri, Nishit Anand, Ashish Seth, Sreyan Ghosh, Dinesh Manocha
| Challenge: | Existing open-vocabulary audio language models struggle to generalize to linguistic variations in textual queries. |
| Approach: | They propose a novel technique to learn audio-language representations agnostic to linguistic variations by reformulating contrastive loss used in CLAP architectures. |
| Outcome: | The proposed approach improves the performance of the open-vocabulary audio language models by 0.8%-13% across benchmarks and enhances robustness to linguistic variation. |
PolyAudio: Advancing Multi-Audio Reasoning in Large Audio Language Models with Interleaved Multi-Audio Contexts (2026.findings-acl)
Copied to clipboard
Sonal Kumar, Sreyan Ghosh, Yueqian Lin, S Sakshi, Ashish Seth, Yiran Chen, Ramani Duraiswami, Dinesh Manocha
| Challenge: | Large Audio Language Models have shown impressive performance on single-clip tasks . however, their ability to reason over interleaved multi-audio contexts remains limited . |
| Approach: | They propose a LALM that targets multi-audio understanding via instruction tuning rather than massive-scale pre-training. |
| Outcome: | The proposed model outperforms baseline models on multi-audio tasks while maintaining robustness. |
DALE: Generative Data Augmentation for Low-Resource Legal NLP (2023.emnlp-main)
Copied to clipboard
Sreyan Ghosh, Chandra Kiran Reddy Evuru, Sonal Kumar, S Ramaneswaran, S Sakshi, Utkarsh Tyagi, Dinesh Manocha
| Challenge: | DALE addresses the challenges existing frameworks pose in generating effective data augmentations of legal documents. |
| Approach: | They propose a generative Data Augmentation framework for low-resource legal NLP that exploits domain-specific language characteristics of templated legal documents to mask collocated spans of text. |
| Outcome: | The proposed framework outperforms baseline frameworks on 13 datasets and 4 low-resource settings. |