Papers by Sonal Kumar
Do Vision-Language Models Understand Compound Nouns? (2024.naacl-short)
Copied to clipboard
| Challenge: | Open-vocabulary vision-language models (CLIP) are emerging as a promising new paradigm for text-to-image retrieval. |
| Approach: | They propose a benchmark to evaluate the effectiveness of open-vocabulary vision-language models (CLIP) for text-to-image retrieval using contrastive loss. |
| Outcome: | The proposed framework improves CN understanding of CLIP by 8.25% on Compun. |
Conversational Semantic Parsing (2020.emnlp-main)
Copied to clipboard
Armen Aghajanyan, Jean Maillard, Akshat Shrivastava, Keith Diedrick, Michael Haeger, Haoran Li, Yashar Mehdad, Veselin Stoyanov, Anuj Kumar, Mike Lewis, Sonal Gupta
| Challenge: | Structured representations for task-oriented assistant systems are limited due to the limitations of the representation. |
| Approach: | They propose a semantic representation for task-oriented conversational systems that can represent co-reference and context carryover. |
| Outcome: | The proposed model improves the best results on ATIS, SNIPS, TOP and DSTC2 by up to 5 points for slot-carryover. |
ASPIRE: Language-Guided Data Augmentation for Improving Robustness Against Spurious Correlations (2024.findings-acl)
Copied to clipboard
Sreyan Ghosh, Chandra Kiran Evuru, Sonal Kumar, Utkarsh Tyagi, S Sakshi, Sanjoy Chowdhury, Dinesh Manocha
| Challenge: | Neural image classifiers often rely on non-predictive features that are spuriously correlated with the class labels in training data. |
| Approach: | They propose a language-guided data augmented with images without spurious correlations that can be used to augment training datasets for robust learning. |
| Outcome: | The proposed model improves the worst-group classification accuracy of prior methods by 1% - 38%. |
EGOILLUSION: Benchmarking Hallucinations in Egocentric Video Understanding (2025.emnlp-main)
Copied to clipboard
Ashish Seth, Utkarsh Tyagi, Ramaneswaran Selvakumar, Nishit Anand, Sonal Kumar, Sreyan Ghosh, Ramani Duraiswami, Chirag Agarwal, Dinesh Manocha
| Challenge: | Multimodal Large Language Models excel at visual perception and reasoning in third-person and egocentric videos, but are prone to hallucinations, generating coherent yet inaccurate responses. |
| Approach: | They propose to use a benchmark to evaluate MLLM hallucinations in egocentric videos. |
| Outcome: | EGOILLUSION comprises 1,400 videos paired with 8,000 human-annotated open and closed-ended questions designed to trigger hallucinations in both visual and auditory cues in egocentric videos. |
CoDa: Constrained Generation based Data Augmentation for Low-Resource NLP (2024.findings-naacl)
Copied to clipboard
| Challenge: | a low-resource dataset is limited in training data, so generating task-specific data is challenging. |
| Approach: | They propose a data augmentation technique that prompts off-the-shelf instruction-following Large Language Models to generate augmentations. |
| Outcome: | The proposed technique outperforms baselines on 11 datasets spanning 3 tasks and 3 low-resource settings. |
ABEX: Data Augmentation for Low-Resource NLU via Expanding Abstract Descriptions (2024.acl-long)
Copied to clipboard
Sreyan Ghosh, Utkarsh Tyagi, Sonal Kumar, Chandra Kiran Evuru, Ramaneswaran S, S Sakshi, Dinesh Manocha
| Challenge: | ABEX is a novel and effective generative data augmentation methodology for low-resource Natural Language Understanding (NLU) tasks. |
| Approach: | They propose a novel generative data augmentation methodology for low-resource Natural Language Understanding (NLU) tasks based on a paradigm for generating diverse forms of an input document . |
| Outcome: | The proposed method outperforms all baselines qualitatively with improvements of 0.04% - 38.8%. |
GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities (2024.emnlp-main)
Copied to clipboard
Sreyan Ghosh, Sonal Kumar, Ashish Seth, Chandra Kiran Evuru, Utkarsh Tyagi, S Sakshi, Oriol Nieto, Ramani Duraiswami, Dinesh Manocha
| Challenge: | We propose a novel large-scale audio-language model with advanced audio understanding and reasoning abilities. |
| Approach: | They propose a general-purpose large audio-language model with advanced audio understanding and reasoning abilities that integrates an LLM with multiple types of audio representations. |
| Outcome: | The proposed model outperforms existing models on audio understanding tasks by 1%-84%. |
PAT: Parameter-Free Audio-Text Aligner to Boost Zero-Shot Audio Classification (2025.naacl-long)
Copied to clipboard
| Challenge: | Audio-Language Models (ALMs) have demonstrated remarkable performance in zero-shot audio classification. |
| Approach: | They propose a training-free method that enhances audio and language representations using mutual feedback. |
| Outcome: | The proposed method outperforms vanilla zero-shot evaluation with significant margins of 0.42%-27.0%. |
CoSyn: Detecting Implicit Hate Speech in Online Conversations Using a Context Synergized Hyperbolic Network (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing work on detecting explicit hate speech has focused on indirect or coded language. |
| Approach: | They propose a context synergized neural network that integrates user- and conversational-contexts for detecting implicit hate speech in online conversations. |
| Outcome: | The proposed framework outperforms baselines on 6 hate speech datasets and shows that it is highly efficient. |
El Volumen Louder Por Favor: Code-switching in Task-oriented Semantic Parsing (2021.eacl-main)
Copied to clipboard
| Challenge: | Code-switching (CS) is the alternation of languages within an utterance or conversation. |
| Approach: | They propose to use translation-and-align and augment with a generation model followed by match-and filter to improve CS generalizability of cross-lingual models when data for only one language is available. |
| Outcome: | The proposed models improve when only English data is available alongside zero or a few CS training instances. |
Semantic Parsing for Task Oriented Dialog using Hierarchical Representations (D18-1)
Copied to clipboard
| Challenge: | Existing work on task oriented dialog systems has limited expressive power to one intent per query and one slot label per token. |
| Approach: | They propose a hierarchical annotation scheme for semantic parsing that allows representation of compositional queries. |
| Outcome: | The proposed representation outperforms sequence-to-sequence approaches on a 44k annotated query dataset. |
MULTIVOX: A Benchmark for Evaluating Voice Assistants for Multimodal Interactions (2025.emnlp-main)
Copied to clipboard
Ramaneswaran Selvakumar, Ashish Seth, Nishit Anand, Utkarsh Tyagi, Sonal Kumar, Sreyan Ghosh, Dinesh Manocha
| Challenge: | omni models lack spoken dialogues, which is essential for assessing conversational and auditory capabilities of voice assistants. |
| Approach: | They propose a benchmark to evaluate the ability of voice assistants to integrate paralinguistic speech features into their models. |
| Outcome: | The multivox voice assistant benchmark evaluates the ability of models to integrate spoken and visual cues including paralinguistic speech features for truly multimodal understanding. |
EH-MAM: Easy-to-Hard Masked Acoustic Modeling for Self-Supervised Speech Representation Learning (2024.emnlp-main)
Copied to clipboard
| Challenge: | EH-MAM is a self-supervised learning approach for speech representation learning . prior methods used random masking schemes to learn speech representations . |
| Approach: | They propose a self-supervised approach that automatically selects hard regions during SSL training and introduces them to the model for reconstruction. |
| Outcome: | The proposed approach outperforms state-of-the-art models across low-resource speech recognition and SUPERB benchmarks by 5%-10%. |
ProSE: Diffusion Priors for Speech Enhancement (2025.naacl-long)
Copied to clipboard
Sonal Kumar, Sreyan Ghosh, Utkarsh Tyagi, Anton Jeran Ratnarajah, Chandra Kiran Reddy Evuru, Ramani Duraiswami, Dinesh Manocha
| Challenge: | deterministic deep learning models have been used for speech enhancement, but generative models have shown promise. |
| Approach: | They propose a method to apply diffusion probabilistic models to speech enhancement using priors in a latent space. |
| Outcome: | The proposed method achieves state-of-the-art performance on synthetic and real-world datasets while consuming less computational costs. |
ACLM: A Selective-Denoising based Generative Data Augmentation Approach for Low-Resource Complex NER (2023.acl-long)
Copied to clipboard
| Challenge: | Named Entity Recognition (NER) is a task of detecting linguistically complex named entities in low-context text. |
| Approach: | They propose a keyword-based augmentation approach to address the context-entity mismatch issue in complex name recognition (NER) they use selective masking to retain the named entities and certain keywords in the input sentence that provide contextually relevant additional knowledge or hints about the named entity. |
| Outcome: | The proposed approach outperforms baseline methods on monolingual, cross-lingual, and multilingual complex NER in various low-resource settings. |
Do Audio-Language Models Understand Linguistic Variations? (2025.naacl-short)
Copied to clipboard
Ramaneswaran Selvakumar, Sonal Kumar, Hemant Kumar Giri, Nishit Anand, Ashish Seth, Sreyan Ghosh, Dinesh Manocha
| Challenge: | Existing open-vocabulary audio language models struggle to generalize to linguistic variations in textual queries. |
| Approach: | They propose a novel technique to learn audio-language representations agnostic to linguistic variations by reformulating contrastive loss used in CLAP architectures. |
| Outcome: | The proposed approach improves the performance of the open-vocabulary audio language models by 0.8%-13% across benchmarks and enhances robustness to linguistic variation. |
PolyAudio: Advancing Multi-Audio Reasoning in Large Audio Language Models with Interleaved Multi-Audio Contexts (2026.findings-acl)
Copied to clipboard
Sonal Kumar, Sreyan Ghosh, Yueqian Lin, S Sakshi, Ashish Seth, Yiran Chen, Ramani Duraiswami, Dinesh Manocha
| Challenge: | Large Audio Language Models have shown impressive performance on single-clip tasks . however, their ability to reason over interleaved multi-audio contexts remains limited . |
| Approach: | They propose a LALM that targets multi-audio understanding via instruction tuning rather than massive-scale pre-training. |
| Outcome: | The proposed model outperforms baseline models on multi-audio tasks while maintaining robustness. |
DALE: Generative Data Augmentation for Low-Resource Legal NLP (2023.emnlp-main)
Copied to clipboard
Sreyan Ghosh, Chandra Kiran Reddy Evuru, Sonal Kumar, S Ramaneswaran, S Sakshi, Utkarsh Tyagi, Dinesh Manocha
| Challenge: | DALE addresses the challenges existing frameworks pose in generating effective data augmentations of legal documents. |
| Approach: | They propose a generative Data Augmentation framework for low-resource legal NLP that exploits domain-specific language characteristics of templated legal documents to mask collocated spans of text. |
| Outcome: | The proposed framework outperforms baseline frameworks on 13 datasets and 4 low-resource settings. |