Papers by Utkarsh Tyagi
Do Vision-Language Models Understand Compound Nouns? (2024.naacl-short)
Copied to clipboard
| Challenge: | Open-vocabulary vision-language models (CLIP) are emerging as a promising new paradigm for text-to-image retrieval. |
| Approach: | They propose a benchmark to evaluate the effectiveness of open-vocabulary vision-language models (CLIP) for text-to-image retrieval using contrastive loss. |
| Outcome: | The proposed framework improves CN understanding of CLIP by 8.25% on Compun. |
ASPIRE: Language-Guided Data Augmentation for Improving Robustness Against Spurious Correlations (2024.findings-acl)
Copied to clipboard
Sreyan Ghosh, Chandra Kiran Evuru, Sonal Kumar, Utkarsh Tyagi, S Sakshi, Sanjoy Chowdhury, Dinesh Manocha
| Challenge: | Neural image classifiers often rely on non-predictive features that are spuriously correlated with the class labels in training data. |
| Approach: | They propose a language-guided data augmented with images without spurious correlations that can be used to augment training datasets for robust learning. |
| Outcome: | The proposed model improves the worst-group classification accuracy of prior methods by 1% - 38%. |
EGOILLUSION: Benchmarking Hallucinations in Egocentric Video Understanding (2025.emnlp-main)
Copied to clipboard
Ashish Seth, Utkarsh Tyagi, Ramaneswaran Selvakumar, Nishit Anand, Sonal Kumar, Sreyan Ghosh, Ramani Duraiswami, Chirag Agarwal, Dinesh Manocha
| Challenge: | Multimodal Large Language Models excel at visual perception and reasoning in third-person and egocentric videos, but are prone to hallucinations, generating coherent yet inaccurate responses. |
| Approach: | They propose to use a benchmark to evaluate MLLM hallucinations in egocentric videos. |
| Outcome: | EGOILLUSION comprises 1,400 videos paired with 8,000 human-annotated open and closed-ended questions designed to trigger hallucinations in both visual and auditory cues in egocentric videos. |
CoDa: Constrained Generation based Data Augmentation for Low-Resource NLP (2024.findings-naacl)
Copied to clipboard
| Challenge: | a low-resource dataset is limited in training data, so generating task-specific data is challenging. |
| Approach: | They propose a data augmentation technique that prompts off-the-shelf instruction-following Large Language Models to generate augmentations. |
| Outcome: | The proposed technique outperforms baselines on 11 datasets spanning 3 tasks and 3 low-resource settings. |
ABEX: Data Augmentation for Low-Resource NLU via Expanding Abstract Descriptions (2024.acl-long)
Copied to clipboard
Sreyan Ghosh, Utkarsh Tyagi, Sonal Kumar, Chandra Kiran Evuru, Ramaneswaran S, S Sakshi, Dinesh Manocha
| Challenge: | ABEX is a novel and effective generative data augmentation methodology for low-resource Natural Language Understanding (NLU) tasks. |
| Approach: | They propose a novel generative data augmentation methodology for low-resource Natural Language Understanding (NLU) tasks based on a paradigm for generating diverse forms of an input document . |
| Outcome: | The proposed method outperforms all baselines qualitatively with improvements of 0.04% - 38.8%. |
GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities (2024.emnlp-main)
Copied to clipboard
Sreyan Ghosh, Sonal Kumar, Ashish Seth, Chandra Kiran Evuru, Utkarsh Tyagi, S Sakshi, Oriol Nieto, Ramani Duraiswami, Dinesh Manocha
| Challenge: | We propose a novel large-scale audio-language model with advanced audio understanding and reasoning abilities. |
| Approach: | They propose a general-purpose large audio-language model with advanced audio understanding and reasoning abilities that integrates an LLM with multiple types of audio representations. |
| Outcome: | The proposed model outperforms existing models on audio understanding tasks by 1%-84%. |
Audio MultiChallenge: A Multi-Turn Evaluation of Spoken Dialogue Systems on Natural Human Interaction (2026.acl-long)
Copied to clipboard
Advait Gosai, Tyler Vuong, Utkarsh Tyagi, Steven Li, Wenjia You, Miheer Bavare, Arda Uçar, Zhongwang Fang, Brian Jang, Bing Liu, Yunzhong He
| Challenge: | End-to-end (E2E) spoken dialogue systems are replacing cascaded pipelines for voice-based human-AI interaction. Existing benchmarks evaluate these systems on synthetic speech and single-turn tasks, leaving multi-turn conversational ability underexplored. |
| Approach: | They propose an open-source benchmark to evaluate spoken dialogue systems under natural multi-turn interaction patterns. |
| Outcome: | The proposed model fails on the highest-performing model with 54.65% pass rate. |
CoSyn: Detecting Implicit Hate Speech in Online Conversations Using a Context Synergized Hyperbolic Network (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing work on detecting explicit hate speech has focused on indirect or coded language. |
| Approach: | They propose a context synergized neural network that integrates user- and conversational-contexts for detecting implicit hate speech in online conversations. |
| Outcome: | The proposed framework outperforms baselines on 6 hate speech datasets and shows that it is highly efficient. |
MULTIVOX: A Benchmark for Evaluating Voice Assistants for Multimodal Interactions (2025.emnlp-main)
Copied to clipboard
Ramaneswaran Selvakumar, Ashish Seth, Nishit Anand, Utkarsh Tyagi, Sonal Kumar, Sreyan Ghosh, Dinesh Manocha
| Challenge: | omni models lack spoken dialogues, which is essential for assessing conversational and auditory capabilities of voice assistants. |
| Approach: | They propose a benchmark to evaluate the ability of voice assistants to integrate paralinguistic speech features into their models. |
| Outcome: | The multivox voice assistant benchmark evaluates the ability of models to integrate spoken and visual cues including paralinguistic speech features for truly multimodal understanding. |
ProSE: Diffusion Priors for Speech Enhancement (2025.naacl-long)
Copied to clipboard
Sonal Kumar, Sreyan Ghosh, Utkarsh Tyagi, Anton Jeran Ratnarajah, Chandra Kiran Reddy Evuru, Ramani Duraiswami, Dinesh Manocha
| Challenge: | deterministic deep learning models have been used for speech enhancement, but generative models have shown promise. |
| Approach: | They propose a method to apply diffusion probabilistic models to speech enhancement using priors in a latent space. |
| Outcome: | The proposed method achieves state-of-the-art performance on synthetic and real-world datasets while consuming less computational costs. |
ACLM: A Selective-Denoising based Generative Data Augmentation Approach for Low-Resource Complex NER (2023.acl-long)
Copied to clipboard
| Challenge: | Named Entity Recognition (NER) is a task of detecting linguistically complex named entities in low-context text. |
| Approach: | They propose a keyword-based augmentation approach to address the context-entity mismatch issue in complex name recognition (NER) they use selective masking to retain the named entities and certain keywords in the input sentence that provide contextually relevant additional knowledge or hints about the named entity. |
| Outcome: | The proposed approach outperforms baseline methods on monolingual, cross-lingual, and multilingual complex NER in various low-resource settings. |
DALE: Generative Data Augmentation for Low-Resource Legal NLP (2023.emnlp-main)
Copied to clipboard
Sreyan Ghosh, Chandra Kiran Reddy Evuru, Sonal Kumar, S Ramaneswaran, S Sakshi, Utkarsh Tyagi, Dinesh Manocha
| Challenge: | DALE addresses the challenges existing frameworks pose in generating effective data augmentations of legal documents. |
| Approach: | They propose a generative Data Augmentation framework for low-resource legal NLP that exploits domain-specific language characteristics of templated legal documents to mask collocated spans of text. |
| Outcome: | The proposed framework outperforms baseline frameworks on 13 datasets and 4 low-resource settings. |