Papers by Utkarsh Tyagi

12 papers
Do Vision-Language Models Understand Compound Nouns? (2024.naacl-short)

Copied to clipboard

Challenge: Open-vocabulary vision-language models (CLIP) are emerging as a promising new paradigm for text-to-image retrieval.
Approach: They propose a benchmark to evaluate the effectiveness of open-vocabulary vision-language models (CLIP) for text-to-image retrieval using contrastive loss.
Outcome: The proposed framework improves CN understanding of CLIP by 8.25% on Compun.
ASPIRE: Language-Guided Data Augmentation for Improving Robustness Against Spurious Correlations (2024.findings-acl)

Copied to clipboard

Challenge: Neural image classifiers often rely on non-predictive features that are spuriously correlated with the class labels in training data.
Approach: They propose a language-guided data augmented with images without spurious correlations that can be used to augment training datasets for robust learning.
Outcome: The proposed model improves the worst-group classification accuracy of prior methods by 1% - 38%.
EGOILLUSION: Benchmarking Hallucinations in Egocentric Video Understanding (2025.emnlp-main)

Copied to clipboard

Challenge: Multimodal Large Language Models excel at visual perception and reasoning in third-person and egocentric videos, but are prone to hallucinations, generating coherent yet inaccurate responses.
Approach: They propose to use a benchmark to evaluate MLLM hallucinations in egocentric videos.
Outcome: EGOILLUSION comprises 1,400 videos paired with 8,000 human-annotated open and closed-ended questions designed to trigger hallucinations in both visual and auditory cues in egocentric videos.
CoDa: Constrained Generation based Data Augmentation for Low-Resource NLP (2024.findings-naacl)

Copied to clipboard

Challenge: a low-resource dataset is limited in training data, so generating task-specific data is challenging.
Approach: They propose a data augmentation technique that prompts off-the-shelf instruction-following Large Language Models to generate augmentations.
Outcome: The proposed technique outperforms baselines on 11 datasets spanning 3 tasks and 3 low-resource settings.
ABEX: Data Augmentation for Low-Resource NLU via Expanding Abstract Descriptions (2024.acl-long)

Copied to clipboard

Challenge: ABEX is a novel and effective generative data augmentation methodology for low-resource Natural Language Understanding (NLU) tasks.
Approach: They propose a novel generative data augmentation methodology for low-resource Natural Language Understanding (NLU) tasks based on a paradigm for generating diverse forms of an input document .
Outcome: The proposed method outperforms all baselines qualitatively with improvements of 0.04% - 38.8%.
GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities (2024.emnlp-main)

Copied to clipboard

Challenge: We propose a novel large-scale audio-language model with advanced audio understanding and reasoning abilities.
Approach: They propose a general-purpose large audio-language model with advanced audio understanding and reasoning abilities that integrates an LLM with multiple types of audio representations.
Outcome: The proposed model outperforms existing models on audio understanding tasks by 1%-84%.
Audio MultiChallenge: A Multi-Turn Evaluation of Spoken Dialogue Systems on Natural Human Interaction (2026.acl-long)

Copied to clipboard

Challenge: End-to-end (E2E) spoken dialogue systems are replacing cascaded pipelines for voice-based human-AI interaction. Existing benchmarks evaluate these systems on synthetic speech and single-turn tasks, leaving multi-turn conversational ability underexplored.
Approach: They propose an open-source benchmark to evaluate spoken dialogue systems under natural multi-turn interaction patterns.
Outcome: The proposed model fails on the highest-performing model with 54.65% pass rate.
CoSyn: Detecting Implicit Hate Speech in Online Conversations Using a Context Synergized Hyperbolic Network (2023.emnlp-main)

Copied to clipboard

Challenge: Existing work on detecting explicit hate speech has focused on indirect or coded language.
Approach: They propose a context synergized neural network that integrates user- and conversational-contexts for detecting implicit hate speech in online conversations.
Outcome: The proposed framework outperforms baselines on 6 hate speech datasets and shows that it is highly efficient.
MULTIVOX: A Benchmark for Evaluating Voice Assistants for Multimodal Interactions (2025.emnlp-main)

Copied to clipboard

Challenge: omni models lack spoken dialogues, which is essential for assessing conversational and auditory capabilities of voice assistants.
Approach: They propose a benchmark to evaluate the ability of voice assistants to integrate paralinguistic speech features into their models.
Outcome: The multivox voice assistant benchmark evaluates the ability of models to integrate spoken and visual cues including paralinguistic speech features for truly multimodal understanding.
ProSE: Diffusion Priors for Speech Enhancement (2025.naacl-long)

Copied to clipboard

Challenge: deterministic deep learning models have been used for speech enhancement, but generative models have shown promise.
Approach: They propose a method to apply diffusion probabilistic models to speech enhancement using priors in a latent space.
Outcome: The proposed method achieves state-of-the-art performance on synthetic and real-world datasets while consuming less computational costs.
ACLM: A Selective-Denoising based Generative Data Augmentation Approach for Low-Resource Complex NER (2023.acl-long)

Copied to clipboard

Challenge: Named Entity Recognition (NER) is a task of detecting linguistically complex named entities in low-context text.
Approach: They propose a keyword-based augmentation approach to address the context-entity mismatch issue in complex name recognition (NER) they use selective masking to retain the named entities and certain keywords in the input sentence that provide contextually relevant additional knowledge or hints about the named entity.
Outcome: The proposed approach outperforms baseline methods on monolingual, cross-lingual, and multilingual complex NER in various low-resource settings.
DALE: Generative Data Augmentation for Low-Resource Legal NLP (2023.emnlp-main)

Copied to clipboard

Challenge: DALE addresses the challenges existing frameworks pose in generating effective data augmentations of legal documents.
Approach: They propose a generative Data Augmentation framework for low-resource legal NLP that exploits domain-specific language characteristics of templated legal documents to mask collocated spans of text.
Outcome: The proposed framework outperforms baseline frameworks on 13 datasets and 4 low-resource settings.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations