Papers by Ashish Seth

8 papers
EGOILLUSION: Benchmarking Hallucinations in Egocentric Video Understanding (2025.emnlp-main)

Copied to clipboard

Challenge: Multimodal Large Language Models excel at visual perception and reasoning in third-person and egocentric videos, but are prone to hallucinations, generating coherent yet inaccurate responses.
Approach: They propose to use a benchmark to evaluate MLLM hallucinations in egocentric videos.
Outcome: EGOILLUSION comprises 1,400 videos paired with 8,000 human-annotated open and closed-ended questions designed to trigger hallucinations in both visual and auditory cues in egocentric videos.
GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities (2024.emnlp-main)

Copied to clipboard

Challenge: We propose a novel large-scale audio-language model with advanced audio understanding and reasoning abilities.
Approach: They propose a general-purpose large audio-language model with advanced audio understanding and reasoning abilities that integrates an LLM with multiple types of audio representations.
Outcome: The proposed model outperforms existing models on audio understanding tasks by 1%-84%.
PAT: Parameter-Free Audio-Text Aligner to Boost Zero-Shot Audio Classification (2025.naacl-long)

Copied to clipboard

Challenge: Audio-Language Models (ALMs) have demonstrated remarkable performance in zero-shot audio classification.
Approach: They propose a training-free method that enhances audio and language representations using mutual feedback.
Outcome: The proposed method outperforms vanilla zero-shot evaluation with significant margins of 0.42%-27.0%.
MULTIVOX: A Benchmark for Evaluating Voice Assistants for Multimodal Interactions (2025.emnlp-main)

Copied to clipboard

Challenge: omni models lack spoken dialogues, which is essential for assessing conversational and auditory capabilities of voice assistants.
Approach: They propose a benchmark to evaluate the ability of voice assistants to integrate paralinguistic speech features into their models.
Outcome: The multivox voice assistant benchmark evaluates the ability of models to integrate spoken and visual cues including paralinguistic speech features for truly multimodal understanding.
EH-MAM: Easy-to-Hard Masked Acoustic Modeling for Self-Supervised Speech Representation Learning (2024.emnlp-main)

Copied to clipboard

Challenge: EH-MAM is a self-supervised learning approach for speech representation learning . prior methods used random masking schemes to learn speech representations .
Approach: They propose a self-supervised approach that automatically selects hard regions during SSL training and introduces them to the model for reconstruction.
Outcome: The proposed approach outperforms state-of-the-art models across low-resource speech recognition and SUPERB benchmarks by 5%-10%.
FIGMA: Towards FIne-Grained Music retrievAl (2026.acl-long)

Copied to clipboard

Challenge: Existing music retrieval models fail to retrieve fine-grained musical attributes when using coarse semantic queries.
Approach: They propose a multi-view contrastive architecture that captures high-level semantic context and fine-grained musical attributes within a unified representation space.
Outcome: The proposed method outperforms existing CLAP-based music retrieval models on multiple benchmarks.
Do Audio-Language Models Understand Linguistic Variations? (2025.naacl-short)

Copied to clipboard

Challenge: Existing open-vocabulary audio language models struggle to generalize to linguistic variations in textual queries.
Approach: They propose a novel technique to learn audio-language representations agnostic to linguistic variations by reformulating contrastive loss used in CLAP architectures.
Outcome: The proposed approach improves the performance of the open-vocabulary audio language models by 0.8%-13% across benchmarks and enhances robustness to linguistic variation.
PolyAudio: Advancing Multi-Audio Reasoning in Large Audio Language Models with Interleaved Multi-Audio Contexts (2026.findings-acl)

Copied to clipboard

Challenge: Large Audio Language Models have shown impressive performance on single-clip tasks . however, their ability to reason over interleaved multi-audio contexts remains limited .
Approach: They propose a LALM that targets multi-audio understanding via instruction tuning rather than massive-scale pre-training.
Outcome: The proposed model outperforms baseline models on multi-audio tasks while maintaining robustness.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations