Papers by Anuj Diwan

5 papers
Why is Winoground Hard? Investigating Failures in Visuolinguistic Compositionality (2022.emnlp-main)

Copied to clipboard

Challenge: Recent visuolinguistic pre-trained models fail miserably on the Winoground dataset, which challenges models to match paired images and English captions.
Approach: They propose to annotate a Winoground dataset that challenges visuolinguistic models to match paired images and English captions with items constructed to overlap lexically but differ in meaning.
Outcome: The proposed dataset challenges models to match paired images and English captions with items constructed to overlap lexically but differ in meaning.
Textless Speech-to-Speech Translation With Limited Parallel Data (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing speech-to-speech translation models either leverage text as an intermediate step or require hundreds of hours of parallel speech data.
Approach: They propose a framework for training textless S2ST models that require dozens of hours of parallel speech data.
Outcome: The proposed model achieves reasonable performance on three domains with single-speaker synthesized speech.
Scaling Rich Style-Prompted Text-to-Speech Datasets (2025.emnlp-main)

Copied to clipboard

Challenge: Existing datasets that only cover basic tags are limited in their scale or coverage of style tags.
Approach: They propose a large-scale dataset that annotates speech utterances with rich style captions.
Outcome: The proposed dataset scales speech utterances with rich style captions for the first time.
When to Use Efficient Self Attention? Profiling Text, Speech and Image Transformer Variants (2023.acl-short)

Copied to clipboard

Challenge: Existing models focus on improving the efficiency of self-attention, but in practice they may be slower, especially given modest input lengths that are typical of many tasks.
Approach: They propose a novel local-attention variant of a self-supervised speech model that uses input length thresholds to identify bottlenecks.
Outcome: The proposed model is based on a self-attention-based model with a high input length threshold.
VoiceCraft-X: Unifying Multilingual, Voice-Cloning Speech Synthesis and Speech Editing (2025.emnlp-main)

Copied to clipboard

Challenge: Autoregressive language model for multilingual speech editing and zero-shot text-to-speech synthesis is available in 11 languages.
Approach: They introduce an autoregressive neural codec language model which unifies multilingual speech editing and zero-shot text-to-speech synthesis across 11 languages.
Outcome: The model generates high-quality, natural-sounding speech, even with limited per-language data . it shows robust performance in diverse linguistic settings, even in limited per language data compared to other models .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations