Papers by Anuj Diwan
Why is Winoground Hard? Investigating Failures in Visuolinguistic Compositionality (2022.emnlp-main)
Copied to clipboard
| Challenge: | Recent visuolinguistic pre-trained models fail miserably on the Winoground dataset, which challenges models to match paired images and English captions. |
| Approach: | They propose to annotate a Winoground dataset that challenges visuolinguistic models to match paired images and English captions with items constructed to overlap lexically but differ in meaning. |
| Outcome: | The proposed dataset challenges models to match paired images and English captions with items constructed to overlap lexically but differ in meaning. |
Textless Speech-to-Speech Translation With Limited Parallel Data (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Existing speech-to-speech translation models either leverage text as an intermediate step or require hundreds of hours of parallel speech data. |
| Approach: | They propose a framework for training textless S2ST models that require dozens of hours of parallel speech data. |
| Outcome: | The proposed model achieves reasonable performance on three domains with single-speaker synthesized speech. |
Scaling Rich Style-Prompted Text-to-Speech Datasets (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing datasets that only cover basic tags are limited in their scale or coverage of style tags. |
| Approach: | They propose a large-scale dataset that annotates speech utterances with rich style captions. |
| Outcome: | The proposed dataset scales speech utterances with rich style captions for the first time. |
When to Use Efficient Self Attention? Profiling Text, Speech and Image Transformer Variants (2023.acl-short)
Copied to clipboard
| Challenge: | Existing models focus on improving the efficiency of self-attention, but in practice they may be slower, especially given modest input lengths that are typical of many tasks. |
| Approach: | They propose a novel local-attention variant of a self-supervised speech model that uses input length thresholds to identify bottlenecks. |
| Outcome: | The proposed model is based on a self-attention-based model with a high input length threshold. |
VoiceCraft-X: Unifying Multilingual, Voice-Cloning Speech Synthesis and Speech Editing (2025.emnlp-main)
Copied to clipboard
Zhisheng Zheng, Puyuan Peng, Anuj Diwan, Cong Phuoc Huynh, Xiaohang Sun, Zhu Liu, Vimal Bhat, David Harwath
| Challenge: | Autoregressive language model for multilingual speech editing and zero-shot text-to-speech synthesis is available in 11 languages. |
| Approach: | They introduce an autoregressive neural codec language model which unifies multilingual speech editing and zero-shot text-to-speech synthesis across 11 languages. |
| Outcome: | The model generates high-quality, natural-sounding speech, even with limited per-language data . it shows robust performance in diverse linguistic settings, even in limited per language data compared to other models . |