Papers with RVQ
Language Model Based Text-to-Audio Generation: Anti-Causally Aligned Collaborative Residual Transformers (2025.emnlp-main)
Copied to clipboard
| Challenge: | Autoregressive language models excel in text-to-audio generation, but lag behind diffusion models by a non-trivial margin. |
| Approach: | They propose a framework that integrates multiple isolated transformers with causal conditioning and anti-causal alignment via reinforcement learning. |
| Outcome: | The proposed framework outperforms existing LM-based and diffusion-based systems in audio synthesis. |
DM-Codec: Distilling Multimodal Representations for Speech Tokenization (2025.findings-emnlp)
Copied to clipboard
Md Mubtasim Ahasan, Md Fahim, Tasnim Mohiuddin, Akmmahbubur Rahman, Aman Chadha, Tariq Iqbal, M Ashraful Amin, Md Mofijul Islam, Amin Ahsan Ali
| Challenge: | Existing speech tokenization models lack contextual representations for speech synthesis . absence of contextual representation results in elevated WER and WIL scores . |
| Approach: | They propose a language model-guided distillation method that incorporates contextual information into a comprehensive speech tokenizer. |
| Outcome: | The proposed method outperforms state-of-the-art tokenization models in reducing WER and WIL scores. |