Papers by Yajun Hu
DiffStyleTTS: Diffusion-based Hierarchical Prosody Modeling for Text-to-Speech with Diverse and Controllable Styles (2025.coling-main)
Copied to clipboard
| Challenge: | Existing models for text-to-speech (TTS) synthesize speech with acoustic features . autoregressive models have problems with word skipping and repeated reading . non-autoregressive acustic models lack probabilistic modeling and unimodal characteristics of Gaussian distribution don't conform to true distribution of aural features, which restricts the diversity of generated prosodic features. |
| Approach: | They propose a multi-speaker acoustic model that hierarchically models speech prosodic features and controls different prosodic styles to guide prosody prediction. |
| Outcome: | The proposed method outperforms baseline models in naturalness and achieves superior synthesis speed compared to baseline models. |