Challenge: Recent advances in text-to-speech (TTS) models have led to improvements in speaker prosody and voices modeling.
Approach: They propose an efficient zero-shot TTS model that leverages distilled time-varying style diffusion to capture diverse speaker identities and prosodies.
Outcome: The proposed model surpasses state-of-the-art models in both naturalness and similarity while reducing inference speed by 90%.

Similar Papers

DiffStyleTTS: Diffusion-based Hierarchical Prosody Modeling for Text-to-Speech with Diverse and Controllable Styles (2025.coling-main)

Copied to clipboard

Challenge: Existing models for text-to-speech (TTS) synthesize speech with acoustic features . autoregressive models have problems with word skipping and repeated reading . non-autoregressive acustic models lack probabilistic modeling and unimodal characteristics of Gaussian distribution don't conform to true distribution of aural features, which restricts the diversity of generated prosodic features.
Approach: They propose a multi-speaker acoustic model that hierarchically models speech prosodic features and controls different prosodic styles to guide prosody prediction.
Outcome: The proposed method outperforms baseline models in naturalness and achieves superior synthesis speed compared to baseline models.
FC-TTS: Style and Timbre Control in Zero-Shot Text-to-Speech with Disentangled Speech Representations (2026.acl-long)

Copied to clipboard

Challenge: Recent advances in text-to-speech (TTS) have enabled accurate imitation of reference speech in terms of both speaking style and speaker timbre.
Approach: They propose a zero-shot text-to-speech framework that enables disentangled control of style and timbre by conditioning on two distinct reference utterances.
Outcome: The proposed framework achieves high-fidelity synthesis and competitive zero-shot naturalness while supporting consistent and independent manipulation of style and timbre.
MultiVerse: Efficient and Expressive Zero-Shot Multi-Task Text-to-Speech (2024.findings-emnlp)

Copied to clipboard

Challenge: Text-to-speech systems that scale up the amount of training data have certain limitations: they require a large amount of data, which increases costs, and overlook prosody similarity.
Approach: They propose a zero-shot multi-task TTS system that can perform TTS or speech style transfer in zero- shot and cross-lingual conditions.
Outcome: The proposed system outperforms other TTS systems trained with the same small amount of data and achieves zero-shot performance comparable to data-driven systems.
ReStyle-TTS: Relative and Continuous Style Control for Zero-Shot Speech Synthesis (2026.findings-acl)

Copied to clipboard

Challenge: Recent controllable zero-shot text-to-speech systems can synthesize speech for unseen speakers from a short reference audio clip, but they also inherit the speaking style present in the reference.
Approach: They propose a framework that enables continuous and reference-relative style control in zero-shot text-to-speech systems by combining style-specific LoRAs with Orthogonal LoRA Fusion.
Outcome: The proposed framework reduces the model's dependence on reference style while preserving text fidelity while maintaining intelligibility and speaker timbre.
ControlSpeech: Towards Simultaneous and Independent Zero-shot Speaker Cloning and Zero-shot Language Style Control (2025.acl-long)

Copied to clipboard

Challenge: Prior zero-shot TTS models only mimic the speaker’s voice without further control and adjustment capabilities while prior controllable TTS systems cannot perform speaker-specific voice generation.
Approach: They propose a style control module that captures codec representations corresponding to timbre, content, and style in a discrete decoupling codec space.
Outcome: The proposed system can fully clone the speaker's voice and perform speech-specific adjustment and control functions.
Prosody-TTS: Improving Prosody with Masked Autoencoder and Conditional Diffusion Model For Expressive Text-to-Speech (2023.findings-acl)

Copied to clipboard

Challenge: Expressive text-to-speech aims to generate high-quality samples with rich prosody . prosodic attributes in highly dynamic voices are difficult to capture and model without intonation .
Approach: They propose a pipeline that enhances prosody modeling and sampling by introducing a self-supervised masked autoencoder and a diffusion model to sample diverse prosodic patterns within the latent space.
Outcome: The proposed pipeline achieves new state-of-the-art in text-to-speech with natural and expressive synthesis.
TCSinger: Zero-Shot Singing Voice Synthesis with Style Transfer and Multi-Level Style Control (2024.emnlp-main)

Copied to clipboard

Challenge: Existing models fail to generate singing voices rich in stylistic nuances for unseen singers due to multifaceted nature of singing styles.
Approach: They propose a zero-shot SVS model for style transfer across cross-lingual speech and singing styles and multi-level style control.
Outcome: Experimental results show that TCSinger outperforms baseline models in synthesis quality, singer similarity, and style controllability.
Low-Resource Multilingual and Zero-Shot Multispeaker TTS (2022.aacl-main)

Copied to clipboard

Challenge: Currently, the amount of data needed for TTS is limited to the vast majority of the spoken languages.
Approach: They propose to use language agnostic meta learning procedure to learn speaking a new language with just 5 minutes of training data while retaining the ability to infer the voice of even unseen speakers.
Outcome: The proposed approach is able to learn speaking a new language using just 5 minutes of training data while retaining the ability to infer the voice of even unseen speakers in the newly learned language.
Scaling Under-Resourced TTS: A Data-Optimized Framework with Advanced Acoustic Modeling for Thai (2025.acl-industry)

Copied to clipboard

Challenge: Text-to-speech (TTS) systems are limited by limited data and linguistic complexities.
Approach: They propose a data-optimized framework with an advanced acoustic model to build high-quality TTS systems for low-resource scenarios.
Outcome: The proposed framework enables zero-shot voice cloning and improved performance across diverse client applications, including finance, healthcare, education, and law.
VoxpopuliTTS: a large-scale multilingual TTS corpus for zero-shot speech generation (2025.coling-main)

Copied to clipboard

Challenge: Existing multilingual TTS datasets are limited in speech generation fields due to lack of quality data.
Approach: They propose to use 30,000 hours of high-quality speech data across 3 languages . they filter out low-quality text-text pairs and concatenate short transcripts .
Outcome: The proposed dataset comprises 30,000 hours of high-quality speech data, across 3 languages with multiple speakers and styles, suitable for various speech tasks such as TTS and ASR.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations