ReStyle-TTS: Relative and Continuous Style Control for Zero-Shot Speech Synthesis (2026.findings-acl)
Copied to clipboard
| Challenge: | Recent controllable zero-shot text-to-speech systems can synthesize speech for unseen speakers from a short reference audio clip, but they also inherit the speaking style present in the reference. |
| Approach: | They propose a framework that enables continuous and reference-relative style control in zero-shot text-to-speech systems by combining style-specific LoRAs with Orthogonal LoRA Fusion. |
| Outcome: | The proposed framework reduces the model's dependence on reference style while preserving text fidelity while maintaining intelligibility and speaker timbre. |
Similar Papers
FC-TTS: Style and Timbre Control in Zero-Shot Text-to-Speech with Disentangled Speech Representations (2026.acl-long)
Copied to clipboard
| Challenge: | Recent advances in text-to-speech (TTS) have enabled accurate imitation of reference speech in terms of both speaking style and speaker timbre. |
| Approach: | They propose a zero-shot text-to-speech framework that enables disentangled control of style and timbre by conditioning on two distinct reference utterances. |
| Outcome: | The proposed framework achieves high-fidelity synthesis and competitive zero-shot naturalness while supporting consistent and independent manipulation of style and timbre. |
ControlSpeech: Towards Simultaneous and Independent Zero-shot Speaker Cloning and Zero-shot Language Style Control (2025.acl-long)
Copied to clipboard
Shengpeng Ji, Qian Chen, Wen Wang, Jialong Zuo, Minghui Fang, Ziyue Jiang, Hai Huang, Zehan Wang, Xize Cheng, Siqi Zheng, Zhou Zhao
| Challenge: | Prior zero-shot TTS models only mimic the speaker’s voice without further control and adjustment capabilities while prior controllable TTS systems cannot perform speaker-specific voice generation. |
| Approach: | They propose a style control module that captures codec representations corresponding to timbre, content, and style in a discrete decoupling codec space. |
| Outcome: | The proposed system can fully clone the speaker's voice and perform speech-specific adjustment and control functions. |
StyleTTS-ZS: Efficient High-Quality Zero-Shot Text-to-Speech Synthesis with Distilled Time-Varying Style Diffusion (2025.naacl-long)
Copied to clipboard
| Challenge: | Recent advances in text-to-speech (TTS) models have led to improvements in speaker prosody and voices modeling. |
| Approach: | They propose an efficient zero-shot TTS model that leverages distilled time-varying style diffusion to capture diverse speaker identities and prosodies. |
| Outcome: | The proposed model surpasses state-of-the-art models in both naturalness and similarity while reducing inference speed by 90%. |
Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment (2025.acl-long)
Copied to clipboard
| Challenge: | Existing zero-shot text-to-speech systems struggle in challenging scenarios such as tongue twisters, repeated words, code-switching, and cross-lingual synthesis. |
| Approach: | They propose a dataset that leverages preference alignment techniques to improve performance . they also extend the Direct Preference Optimization framework to accommodate diverse TTS architectures . |
| Outcome: | The proposed dataset improves intelligibility, similarity, and audio quality for multiple models across domains. |
TCSinger: Zero-Shot Singing Voice Synthesis with Style Transfer and Multi-Level Style Control (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing models fail to generate singing voices rich in stylistic nuances for unseen singers due to multifaceted nature of singing styles. |
| Approach: | They propose a zero-shot SVS model for style transfer across cross-lingual speech and singing styles and multi-level style control. |
| Outcome: | Experimental results show that TCSinger outperforms baseline models in synthesis quality, singer similarity, and style controllability. |
Low-Resource Multilingual and Zero-Shot Multispeaker TTS (2022.aacl-main)
Copied to clipboard
| Challenge: | Currently, the amount of data needed for TTS is limited to the vast majority of the spoken languages. |
| Approach: | They propose to use language agnostic meta learning procedure to learn speaking a new language with just 5 minutes of training data while retaining the ability to infer the voice of even unseen speakers. |
| Outcome: | The proposed approach is able to learn speaking a new language using just 5 minutes of training data while retaining the ability to infer the voice of even unseen speakers in the newly learned language. |
MultiVerse: Efficient and Expressive Zero-Shot Multi-Task Text-to-Speech (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Text-to-speech systems that scale up the amount of training data have certain limitations: they require a large amount of data, which increases costs, and overlook prosody similarity. |
| Approach: | They propose a zero-shot multi-task TTS system that can perform TTS or speech style transfer in zero- shot and cross-lingual conditions. |
| Outcome: | The proposed system outperforms other TTS systems trained with the same small amount of data and achieves zero-shot performance comparable to data-driven systems. |
kNN Retrieval for Simple and Effective Zero-Shot Multi-speaker Text-to-Speech (2025.naacl-short)
Copied to clipboard
| Challenge: | Neural text-to-speech (TTS) models typically rely on extensive transcribed speech datasets and intricate training pipelines. |
| Approach: | They propose a framework for zero-shot multi-speaker text-to-speech using retrieval methods which leverage the linear relationships between SSL features. |
| Outcome: | The proposed framework achieves comparable performance to state-of-the-art models trained on large training datasets. |
BnTTS: Few-Shot Speaker Adaptation in Low-Resource Setting (2025.findings-naacl)
Copied to clipboard
Mohammad Jahid Ibna Basher, Md Kowsher, Md Saiful Islam, Rabindra Nath Nandi, Nusrat Jahan Prottasha, Mehadi Hasan Menon, Tareq Al Muntasir, Shammur Absar Chowdhury, Firoj Alam, Niloofar Yousefi, Ozlem Garibay
| Challenge: | Empirical evaluations in few-shot settings show that BnTTS significantly improves the naturalness, intelligibility, and speaker fidelity of synthesized Bangla speech. |
| Approach: | They propose to integrate Bangla into a multilingual TTS pipeline with modifications to account for the phonetic and linguistic characteristics of the language. |
| Outcome: | The proposed framework improves the naturalness, intelligibility, and speaker fidelity of synthesized Bangla speech compared to state-of-the-art systems. |
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation (2026.findings-acl)
Copied to clipboard
| Challenge: | Neural codec language models (NCLMs) lack fine-grained controllability and inability to extrapolate to sequence lengths much longer than those seen during training. |
| Approach: | They propose a novel autoregressive encoder-decoder neural codec language model that can be trained with a Continuation-Prompt Mixed training system. |
| Outcome: | The proposed model outperforms or is on par with current state-of-the-art models on short-form benchmarks such as LibriSpeech and Seed-TTS in terms of intelligibility and naturalness. |