Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment (2025.acl-long)
Copied to clipboard
| Challenge: | Existing zero-shot text-to-speech systems struggle in challenging scenarios such as tongue twisters, repeated words, code-switching, and cross-lingual synthesis. |
| Approach: | They propose a dataset that leverages preference alignment techniques to improve performance . they also extend the Direct Preference Optimization framework to accommodate diverse TTS architectures . |
| Outcome: | The proposed dataset improves intelligibility, similarity, and audio quality for multiple models across domains. |
Similar Papers
Data-efficient Targeted Token-level Preference Optimization for LLM-based Text-to-Speech (2026.acl-short)
Copied to clipboard
| Challenge: | Recent work has shown that text-to-speech (TTS) models generate contextaware pronunciations from raw text without morphological analysis. |
| Approach: | They propose a preference optimization algorithm that aligns text-to-speech (TTS) outputs with human feedback. |
| Outcome: | The proposed method improves the challenging Japanese pronunciation accuracy by 39% and reduces CER by 54%. |
FC-TTS: Style and Timbre Control in Zero-Shot Text-to-Speech with Disentangled Speech Representations (2026.acl-long)
Copied to clipboard
| Challenge: | Recent advances in text-to-speech (TTS) have enabled accurate imitation of reference speech in terms of both speaking style and speaker timbre. |
| Approach: | They propose a zero-shot text-to-speech framework that enables disentangled control of style and timbre by conditioning on two distinct reference utterances. |
| Outcome: | The proposed framework achieves high-fidelity synthesis and competitive zero-shot naturalness while supporting consistent and independent manipulation of style and timbre. |
VoxpopuliTTS: a large-scale multilingual TTS corpus for zero-shot speech generation (2025.coling-main)
Copied to clipboard
Wenrui Liu, Jionghao Bai, Xize Cheng, Jialong Zuo, Ziyue Jiang, Shengpeng Ji, Minghui Fang, Xiaoda Yang, Qian Yang, Zhou Zhao
| Challenge: | Existing multilingual TTS datasets are limited in speech generation fields due to lack of quality data. |
| Approach: | They propose to use 30,000 hours of high-quality speech data across 3 languages . they filter out low-quality text-text pairs and concatenate short transcripts . |
| Outcome: | The proposed dataset comprises 30,000 hours of high-quality speech data, across 3 languages with multiple speakers and styles, suitable for various speech tasks such as TTS and ASR. |
Koel-TTS: Enhancing LLM based Speech Generation with Preference Alignment and Classifier Free Guidance (2025.emnlp-main)
Copied to clipboard
Shehzeen Samarah Hussain, Paarth Neekhara, Xuesong Yang, Edresson Casanova, Subhankar Ghosh, Roy Fejgin, Mikyas T. Desta, Rafael Valle, Jason Li
| Challenge: | Autoregressive speech token generation models suffer from hallucinations and undesired vocalizations that do not conform to conditioning inputs. |
| Approach: | They propose an encoder-decoder transformer model that improves contextual adherence of speech token generation LLMs through preference alignment and classifier-free guidance. |
| Outcome: | The proposed model outperforms previous LLM-based models on intelligibility, speaker similarity and naturalness. |
Low-Resource Multilingual and Zero-Shot Multispeaker TTS (2022.aacl-main)
Copied to clipboard
| Challenge: | Currently, the amount of data needed for TTS is limited to the vast majority of the spoken languages. |
| Approach: | They propose to use language agnostic meta learning procedure to learn speaking a new language with just 5 minutes of training data while retaining the ability to infer the voice of even unseen speakers. |
| Outcome: | The proposed approach is able to learn speaking a new language using just 5 minutes of training data while retaining the ability to infer the voice of even unseen speakers in the newly learned language. |
Scaling Under-Resourced TTS: A Data-Optimized Framework with Advanced Acoustic Modeling for Thai (2025.acl-industry)
Copied to clipboard
| Challenge: | Text-to-speech (TTS) systems are limited by limited data and linguistic complexities. |
| Approach: | They propose a data-optimized framework with an advanced acoustic model to build high-quality TTS systems for low-resource scenarios. |
| Outcome: | The proposed framework enables zero-shot voice cloning and improved performance across diverse client applications, including finance, healthcare, education, and law. |
MultiVerse: Efficient and Expressive Zero-Shot Multi-Task Text-to-Speech (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Text-to-speech systems that scale up the amount of training data have certain limitations: they require a large amount of data, which increases costs, and overlook prosody similarity. |
| Approach: | They propose a zero-shot multi-task TTS system that can perform TTS or speech style transfer in zero- shot and cross-lingual conditions. |
| Outcome: | The proposed system outperforms other TTS systems trained with the same small amount of data and achieves zero-shot performance comparable to data-driven systems. |
DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization (2025.findings-acl)
Copied to clipboard
Amitava Das, Suranjana Trivedy, Danush Khanna, Yaswanth Narsupalli, Basab Ghosh, Rajarshi Roy, Gurpreet Singh, Vinija Jain, Vasu Sharma, Aishwarya Naresh Reganti, Aman Chadha
| Challenge: | Direct Preference Optimization (DPO) is a cornerstone for preference alignment but is constrained by fixed divergence measures and limited feature transformations. |
| Approach: | They propose a new enhancement of Direct Preference Optimization that integrates kernel methods to overcome these challenges. |
| Outcome: | The proposed model improves divergence measures and features by using kernels . the proposed model achieves state-of-the-art generalization in factuality, safety, reasoning, and instruction following . |
A Simple and Effective Method to Improve Zero-Shot Cross-Lingual Transfer Learning (2022.coling-1)
Copied to clipboard
Kunbo Ding, Weijie Liu, Yuejian Fang, Weiquan Mao, Zhe Zhao, Tao Zhu, Haoyan Liu, Rong Tian, Yiren Chen
| Challenge: | Existing zero-shot cross-lingual transfer methods rely on parallel corpora or bilingual dictionaries . however, its effect is limited by the gap between embedding clusters of different languages . |
| Approach: | They propose Embedding-Push, Attention-Pull, and Robust targets to transfer English embeddings to virtual multilingual embedders without semantic loss. |
| Outcome: | Experimental results show that the proposed method outperforms existing methods on cross-lingual tasks and can achieve a better multilingual alignment. |
Takin-VC: Expressive Zero-Shot Voice Conversion via Adaptive Hybrid Content Encoding and Enhanced Timbre Modeling (2025.acl-long)
Copied to clipboard
Yang Yuguang, Yu Pan, Jixun Yao, Xiang Zhang, Jianhao Ye, Hongbin Zhou, Lei Xie, Lei Ma, Jianjun Zhao
| Challenge: | Expressive zero-shot voice conversion (VC) aims to modify source timbre to match unseen speaker . existing zero- shot VC systems struggle to reproduce paralinguistic information in highly expressive speech . |
| Approach: | They propose a framework for expressive zero-shot voice conversion that uses hybrid content encoding and memory-augmented context-aware timbre modeling. |
| Outcome: | The proposed framework surpasses state-of-the-art VC systems in speech naturalness, speaker similarity, and speaker similarness. |