| Challenge: | a text-to-speech voice building tool is available for low-resourced languages . the tool allows researchers to run voice training experiments and listen to the resulting voice . |
| Approach: | They propose an opensource text-to-speech (TTS) voice building tool that focuses on simplicity, flexibility, and collaboration. |
| Outcome: | The proposed tool can help improve TTS research especially for low-resourced languages . it can be used to run voice training experiments and listen to the resulting synthesized voice . |
Similar Papers
SpeechWeave: Diverse Multilingual Synthetic Text & Audio Data Generation Pipeline for Training Text to Speech Models (2025.acl-industry)
Copied to clipboard
| Challenge: | Text-to-Speech (TTS) training requires extensive and diverse text and speech data. |
| Approach: | They propose a synthetic speech data generation pipeline that generates multilingual, domain-specific datasets for TTS training. |
| Outcome: | The proposed pipeline generates data that is 10–48% more diverse than baseline across various linguistic and phonetic metrics, along with speaker-standardized speech audio while generating approximately 97% correctly normalized text. |
Scaling Under-Resourced TTS: A Data-Optimized Framework with Advanced Acoustic Modeling for Thai (2025.acl-industry)
Copied to clipboard
| Challenge: | Text-to-speech (TTS) systems are limited by limited data and linguistic complexities. |
| Approach: | They propose a data-optimized framework with an advanced acoustic model to build high-quality TTS systems for low-resource scenarios. |
| Outcome: | The proposed framework enables zero-shot voice cloning and improved performance across diverse client applications, including finance, healthcare, education, and law. |
Creating New Language and Voice Components for the Updated MaryTTS Text-to-Speech Synthesis Platform (L18-1)
Copied to clipboard
| Challenge: | a reboot of the MaryTTS system became unavoidable due to the number of people who have contributed to its development over the years. |
| Approach: | They propose a workflow to create components for the MaryTTS text-to-speech synthesis platform. |
| Outcome: | The proposed workflow is compatible with the updated MaryTTS architecture, enabling new features and state-of-the-art paradigms such as synthesis based on deep neural networks (DNNs). |
Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey (2025.emnlp-main)
Copied to clipboard
| Challenge: | Text-to-speech (TTS) has advanced from generating natural-sounding speech to enabling fine-grained control over speech attributes. |
| Approach: | They provide a review of controllable TTS methods from traditional control techniques to emerging approaches using natural language prompts. |
| Outcome: | The proposed methods are based on models, strategies, and features, and summarize challenges, datasets, and evaluations. |
VoiceCraft-X: Unifying Multilingual, Voice-Cloning Speech Synthesis and Speech Editing (2025.emnlp-main)
Copied to clipboard
Zhisheng Zheng, Puyuan Peng, Anuj Diwan, Cong Phuoc Huynh, Xiaohang Sun, Zhu Liu, Vimal Bhat, David Harwath
| Challenge: | Autoregressive language model for multilingual speech editing and zero-shot text-to-speech synthesis is available in 11 languages. |
| Approach: | They introduce an autoregressive neural codec language model which unifies multilingual speech editing and zero-shot text-to-speech synthesis across 11 languages. |
| Outcome: | The model generates high-quality, natural-sounding speech, even with limited per-language data . it shows robust performance in diverse linguistic settings, even in limited per language data compared to other models . |
ProsodyFlow: High-fidelity Text-to-Speech through Conditional Flow Matching and Prosody Modeling with Large Speech Language Models (2025.coling-main)
Copied to clipboard
| Challenge: | Text-to-speech (TTS) models have been developed to generate high-quality speech. |
| Approach: | They propose an end-to-end TTS model that integrates large self-supervised speech models and conditional flow matching to model prosodic features effectively. |
| Outcome: | The proposed model improves synthesis quality and efficiency compared to existing models, showing that it generates more prosodic and expressive speech synthesizing. |
Low-Resource Multilingual and Zero-Shot Multispeaker TTS (2022.aacl-main)
Copied to clipboard
| Challenge: | Currently, the amount of data needed for TTS is limited to the vast majority of the spoken languages. |
| Approach: | They propose to use language agnostic meta learning procedure to learn speaking a new language with just 5 minutes of training data while retaining the ability to infer the voice of even unseen speakers. |
| Outcome: | The proposed approach is able to learn speaking a new language using just 5 minutes of training data while retaining the ability to infer the voice of even unseen speakers in the newly learned language. |
An Automated End-to-End Open-Source Software for High-Quality Text-to-Speech Dataset Generation (2024.lrec-main)
Copied to clipboard
Ahmet Gunduz, Kamer Ali Yuksel, Kareem Darwish, Golara Javadi, Fabio Minazzi, Nicola Sobieski, Sébastien Bratières
| Challenge: | Text-to-speech (TTS) models require data availability and quality of training data. |
| Approach: | They propose an end-to-end tool to generate high-quality datasets for text-to speech models . language-specific phoneme distribution is integrated into sample selection, they argue . |
| Outcome: | The proposed tool aims to streamline the dataset creation process for voice-based technologies by integrating language-specific phonemes into sample selection and quality assurance of recordings. |
Tutorial: End-to-End Speech Translation (2021.eacl-tutorials)
Copied to clipboard
| Challenge: | Speech translation is the translation of speech in one language typically to text in another, traditionally accomplished through a combination of automatic speech recognition and machine translation. |
| Approach: | This tutorial introduces the techniques used in cutting-edge research on speech translation. |
| Outcome: | The proposed models achieve state-of-the-art performance with end-to-end speech translation for both high- and low-resource languages. |
Make-A-Voice: Revisiting Voice Large Language Models as Scalable Multilingual and Multitask Learners (2024.acl-long)
Copied to clipboard
Rongjie Huang, Chunlei Zhang, Yongqi Wang, Dongchao Yang, Jinchuan Tian, Zhenhui Ye, Luping Liu, Zehan Wang, Ziyue Jiang, Xuankai Chang, Jiatong Shi, Chao Weng, Zhou Zhao, Dong Yu
| Challenge: | Large language models (LLMs) have been used for general-purpose interfaces across multiple tasks and languages. |
| Approach: | They propose to use large language models as a general-purpose interface across multiple tasks and languages. |
| Outcome: | The proposed model performs better on 200K hours of 6-language data for voice generation applications. |