Voice Builder: A Tool for Building Text-To-Speech Voices (L18-1)

Copied to clipboard

Challenge: a text-to-speech voice building tool is available for low-resourced languages . the tool allows researchers to run voice training experiments and listen to the resulting voice .
Approach: They propose an opensource text-to-speech (TTS) voice building tool that focuses on simplicity, flexibility, and collaboration.
Outcome: The proposed tool can help improve TTS research especially for low-resourced languages . it can be used to run voice training experiments and listen to the resulting synthesized voice .

Similar Papers

SpeechWeave: Diverse Multilingual Synthetic Text & Audio Data Generation Pipeline for Training Text to Speech Models (2025.acl-industry)

Copied to clipboard

Challenge: Text-to-Speech (TTS) training requires extensive and diverse text and speech data.
Approach: They propose a synthetic speech data generation pipeline that generates multilingual, domain-specific datasets for TTS training.
Outcome: The proposed pipeline generates data that is 10–48% more diverse than baseline across various linguistic and phonetic metrics, along with speaker-standardized speech audio while generating approximately 97% correctly normalized text.
Scaling Under-Resourced TTS: A Data-Optimized Framework with Advanced Acoustic Modeling for Thai (2025.acl-industry)

Copied to clipboard

Challenge: Text-to-speech (TTS) systems are limited by limited data and linguistic complexities.
Approach: They propose a data-optimized framework with an advanced acoustic model to build high-quality TTS systems for low-resource scenarios.
Outcome: The proposed framework enables zero-shot voice cloning and improved performance across diverse client applications, including finance, healthcare, education, and law.
Creating New Language and Voice Components for the Updated MaryTTS Text-to-Speech Synthesis Platform (L18-1)

Copied to clipboard

Challenge: a reboot of the MaryTTS system became unavoidable due to the number of people who have contributed to its development over the years.
Approach: They propose a workflow to create components for the MaryTTS text-to-speech synthesis platform.
Outcome: The proposed workflow is compatible with the updated MaryTTS architecture, enabling new features and state-of-the-art paradigms such as synthesis based on deep neural networks (DNNs).
Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey (2025.emnlp-main)

Copied to clipboard

Challenge: Text-to-speech (TTS) has advanced from generating natural-sounding speech to enabling fine-grained control over speech attributes.
Approach: They provide a review of controllable TTS methods from traditional control techniques to emerging approaches using natural language prompts.
Outcome: The proposed methods are based on models, strategies, and features, and summarize challenges, datasets, and evaluations.
VoiceCraft-X: Unifying Multilingual, Voice-Cloning Speech Synthesis and Speech Editing (2025.emnlp-main)

Copied to clipboard

Challenge: Autoregressive language model for multilingual speech editing and zero-shot text-to-speech synthesis is available in 11 languages.
Approach: They introduce an autoregressive neural codec language model which unifies multilingual speech editing and zero-shot text-to-speech synthesis across 11 languages.
Outcome: The model generates high-quality, natural-sounding speech, even with limited per-language data . it shows robust performance in diverse linguistic settings, even in limited per language data compared to other models .
ProsodyFlow: High-fidelity Text-to-Speech through Conditional Flow Matching and Prosody Modeling with Large Speech Language Models (2025.coling-main)

Copied to clipboard

Challenge: Text-to-speech (TTS) models have been developed to generate high-quality speech.
Approach: They propose an end-to-end TTS model that integrates large self-supervised speech models and conditional flow matching to model prosodic features effectively.
Outcome: The proposed model improves synthesis quality and efficiency compared to existing models, showing that it generates more prosodic and expressive speech synthesizing.
Low-Resource Multilingual and Zero-Shot Multispeaker TTS (2022.aacl-main)

Copied to clipboard

Challenge: Currently, the amount of data needed for TTS is limited to the vast majority of the spoken languages.
Approach: They propose to use language agnostic meta learning procedure to learn speaking a new language with just 5 minutes of training data while retaining the ability to infer the voice of even unseen speakers.
Outcome: The proposed approach is able to learn speaking a new language using just 5 minutes of training data while retaining the ability to infer the voice of even unseen speakers in the newly learned language.
An Automated End-to-End Open-Source Software for High-Quality Text-to-Speech Dataset Generation (2024.lrec-main)

Copied to clipboard

Challenge: Text-to-speech (TTS) models require data availability and quality of training data.
Approach: They propose an end-to-end tool to generate high-quality datasets for text-to speech models . language-specific phoneme distribution is integrated into sample selection, they argue .
Outcome: The proposed tool aims to streamline the dataset creation process for voice-based technologies by integrating language-specific phonemes into sample selection and quality assurance of recordings.
Tutorial: End-to-End Speech Translation (2021.eacl-tutorials)

Copied to clipboard

Challenge: Speech translation is the translation of speech in one language typically to text in another, traditionally accomplished through a combination of automatic speech recognition and machine translation.
Approach: This tutorial introduces the techniques used in cutting-edge research on speech translation.
Outcome: The proposed models achieve state-of-the-art performance with end-to-end speech translation for both high- and low-resource languages.
Make-A-Voice: Revisiting Voice Large Language Models as Scalable Multilingual and Multitask Learners (2024.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) have been used for general-purpose interfaces across multiple tasks and languages.
Approach: They propose to use large language models as a general-purpose interface across multiple tasks and languages.
Outcome: The proposed model performs better on 200K hours of 6-language data for voice generation applications.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations