Challenge: Recent advances in text-to-speech synthesis have enabled researchers to generate high-quality speech with a wide range of prosodic variations.
Approach: They propose a model that can synthesise newscaster-style speech with a few hours of data . they propose to factor in contextual word embeddings and evaluate it against neutral synthesis .
Outcome: The proposed model can synthesise newscaster-style speech with just a few hours of data.

Similar Papers

FillerSpeech: Towards Human-Like Text-to-Speech Synthesis with Filler Insertion and Filler Style Control (2025.emnlp-main)

Copied to clipboard

Challenge: Recent advances in speech synthesis have improved audio quality and pronunciation . fillers are an integral part of natural human conversation, but achieving human-like conversational speech remains a challenge.
Approach: They propose a speech synthesis framework that enables natural filler insertion and style control . they propose 'filler-inclusive' speech data that includes fillers with pitch and duration information .
Outcome: The proposed framework enables natural filler insertion and style control . the proposed framework is validated and can be used to predict filler style .
Effectiveness of Data Augmentation and Pretraining for Improving Neural Headline Generation in Low-Resource Settings (2022.lrec-1)

Copied to clipboard

Challenge: Neural approaches for natural language generation (NLG) have mushroomed due to large textual resources.
Approach: They propose to use a pretrained multilingual encoder-decoder model and a combination of two pretrained language models to train a model in a low-resource setting.
Outcome: The proposed model outperforms the previous model on English and on a small subset of the same data.
Inference Time Style Control for Summarization (2021.naacl-main)

Copied to clipboard

Challenge: Existing methods to generate summaries of different styles without training separate models are lacking parallel data and expensive (re)training.
Approach: They propose two methods that can be deployed during summary decoding on any pre-trained Transformer-based summarization model.
Outcome: The proposed methods generate news headlines with various ideological leanings while still informative.
Audiobook Dialogues as Training Data for Conversational Style Synthetic Voices (2022.lrec-1)

Copied to clipboard

Challenge: Synthetic voices are increasingly used in applications that require a conversational speaking style.
Approach: They compare voices trained on audiobook character speech corpus, audiobook narrator speech corpu and neutral-style sentence-based corpus . they conclude that the character speech and neutral style corpus are more suitable .
Outcome: The evaluation of voices trained on three corpora of equal size was conducted by voice chatbots . the results may have been confounded by the greater acoustic variability and poorer phonemic coverage .
StyleTTS-ZS: Efficient High-Quality Zero-Shot Text-to-Speech Synthesis with Distilled Time-Varying Style Diffusion (2025.naacl-long)

Copied to clipboard

Challenge: Recent advances in text-to-speech (TTS) models have led to improvements in speaker prosody and voices modeling.
Approach: They propose an efficient zero-shot TTS model that leverages distilled time-varying style diffusion to capture diverse speaker identities and prosodies.
Outcome: The proposed model surpasses state-of-the-art models in both naturalness and similarity while reducing inference speed by 90%.
Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey (2025.emnlp-main)

Copied to clipboard

Challenge: Text-to-speech (TTS) has advanced from generating natural-sounding speech to enabling fine-grained control over speech attributes.
Approach: They provide a review of controllable TTS methods from traditional control techniques to emerging approaches using natural language prompts.
Outcome: The proposed methods are based on models, strategies, and features, and summarize challenges, datasets, and evaluations.
Text-Free Image-to-Speech Synthesis Using Learned Segmental Units (2021.acl-long)

Copied to clipboard

Challenge: Existing models for synthesising fluent, natural-sounding spoken audio captions do not require natural language text as an intermediate representation or source of supervision.
Approach: They propose a model for directly synthesizing fluent, natural-sounding spoken audio captions for images that does not require natural language text as an intermediate representation or source of supervision.
Outcome: The proposed model captures diverse visual semantics of images and can replace text with a set of discrete, sub-word speech units.
ProsodyFlow: High-fidelity Text-to-Speech through Conditional Flow Matching and Prosody Modeling with Large Speech Language Models (2025.coling-main)

Copied to clipboard

Challenge: Text-to-speech (TTS) models have been developed to generate high-quality speech.
Approach: They propose an end-to-end TTS model that integrates large self-supervised speech models and conditional flow matching to model prosodic features effectively.
Outcome: The proposed model improves synthesis quality and efficiency compared to existing models, showing that it generates more prosodic and expressive speech synthesizing.
Sentence-Level Content Planning and Style Specification for Neural Text Generation (D19-1)

Copied to clipboard

Challenge: Recent advances in text generation systems often produce incoherent and unfaithful outputs . a novel automated text generation system takes into account content selection, text planning, and surface realization.
Approach: They propose an end-to-end trained two-step text generation model that considers sentence-level content planners and language styles.
Outcome: The proposed model outperforms competing models in three domains with diverse topics and varying language styles.
A Case Study on Neural Headline Generation for Editing Support (N19-2)

Copied to clipboard

Challenge: a news-aggregator is a website or mobile application that aggregates web content . dozens of professional editors manually create their headlines, which are much shorter than the original headlines.
Approach: They propose a neural headline generation model that automatically generates short headlines from news articles.
Outcome: The proposed model is deployed to an editing support tool and compares editors' behavior before and after the release.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations