Challenge: Audiobook readers play with their voices to emphasize some text passages, highlight discourse changes or significant events, or in order to make listening easier and entertaining.
Approach: They propose to modify the narrator’s voice to fit the context of the story, such as the character who is speaking, using voice conversion.
Outcome: The proposed method improves the quality of the voice conversion system and the speaker similarity.

Similar Papers

Audiobook Dialogues as Training Data for Conversational Style Synthetic Voices (2022.lrec-1)

Copied to clipboard

Challenge: Synthetic voices are increasingly used in applications that require a conversational speaking style.
Approach: They compare voices trained on audiobook character speech corpus, audiobook narrator speech corpu and neutral-style sentence-based corpus . they conclude that the character speech and neutral style corpus are more suitable .
Outcome: The evaluation of voices trained on three corpora of equal size was conducted by voice chatbots . the results may have been confounded by the greater acoustic variability and poorer phonemic coverage .
Computational Narrative Understanding for Expressive Text-to-Speech (2026.findings-acl)

Copied to clipboard

Challenge: Recent advances in text-to-speech systems have been driven by large, multi-domain speech corpora.
Approach: They propose a large-scale 5.3K hours of expressive speech drawn from character quotations . they fine-tune a flow-matching model and train from scratch .
Outcome: The proposed model improves expressivity and intelligibility while training from scratch improves expressiveness of an autoregressive model.
What does the sea say to the shore? A BERT based DST style approach for speaker to dialogue attribution in novels (2022.acl-long)

Copied to clipboard

Challenge: Existing systems can perform the first two tasks accurately, but attributing characters to direct speech is a challenging problem due to the narrator’s lack of explicit character mentions and the frequent use of nominal and pronominal coreference when such explicit mentions are made.
Approach: They propose a pipeline to extract characters and link them to their direct-speech utterances by using a novel's list of characters and a list of attributing them to the speaking characters.
Outcome: The proposed pipeline improves state-of-the-art models by 50% in F1-score compared with existing models .
SynPaFlex-Corpus: An Expressive French Audiobooks Corpus dedicated to expressive speech synthesis. (L18-1)

Copied to clipboard

Challenge: a French audiobooks corpus contains 87 hours of good audio quality speech . audiobooks provide mono-genre and multi-speaker speech whereas audiobooks usually provide a few hours of mono- and multispeakers .
Approach: They present an expressive French audiobooks corpus containing eighty seven hours of speech . the corpus is annotated automatically and provides information as phone labels, phone boundaries, syllables, words or morpho-syntactic tagging.
Outcome: The proposed corpus contains 87 hours of speech recorded by a single speaker . the data will allow developing models to better control expressiveness in speech synthesis .
Reducing Sensitivity on Speaker Names for Text Generation from Dialogues (2023.findings-acl)

Copied to clipboard

Challenge: Pre-trained language models are sensitive to nuances, resulting in unfairness in real-world applications.
Approach: They propose to quantitatively measure a model's sensitivity on speaker names and comprehensively evaluate a number of known methods for reducing speaker name sensitivity.
Outcome: The proposed approach reduces speaker name sensitivity and improves quality of generation.
Enhancing Polyglot Voices by Leveraging Cross-Lingual Fine-Tuning in Any-to-One Voice Conversion (2024.findings-emnlp)

Copied to clipboard

Challenge: Recent advances in speech synthesis have improved the quality of polyglot voices.
Approach: They propose a cross-lingual any-to-one voice conversion system that preserves the source accent without multilingual data from the target speaker.
Outcome: The proposed system preserves source accent without multilingual data from target speaker and reduces training data requirements.
Exploring Cross-Lingual Voice Conversion Methods for Anonymizing Low-Resource Text-to-Speech (2026.eacl-short)

Copied to clipboard

Challenge: a growing number of speech synthesis systems clone a person's voice, a new study finds . a variety of voice conversion techniques can mask speaker identities in low-resource text-to-speech systems.
Approach: They compare voice conversion techniques to mask speaker identities in text-to-speech systems . they build and evaluate speaker-anonymized systems for two Canadian Indigenous languages .
Outcome: The proposed methods are compared with other approaches for using voice conversion to mask speaker identities in low-resource text-to-speech systems.
Dubbing in Practice: A Large Scale Study of Human Localization With Insights for Automatic Dubbing (2023.tacl-1)

Copied to clipboard

Challenge: a large-scale study of human dubbing in practice is lacking in qualitative literature on human dubs . authors argue for vocal naturalness and translation quality over isometric constraints . a data-driven examination of the way humans perform this task is needed .
Approach: They analyze 319.57 hours of video from 54 professionally produced titles . they argue for vocal naturalness and translation quality over isometric constraints . authors say they need to preserve speech characteristics and transfer of semantic properties .
Outcome: The study challenges assumptions in qualitative and machine-learning literature on dubbing . it also finds that source-side audio influences human dubbing through other channels .
Improve Speech Translation Through Text Rewrite (2025.coling-industry)

Copied to clipboard

Challenge: Recent advances in speech translation (ST) research have focused on the unique characteristics of spontaneous speech, including accents and presentation quality.
Approach: They propose to transform transcribed speech into a cleaner style more in line with the expectations of translation models built from written text.
Outcome: Experiments on public and in-house translation models show that the proposed model can be effectively distilled into a standalone translation model.
Unsupervised Multi-scale Expressive Speaking Style Modeling with Hierarchical Context Information for Audiobook Speech Synthesis (2022.coling-1)

Copied to clipboard

Challenge: a recent study has shown that expressiveness of audiobooks is limited by the averaged global-scale speaking style representation.
Approach: They propose an unsupervised multi-scale context-sensitive text-to-speech model for audiobooks . they use hierarchical context encoder to predict global-scale contextual style embeddings .
Outcome: The proposed model outperforms existing models on a real-world Mandarin audio dataset.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations