Investigating Inter- and Intra-speaker Voice Conversion using Audiobooks (2022.lrec-1)
Copied to clipboard
| Challenge: | Audiobook readers play with their voices to emphasize some text passages, highlight discourse changes or significant events, or in order to make listening easier and entertaining. |
| Approach: | They propose to modify the narrator’s voice to fit the context of the story, such as the character who is speaking, using voice conversion. |
| Outcome: | The proposed method improves the quality of the voice conversion system and the speaker similarity. |
Similar Papers
Audiobook Dialogues as Training Data for Conversational Style Synthetic Voices (2022.lrec-1)
Copied to clipboard
Liisi Piits, Hille Pajupuu, Heete Sahkai, Rene Altrov, Liis Ermus, Kairi Tamuri, Indrek Hein, Meelis Mihkla, Indrek Kiissel, Egert Männisalu, Kristjan Suluste, Jaan Pajupuu
| Challenge: | Synthetic voices are increasingly used in applications that require a conversational speaking style. |
| Approach: | They compare voices trained on audiobook character speech corpus, audiobook narrator speech corpu and neutral-style sentence-based corpus . they conclude that the character speech and neutral style corpus are more suitable . |
| Outcome: | The evaluation of voices trained on three corpora of equal size was conducted by voice chatbots . the results may have been confounded by the greater acoustic variability and poorer phonemic coverage . |
Computational Narrative Understanding for Expressive Text-to-Speech (2026.findings-acl)
Copied to clipboard
| Challenge: | Recent advances in text-to-speech systems have been driven by large, multi-domain speech corpora. |
| Approach: | They propose a large-scale 5.3K hours of expressive speech drawn from character quotations . they fine-tune a flow-matching model and train from scratch . |
| Outcome: | The proposed model improves expressivity and intelligibility while training from scratch improves expressiveness of an autoregressive model. |
What does the sea say to the shore? A BERT based DST style approach for speaker to dialogue attribution in novels (2022.acl-long)
Copied to clipboard
| Challenge: | Existing systems can perform the first two tasks accurately, but attributing characters to direct speech is a challenging problem due to the narrator’s lack of explicit character mentions and the frequent use of nominal and pronominal coreference when such explicit mentions are made. |
| Approach: | They propose a pipeline to extract characters and link them to their direct-speech utterances by using a novel's list of characters and a list of attributing them to the speaking characters. |
| Outcome: | The proposed pipeline improves state-of-the-art models by 50% in F1-score compared with existing models . |
SynPaFlex-Corpus: An Expressive French Audiobooks Corpus dedicated to expressive speech synthesis. (L18-1)
Copied to clipboard
| Challenge: | a French audiobooks corpus contains 87 hours of good audio quality speech . audiobooks provide mono-genre and multi-speaker speech whereas audiobooks usually provide a few hours of mono- and multispeakers . |
| Approach: | They present an expressive French audiobooks corpus containing eighty seven hours of speech . the corpus is annotated automatically and provides information as phone labels, phone boundaries, syllables, words or morpho-syntactic tagging. |
| Outcome: | The proposed corpus contains 87 hours of speech recorded by a single speaker . the data will allow developing models to better control expressiveness in speech synthesis . |
Reducing Sensitivity on Speaker Names for Text Generation from Dialogues (2023.findings-acl)
Copied to clipboard
| Challenge: | Pre-trained language models are sensitive to nuances, resulting in unfairness in real-world applications. |
| Approach: | They propose to quantitatively measure a model's sensitivity on speaker names and comprehensively evaluate a number of known methods for reducing speaker name sensitivity. |
| Outcome: | The proposed approach reduces speaker name sensitivity and improves quality of generation. |
Enhancing Polyglot Voices by Leveraging Cross-Lingual Fine-Tuning in Any-to-One Voice Conversion (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Recent advances in speech synthesis have improved the quality of polyglot voices. |
| Approach: | They propose a cross-lingual any-to-one voice conversion system that preserves the source accent without multilingual data from the target speaker. |
| Outcome: | The proposed system preserves source accent without multilingual data from target speaker and reduces training data requirements. |
Exploring Cross-Lingual Voice Conversion Methods for Anonymizing Low-Resource Text-to-Speech (2026.eacl-short)
Copied to clipboard
| Challenge: | a growing number of speech synthesis systems clone a person's voice, a new study finds . a variety of voice conversion techniques can mask speaker identities in low-resource text-to-speech systems. |
| Approach: | They compare voice conversion techniques to mask speaker identities in text-to-speech systems . they build and evaluate speaker-anonymized systems for two Canadian Indigenous languages . |
| Outcome: | The proposed methods are compared with other approaches for using voice conversion to mask speaker identities in low-resource text-to-speech systems. |
Dubbing in Practice: A Large Scale Study of Human Localization With Insights for Automatic Dubbing (2023.tacl-1)
Copied to clipboard
| Challenge: | a large-scale study of human dubbing in practice is lacking in qualitative literature on human dubs . authors argue for vocal naturalness and translation quality over isometric constraints . a data-driven examination of the way humans perform this task is needed . |
| Approach: | They analyze 319.57 hours of video from 54 professionally produced titles . they argue for vocal naturalness and translation quality over isometric constraints . authors say they need to preserve speech characteristics and transfer of semantic properties . |
| Outcome: | The study challenges assumptions in qualitative and machine-learning literature on dubbing . it also finds that source-side audio influences human dubbing through other channels . |
Improve Speech Translation Through Text Rewrite (2025.coling-industry)
Copied to clipboard
| Challenge: | Recent advances in speech translation (ST) research have focused on the unique characteristics of spontaneous speech, including accents and presentation quality. |
| Approach: | They propose to transform transcribed speech into a cleaner style more in line with the expectations of translation models built from written text. |
| Outcome: | Experiments on public and in-house translation models show that the proposed model can be effectively distilled into a standalone translation model. |
Unsupervised Multi-scale Expressive Speaking Style Modeling with Hierarchical Context Information for Audiobook Speech Synthesis (2022.coling-1)
Copied to clipboard
| Challenge: | a recent study has shown that expressiveness of audiobooks is limited by the averaged global-scale speaking style representation. |
| Approach: | They propose an unsupervised multi-scale context-sensitive text-to-speech model for audiobooks . they use hierarchical context encoder to predict global-scale contextual style embeddings . |
| Outcome: | The proposed model outperforms existing models on a real-world Mandarin audio dataset. |