Audiobook Dialogues as Training Data for Conversational Style Synthetic Voices (2022.lrec-1)
Copied to clipboard
Liisi Piits, Hille Pajupuu, Heete Sahkai, Rene Altrov, Liis Ermus, Kairi Tamuri, Indrek Hein, Meelis Mihkla, Indrek Kiissel, Egert Männisalu, Kristjan Suluste, Jaan Pajupuu
| Challenge: | Synthetic voices are increasingly used in applications that require a conversational speaking style. |
| Approach: | They compare voices trained on audiobook character speech corpus, audiobook narrator speech corpu and neutral-style sentence-based corpus . they conclude that the character speech and neutral style corpus are more suitable . |
| Outcome: | The evaluation of voices trained on three corpora of equal size was conducted by voice chatbots . the results may have been confounded by the greater acoustic variability and poorer phonemic coverage . |
Similar Papers
Investigating Inter- and Intra-speaker Voice Conversion using Audiobooks (2022.lrec-1)
Copied to clipboard
| Challenge: | Audiobook readers play with their voices to emphasize some text passages, highlight discourse changes or significant events, or in order to make listening easier and entertaining. |
| Approach: | They propose to modify the narrator’s voice to fit the context of the story, such as the character who is speaking, using voice conversion. |
| Outcome: | The proposed method improves the quality of the voice conversion system and the speaker similarity. |
Lightweight Transformers for Conversational AI (2022.naacl-industry)
Copied to clipboard
| Challenge: | Commercial dialogue systems typically require a small footprint and fast execution time, but recent trends are in the other direction, resulting in difficulties in model deployment. |
| Approach: | They build Transformer-based Language Models from scratch on large corpora of conversational data and compare their performance against BERT and other strong baselines on dialogue probing tasks. |
| Outcome: | The proposed model outperforms existing models on dialogue probing tasks and can be fine-tuned on a single consumer GPU card. |
SynPaFlex-Corpus: An Expressive French Audiobooks Corpus dedicated to expressive speech synthesis. (L18-1)
Copied to clipboard
| Challenge: | a French audiobooks corpus contains 87 hours of good audio quality speech . audiobooks provide mono-genre and multi-speaker speech whereas audiobooks usually provide a few hours of mono- and multispeakers . |
| Approach: | They present an expressive French audiobooks corpus containing eighty seven hours of speech . the corpus is annotated automatically and provides information as phone labels, phone boundaries, syllables, words or morpho-syntactic tagging. |
| Outcome: | The proposed corpus contains 87 hours of speech recorded by a single speaker . the data will allow developing models to better control expressiveness in speech synthesis . |
Book2Dial: Generating Teacher Student Interactions from Textbooks for Cost-Effective Development of Educational Chatbots (2024.findings-acl)
Copied to clipboard
| Challenge: | Educational chatbots are a promising tool for assisting student learning, but high-quality data is difficult to obtain due to privacy concerns. |
| Approach: | They propose a framework for generating synthetic teacher-student interactions grounded in a set of textbooks and propose to open-source their results. |
| Outcome: | The proposed framework captures a key aspect of learning interactions where curious students with partial knowledge ask teachers questions about the material in the textbook. |
Faithful Persona-based Conversational Dataset Generation with Large Language Models (2024.findings-acl)
Copied to clipboard
| Challenge: | Existing datasets for training conversational AI models do not sufficiently model their users. |
| Approach: | They propose a generator-critic architecture framework to expand the initial dataset while improving the quality of its conversations. |
| Outcome: | The proposed framework expands the initial dataset while improving the quality of its conversations. |
On Generative Spoken Language Modeling from Raw Audio (2021.tacl-1)
Copied to clipboard
Kushal Lakhotia, Eugene Kharitonov, Wei-Ning Hsu, Yossi Adi, Adam Polyak, Benjamin Bolte, Tu-Anh Nguyen, Jade Copet, Alexei Baevski, Abdelrahman Mohamed, Emmanuel Dupoux
| Challenge: | Using a set of metrics to evaluate the learned representations, we aim to create a system that learns from natural interactions as infants learn their first language. |
| Approach: | They propose a task of learning acoustic and linguistic characteristics from raw audio and a set of metrics to evaluate the learned representations at acustic, linguistic and encoding levels. |
| Outcome: | The proposed models evaluate the learned representations at acoustic and linguistic levels for both encoding and generation. |
Towards Neural Speaker Modeling in Multi-Party Conversation: The Task, Dataset, and Models (L18-1)
Copied to clipboard
| Challenge: | Existing methods for speaker modeling are based on hand-crafted statistics and ad hoc to a certain application. |
| Approach: | They propose to use speaker classification as a surrogate task for general speaker modeling and collect massive data to facilitate research in this direction. |
| Outcome: | The proposed models outperform the existing models and are feasible with speaker identity information. |
A Dynamic Speaker Model for Conversational Interactions (N19-1)
Copied to clipboard
| Challenge: | a neural model for characterizing individual differences in speakers is shown to be useful in human-computer interaction and dialog act prediction. |
| Approach: | They propose a neural model for learning a dynamically updated speaker embedding in a conversational context. |
| Outcome: | The proposed model is used for content ranking and dialog act prediction in human-human conversations. |
DIRECT: Direct and Indirect Responses in Conversational Text Corpus (2021.findings-emnlp)
Copied to clipboard
| Challenge: | Neural conversation models have been able to generate fluent responses through training on a dialogue corpus, but they lack the ability to reveal the implied intentions of users. |
| Approach: | They propose to train neural conversation models on a dialogue corpus that provides pragmatic paraphrases to advance techniques for natural language understanding in dialogue systems. |
| Outcome: | The proposed corpus provides 71,498 pairs of indirect–direct utterance pairs accompanied by a multi-turn dialogue history extracted from the MultiWoZ dataset. |
Finding A Voice: Exploring the Potential of African American Dialect and Voice Generation for Chatbots (2025.acl-long)
Copied to clipboard
| Challenge: | This study examines how linguistic similarity affects chatbot performance, focusing on integrating African American English (AAE) into virtual agents to better serve the African American community. |
| Approach: | They develop text-based and spoken chatbots using large language models and text-to-speech technology and evaluate them with AAE speakers to better serve the African American community. |
| Outcome: | The proposed language-based chatbots with African American English speakers outperform standard English chatbot models and show that spoken chatbot features improve performance and preference. |