Challenge: Synthetic voices are increasingly used in applications that require a conversational speaking style.
Approach: They compare voices trained on audiobook character speech corpus, audiobook narrator speech corpu and neutral-style sentence-based corpus . they conclude that the character speech and neutral style corpus are more suitable .
Outcome: The evaluation of voices trained on three corpora of equal size was conducted by voice chatbots . the results may have been confounded by the greater acoustic variability and poorer phonemic coverage .

Similar Papers

Investigating Inter- and Intra-speaker Voice Conversion using Audiobooks (2022.lrec-1)

Copied to clipboard

Challenge: Audiobook readers play with their voices to emphasize some text passages, highlight discourse changes or significant events, or in order to make listening easier and entertaining.
Approach: They propose to modify the narrator’s voice to fit the context of the story, such as the character who is speaking, using voice conversion.
Outcome: The proposed method improves the quality of the voice conversion system and the speaker similarity.
Lightweight Transformers for Conversational AI (2022.naacl-industry)

Copied to clipboard

Challenge: Commercial dialogue systems typically require a small footprint and fast execution time, but recent trends are in the other direction, resulting in difficulties in model deployment.
Approach: They build Transformer-based Language Models from scratch on large corpora of conversational data and compare their performance against BERT and other strong baselines on dialogue probing tasks.
Outcome: The proposed model outperforms existing models on dialogue probing tasks and can be fine-tuned on a single consumer GPU card.
SynPaFlex-Corpus: An Expressive French Audiobooks Corpus dedicated to expressive speech synthesis. (L18-1)

Copied to clipboard

Challenge: a French audiobooks corpus contains 87 hours of good audio quality speech . audiobooks provide mono-genre and multi-speaker speech whereas audiobooks usually provide a few hours of mono- and multispeakers .
Approach: They present an expressive French audiobooks corpus containing eighty seven hours of speech . the corpus is annotated automatically and provides information as phone labels, phone boundaries, syllables, words or morpho-syntactic tagging.
Outcome: The proposed corpus contains 87 hours of speech recorded by a single speaker . the data will allow developing models to better control expressiveness in speech synthesis .
Book2Dial: Generating Teacher Student Interactions from Textbooks for Cost-Effective Development of Educational Chatbots (2024.findings-acl)

Copied to clipboard

Challenge: Educational chatbots are a promising tool for assisting student learning, but high-quality data is difficult to obtain due to privacy concerns.
Approach: They propose a framework for generating synthetic teacher-student interactions grounded in a set of textbooks and propose to open-source their results.
Outcome: The proposed framework captures a key aspect of learning interactions where curious students with partial knowledge ask teachers questions about the material in the textbook.
Faithful Persona-based Conversational Dataset Generation with Large Language Models (2024.findings-acl)

Copied to clipboard

Challenge: Existing datasets for training conversational AI models do not sufficiently model their users.
Approach: They propose a generator-critic architecture framework to expand the initial dataset while improving the quality of its conversations.
Outcome: The proposed framework expands the initial dataset while improving the quality of its conversations.
On Generative Spoken Language Modeling from Raw Audio (2021.tacl-1)

Copied to clipboard

Challenge: Using a set of metrics to evaluate the learned representations, we aim to create a system that learns from natural interactions as infants learn their first language.
Approach: They propose a task of learning acoustic and linguistic characteristics from raw audio and a set of metrics to evaluate the learned representations at acustic, linguistic and encoding levels.
Outcome: The proposed models evaluate the learned representations at acoustic and linguistic levels for both encoding and generation.
Towards Neural Speaker Modeling in Multi-Party Conversation: The Task, Dataset, and Models (L18-1)

Copied to clipboard

Challenge: Existing methods for speaker modeling are based on hand-crafted statistics and ad hoc to a certain application.
Approach: They propose to use speaker classification as a surrogate task for general speaker modeling and collect massive data to facilitate research in this direction.
Outcome: The proposed models outperform the existing models and are feasible with speaker identity information.
A Dynamic Speaker Model for Conversational Interactions (N19-1)

Copied to clipboard

Challenge: a neural model for characterizing individual differences in speakers is shown to be useful in human-computer interaction and dialog act prediction.
Approach: They propose a neural model for learning a dynamically updated speaker embedding in a conversational context.
Outcome: The proposed model is used for content ranking and dialog act prediction in human-human conversations.
DIRECT: Direct and Indirect Responses in Conversational Text Corpus (2021.findings-emnlp)

Copied to clipboard

Challenge: Neural conversation models have been able to generate fluent responses through training on a dialogue corpus, but they lack the ability to reveal the implied intentions of users.
Approach: They propose to train neural conversation models on a dialogue corpus that provides pragmatic paraphrases to advance techniques for natural language understanding in dialogue systems.
Outcome: The proposed corpus provides 71,498 pairs of indirect–direct utterance pairs accompanied by a multi-turn dialogue history extracted from the MultiWoZ dataset.
Finding A Voice: Exploring the Potential of African American Dialect and Voice Generation for Chatbots (2025.acl-long)

Copied to clipboard

Challenge: This study examines how linguistic similarity affects chatbot performance, focusing on integrating African American English (AAE) into virtual agents to better serve the African American community.
Approach: They develop text-based and spoken chatbots using large language models and text-to-speech technology and evaluate them with AAE speakers to better serve the African American community.
Outcome: The proposed language-based chatbots with African American English speakers outperform standard English chatbot models and show that spoken chatbot features improve performance and preference.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations