Challenge: a method to control affective prosody of text-to-speech systems is proposed to use phoneme-level intermediate features as levers . DS is used to disentangle features relating to affective proody from those due to acoustics conditions and speaker identity .
Approach: They propose a method to control the emotional prosody of Text to Speech systems by using phoneme-level intermediate features as levers.
Outcome: The proposed method improves over the prior art in emulating emotion in speech . it adds the much-coveted "human touch" in machine dialogue, the authors say .

Similar Papers

Prosody-TTS: Improving Prosody with Masked Autoencoder and Conditional Diffusion Model For Expressive Text-to-Speech (2023.findings-acl)

Copied to clipboard

Challenge: Expressive text-to-speech aims to generate high-quality samples with rich prosody . prosodic attributes in highly dynamic voices are difficult to capture and model without intonation .
Approach: They propose a pipeline that enhances prosody modeling and sampling by introducing a self-supervised masked autoencoder and a diffusion model to sample diverse prosodic patterns within the latent space.
Outcome: The proposed pipeline achieves new state-of-the-art in text-to-speech with natural and expressive synthesis.
DPP-TTS: Diversifying prosodic features of speech via determinantal point processes (2023.emnlp-main)

Copied to clipboard

Challenge: Recent advances in deep generative models have succeeded in synthesizing human-like speech.
Approach: They propose a text-to-speech model with a prosody diversifying module that considers perceptual diversity in each sample and among multiple samples.
Outcome: The proposed model generates speech samples with more diversified prosody than baselines in the side-by-side comparison test considering the naturalness of speech at the same time.
Emp-RFT: Empathetic Response Generation via Recognizing Feature Transitions between Utterances (2022.naacl-main)

Copied to clipboard

Challenge: Existing approaches for recognizing feature transitions between utterances extract features for the context at the coarse-grained level.
Approach: They propose a method to recognize feature transitions between utterances that helps understand dialogue flow . they propose empathetic response generation strategy to focus on emotion and keywords related to appropriate features when generating responses.
Outcome: The proposed approach outperforms baseline approaches and improves on multi-turn dialogues.
TED-TTS: Training-Free Intra-Utterance Emotion and Duration Control for Text-to-Speech Synthesis (2026.acl-long)

Copied to clipboard

Challenge: Existing controllable Text-to-Speech methods limited to inter-utterance-level control . utterance expressiveness remains a challenge in building human-like TTS synthesis systems .
Approach: They propose a training-free controllable framework for pretrained zero-shot TTS to enable intra-utterance emotion and duration expression.
Outcome: The proposed framework achieves state-of-the-art intra-utterance consistency while maintaining baseline-level speech quality.
ES4R: Speech Encoding Based on Prepositive Affective Modeling for Empathetic Response Generation (2026.acl-long)

Copied to clipboard

Challenge: Existing speech-to-speech large language models rely on ASR transcription or use encoders to extract latent representations, weakening affective information and contextual coherence in multi-turn dialogues.
Approach: They propose a framework for speech-based empathetic response generation that captures turn-level affective states and dialogue-level emotional dynamics.
Outcome: The proposed framework outperforms baselines in automatic and human evaluations and remains robust across different Large Language Model (LLM) backbones.
Emotags: Computer-Assisted Verbal Labelling of Expressive Audiovisual Utterances for Expressive Multimodal TTS (2024.lrec-main)

Copied to clipboard

Challenge: We show that ascribing verbal descriptions to expressive audiovisual utterances is efficient and efficient.
Approach: They propose a web app for ascribing verbal descriptions to expressive audiovisual utterances.
Outcome: The proposed system can be deployed at a large scale to efficiently collect relevant verbal descriptions.
From Personas to Talks: Revisiting the Impact of Personas on LLM-Synthesized Emotional Support Conversations (2025.emnlp-main)

Copied to clipboard

Challenge: Experimental results show that LLMs can infer persona traits and subtle shifts in emotionality and extraversion occur . scalable solutions with reduced costs and enhanced data privacy are needed .
Approach: They explore the role of personas in the creation of emotional support conversations by LLMs.
Outcome: The proposed model can infer persona traits and maintain key persona characteristics while revealing shifts in emotionality and extraversion.
Prosody in Cascade and Direct Speech-to-Text Translation: a case study on Korean Wh-Phrases (2024.findings-eacl)

Copied to clipboard

Challenge: Existing direct S2TT systems have been unable to disambiguate utterances where prosody plays a crucial role.
Approach: They propose to use contrastive evaluation to quantitatively measure the ability of direct S2TT systems to disambiguate utterances where prosody plays a crucial role.
Outcome: The proposed system improves overall accuracy 12.9% and improves intent scores 15.6%.
ProsodyFlow: High-fidelity Text-to-Speech through Conditional Flow Matching and Prosody Modeling with Large Speech Language Models (2025.coling-main)

Copied to clipboard

Challenge: Text-to-speech (TTS) models have been developed to generate high-quality speech.
Approach: They propose an end-to-end TTS model that integrates large self-supervised speech models and conditional flow matching to model prosodic features effectively.
Outcome: The proposed model improves synthesis quality and efficiency compared to existing models, showing that it generates more prosodic and expressive speech synthesizing.
MoEL: Mixture of Empathetic Listeners (D19-1)

Copied to clipboard

Challenge: Neural network approaches for conversation models have shown to be successful in generating fluent and relevant responses.
Approach: They propose a novel end-to-end approach for modeling empathy in dialogue systems by using Mixture of Empathetic Listeners (MoEL).
Outcome: The proposed model outperforms multitask training baseline in terms of empathy, relevance, and fluency.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations