Empathic Machines: Using Intermediate Features as Levers to Emulate Emotions in Text-To-Speech Systems (2022.naacl-main)
Copied to clipboard
| Challenge: | a method to control affective prosody of text-to-speech systems is proposed to use phoneme-level intermediate features as levers . DS is used to disentangle features relating to affective proody from those due to acoustics conditions and speaker identity . |
| Approach: | They propose a method to control the emotional prosody of Text to Speech systems by using phoneme-level intermediate features as levers. |
| Outcome: | The proposed method improves over the prior art in emulating emotion in speech . it adds the much-coveted "human touch" in machine dialogue, the authors say . |
Similar Papers
Prosody-TTS: Improving Prosody with Masked Autoencoder and Conditional Diffusion Model For Expressive Text-to-Speech (2023.findings-acl)
Copied to clipboard
| Challenge: | Expressive text-to-speech aims to generate high-quality samples with rich prosody . prosodic attributes in highly dynamic voices are difficult to capture and model without intonation . |
| Approach: | They propose a pipeline that enhances prosody modeling and sampling by introducing a self-supervised masked autoencoder and a diffusion model to sample diverse prosodic patterns within the latent space. |
| Outcome: | The proposed pipeline achieves new state-of-the-art in text-to-speech with natural and expressive synthesis. |
DPP-TTS: Diversifying prosodic features of speech via determinantal point processes (2023.emnlp-main)
Copied to clipboard
| Challenge: | Recent advances in deep generative models have succeeded in synthesizing human-like speech. |
| Approach: | They propose a text-to-speech model with a prosody diversifying module that considers perceptual diversity in each sample and among multiple samples. |
| Outcome: | The proposed model generates speech samples with more diversified prosody than baselines in the side-by-side comparison test considering the naturalness of speech at the same time. |
Emp-RFT: Empathetic Response Generation via Recognizing Feature Transitions between Utterances (2022.naacl-main)
Copied to clipboard
| Challenge: | Existing approaches for recognizing feature transitions between utterances extract features for the context at the coarse-grained level. |
| Approach: | They propose a method to recognize feature transitions between utterances that helps understand dialogue flow . they propose empathetic response generation strategy to focus on emotion and keywords related to appropriate features when generating responses. |
| Outcome: | The proposed approach outperforms baseline approaches and improves on multi-turn dialogues. |
TED-TTS: Training-Free Intra-Utterance Emotion and Duration Control for Text-to-Speech Synthesis (2026.acl-long)
Copied to clipboard
| Challenge: | Existing controllable Text-to-Speech methods limited to inter-utterance-level control . utterance expressiveness remains a challenge in building human-like TTS synthesis systems . |
| Approach: | They propose a training-free controllable framework for pretrained zero-shot TTS to enable intra-utterance emotion and duration expression. |
| Outcome: | The proposed framework achieves state-of-the-art intra-utterance consistency while maintaining baseline-level speech quality. |
ES4R: Speech Encoding Based on Prepositive Affective Modeling for Empathetic Response Generation (2026.acl-long)
Copied to clipboard
| Challenge: | Existing speech-to-speech large language models rely on ASR transcription or use encoders to extract latent representations, weakening affective information and contextual coherence in multi-turn dialogues. |
| Approach: | They propose a framework for speech-based empathetic response generation that captures turn-level affective states and dialogue-level emotional dynamics. |
| Outcome: | The proposed framework outperforms baselines in automatic and human evaluations and remains robust across different Large Language Model (LLM) backbones. |
Emotags: Computer-Assisted Verbal Labelling of Expressive Audiovisual Utterances for Expressive Multimodal TTS (2024.lrec-main)
Copied to clipboard
| Challenge: | We show that ascribing verbal descriptions to expressive audiovisual utterances is efficient and efficient. |
| Approach: | They propose a web app for ascribing verbal descriptions to expressive audiovisual utterances. |
| Outcome: | The proposed system can be deployed at a large scale to efficiently collect relevant verbal descriptions. |
From Personas to Talks: Revisiting the Impact of Personas on LLM-Synthesized Emotional Support Conversations (2025.emnlp-main)
Copied to clipboard
| Challenge: | Experimental results show that LLMs can infer persona traits and subtle shifts in emotionality and extraversion occur . scalable solutions with reduced costs and enhanced data privacy are needed . |
| Approach: | They explore the role of personas in the creation of emotional support conversations by LLMs. |
| Outcome: | The proposed model can infer persona traits and maintain key persona characteristics while revealing shifts in emotionality and extraversion. |
Prosody in Cascade and Direct Speech-to-Text Translation: a case study on Korean Wh-Phrases (2024.findings-eacl)
Copied to clipboard
| Challenge: | Existing direct S2TT systems have been unable to disambiguate utterances where prosody plays a crucial role. |
| Approach: | They propose to use contrastive evaluation to quantitatively measure the ability of direct S2TT systems to disambiguate utterances where prosody plays a crucial role. |
| Outcome: | The proposed system improves overall accuracy 12.9% and improves intent scores 15.6%. |
ProsodyFlow: High-fidelity Text-to-Speech through Conditional Flow Matching and Prosody Modeling with Large Speech Language Models (2025.coling-main)
Copied to clipboard
| Challenge: | Text-to-speech (TTS) models have been developed to generate high-quality speech. |
| Approach: | They propose an end-to-end TTS model that integrates large self-supervised speech models and conditional flow matching to model prosodic features effectively. |
| Outcome: | The proposed model improves synthesis quality and efficiency compared to existing models, showing that it generates more prosodic and expressive speech synthesizing. |
MoEL: Mixture of Empathetic Listeners (D19-1)
Copied to clipboard
| Challenge: | Neural network approaches for conversation models have shown to be successful in generating fluent and relevant responses. |
| Approach: | They propose a novel end-to-end approach for modeling empathy in dialogue systems by using Mixture of Empathetic Listeners (MoEL). |
| Outcome: | The proposed model outperforms multitask training baseline in terms of empathy, relevance, and fluency. |