Emotags: Computer-Assisted Verbal Labelling of Expressive Audiovisual Utterances for Expressive Multimodal TTS (2024.lrec-main)
Copied to clipboard
| Challenge: | We show that ascribing verbal descriptions to expressive audiovisual utterances is efficient and efficient. |
| Approach: | They propose a web app for ascribing verbal descriptions to expressive audiovisual utterances. |
| Outcome: | The proposed system can be deployed at a large scale to efficiently collect relevant verbal descriptions. |
Similar Papers
Computational Narrative Understanding for Expressive Text-to-Speech (2026.findings-acl)
Copied to clipboard
| Challenge: | Recent advances in text-to-speech systems have been driven by large, multi-domain speech corpora. |
| Approach: | They propose a large-scale 5.3K hours of expressive speech drawn from character quotations . they fine-tune a flow-matching model and train from scratch . |
| Outcome: | The proposed model improves expressivity and intelligibility while training from scratch improves expressiveness of an autoregressive model. |
When Words Smile: Generating Diverse Emotional Facial Expressions from Text (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing systems that generate only coarse facial expressions ignore the rich and dynamic nature of face-to-face communication. |
| Approach: | They propose an end-to-end text-to expression model that explicitly focuses on emotional dynamics. |
| Outcome: | The proposed model outperforms baselines on 15,000 text–3D expression pairs on a large-scale dataset. |
Prompt-Guided Selective Masking Loss for Context-Aware Emotive Text-to-Speech (2025.findings-naacl)
Copied to clipboard
| Challenge: | Emotional dialogue speech synthesis (EDSS) aims to generate expressive speech by leveraging the dialogue context between interlocutors. |
| Approach: | They propose a large language model to generate holistic emotion tags based on prior dialogue context and pinpoint key words in the target utterance that align with the predicted emotional state. |
| Outcome: | The proposed method improves emotional expressiveness and facilitates automatic emotion speech generation during inference. |
Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey (2025.emnlp-main)
Copied to clipboard
| Challenge: | Text-to-speech (TTS) has advanced from generating natural-sounding speech to enabling fine-grained control over speech attributes. |
| Approach: | They provide a review of controllable TTS methods from traditional control techniques to emerging approaches using natural language prompts. |
| Outcome: | The proposed methods are based on models, strategies, and features, and summarize challenges, datasets, and evaluations. |
Praaline: An Open-Source System for Managing, Annotating, Visualising and Analysing Speech Corpora (P18-4)
Copied to clipboard
| Challenge: | Praaline is an open-source software system for constituting and managing spoken language and multimodal corpora. |
| Approach: | They present the latest developments of Praaline, an open-source software system for constituting and managing spoken language and multimodal corpora. |
| Outcome: | The proposed system can be used for creating, managing, visualising and analysing spoken language and multimodal corpora. |
Emosical: An Emotion-Annotated Musical Theatre Dataset (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Emosical provides rich emotion annotations for musical films by inferring the background story of the characters. |
| Approach: | They propose to use a multimodal dataset of musical films to generate annotated emotion tags for each sample by inferring the background story of the characters. |
| Outcome: | The proposed dataset provides rich emotion annotations for musical films by inferring the background story of the characters. |
End-to-end ASR to jointly predict transcriptions and linguistic annotations (2021.naacl-main)
Copied to clipboard
| Challenge: | Existing models generate audio transcripts by sequentially producing likely graphemes, or multi-graphemic units, from which lexical items of a language can be recovered. |
| Approach: | They propose a Transformer-based sequence-to-sequence model for automatic speech recognition that can produce high-quality transcriptions and linguistic annotations. |
| Outcome: | The proposed model can produce high-quality transcriptions and linguistic annotations on Japanese and English audio datasets. |
Beyond Sentence-level Labels: Integrating Conversational Context and Personal Experience for Natural Emotional Expression (2026.findings-acl)
Copied to clipboard
Haiyang Sun, Chenyang Le, Wei Wang, Leying Zhang, Chuang Li, Bing Han, Chenda Li, Mengxiao Bi, Yanmin Qian
| Challenge: | Existing systems rely on sentence-level labels, which fails to capture the subtle nuances of human affect. |
| Approach: | They propose to use a large-scale, context-aware speech corpus derived from multi-speaker audiobooks to generate a speech that is human-like. |
| Outcome: | The proposed model outperforms existing methods in terms of emotional expression accuracy and naturalness. |
ES4R: Speech Encoding Based on Prepositive Affective Modeling for Empathetic Response Generation (2026.acl-long)
Copied to clipboard
| Challenge: | Existing speech-to-speech large language models rely on ASR transcription or use encoders to extract latent representations, weakening affective information and contextual coherence in multi-turn dialogues. |
| Approach: | They propose a framework for speech-based empathetic response generation that captures turn-level affective states and dialogue-level emotional dynamics. |
| Outcome: | The proposed framework outperforms baselines in automatic and human evaluations and remains robust across different Large Language Model (LLM) backbones. |
Llama-VITS: Enhancing TTS Synthesis with Semantic Awareness (2024.lrec-main)
Copied to clipboard
| Challenge: | Recent advances in Natural Language Processing (NLP) have seen Large-scale Language Models excel at producing high-quality text for various purposes. |
| Approach: | They propose a language model that enriches semantic content of text using Llama2 . their method enhances emotive expressiveness on a dataset . |
| Outcome: | The proposed model matches the naturalness of the original VITS and incorporates BERT (BERT-VITS) on the LJSpeech dataset, highlighting its potential to generate emotive speech. |