Challenge: We show that ascribing verbal descriptions to expressive audiovisual utterances is efficient and efficient.
Approach: They propose a web app for ascribing verbal descriptions to expressive audiovisual utterances.
Outcome: The proposed system can be deployed at a large scale to efficiently collect relevant verbal descriptions.

Similar Papers

Computational Narrative Understanding for Expressive Text-to-Speech (2026.findings-acl)

Copied to clipboard

Challenge: Recent advances in text-to-speech systems have been driven by large, multi-domain speech corpora.
Approach: They propose a large-scale 5.3K hours of expressive speech drawn from character quotations . they fine-tune a flow-matching model and train from scratch .
Outcome: The proposed model improves expressivity and intelligibility while training from scratch improves expressiveness of an autoregressive model.
When Words Smile: Generating Diverse Emotional Facial Expressions from Text (2025.emnlp-main)

Copied to clipboard

Challenge: Existing systems that generate only coarse facial expressions ignore the rich and dynamic nature of face-to-face communication.
Approach: They propose an end-to-end text-to expression model that explicitly focuses on emotional dynamics.
Outcome: The proposed model outperforms baselines on 15,000 text–3D expression pairs on a large-scale dataset.
Prompt-Guided Selective Masking Loss for Context-Aware Emotive Text-to-Speech (2025.findings-naacl)

Copied to clipboard

Challenge: Emotional dialogue speech synthesis (EDSS) aims to generate expressive speech by leveraging the dialogue context between interlocutors.
Approach: They propose a large language model to generate holistic emotion tags based on prior dialogue context and pinpoint key words in the target utterance that align with the predicted emotional state.
Outcome: The proposed method improves emotional expressiveness and facilitates automatic emotion speech generation during inference.
Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey (2025.emnlp-main)

Copied to clipboard

Challenge: Text-to-speech (TTS) has advanced from generating natural-sounding speech to enabling fine-grained control over speech attributes.
Approach: They provide a review of controllable TTS methods from traditional control techniques to emerging approaches using natural language prompts.
Outcome: The proposed methods are based on models, strategies, and features, and summarize challenges, datasets, and evaluations.
Praaline: An Open-Source System for Managing, Annotating, Visualising and Analysing Speech Corpora (P18-4)

Copied to clipboard

Challenge: Praaline is an open-source software system for constituting and managing spoken language and multimodal corpora.
Approach: They present the latest developments of Praaline, an open-source software system for constituting and managing spoken language and multimodal corpora.
Outcome: The proposed system can be used for creating, managing, visualising and analysing spoken language and multimodal corpora.
Emosical: An Emotion-Annotated Musical Theatre Dataset (2024.findings-emnlp)

Copied to clipboard

Challenge: Emosical provides rich emotion annotations for musical films by inferring the background story of the characters.
Approach: They propose to use a multimodal dataset of musical films to generate annotated emotion tags for each sample by inferring the background story of the characters.
Outcome: The proposed dataset provides rich emotion annotations for musical films by inferring the background story of the characters.
End-to-end ASR to jointly predict transcriptions and linguistic annotations (2021.naacl-main)

Copied to clipboard

Challenge: Existing models generate audio transcripts by sequentially producing likely graphemes, or multi-graphemic units, from which lexical items of a language can be recovered.
Approach: They propose a Transformer-based sequence-to-sequence model for automatic speech recognition that can produce high-quality transcriptions and linguistic annotations.
Outcome: The proposed model can produce high-quality transcriptions and linguistic annotations on Japanese and English audio datasets.
Beyond Sentence-level Labels: Integrating Conversational Context and Personal Experience for Natural Emotional Expression (2026.findings-acl)

Copied to clipboard

Challenge: Existing systems rely on sentence-level labels, which fails to capture the subtle nuances of human affect.
Approach: They propose to use a large-scale, context-aware speech corpus derived from multi-speaker audiobooks to generate a speech that is human-like.
Outcome: The proposed model outperforms existing methods in terms of emotional expression accuracy and naturalness.
ES4R: Speech Encoding Based on Prepositive Affective Modeling for Empathetic Response Generation (2026.acl-long)

Copied to clipboard

Challenge: Existing speech-to-speech large language models rely on ASR transcription or use encoders to extract latent representations, weakening affective information and contextual coherence in multi-turn dialogues.
Approach: They propose a framework for speech-based empathetic response generation that captures turn-level affective states and dialogue-level emotional dynamics.
Outcome: The proposed framework outperforms baselines in automatic and human evaluations and remains robust across different Large Language Model (LLM) backbones.
Llama-VITS: Enhancing TTS Synthesis with Semantic Awareness (2024.lrec-main)

Copied to clipboard

Challenge: Recent advances in Natural Language Processing (NLP) have seen Large-scale Language Models excel at producing high-quality text for various purposes.
Approach: They propose a language model that enriches semantic content of text using Llama2 . their method enhances emotive expressiveness on a dataset .
Outcome: The proposed model matches the naturalness of the original VITS and incorporates BERT (BERT-VITS) on the LJSpeech dataset, highlighting its potential to generate emotive speech.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations