Challenge: Prosody—the melody of speech—conveys critical information often not captured by the words or text of a message.
Approach: They propose an information-theoretic approach to quantify how much is conveyed by prosody that is not recoverable from text alone.
Outcome: The proposed framework can quantify how much is conveyed by prosody that is not recoverable from text alone and crucially, what prosody conveys.

Similar Papers

Quantifying the redundancy between prosody and text (2023.emnlp-main)

Copied to clipboard

Challenge: Existing studies suggest partial redundancy between prosody and linguistic information.
Approach: They use large language models to estimate how much information is redundant between prosody and the words themselves.
Outcome: The proposed model can predict prosodic features across prosodic features, including intensity, duration, pauses, and pitch contours.
The time scale of redundancy between prosody and linguistic context (2025.acl-long)

Copied to clipboard

Challenge: Prior work has shown that the information carried by prosodic features is substantially redundant with that carried by the surrounding words.
Approach: They examine the time scale of this relationship, studying how it varies with the length of past and future contexts.
Outcome: The results show that prosody features show some redundancy with future words, but only with a short scale of 1-2 words, consistent with reports of incremental short-term planning in language production.
Prosody: Models, Methods, and Applications (2021.acl-tutorials)

Copied to clipboard

Challenge: This tutorial will overview the computational modeling of prosody.
Approach: This tutorial will overview the computational modeling of prosody . it will discuss the latest advances in prosody and diverse applications .
Outcome: This tutorial will overview the computational modeling of prosody.
Using Information Theory to Characterize Prosodic Typology: The Case of Tone, Pitch-Accent and Stress-Accent (2025.acl-long)

Copied to clipboard

Challenge: lexical identity and prosody are well-studied parameters of linguistic variation, but they are difficult to predict in tonal languages.
Approach: They propose to characterize the relationship between lexical identity and prosody using information theory to estimate mutual information between the text and pitch curves.
Outcome: The proposed hypothesis supports perspectives that view linguistic typology as gradient, rather than categorical.
The Role of Prosody in Spoken Question Answering (2025.findings-naacl)

Copied to clipboard

Challenge: lexical information is not available in most models, but prosody is important in understanding spoken language.
Approach: They investigate the role of prosody in the process of answering a spoken question by isolating prosodic and lexical information from a natural speech dataset.
Outcome: The proposed models can perform reasonably well on the SLUE-SQA-5 dataset, but when lexical information is available, models tend to predominantly rely on it.
Compilation of Corpora for the Study of the Information Structure–Prosody Interface (L18-1)

Copied to clipboard

Challenge: empirical studies on the Information Structure-prosody interface are scarce . thematicity defines how content is packaged in terms of "what is being talked about" a different view on thematicality is advocated by I. Mel'uk in the context of the MTT.
Approach: They propose a method for the compilation of annotated corpora to study the correspondence between Information Structure and prosody.
Outcome: The proposed method is applied to a corpus of read speech in English annotated with hierarchical thematicity and automatically extracted prosodic parameters.
The Prosody of Emojis (2026.acl-long)

Copied to clipboard

Challenge: emojis are useful in spoken communication because they add affective and pragmatic nuance to textual cues.
Approach: They analyze human speech data to find prosodic features that are important in spoken communication.
Outcome: The proposed model shows that speakers adapt prosody based on emoji cues, and that listeners can recover intended meanings significantly above chance.
Prosody in Cascade and Direct Speech-to-Text Translation: a case study on Korean Wh-Phrases (2024.findings-eacl)

Copied to clipboard

Challenge: Existing direct S2TT systems have been unable to disambiguate utterances where prosody plays a crucial role.
Approach: They propose to use contrastive evaluation to quantitatively measure the ability of direct S2TT systems to disambiguate utterances where prosody plays a crucial role.
Outcome: The proposed system improves overall accuracy 12.9% and improves intent scores 15.6%.
ProsodyFlow: High-fidelity Text-to-Speech through Conditional Flow Matching and Prosody Modeling with Large Speech Language Models (2025.coling-main)

Copied to clipboard

Challenge: Text-to-speech (TTS) models have been developed to generate high-quality speech.
Approach: They propose an end-to-end TTS model that integrates large self-supervised speech models and conditional flow matching to model prosodic features effectively.
Outcome: The proposed model improves synthesis quality and efficiency compared to existing models, showing that it generates more prosodic and expressive speech synthesizing.
Computational Narrative Understanding for Expressive Text-to-Speech (2026.findings-acl)

Copied to clipboard

Challenge: Recent advances in text-to-speech systems have been driven by large, multi-domain speech corpora.
Approach: They propose a large-scale 5.3K hours of expressive speech drawn from character quotations . they fine-tune a flow-matching model and train from scratch .
Outcome: The proposed model improves expressivity and intelligibility while training from scratch improves expressiveness of an autoregressive model.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations