What Do Prosody and Text Convey? Characterizing How Meaningful Information is Distributed Across Multiple Channels (2026.acl-long)
Copied to clipboard
| Challenge: | Prosody—the melody of speech—conveys critical information often not captured by the words or text of a message. |
| Approach: | They propose an information-theoretic approach to quantify how much is conveyed by prosody that is not recoverable from text alone. |
| Outcome: | The proposed framework can quantify how much is conveyed by prosody that is not recoverable from text alone and crucially, what prosody conveys. |
Similar Papers
Quantifying the redundancy between prosody and text (2023.emnlp-main)
Copied to clipboard
Lukas Wolf, Tiago Pimentel, Evelina Fedorenko, Ryan Cotterell, Alex Warstadt, Ethan Wilcox, Tamar Regev
| Challenge: | Existing studies suggest partial redundancy between prosody and linguistic information. |
| Approach: | They use large language models to estimate how much information is redundant between prosody and the words themselves. |
| Outcome: | The proposed model can predict prosodic features across prosodic features, including intensity, duration, pauses, and pitch contours. |
The time scale of redundancy between prosody and linguistic context (2025.acl-long)
Copied to clipboard
Tamar I Regev, Chiebuka Ohams, Shaylee Xie, Lukas Wolf, Evelina Fedorenko, Alex Warstadt, Ethan Wilcox, Tiago Pimentel
| Challenge: | Prior work has shown that the information carried by prosodic features is substantially redundant with that carried by the surrounding words. |
| Approach: | They examine the time scale of this relationship, studying how it varies with the length of past and future contexts. |
| Outcome: | The results show that prosody features show some redundancy with future words, but only with a short scale of 1-2 words, consistent with reports of incremental short-term planning in language production. |
Prosody: Models, Methods, and Applications (2021.acl-tutorials)
Copied to clipboard
| Challenge: | This tutorial will overview the computational modeling of prosody. |
| Approach: | This tutorial will overview the computational modeling of prosody . it will discuss the latest advances in prosody and diverse applications . |
| Outcome: | This tutorial will overview the computational modeling of prosody. |
Using Information Theory to Characterize Prosodic Typology: The Case of Tone, Pitch-Accent and Stress-Accent (2025.acl-long)
Copied to clipboard
| Challenge: | lexical identity and prosody are well-studied parameters of linguistic variation, but they are difficult to predict in tonal languages. |
| Approach: | They propose to characterize the relationship between lexical identity and prosody using information theory to estimate mutual information between the text and pitch curves. |
| Outcome: | The proposed hypothesis supports perspectives that view linguistic typology as gradient, rather than categorical. |
The Role of Prosody in Spoken Question Answering (2025.findings-naacl)
Copied to clipboard
| Challenge: | lexical information is not available in most models, but prosody is important in understanding spoken language. |
| Approach: | They investigate the role of prosody in the process of answering a spoken question by isolating prosodic and lexical information from a natural speech dataset. |
| Outcome: | The proposed models can perform reasonably well on the SLUE-SQA-5 dataset, but when lexical information is available, models tend to predominantly rely on it. |
Compilation of Corpora for the Study of the Information Structure–Prosody Interface (L18-1)
Copied to clipboard
| Challenge: | empirical studies on the Information Structure-prosody interface are scarce . thematicity defines how content is packaged in terms of "what is being talked about" a different view on thematicality is advocated by I. Mel'uk in the context of the MTT. |
| Approach: | They propose a method for the compilation of annotated corpora to study the correspondence between Information Structure and prosody. |
| Outcome: | The proposed method is applied to a corpus of read speech in English annotated with hierarchical thematicity and automatically extracted prosodic parameters. |
The Prosody of Emojis (2026.acl-long)
Copied to clipboard
| Challenge: | emojis are useful in spoken communication because they add affective and pragmatic nuance to textual cues. |
| Approach: | They analyze human speech data to find prosodic features that are important in spoken communication. |
| Outcome: | The proposed model shows that speakers adapt prosody based on emoji cues, and that listeners can recover intended meanings significantly above chance. |
Prosody in Cascade and Direct Speech-to-Text Translation: a case study on Korean Wh-Phrases (2024.findings-eacl)
Copied to clipboard
| Challenge: | Existing direct S2TT systems have been unable to disambiguate utterances where prosody plays a crucial role. |
| Approach: | They propose to use contrastive evaluation to quantitatively measure the ability of direct S2TT systems to disambiguate utterances where prosody plays a crucial role. |
| Outcome: | The proposed system improves overall accuracy 12.9% and improves intent scores 15.6%. |
ProsodyFlow: High-fidelity Text-to-Speech through Conditional Flow Matching and Prosody Modeling with Large Speech Language Models (2025.coling-main)
Copied to clipboard
| Challenge: | Text-to-speech (TTS) models have been developed to generate high-quality speech. |
| Approach: | They propose an end-to-end TTS model that integrates large self-supervised speech models and conditional flow matching to model prosodic features effectively. |
| Outcome: | The proposed model improves synthesis quality and efficiency compared to existing models, showing that it generates more prosodic and expressive speech synthesizing. |
Computational Narrative Understanding for Expressive Text-to-Speech (2026.findings-acl)
Copied to clipboard
| Challenge: | Recent advances in text-to-speech systems have been driven by large, multi-domain speech corpora. |
| Approach: | They propose a large-scale 5.3K hours of expressive speech drawn from character quotations . they fine-tune a flow-matching model and train from scratch . |
| Outcome: | The proposed model improves expressivity and intelligibility while training from scratch improves expressiveness of an autoregressive model. |