Quantifying the redundancy between prosody and text (2023.emnlp-main)

Copied to clipboard

Challenge: Existing studies suggest partial redundancy between prosody and linguistic information.
Approach: They use large language models to estimate how much information is redundant between prosody and the words themselves.
Outcome: The proposed model can predict prosodic features across prosodic features, including intensity, duration, pauses, and pitch contours.

Similar Papers

The time scale of redundancy between prosody and linguistic context (2025.acl-long)

Copied to clipboard

Challenge: Prior work has shown that the information carried by prosodic features is substantially redundant with that carried by the surrounding words.
Approach: They examine the time scale of this relationship, studying how it varies with the length of past and future contexts.
Outcome: The results show that prosody features show some redundancy with future words, but only with a short scale of 1-2 words, consistent with reports of incremental short-term planning in language production.
What Do Prosody and Text Convey? Characterizing How Meaningful Information is Distributed Across Multiple Channels (2026.acl-long)

Copied to clipboard

Challenge: Prosody—the melody of speech—conveys critical information often not captured by the words or text of a message.
Approach: They propose an information-theoretic approach to quantify how much is conveyed by prosody that is not recoverable from text alone.
Outcome: The proposed framework can quantify how much is conveyed by prosody that is not recoverable from text alone and crucially, what prosody conveys.
Compilation of Corpora for the Study of the Information Structure–Prosody Interface (L18-1)

Copied to clipboard

Challenge: empirical studies on the Information Structure-prosody interface are scarce . thematicity defines how content is packaged in terms of "what is being talked about" a different view on thematicality is advocated by I. Mel'uk in the context of the MTT.
Approach: They propose a method for the compilation of annotated corpora to study the correspondence between Information Structure and prosody.
Outcome: The proposed method is applied to a corpus of read speech in English annotated with hierarchical thematicity and automatically extracted prosodic parameters.
Giving Attention to the Unexpected: Using Prosody Innovations in Disfluency Detection (N19-1)

Copied to clipboard

Challenge: Disfluencies in spontaneous speech are associated with prosodic disruptions.
Approach: They propose a method to extract acoustic-prosodic cues from word transcripts . they explore early and late fusion techniques for integrating text and prosody .
Outcome: The proposed approach shows gains over a high-accuracy text-only model.
The Prosody of Emojis (2026.acl-long)

Copied to clipboard

Challenge: emojis are useful in spoken communication because they add affective and pragmatic nuance to textual cues.
Approach: They analyze human speech data to find prosodic features that are important in spoken communication.
Outcome: The proposed model shows that speakers adapt prosody based on emoji cues, and that listeners can recover intended meanings significantly above chance.
When Large Language Models Meet Speech: A Survey on Integration Approaches (2025.findings-acl)

Copied to clipboard

Challenge: Recent advances in large language models have spurred interest in expanding their application beyond text-based tasks.
Approach: They propose to categorize the integration of speech with LLMs into three main approaches . they demonstrate how these methods are applied across various speech-related applications .
Outcome: The proposed methods are applied across speech-related applications and highlight the challenges in this field to offer inspiration for future research.
Pragmatics in the Era of Large Language Models: A Survey on Datasets, Evaluation, Opportunities and Challenges (2025.acl-long)

Copied to clipboard

Challenge: linguistics studies how context influences meaning of language and how people use it to convey implied meanings, emotions, and intentions.
Approach: They analyze task designs, data collection methods, evaluation approaches and their relevance to real-world applications.
Outcome: The findings highlight emerging trends, challenges, and gaps in existing benchmarks . the findings will contribute to more nuanced and context-aware NLP models .
ProsodyFlow: High-fidelity Text-to-Speech through Conditional Flow Matching and Prosody Modeling with Large Speech Language Models (2025.coling-main)

Copied to clipboard

Challenge: Text-to-speech (TTS) models have been developed to generate high-quality speech.
Approach: They propose an end-to-end TTS model that integrates large self-supervised speech models and conditional flow matching to model prosodic features effectively.
Outcome: The proposed model improves synthesis quality and efficiency compared to existing models, showing that it generates more prosodic and expressive speech synthesizing.
Low-Perplexity LLM-Generated Sequences and Where To Find Them (2025.acl-srw)

Copied to clipboard

Challenge: Large Language Models (LLMs) are increasingly applied across various domains, but the ways they leverage their training data during inference remains only partially understood.
Approach: They propose a systematic approach that analyzes low-perplexity sequences and traces them back to their sources in the training data.
Outcome: The proposed pipeline extracts low-perplexity sequences across diverse topics while avoiding degeneration, then trace them back to their sources in the training data.
The Role of Prosody in Spoken Question Answering (2025.findings-naacl)

Copied to clipboard

Challenge: lexical information is not available in most models, but prosody is important in understanding spoken language.
Approach: They investigate the role of prosody in the process of answering a spoken question by isolating prosodic and lexical information from a natural speech dataset.
Outcome: The proposed models can perform reasonably well on the SLUE-SQA-5 dataset, but when lexical information is available, models tend to predominantly rely on it.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations