Challenge: a study aims to understand the role of disfluencies in speech production . speakers tend to lessen cognitive load for upcoming difficulties .
Approach: They examine the role of three influential theories of language processing in predicting disfluencies in speech production.
Outcome: The proposed classifiers predict disfluencies in English conversational speech . the classifier features lexical surprisal, word duration and DLT integration costs .

Similar Papers

Disfluent Cues for Enhanced Speech Understanding in Large Language Models (2023.findings-emnlp)

Copied to clipboard

Challenge: a large number of language models struggle to handle disfluencies, authors say . when a speaker hesitates, interrupts themselves, repeats or corrects words, or abandons phrases, it can make their speech fragmented.
Approach: They propose to use disfluent queries to “clean” spontaneous speech . they propose to apply disfluencies to models that use different types of speech repairs .
Outcome: The proposed model improves on a reading comprehension task using disfluent queries . the results suggest that disfluencies can improve model performance, rather than their removal .
Revisiting the Uniform Information Density Hypothesis (2021.emnlp-main)

Copied to clipboard

Challenge: The uniform information density hypothesis posits a preference among language users for utterances structured such that information is distributed uniformly across a signal.
Approach: They propose to test the hypothesis by using reading time and acceptability data to examine the effect of surprisal on language comprehension and acceptabilities.
Outcome: The proposed hypothesis makes predictions about language comprehension and linguistic acceptability .
Surprise! Uniform Information Density Isn’t the Whole Story: Predicting Surprisal Contours in Long-form Discourse (2024.emnlp-main)

Copied to clipboard

Challenge: Uniform Information Density (UID) hypothesis posits that speakers tend to distribute information evenly across linguistic units to achieve efficient communication.
Approach: They propose a functional pressure that speakers modulate information rate based on location within a hierarchically-structured model of discourse.
Outcome: The proposed hypothesis posits that speakers tend to distribute information evenly across linguistic units to achieve efficient communication.
Rethinking the Idiomaticity Decomposability Hypothesis: Evidence from Distributional Learning (2026.acl-long)

Copied to clipboard

Challenge: Decomposability is thought to predict syntactic flexibility, but is not attributed to distributional experience.
Approach: They propose a model-internal measure of decomposability and relate it to human ratings, syntactic flexibility, and predictability while tracking idiom learning during pretraining.
Outcome: The proposed model-internal measure correlates weakly with human judgments and shows a small but consistent negative relationship with syntactic flexibility.
Do Language Models Exhibit Human-like Structural Priming Effects? (2024.findings-acl)

Copied to clipboard

Challenge: a recent exposure to a structure facilitates processing of the same structure, a study finds . structural priming is well attested in humans, for both language production and comprehension .
Approach: They use the structural priming paradigm to investigate where priming effects manifest . they find that rarer elements within a prime increase priming effect .
Outcome: The findings provide an important piece in the puzzle of understanding how properties within their context affect structural prediction in language models.
The importance of fillers for text representations of speech transcripts (2020.emnlp-main)

Copied to clipboard

Challenge: Fillers are a type of disfluency that can be a sound ("um" or "uh") filling a pause in an utterance or conversation.
Approach: They propose to represent fillers with deep contextualised embeddings to improve modelling of spoken language and two downstream tasks .
Outcome: The proposed representations improve modelling of spoken language and two downstream tasks, predicting a speaker’s stance and expressed confidence.
Expect the Unexpected? Testing the Surprisal of Salient Entities (2026.acl-long)

Copied to clipboard

Challenge: Existing work on the Uniform Information Density hypothesis has neglected the relative salience of discourse participants.
Approach: They propose to use an annotated text to examine how overall salience of entities in discourse relates to surprisal.
Outcome: The proposed method shows that global salience is a mechanism shaping information distribution in discourse.
Investigating the Role and Impact of Disfluency on Summarization (2023.emnlp-industry)

Copied to clipboard

Challenge: Existing studies have focused on disfluency detection and removal, with limited studies into its impact on downstream tasks.
Approach: They propose to incorporate disfluency in summarization models to reduce the impact of replacement disfluencies on natural language processing tasks.
Outcome: The proposed model improves on both public and real-life datasets and shows that it can handle disfluent data with up to 6.99-point degradation in Rouge-L score and replacement disfluencies have the highest negative impact.
Quantifying the Impact of Disfluency on Spoken Content Summarization (2024.lrec-main)

Copied to clipboard

Challenge: a recent study has found that disfluencies negatively impact spoken content summarization .
Approach: They aim to quantify the impact of disfluency on spoken content summarization . they also investigate two methods towards improving summarizing in the presence of disflouencies .
Outcome: The proposed methods improve summarization quality in the presence of disfluencies.
Why Does Surprisal From Larger Transformer-Based Language Models Provide a Poorer Fit to Human Reading Times? (2023.tacl-1)

Copied to clipboard

Challenge: Existing studies have shown that larger pre-trained language models with more parameters and lower perplexity are less predictive of human reading times.
Approach: They propose to use a transformer-based model with more parameters and lower perplexity to investigate why these models are less predictive of human reading times.
Outcome: The results show that the larger models with more parameters and lower perplexity are less predictive of human reading times and eye-gaze durations collected during naturalistic reading.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations