Challenge: Existing studies show that optional reductions are sensitive to contextual predictability . unclear whether speaker choices are driven by audience design or to facilitate production .
Approach: They argue that Uniform Information Density and availability-based production make opposite predictions regarding the predictability of upcoming material and speaker choices.
Outcome: The proposed model shows that speaker choices support availability-based production account, not the UID hypothesis.

Similar Papers

Surprisal Minimisation over Goal-directed Alternatives Predicts Production Choice in Dialogue (2026.acl-long)

Copied to clipboard

Challenge: a method to model utterance production is based on information-theoretic notions of cost . a technique to generate alternative sets of utterables is proposed .
Approach: They propose a procedure to generate both types of alternative sets using language models.
Outcome: The proposed procedure allows for speaker- and listener-oriented interpretations of different cost measures.
That’s Optional: A Contemporary Exploration of “that” Omission in English Subordinate Clauses (2024.acl-short)

Copied to clipboard

Challenge: Uniform information density (UID) hypothesis posits speakers optimize the communicative properties of their utterances by avoiding spikes in information.
Approach: They propose to extend the information-uniformity principles by the notion of entropy to estimate the UID manifestations in the usecase of syntactic reduction choices.
Outcome: The proposed hypothesis is based on the optional omission of the connector "that" in English subordinate clauses.
How Relevant Are Selectional Preferences for Transformer-based Language Models? (2020.coling-main)

Copied to clipboard

Challenge: Selectional preference is defined as the tendency of a predicate to favor particular arguments within a certain linguistic context and reject others that result in conflicting or implausible meanings.
Approach: They examine the probability that Bert assigns a dependent word given the presence of a head word in a sentence to determine whether selectional preference exists.
Outcome: The proposed model is based on the SP-10K corpus of selectional preference and the ukWaC corpus.
Mandarin classifier systems optimize to accommodate communicative pressures (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing studies suggest that gendered languages are inherently optimized to accommodate communicative pressures on language learning and processing.
Approach: They propose to use grammatical or probabilistic modifiers to smooth the entropy of nouns in context to find the same frequency, similarity, and co-occurrence interactions that structure gender systems.
Outcome: The proposed noun classification device is sensitive to frequency, similarity, and co-occurrence interactions that structure gender systems.
Discourse Context Predictability Effects in Hindi Word Order (2022.emnlp-main)

Copied to clipboard

Challenge: Prior work has shown that information status, dependency length, and syntactic surprisal influence word order preferences, but the role of discourse predictability is underexplored in the literature.
Approach: They propose to use Hindi-Urdu Treebank corpus to build a classifier to predict which sentences actually occurred in the corpus against artificially generated distractors.
Outcome: The proposed classifier predicts which sentences occur in the Hindi-Urdu Treebank corpus against artificial distractors.
Language Models at the Syntax-Semantics Interface: A Case Study of the Long-Distance Binding of Chinese Reflexive Ziji (2025.coling-main)

Copied to clipboard

Challenge: Existing language models tend to rely heavily on sequential cues, but not always favoring the closest strings.
Approach: They construct a dataset of 320 synthetic sentences and 360 natural sentences from the BCC corpus . they evaluate 21 language models against this dataset and compare their performance to native Mandarin speakers .
Outcome: The proposed models do not replicate human-like judgments in Mandarin Chinese . the results show that existing models tend to rely heavily on sequential cues .
The Role of Abstract Representations and Observed Preferences in the Ordering of Binomials in Large Language Models (2025.acl-short)

Copied to clipboard

Challenge: Using binomial ordering preferences, large language models learn abstract representations versus more superficial aspects of their training corpora.
Approach: They examine binomial ordering preferences involving two conjoined nouns in English and examine whether large language models rely on observed binomialisms or on abstract ordering preferences.
Outcome: The proposed model learning is based on the observed binomial ordering preferences in English, and not on human linguis-tic input.
Speechworthy Instruction-tuned Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: Current instruction tuned language models are trained on textual preference data and therefore not aligned to speech domain.
Approach: They propose to use radio-industry best practices to prompt and learn speech-based preference data to improve speech-suitability of popular instruction tuned language models.
Outcome: The proposed methods achieve the best win rates in head-to-head comparisons, resulting in preferred or tied to the base model in 76.2% of comparisons on average.
A Cognitive Regularizer for Language Modeling (2021.acl-long)

Copied to clipboard

Challenge: a uniform information density hypothesis is used to explain certain linguistic phenomena . a regularizer that encodes the UID hypothesis can be used for language training .
Approach: They propose to augment the canonical MLE objective with a regularizer that encodes UID . they find that regularization consistently improves perplexity in language models .
Outcome: The proposed hypothesis can be operationalized as an inductive bias for language modeling.
Revisiting the Uniform Information Density Hypothesis (2021.emnlp-main)

Copied to clipboard

Challenge: The uniform information density hypothesis posits a preference among language users for utterances structured such that information is distributed uniformly across a signal.
Approach: They propose to test the hypothesis by using reading time and acceptability data to examine the effect of surprisal on language comprehension and acceptabilities.
Outcome: The proposed hypothesis makes predictions about language comprehension and linguistic acceptability .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations