Challenge: Existing studies have examined Zipf's meaning-frequency law as a relationship between word frequency and the number of meanings based on contextualized word vectors .
Approach: They propose to use word frequency as a relationship between word frequency and contextual diversity to examine Zipf's meaning-frequency law for a wider variety of words and corpora than previous studies have shown.
Outcome: The proposed formulation gives a new interpretation of Zipf's meaning-frequency law and enables us to examine it for a wider variety of words and corpora than previous studies have shown.

Similar Papers

More frequent verbs are associated with more diverse valency frames: Efficient principles at the lexicon-grammar interface (2024.acl-long)

Copied to clipboard

Challenge: Existing evidence has focused on word-internal properties, such as Zipf's observation that more frequent words are optimized in form to minimize communicative cost.
Approach: They propose to examine the hypothesis that efficient lexicon organization is also reflected in valency, or the combinations and orders of additional words and phrases a verb selects for in a sentence.
Outcome: The proposed hypothesis is consistent with communicative efficiency principles.
A Three-Parameter Rank-Frequency Relation in Natural Languages (2020.acl-main)

Copied to clipboard

Challenge: Existing empirical law to form rank-frequency relation in textual data is Zipf's/power law .
Approach: They propose a rank-frequency relation that follows f r-(r+)- in textual data.
Outcome: The proposed formulation is the power law when =0 and the Zipf–Mandelbrot law when=1 .
Adam’s Law: Textual Frequency Law on Large Language Models (2026.acl-long)

Copied to clipboard

Challenge: Textual frequency is a topic of understudied research, but its relevance to Large Language Models is not well understood.
Approach: They propose a framework to estimate textual data frequency using a paraphraser and a textual distillation method to refine LLMs.
Outcome: The proposed framework can be used to estimate sentence-level frequency with word-level frequencies.
Contextual Diversity Measure (CDM) for Controllable Story Generation in Large Language Models (2026.acl-srw)

Copied to clipboard

Challenge: Existing studies on controllable text generation focus on controlling attributes such as sentiment, writing style, and writing style.
Approach: They introduce a metric that quantifies semantic diversity for scenario generation under fixed abstract semantic constraints and validate it through controlled experiments.
Outcome: The proposed metric achieves excellent discrimination accuracy (100% and 91.9%, respectively), with discriminative power up to 5.5 greater than the best baseline.
How (Non-)Optimal is the Lexicon? (2021.naacl-main)

Copied to clipboard

Challenge: lexical meanings are mapped to wordforms by usage pressures and constraints on sequences of symbols.
Approach: They propose a coding-theoretic view of the lexicon and a novel generative statistical model to quantify its compressibility under various constraints.
Outcome: The proposed model shows that (compositional) morphology and graphotactics can account for most of the complexity of natural codes—as measured by code length.
Exploring the Representation of Word Meanings in Context: A Case Study on Homonymy and Synonymy (2021.acl-long)

Copied to clipboard

Challenge: Existing models that represent different senses of words in context are not accurate for polysemous words.
Approach: They propose a multilingual dataset that evaluates the ability of models to accurately represent different lexical-semantic relations such as homonymy and synonymy.
Outcome: The proposed models can disambiguate homonyms in context, but fail to represent words with different senses when occurring in similar sentences.
A large-scale study of the effects of word frequency and predictability in naturalistic reading (N19-1)

Copied to clipboard

Challenge: Recent studies have shown separable effects of word frequency and predictability on human sentence processing . other theories hold that apparent effects of frequency are underlyingly effects of predictability .
Approach: They examine the generalizability of this finding to more realistic conditions of sentence processing by studying effects of frequency and predictability in three large-scale naturalistic reading corpora.
Outcome: The results show that word frequency and predictability are significant in isolation but not over and above predictability, and raise doubts about the existence of such a distinction in everyday sentence comprehension.
Patterns of Polysemy and Homonymy in Contextualised Language Models (2021.findings-emnlp)

Copied to clipboard

Challenge: a recent study has focused on homonymy, a variety of multiplicity of meanings exemplified by word forms with unrelated meanings.
Approach: They investigate the extent to which contextualised embeddings reflect traditional distinctions of polysemy and homonymy.
Outcome: The proposed model shows that it can distinguish between polysemy and homonymy . it shows that the model fails to replicate the results of the human-annotated dataset .
Why is penguin more similar to polar bear than to sea gull? Analyzing conceptual knowledge in distributional models (2020.acl-srw)

Copied to clipboard

Challenge: Several analysis methods have been shown to be limited and are not well understood . thesis aims to understand distributional semantic representations based on linguistic data .
Approach: They propose a framework for investigating the information encoded in distributional semantic models . they combine observations made on corpora with insights obtained from data manipulation experiments .
Outcome: The proposed framework pairs observations made on corpora with insights obtained from data manipulation experiments.
Frequency & Compositionality in Emergent Communication (2025.emnlp-main)

Copied to clipboard

Challenge: Natural languages exhibit a universal tendency to resist regular patterns, developing idiosyncratic forms.
Approach: They investigate the relationship between frequency and compositionality in emergent languages . they use a referential game setting to manipulate input frequency through Zipfian distributions .
Outcome: The proposed method shows that the frequency distributions of the most frequent words resist regular patterns, resulting in less compositional structure.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations