A New Formulation of Zipf’s Meaning-Frequency Law through Contextual Diversity (2025.acl-long)
Copied to clipboard
| Challenge: | Existing studies have examined Zipf's meaning-frequency law as a relationship between word frequency and the number of meanings based on contextualized word vectors . |
| Approach: | They propose to use word frequency as a relationship between word frequency and contextual diversity to examine Zipf's meaning-frequency law for a wider variety of words and corpora than previous studies have shown. |
| Outcome: | The proposed formulation gives a new interpretation of Zipf's meaning-frequency law and enables us to examine it for a wider variety of words and corpora than previous studies have shown. |
Similar Papers
More frequent verbs are associated with more diverse valency frames: Efficient principles at the lexicon-grammar interface (2024.acl-long)
Copied to clipboard
| Challenge: | Existing evidence has focused on word-internal properties, such as Zipf's observation that more frequent words are optimized in form to minimize communicative cost. |
| Approach: | They propose to examine the hypothesis that efficient lexicon organization is also reflected in valency, or the combinations and orders of additional words and phrases a verb selects for in a sentence. |
| Outcome: | The proposed hypothesis is consistent with communicative efficiency principles. |
A Three-Parameter Rank-Frequency Relation in Natural Languages (2020.acl-main)
Copied to clipboard
| Challenge: | Existing empirical law to form rank-frequency relation in textual data is Zipf's/power law . |
| Approach: | They propose a rank-frequency relation that follows f r-(r+)- in textual data. |
| Outcome: | The proposed formulation is the power law when =0 and the Zipf–Mandelbrot law when=1 . |
Adam’s Law: Textual Frequency Law on Large Language Models (2026.acl-long)
Copied to clipboard
| Challenge: | Textual frequency is a topic of understudied research, but its relevance to Large Language Models is not well understood. |
| Approach: | They propose a framework to estimate textual data frequency using a paraphraser and a textual distillation method to refine LLMs. |
| Outcome: | The proposed framework can be used to estimate sentence-level frequency with word-level frequencies. |
Contextual Diversity Measure (CDM) for Controllable Story Generation in Large Language Models (2026.acl-srw)
Copied to clipboard
| Challenge: | Existing studies on controllable text generation focus on controlling attributes such as sentiment, writing style, and writing style. |
| Approach: | They introduce a metric that quantifies semantic diversity for scenario generation under fixed abstract semantic constraints and validate it through controlled experiments. |
| Outcome: | The proposed metric achieves excellent discrimination accuracy (100% and 91.9%, respectively), with discriminative power up to 5.5 greater than the best baseline. |
How (Non-)Optimal is the Lexicon? (2021.naacl-main)
Copied to clipboard
| Challenge: | lexical meanings are mapped to wordforms by usage pressures and constraints on sequences of symbols. |
| Approach: | They propose a coding-theoretic view of the lexicon and a novel generative statistical model to quantify its compressibility under various constraints. |
| Outcome: | The proposed model shows that (compositional) morphology and graphotactics can account for most of the complexity of natural codes—as measured by code length. |
Exploring the Representation of Word Meanings in Context: A Case Study on Homonymy and Synonymy (2021.acl-long)
Copied to clipboard
| Challenge: | Existing models that represent different senses of words in context are not accurate for polysemous words. |
| Approach: | They propose a multilingual dataset that evaluates the ability of models to accurately represent different lexical-semantic relations such as homonymy and synonymy. |
| Outcome: | The proposed models can disambiguate homonyms in context, but fail to represent words with different senses when occurring in similar sentences. |
A large-scale study of the effects of word frequency and predictability in naturalistic reading (N19-1)
Copied to clipboard
| Challenge: | Recent studies have shown separable effects of word frequency and predictability on human sentence processing . other theories hold that apparent effects of frequency are underlyingly effects of predictability . |
| Approach: | They examine the generalizability of this finding to more realistic conditions of sentence processing by studying effects of frequency and predictability in three large-scale naturalistic reading corpora. |
| Outcome: | The results show that word frequency and predictability are significant in isolation but not over and above predictability, and raise doubts about the existence of such a distinction in everyday sentence comprehension. |
Patterns of Polysemy and Homonymy in Contextualised Language Models (2021.findings-emnlp)
Copied to clipboard
| Challenge: | a recent study has focused on homonymy, a variety of multiplicity of meanings exemplified by word forms with unrelated meanings. |
| Approach: | They investigate the extent to which contextualised embeddings reflect traditional distinctions of polysemy and homonymy. |
| Outcome: | The proposed model shows that it can distinguish between polysemy and homonymy . it shows that the model fails to replicate the results of the human-annotated dataset . |
Why is penguin more similar to polar bear than to sea gull? Analyzing conceptual knowledge in distributional models (2020.acl-srw)
Copied to clipboard
| Challenge: | Several analysis methods have been shown to be limited and are not well understood . thesis aims to understand distributional semantic representations based on linguistic data . |
| Approach: | They propose a framework for investigating the information encoded in distributional semantic models . they combine observations made on corpora with insights obtained from data manipulation experiments . |
| Outcome: | The proposed framework pairs observations made on corpora with insights obtained from data manipulation experiments. |
Frequency & Compositionality in Emergent Communication (2025.emnlp-main)
Copied to clipboard
| Challenge: | Natural languages exhibit a universal tendency to resist regular patterns, developing idiosyncratic forms. |
| Approach: | They investigate the relationship between frequency and compositionality in emergent languages . they use a referential game setting to manipulate input frequency through Zipfian distributions . |
| Outcome: | The proposed method shows that the frequency distributions of the most frequent words resist regular patterns, resulting in less compositional structure. |