Challenge: Languages vary in how meanings map to word forms, but this theory does not account for systematic relations within word forms.
Approach: They propose a model that measures the learnability of meaning-to-form mappings by inverse of simplicity.
Outcome: The proposed model captures fine-grained regularities in linguistic form, allowing better discrimination between attested and unattested systems.

Similar Papers

Meaning to Form: Measuring Systematicity as Information (P19-1)

Copied to clipboard

Challenge: A longstanding debate in semiotics centers on the relationship between linguistic signs and their corresponding semantics: is there an arbitrary relationship between word forms and their meaning, or does some systematic phenomenon pervade?
Approach: They propose to quantify the systematicity of the sign using mutual information and recurrent neural networks to examine 106 languages.
Outcome: The proposed model reduces entropy in a word form conditioned on its semantic representation and recovers English examples of systematic affixes.
Recursive numeral systems are highly regular and easy to process (2026.eacl-long)

Copied to clipboard

Challenge: Existing studies on linguistic efficiency have focused on the systematicity of forms, a key property of natural language.
Approach: They propose to incorporate regularity across sets of forms in studies of efficiency in language . they use the Minimum Description Length approach to measure regularity and processing complexity .
Outcome: The proposed method shows that recursive numeral systems are more efficient with respect to regularity and processing complexity.
Finding Concept-specific Biases in Form–Meaning Associations (2021.naacl-main)

Copied to clipboard

Challenge: Existing methods to detect cross-linguistic associations are not effective, but their effects are minor.
Approach: They propose a method to measure cross-linguistic associations by controlling for the influence of language family and geographic proximity within a large concept-aligned, cross-lingual lexicon.
Outcome: The proposed method shows that it is small, but it is unsurprisingly small (less than 0.5% on average).
AND does not mean OR: Using Formal Languages to Study Language Models’ Representations (2021.acl-short)

Copied to clipboard

Challenge: A current open question in natural language processing is to what extent language models are able to capture the meaning of language.
Approach: They propose to simulate a distributional language model’s ability to differentiate logical symbols using motivated constraints and motivated constraints.
Outcome: The results show that the proposed models are unable to differentiate meaningfully different symbols, suggesting a limitation to the types of semantic signals that current models are capable of exploiting.
How (Non-)Optimal is the Lexicon? (2021.naacl-main)

Copied to clipboard

Challenge: lexical meanings are mapped to wordforms by usage pressures and constraints on sequences of symbols.
Approach: They propose a coding-theoretic view of the lexicon and a novel generative statistical model to quantify its compressibility under various constraints.
Outcome: The proposed model shows that (compositional) morphology and graphotactics can account for most of the complexity of natural codes—as measured by code length.
A Survey of Meaning Representations – From Theory to Practical Utility (2024.naacl-long)

Copied to clipboard

Challenge: Symbolic meaning representations of natural language text have been studied since at least the 1960s . with the availability of large annotated corpora, the field has recently seen several new developments .
Approach: They propose a framework for expressing meaning in natural language text using annotated corpora and a set of tools for machine learning.
Outcome: The frameworks are based on a set of theoretical and practical problems and their applications.
What Can String Probability Tell Us About Grammaticality? (2026.tacl-1)

Copied to clipboard

Challenge: linguistic theories have argued that language models have largely achieved grammatical competence, but they will assign non-zero probability to all strings.
Approach: They propose a theoretical framework for analyzing string probabilities in linguistics based on simple assumptions about the generative process of corpus data.
Outcome: The proposed framework makes three predictions using 280K sentence pairs in English and Chinese.
A Tour of Explicit Multilingual Semantics: Word Sense Disambiguation, Semantic Role Labeling and Semantic Parsing (2022.aacl-tutorials)

Copied to clipboard

Challenge: a recent advent of pretrained language models has sparked a revolution in NLP . but, there are still questions about whether current approaches capture explicit, symbolic meaning . this tutorial will review efforts to tackle three key open problems in lexical and sentence-level semantics .
Approach: This tutorial reviews recent efforts to shed light on meaning in NLP . it will focus on three key open problems in lexical and sentence-level semantics .
Outcome: This tutorial reviews recent efforts to shed light on meaning in NLP . it focuses on three key open problems in lexical and sentence-level semantics .
Mandarin classifier systems optimize to accommodate communicative pressures (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing studies suggest that gendered languages are inherently optimized to accommodate communicative pressures on language learning and processing.
Approach: They propose to use grammatical or probabilistic modifiers to smooth the entropy of nouns in context to find the same frequency, similarity, and co-occurrence interactions that structure gender systems.
Outcome: The proposed noun classification device is sensitive to frequency, similarity, and co-occurrence interactions that structure gender systems.
The Linguistic Connectivities Within Large Language Models (2025.findings-acl)

Copied to clipboard

Challenge: Recent studies have discovered notable disparities in their performance across different languages.
Approach: They conduct a systematic investigation into the behaviors of large language models across 27 different languages on 3 different scenarios and reveals a Linguistic Map correlates with the richness of available resources and linguistic family relations.
Outcome: The proposed model demonstrates that there are significant disparities in performance across languages across 27 different languages on 3 different scenarios.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations