Challenge: a standard assumption in phonology is that finding a minimal feature specification is an automatic part of acquisition and generalization.
Approach: They analyze the problem of determining whether a set of phonemes forms a natural class and find the minimal feature specification for the class.
Outcome: The proposed model is based on a greedy algorithm that fails to find minimal features . the proposed model can be used to find features that are universal across languages .

Similar Papers

Phonotactic Complexity and Its Trade-offs (2020.tacl-1)

Copied to clipboard

Challenge: Existing measures of linguistic complexity are relatively coarse-see, for example, Moran and Blasi (2014) and 2 below for reviews.
Approach: They propose to measure bits per phoneme using the negative log-probability of a word in a language model and a collection of 1016 basic concept words across 106 languages.
Outcome: The proposed measure allows a cross-linguistic comparison of phonotactic complexity across languages.
Phonotactic Complexity across Dialects (2024.lrec-main)

Copied to clipboard

Challenge: Recent studies show a moderate negative correlation between phonotactic complexity and word length in 106 languages.
Approach: They propose to use a phone-level language model to measure phonotactic complexity . they find a tradeoff between word length and phonomactic complex .
Outcome: The proposed model shows that low phonotactic complexity dialects concentrate around capital regions.
What Do Neural Speech Models Know About Phonology? Evidence from Structured Phoneme Confusions (2026.findings-acl)

Copied to clipboard

Challenge: acoustic and phonological models of speech recognition are often limited to the phoneme level . a recent study has shown that phoneme confusions are strongly structured in phonology space .
Approach: They adopt a featural representation of phonemes grounded in phonological theory which models speech sounds as structured bundles of distinctive articulatory and acoustic properties.
Outcome: The proposed model allows us to analyse phoneme confusions at a finer granularity and to investigate whether certain phonological features are more vulnerable than others.
Constructions Are So Difficult That Even Large Language Models Get Them Right for the Wrong Reasons (2024.lrec-main)

Copied to clipboard

Challenge: In this paper, we examine the ability of large language models (LLMs) to identify different meanings in sentences that are superficially similar.
Approach: They propose a challenge dataset for NLP with large lexical overlap which minimises the possibility of models discerning entailment solely based on token distinctions.
Outcome: The proposed model fails to distinguish between constructions with three classes of adjectives which cannot be distinguished by surface features.
Geometric Signatures of Compositionality Across a Language Model’s Lifetime (2025.acl-long)

Copied to clipboard

Challenge: linguistic compositionality allows atoms to locally combine to create global meaning . a rich array of meanings at the level of a phrase may be explained by simple rules of composition.
Approach: They propose to relate the degree of compositionality in a dataset to the intrinsic dimension of its representations under an LM, a measure of feature complexity.
Outcome: The proposed model is based on a geometric view of the compositionality of a dataset and the intrinsic dimension of its representations under an LM.
How Tokenization Limits Phonological Knowledge Representation in Language Models and How to Improve Them (2026.acl-long)

Copied to clipboard

Challenge: Tokenization is the first step in every language model (LM), yet it never takes the sounds of words into account.
Approach: They propose a lightweight IPA-based fine-tuning method that infuses phonological awareness into LMs.
Outcome: The proposed method improves phonological awareness across three phonology-related tasks while preserving math and general reasoning ability.
A Regex Minimization Benchmark: A PSPACE-Complete Challenge for Language Models (2026.eacl-long)

Copied to clipboard

Challenge: Language models (LMs) have demonstrated impressive reasoning capabilities across domains . but their ability to handle PSPACE-complete problems remains underexplored . a new benchmark for regex minimization is proposed to evaluate LMs' reasoning capabilities .
Approach: They propose a benchmark for regex minimization to evaluate LMs' reasoning power . they use a million regexes paired with their minimal equivalents to evaluate their performance .
Outcome: The proposed model can solve NP-complete problems, but their ability to handle PSPACE-complete ones remains underexplored.
Small Language Models Also Work With Small Vocabularies: Probing the Linguistic Abilities of Grapheme- and Phoneme-Based Baby Llamas (2025.coling-main)

Copied to clipboard

Challenge: Existing studies on LMs have focused on linguistic generalizations and representations from developmentally plausible data.
Approach: They propose to use phoneme- and grapheme-based language models to learn linguistic units at and below the word level.
Outcome: The proposed models can achieve strong performance on syntactic and novel benchmarks and match grapheme-based models in standard tasks and novel evaluations.
How (Non-)Optimal is the Lexicon? (2021.naacl-main)

Copied to clipboard

Challenge: lexical meanings are mapped to wordforms by usage pressures and constraints on sequences of symbols.
Approach: They propose a coding-theoretic view of the lexicon and a novel generative statistical model to quantify its compressibility under various constraints.
Outcome: The proposed model shows that (compositional) morphology and graphotactics can account for most of the complexity of natural codes—as measured by code length.
Optimizing over subsequences generates context-sensitive languages (2021.tacl-1)

Copied to clipboard

Challenge: Optimality Theory is a framework that is commonly used to model phonology but it is known to generate non-finite-state mappings and languages.
Approach: They propose to use Optimality Theory to generate non-context-free languages using constraints defined over subsequences to demonstrate its generative capacity.
Outcome: The proposed framework is capable of generating non-context-free languages with minimal modification as it is standardly employed.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations