The Computational Complexity of Distinctive Feature Minimization in Phonology (N18-2)
Copied to clipboard
| Challenge: | a standard assumption in phonology is that finding a minimal feature specification is an automatic part of acquisition and generalization. |
| Approach: | They analyze the problem of determining whether a set of phonemes forms a natural class and find the minimal feature specification for the class. |
| Outcome: | The proposed model is based on a greedy algorithm that fails to find minimal features . the proposed model can be used to find features that are universal across languages . |
Similar Papers
Phonotactic Complexity and Its Trade-offs (2020.tacl-1)
Copied to clipboard
| Challenge: | Existing measures of linguistic complexity are relatively coarse-see, for example, Moran and Blasi (2014) and 2 below for reviews. |
| Approach: | They propose to measure bits per phoneme using the negative log-probability of a word in a language model and a collection of 1016 basic concept words across 106 languages. |
| Outcome: | The proposed measure allows a cross-linguistic comparison of phonotactic complexity across languages. |
Phonotactic Complexity across Dialects (2024.lrec-main)
Copied to clipboard
| Challenge: | Recent studies show a moderate negative correlation between phonotactic complexity and word length in 106 languages. |
| Approach: | They propose to use a phone-level language model to measure phonotactic complexity . they find a tradeoff between word length and phonomactic complex . |
| Outcome: | The proposed model shows that low phonotactic complexity dialects concentrate around capital regions. |
What Do Neural Speech Models Know About Phonology? Evidence from Structured Phoneme Confusions (2026.findings-acl)
Copied to clipboard
| Challenge: | acoustic and phonological models of speech recognition are often limited to the phoneme level . a recent study has shown that phoneme confusions are strongly structured in phonology space . |
| Approach: | They adopt a featural representation of phonemes grounded in phonological theory which models speech sounds as structured bundles of distinctive articulatory and acoustic properties. |
| Outcome: | The proposed model allows us to analyse phoneme confusions at a finer granularity and to investigate whether certain phonological features are more vulnerable than others. |
Constructions Are So Difficult That Even Large Language Models Get Them Right for the Wrong Reasons (2024.lrec-main)
Copied to clipboard
| Challenge: | In this paper, we examine the ability of large language models (LLMs) to identify different meanings in sentences that are superficially similar. |
| Approach: | They propose a challenge dataset for NLP with large lexical overlap which minimises the possibility of models discerning entailment solely based on token distinctions. |
| Outcome: | The proposed model fails to distinguish between constructions with three classes of adjectives which cannot be distinguished by surface features. |
Geometric Signatures of Compositionality Across a Language Model’s Lifetime (2025.acl-long)
Copied to clipboard
| Challenge: | linguistic compositionality allows atoms to locally combine to create global meaning . a rich array of meanings at the level of a phrase may be explained by simple rules of composition. |
| Approach: | They propose to relate the degree of compositionality in a dataset to the intrinsic dimension of its representations under an LM, a measure of feature complexity. |
| Outcome: | The proposed model is based on a geometric view of the compositionality of a dataset and the intrinsic dimension of its representations under an LM. |
How Tokenization Limits Phonological Knowledge Representation in Language Models and How to Improve Them (2026.acl-long)
Copied to clipboard
| Challenge: | Tokenization is the first step in every language model (LM), yet it never takes the sounds of words into account. |
| Approach: | They propose a lightweight IPA-based fine-tuning method that infuses phonological awareness into LMs. |
| Outcome: | The proposed method improves phonological awareness across three phonology-related tasks while preserving math and general reasoning ability. |
A Regex Minimization Benchmark: A PSPACE-Complete Challenge for Language Models (2026.eacl-long)
Copied to clipboard
| Challenge: | Language models (LMs) have demonstrated impressive reasoning capabilities across domains . but their ability to handle PSPACE-complete problems remains underexplored . a new benchmark for regex minimization is proposed to evaluate LMs' reasoning capabilities . |
| Approach: | They propose a benchmark for regex minimization to evaluate LMs' reasoning power . they use a million regexes paired with their minimal equivalents to evaluate their performance . |
| Outcome: | The proposed model can solve NP-complete problems, but their ability to handle PSPACE-complete ones remains underexplored. |
Small Language Models Also Work With Small Vocabularies: Probing the Linguistic Abilities of Grapheme- and Phoneme-Based Baby Llamas (2025.coling-main)
Copied to clipboard
| Challenge: | Existing studies on LMs have focused on linguistic generalizations and representations from developmentally plausible data. |
| Approach: | They propose to use phoneme- and grapheme-based language models to learn linguistic units at and below the word level. |
| Outcome: | The proposed models can achieve strong performance on syntactic and novel benchmarks and match grapheme-based models in standard tasks and novel evaluations. |
How (Non-)Optimal is the Lexicon? (2021.naacl-main)
Copied to clipboard
| Challenge: | lexical meanings are mapped to wordforms by usage pressures and constraints on sequences of symbols. |
| Approach: | They propose a coding-theoretic view of the lexicon and a novel generative statistical model to quantify its compressibility under various constraints. |
| Outcome: | The proposed model shows that (compositional) morphology and graphotactics can account for most of the complexity of natural codes—as measured by code length. |
Optimizing over subsequences generates context-sensitive languages (2021.tacl-1)
Copied to clipboard
| Challenge: | Optimality Theory is a framework that is commonly used to model phonology but it is known to generate non-finite-state mappings and languages. |
| Approach: | They propose to use Optimality Theory to generate non-context-free languages using constraints defined over subsequences to demonstrate its generative capacity. |
| Outcome: | The proposed framework is capable of generating non-context-free languages with minimal modification as it is standardly employed. |