| Challenge: | Recent studies show a moderate negative correlation between phonotactic complexity and word length in 106 languages. |
| Approach: | They propose to use a phone-level language model to measure phonotactic complexity . they find a tradeoff between word length and phonomactic complex . |
| Outcome: | The proposed model shows that low phonotactic complexity dialects concentrate around capital regions. |
Similar Papers
Phonotactic Complexity and Its Trade-offs (2020.tacl-1)
Copied to clipboard
| Challenge: | Existing measures of linguistic complexity are relatively coarse-see, for example, Moran and Blasi (2014) and 2 below for reviews. |
| Approach: | They propose to measure bits per phoneme using the negative log-probability of a word in a language model and a collection of 1016 basic concept words across 106 languages. |
| Outcome: | The proposed measure allows a cross-linguistic comparison of phonotactic complexity across languages. |
Quantifying the Dialect Gap and its Correlates Across Languages (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Historically, studies investigating minority variants of languages have been limited to a select few languages. |
| Approach: | They evaluate state-of-the-art large language models for regional dialects of several high- and low-resource languages and analyze how regional dialect gap is correlated with economic, social, and linguistic factors. |
| Outcome: | The proposed model is compared with two high-use applications and shows that it can solve the regional dialect gap. |
Dialetto, ma Quanto Dialetto? Transcribing and Evaluating Dialects on a Continuum (2025.findings-naacl)
Copied to clipboard
| Challenge: | Recent work on dialect variation in NLP treats dialects as discrete categories . dialect variation is a focus of increasing interest in the field . |
| Approach: | They examine performance differences between Italian dialects by incorporating performance data from different regions of the world. |
| Outcome: | The results show that performance disparities are due to dialects that are more similar to the standard variety. |
Word Complexity is in the Eye of the Beholder (2021.naacl-main)
Copied to clipboard
| Challenge: | Lexical complexity is a subjective notion, yet it is often neglected in lexical simplification and readability systems which use a ”one-size-fits-all” approach. |
| Approach: | They propose to use a dataset of complex words annotated by readers with different backgrounds to investigate which aspects contribute to the notion of lexical complexity. |
| Outcome: | The proposed approach can be replicated in a dataset of complex words annotated by readers with different backgrounds. |
Are All Languages Equally Hard to Language-Model? (N18-2)
Copied to clipboard
| Challenge: | a fair comparison of language models is tricky because of the size of the corpora and the variability of orthographic systems. |
| Approach: | They propose a framework for fair cross-linguistic comparison of language models . they show that in some languages, textual expression is harder to predict with n-gram models compared to LSTM models based on translated text . |
| Outcome: | The proposed framework is based on translated text and language models on 21 languages. |
The Computational Complexity of Distinctive Feature Minimization in Phonology (N18-2)
Copied to clipboard
| Challenge: | a standard assumption in phonology is that finding a minimal feature specification is an automatic part of acquisition and generalization. |
| Approach: | They analyze the problem of determining whether a set of phonemes forms a natural class and find the minimal feature specification for the class. |
| Outcome: | The proposed model is based on a greedy algorithm that fails to find minimal features . the proposed model can be used to find features that are universal across languages . |
Large Language Models Discriminate Against Speakers of German Dialects (2025.emnlp-main)
Copied to clipboard
| Challenge: | In Germany, more than 40% of the population speaks a regional dialect . however, dialect speakers face negative societal stereotypes . |
| Approach: | They construct a corpus that pairs sentences from seven regional German dialects with their standard German counterparts to assess their dialect usage bias. |
| Outcome: | The proposed model reproduces dialect usage bias in association task and decision task. |
Revisiting the Hierarchical Multiscale LSTM (C18-1)
Copied to clipboard
| Challenge: | Hierarchical Multiscale LSTM model learns structure from character input . high complexity of architecture, training and implementations might hinder its applicability . |
| Approach: | They propose to reproduce and ablate hierarchical multiscale LSTM language model and show that simplifying certain aspects of the architecture can improve its performance. |
| Outcome: | The proposed model performs better when simplified and linguistic units are learned by different levels of the model. |
Quantifying Language Variation Acoustically with Few Resources (2022.naacl-main)
Copied to clipboard
| Challenge: | acoustic models represent linguistic information based on massive amounts of data. |
| Approach: | They examine the model's ability to distinguish low-resource (Dutch) regional varieties by extracting embeddings from hidden layers and dynamic time warping. |
| Outcome: | The proposed model outperforms transcription-based models without phonetic transcriptions on the basis of only six seconds of speech. |
What Kind of Language Is Hard to Language-Model? (P19-1)
Copied to clipboard
| Challenge: | a recent study suggests that language models perform poorly across languages. |
| Approach: | They propose a model that fits a paired-sample multiplicative mixed-effects model to obtain language difficulty coefficients from at least-pairwise parallel corpora. |
| Outcome: | The proposed model is able to handle missing data and is aware of inter-sentence variation. |