Phonotactic Complexity across Dialects (2024.lrec-main)

Copied to clipboard

Challenge: Recent studies show a moderate negative correlation between phonotactic complexity and word length in 106 languages.
Approach: They propose to use a phone-level language model to measure phonotactic complexity . they find a tradeoff between word length and phonomactic complex .
Outcome: The proposed model shows that low phonotactic complexity dialects concentrate around capital regions.

Similar Papers

Phonotactic Complexity and Its Trade-offs (2020.tacl-1)

Copied to clipboard

Challenge: Existing measures of linguistic complexity are relatively coarse-see, for example, Moran and Blasi (2014) and 2 below for reviews.
Approach: They propose to measure bits per phoneme using the negative log-probability of a word in a language model and a collection of 1016 basic concept words across 106 languages.
Outcome: The proposed measure allows a cross-linguistic comparison of phonotactic complexity across languages.
Quantifying the Dialect Gap and its Correlates Across Languages (2023.findings-emnlp)

Copied to clipboard

Challenge: Historically, studies investigating minority variants of languages have been limited to a select few languages.
Approach: They evaluate state-of-the-art large language models for regional dialects of several high- and low-resource languages and analyze how regional dialect gap is correlated with economic, social, and linguistic factors.
Outcome: The proposed model is compared with two high-use applications and shows that it can solve the regional dialect gap.
Dialetto, ma Quanto Dialetto? Transcribing and Evaluating Dialects on a Continuum (2025.findings-naacl)

Copied to clipboard

Challenge: Recent work on dialect variation in NLP treats dialects as discrete categories . dialect variation is a focus of increasing interest in the field .
Approach: They examine performance differences between Italian dialects by incorporating performance data from different regions of the world.
Outcome: The results show that performance disparities are due to dialects that are more similar to the standard variety.
Word Complexity is in the Eye of the Beholder (2021.naacl-main)

Copied to clipboard

Challenge: Lexical complexity is a subjective notion, yet it is often neglected in lexical simplification and readability systems which use a ”one-size-fits-all” approach.
Approach: They propose to use a dataset of complex words annotated by readers with different backgrounds to investigate which aspects contribute to the notion of lexical complexity.
Outcome: The proposed approach can be replicated in a dataset of complex words annotated by readers with different backgrounds.
Are All Languages Equally Hard to Language-Model? (N18-2)

Copied to clipboard

Challenge: a fair comparison of language models is tricky because of the size of the corpora and the variability of orthographic systems.
Approach: They propose a framework for fair cross-linguistic comparison of language models . they show that in some languages, textual expression is harder to predict with n-gram models compared to LSTM models based on translated text .
Outcome: The proposed framework is based on translated text and language models on 21 languages.
The Computational Complexity of Distinctive Feature Minimization in Phonology (N18-2)

Copied to clipboard

Challenge: a standard assumption in phonology is that finding a minimal feature specification is an automatic part of acquisition and generalization.
Approach: They analyze the problem of determining whether a set of phonemes forms a natural class and find the minimal feature specification for the class.
Outcome: The proposed model is based on a greedy algorithm that fails to find minimal features . the proposed model can be used to find features that are universal across languages .
Large Language Models Discriminate Against Speakers of German Dialects (2025.emnlp-main)

Copied to clipboard

Challenge: In Germany, more than 40% of the population speaks a regional dialect . however, dialect speakers face negative societal stereotypes .
Approach: They construct a corpus that pairs sentences from seven regional German dialects with their standard German counterparts to assess their dialect usage bias.
Outcome: The proposed model reproduces dialect usage bias in association task and decision task.
Revisiting the Hierarchical Multiscale LSTM (C18-1)

Copied to clipboard

Challenge: Hierarchical Multiscale LSTM model learns structure from character input . high complexity of architecture, training and implementations might hinder its applicability .
Approach: They propose to reproduce and ablate hierarchical multiscale LSTM language model and show that simplifying certain aspects of the architecture can improve its performance.
Outcome: The proposed model performs better when simplified and linguistic units are learned by different levels of the model.
Quantifying Language Variation Acoustically with Few Resources (2022.naacl-main)

Copied to clipboard

Challenge: acoustic models represent linguistic information based on massive amounts of data.
Approach: They examine the model's ability to distinguish low-resource (Dutch) regional varieties by extracting embeddings from hidden layers and dynamic time warping.
Outcome: The proposed model outperforms transcription-based models without phonetic transcriptions on the basis of only six seconds of speech.
What Kind of Language Is Hard to Language-Model? (P19-1)

Copied to clipboard

Challenge: a recent study suggests that language models perform poorly across languages.
Approach: They propose a model that fits a paired-sample multiplicative mixed-effects model to obtain language difficulty coefficients from at least-pairwise parallel corpora.
Outcome: The proposed model is able to handle missing data and is aware of inter-sentence variation.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations