What is it? Towards a Generalizable Native American Language Identification System (2025.naacl-srw)
Copied to clipboard
| Challenge: | Despite their cultural and historical significance, Native American languages remain unsupported by major commercial language identification systems. |
| Approach: | They propose to curate linguistic resources across all Native American languages for robust training and tailor data augmentation to generate synthetic yet linguistically coherent training samples. |
| Outcome: | The proposed system would be generalizable across all Native American languages . it would also generate coherent training samples for low-resource languages based on Plains Apache . |
Similar Papers
Visibility as Survival: Generalizing NLP for Native Alaskan Language Identification (2025.findings-acl)
Copied to clipboard
| Challenge: | Indigenous languages are largely invisible in commercial language identification systems, a stark reality exemplified by Google Translate’s LangID tool, which excludes all 150 Indigenous languages of North America. |
| Approach: | They propose a framework that shows how large language models and specialized classifiers can effectively identify these languages with minimal data. |
| Outcome: | The proposed framework shows that large language models and specialized classifiers can effectively identify these languages with minimal data. |
Is It Navajo? Accurate Language Detection for Endangered Athabaskan Languages (2025.naacl-short)
Copied to clipboard
| Challenge: | Endangered languages are significantly underrepresented in contemporary language technologies, exacerbating the challenges of their preservation and revitalization. |
| Approach: | They propose a random forest classifier trained on Navajo and 20 erroneously suggested languages by Google's Language Identification tool. |
| Outcome: | The proposed classifier achieves near-perfect accuracy across other Athabaskan languages suggesting its potential for broader application. |
Native Language Identification with User Generated Content (D18-1)
Copied to clipboard
| Challenge: | Using both linguistically-motivated features and the characteristics of the social media outlet, we obtain high accuracy on this challenging task. |
| Approach: | They propose to use linguistically-motivated features and social media characteristics to obtain high accuracy on this task. |
| Outcome: | The proposed method is highly accurate on a social media content where authors are highly-fluent nonnative speakers. |
A Deep Generative Approach to Native Language Identification (2020.coling-main)
Copied to clipboard
| Challenge: | Native language identification (NLI) is a multi-class classification task involving multiple features that capture the systematic fingerprints of the first language in the second language writing. |
| Approach: | They propose a deep generative language modelling approach to NLI that fine-tunes a GPT-2 model separately on texts written by the authors with the same L1 and assigns n-grams to an unseen text. |
| Outcome: | The proposed method outperforms traditional machine learning approaches and currently achieves the best results on the benchmark NLI datasets. |
Challenges of language technologies for the indigenous languages of the Americas (C18-1)
Copied to clipboard
| Challenge: | Indigenous languages of the American continent are highly diverse, but have received little attention from the technological perspective. |
| Approach: | They review the research, the digital resources and the available NLP systems for indigenous languages of the American continent . they stress the need of developing language resources and NLP tools for these languages . |
| Outcome: | The authors review the research and the available NLP systems on indigenous languages of the Americas . they argue that the lack of resources and tools can have a negative impact on the communities which depend on these languages . |
Evaluating the Efficacy of Large Acoustic Model for Documenting Non-Orthographic Tribal Languages in India (2024.lrec-main)
Copied to clipboard
| Challenge: | Pre-trained Large Acoustic Models have been shown to improve performance in spoken languages . however, their potential for novel under-resourced languages is not fully known . |
| Approach: | They propose to use pre-trained Large Acoustic Models to document under-resourced languages . they use scripts from languages that hold a prominent presence in the geographical regions . |
| Outcome: | The proposed model can document under-resourced languages in the electronic domain . the model can be used to document languages with a written script . |
Evaluating Self-Supervised Speech Representations for Indigenous American Languages (2024.lrec-main)
Copied to clipboard
| Challenge: | a recent study focused on the use of self-supervised learning to learn speech representations for indigenous languages . aaron e. scott: the vast linguistic diversity represented by indigenous languages remains unexplored . by expanding the scope of language processing to include indigenous languages, we can foster linguistic inclusivity, he says . |
| Approach: | They benchmark the efficacy of large-scale self-supervised learning models on indigenous American languages. |
| Outcome: | The proposed model can generalize to real-world data, showing strong performance . evaluators found that the model performed better than monolingual models on indigenous languages . |
Not always about you: Prioritizing community needs when developing endangered language technology (2022.acl-long)
Copied to clipboard
| Challenge: | low-resource languages lack the quantity of data needed to train statistical and machine learning tools and models. |
| Approach: | They propose to use language technology to support endangered languages' revitalization . they propose to work with indigenous speakers to develop technology for such training . |
| Outcome: | The authors discuss the challenges that researchers and indigenous speech community members face when working together to develop language technology to support endangered languages. |
How can NLP Help Revitalize Endangered Languages? A Case Study and Roadmap for the Cherokee Language (2022.acl-long)
Copied to clipboard
| Challenge: | There are an estimated 6000 to 7000 spoken languages in the world, and at least 43% of them are endangered. |
| Approach: | They propose three principles that may help NLP practitioners foster mutual understanding and collaboration with language communities and three ways in which NLP can potentially assist in language education. |
| Outcome: | The proposed methods can be used to enrich Cherokee language resources with machine-in-the-loop processing and to provide language education. |
BigNLI: Native Language Identification with Big Bird Embeddings (2024.lrec-main)
Copied to clipboard
| Challenge: | Native Language Identification (NLI) is a task that relies on time-consuming linguistic feature engineering and current transformer models are limited by input size. |
| Approach: | They propose to train a logistic regression classifier which only uses Big Bird embeddings to overcome this limitation. |
| Outcome: | The proposed method outperforms linguistic feature engineering models on the Reddit-L2 dataset and shows consistent out-of-sample and out-off-domain performance. |