Challenge: Despite their cultural and historical significance, Native American languages remain unsupported by major commercial language identification systems.
Approach: They propose to curate linguistic resources across all Native American languages for robust training and tailor data augmentation to generate synthetic yet linguistically coherent training samples.
Outcome: The proposed system would be generalizable across all Native American languages . it would also generate coherent training samples for low-resource languages based on Plains Apache .

Similar Papers

Visibility as Survival: Generalizing NLP for Native Alaskan Language Identification (2025.findings-acl)

Copied to clipboard

Challenge: Indigenous languages are largely invisible in commercial language identification systems, a stark reality exemplified by Google Translate’s LangID tool, which excludes all 150 Indigenous languages of North America.
Approach: They propose a framework that shows how large language models and specialized classifiers can effectively identify these languages with minimal data.
Outcome: The proposed framework shows that large language models and specialized classifiers can effectively identify these languages with minimal data.
Is It Navajo? Accurate Language Detection for Endangered Athabaskan Languages (2025.naacl-short)

Copied to clipboard

Challenge: Endangered languages are significantly underrepresented in contemporary language technologies, exacerbating the challenges of their preservation and revitalization.
Approach: They propose a random forest classifier trained on Navajo and 20 erroneously suggested languages by Google's Language Identification tool.
Outcome: The proposed classifier achieves near-perfect accuracy across other Athabaskan languages suggesting its potential for broader application.
Native Language Identification with User Generated Content (D18-1)

Copied to clipboard

Challenge: Using both linguistically-motivated features and the characteristics of the social media outlet, we obtain high accuracy on this challenging task.
Approach: They propose to use linguistically-motivated features and social media characteristics to obtain high accuracy on this task.
Outcome: The proposed method is highly accurate on a social media content where authors are highly-fluent nonnative speakers.
A Deep Generative Approach to Native Language Identification (2020.coling-main)

Copied to clipboard

Challenge: Native language identification (NLI) is a multi-class classification task involving multiple features that capture the systematic fingerprints of the first language in the second language writing.
Approach: They propose a deep generative language modelling approach to NLI that fine-tunes a GPT-2 model separately on texts written by the authors with the same L1 and assigns n-grams to an unseen text.
Outcome: The proposed method outperforms traditional machine learning approaches and currently achieves the best results on the benchmark NLI datasets.
Challenges of language technologies for the indigenous languages of the Americas (C18-1)

Copied to clipboard

Challenge: Indigenous languages of the American continent are highly diverse, but have received little attention from the technological perspective.
Approach: They review the research, the digital resources and the available NLP systems for indigenous languages of the American continent . they stress the need of developing language resources and NLP tools for these languages .
Outcome: The authors review the research and the available NLP systems on indigenous languages of the Americas . they argue that the lack of resources and tools can have a negative impact on the communities which depend on these languages .
Evaluating the Efficacy of Large Acoustic Model for Documenting Non-Orthographic Tribal Languages in India (2024.lrec-main)

Copied to clipboard

Challenge: Pre-trained Large Acoustic Models have been shown to improve performance in spoken languages . however, their potential for novel under-resourced languages is not fully known .
Approach: They propose to use pre-trained Large Acoustic Models to document under-resourced languages . they use scripts from languages that hold a prominent presence in the geographical regions .
Outcome: The proposed model can document under-resourced languages in the electronic domain . the model can be used to document languages with a written script .
Evaluating Self-Supervised Speech Representations for Indigenous American Languages (2024.lrec-main)

Copied to clipboard

Challenge: a recent study focused on the use of self-supervised learning to learn speech representations for indigenous languages . aaron e. scott: the vast linguistic diversity represented by indigenous languages remains unexplored . by expanding the scope of language processing to include indigenous languages, we can foster linguistic inclusivity, he says .
Approach: They benchmark the efficacy of large-scale self-supervised learning models on indigenous American languages.
Outcome: The proposed model can generalize to real-world data, showing strong performance . evaluators found that the model performed better than monolingual models on indigenous languages .
Not always about you: Prioritizing community needs when developing endangered language technology (2022.acl-long)

Copied to clipboard

Challenge: low-resource languages lack the quantity of data needed to train statistical and machine learning tools and models.
Approach: They propose to use language technology to support endangered languages' revitalization . they propose to work with indigenous speakers to develop technology for such training .
Outcome: The authors discuss the challenges that researchers and indigenous speech community members face when working together to develop language technology to support endangered languages.
How can NLP Help Revitalize Endangered Languages? A Case Study and Roadmap for the Cherokee Language (2022.acl-long)

Copied to clipboard

Challenge: There are an estimated 6000 to 7000 spoken languages in the world, and at least 43% of them are endangered.
Approach: They propose three principles that may help NLP practitioners foster mutual understanding and collaboration with language communities and three ways in which NLP can potentially assist in language education.
Outcome: The proposed methods can be used to enrich Cherokee language resources with machine-in-the-loop processing and to provide language education.
BigNLI: Native Language Identification with Big Bird Embeddings (2024.lrec-main)

Copied to clipboard

Challenge: Native Language Identification (NLI) is a task that relies on time-consuming linguistic feature engineering and current transformer models are limited by input size.
Approach: They propose to train a logistic regression classifier which only uses Big Bird embeddings to overcome this limitation.
Outcome: The proposed method outperforms linguistic feature engineering models on the Reddit-L2 dataset and shows consistent out-of-sample and out-off-domain performance.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations