Visibility as Survival: Generalizing NLP for Native Alaskan Language Identification (2025.findings-acl)
Copied to clipboard
| Challenge: | Indigenous languages are largely invisible in commercial language identification systems, a stark reality exemplified by Google Translate’s LangID tool, which excludes all 150 Indigenous languages of North America. |
| Approach: | They propose a framework that shows how large language models and specialized classifiers can effectively identify these languages with minimal data. |
| Outcome: | The proposed framework shows that large language models and specialized classifiers can effectively identify these languages with minimal data. |
Similar Papers
Is It Navajo? Accurate Language Detection for Endangered Athabaskan Languages (2025.naacl-short)
Copied to clipboard
| Challenge: | Endangered languages are significantly underrepresented in contemporary language technologies, exacerbating the challenges of their preservation and revitalization. |
| Approach: | They propose a random forest classifier trained on Navajo and 20 erroneously suggested languages by Google's Language Identification tool. |
| Outcome: | The proposed classifier achieves near-perfect accuracy across other Athabaskan languages suggesting its potential for broader application. |
What is it? Towards a Generalizable Native American Language Identification System (2025.naacl-srw)
Copied to clipboard
| Challenge: | Despite their cultural and historical significance, Native American languages remain unsupported by major commercial language identification systems. |
| Approach: | They propose to curate linguistic resources across all Native American languages for robust training and tailor data augmentation to generate synthetic yet linguistically coherent training samples. |
| Outcome: | The proposed system would be generalizable across all Native American languages . it would also generate coherent training samples for low-resource languages based on Plains Apache . |
How can NLP Help Revitalize Endangered Languages? A Case Study and Roadmap for the Cherokee Language (2022.acl-long)
Copied to clipboard
| Challenge: | There are an estimated 6000 to 7000 spoken languages in the world, and at least 43% of them are endangered. |
| Approach: | They propose three principles that may help NLP practitioners foster mutual understanding and collaboration with language communities and three ways in which NLP can potentially assist in language education. |
| Outcome: | The proposed methods can be used to enrich Cherokee language resources with machine-in-the-loop processing and to provide language education. |
Challenges of language technologies for the indigenous languages of the Americas (C18-1)
Copied to clipboard
| Challenge: | Indigenous languages of the American continent are highly diverse, but have received little attention from the technological perspective. |
| Approach: | They review the research, the digital resources and the available NLP systems for indigenous languages of the American continent . they stress the need of developing language resources and NLP tools for these languages . |
| Outcome: | The authors review the research and the available NLP systems on indigenous languages of the Americas . they argue that the lack of resources and tools can have a negative impact on the communities which depend on these languages . |
Revitalization of Indigenous Languages through Pre-processing and Neural Machine Translation: The case of Inuktitut (2020.coling-main)
Copied to clipboard
| Challenge: | Indigenous languages have been considered low-resource and/or endangered . authors propose a method to revitalize the language spoken in northern canada . |
| Approach: | They propose to revitalize the Inuktitut language through pre-processing and neural machine translation . they propose to use this technique to perform morphological analysis and neural translation tasks . |
| Outcome: | The proposed approach improves the Inuktitut language compared to the state-of-the-art . the proposed approach is based on preprocessing and neural machine translation . |
An Open Dataset and Model for Language Identification (2023.acl-short)
Copied to clipboard
| Challenge: | Existing LID systems perform poorly on low-resource languages, causing 'representation washing', where the community is given a false view of the actual progress of low-source NLP. |
| Approach: | They propose a model which achieves a macro-average F1 score of 0.93 and a false positive rate of 0.033% across 201 languages, outperforming previous work. |
| Outcome: | The proposed model outperforms existing models and datasets on 201 languages and a false positive rate of 0.033%. |
FormosanBench: Benchmarking Low-Resource Austronesian Languages in the Era of Large Language Models (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Existing LLMs consistently underperform across all tasks, with 10-shot learning and fine-tuning offering only limited improvements. |
| Approach: | They introduce FormosanBench, a benchmark for evaluating LLMs on low-resource Austronesian languages. |
| Outcome: | The proposed benchmark covers three endangered Formosan languages: Atayal, Amis, and Paiwan . existing LLMs consistently underperform across all tasks, with 10-shot learning and fine-tuning offering only limited improvements. |
One Country, 700+ Languages: NLP Challenges for Underrepresented Languages and Dialects in Indonesia (2022.acl-long)
Copied to clipboard
Alham Fikri Aji, Genta Indra Winata, Fajri Koto, Samuel Cahyawijaya, Ade Romadhony, Rahmad Mahendra, Kemal Kurniawan, David Moeljadi, Radityo Eko Prasojo, Timothy Baldwin, Jey Han Lau, Sebastian Ruder
| Challenge: | There are more than 700 languages spoken in Indonesia, equal to 10% of the world's languages, second only to Papua New Guinea. |
| Approach: | They focus on the languages spoken in Indonesia, the world's second most linguistically diverse nation, and the fourth most populous nation of the world. |
| Outcome: | The proposed model is based on the languages spoken in Indonesia, the world's second-most linguistically diverse nation, with 273 million people spread over 17,508 islands. |
NLP Progress in Indigenous Latin American Languages (2024.naacl-long)
Copied to clipboard
Atnafu Tonja, Fazlourrahman Balouchzahi, Sabur Butt, Olga Kolesnikova, Hector Ceballos, Alexander Gelbukh, Thamar Solorio
| Challenge: | a new study examines the marginalization of indigenous languages in the face of rapid technological advancements. |
| Approach: | They highlight the cultural richness of indigenous languages and the risk they face of being overlooked in the realm of natural language processing. |
| Outcome: | The authors highlight the cultural richness of indigenous languages and their risk of being overlooked in the realm of natural language processing. |
Language ID in the Wild: Unexpected Challenges on the Path to a Thousand-Language Web Text Corpus (2020.coling-main)
Copied to clipboard
| Challenge: | Large text corpora are increasingly important for a wide variety of NLP tasks. |
| Approach: | They propose to train automatic language identification models on up to 1,629 languages . they find that human-judged accuracy for web-crawl text corpora is only around 5% for many lower-resource languages. |
| Outcome: | The proposed models achieve over 90% average F1 on 1,629 languages . human-judged accuracy for web-crawl text corpora is only around 5% for many lower-resource languages - suggesting a need for more robust evaluation. |