Papers with LangID
Is It Navajo? Accurate Language Detection for Endangered Athabaskan Languages (2025.naacl-short)
Copied to clipboard
| Challenge: | Endangered languages are significantly underrepresented in contemporary language technologies, exacerbating the challenges of their preservation and revitalization. |
| Approach: | They propose a random forest classifier trained on Navajo and 20 erroneously suggested languages by Google's Language Identification tool. |
| Outcome: | The proposed classifier achieves near-perfect accuracy across other Athabaskan languages suggesting its potential for broader application. |
Language ID in the Wild: Unexpected Challenges on the Path to a Thousand-Language Web Text Corpus (2020.coling-main)
Copied to clipboard
| Challenge: | Large text corpora are increasingly important for a wide variety of NLP tasks. |
| Approach: | They propose to train automatic language identification models on up to 1,629 languages . they find that human-judged accuracy for web-crawl text corpora is only around 5% for many lower-resource languages. |
| Outcome: | The proposed models achieve over 90% average F1 on 1,629 languages . human-judged accuracy for web-crawl text corpora is only around 5% for many lower-resource languages - suggesting a need for more robust evaluation. |
NusaAksara: A Multimodal and Multilingual Benchmark for Preserving Indonesian Indigenous Scripts (2025.acl-long)
Copied to clipboard
Muhammad Farid Adilazuarda, Musa Izzanardi Wijanarko, Lucky Susanto, Khumaisa Nur’aini, Derry Tanti Wijaya, Alham Fikri Aji
| Challenge: | NusaAksara covers 8 scripts across 7 languages, including low-resource languages not commonly seen in NLP benchmarks. |
| Approach: | They propose a benchmark for Indonesian scripts that includes their original scripts and a dataset that includes 8 scripts across 7 languages. |
| Outcome: | The proposed benchmark covers 8 scripts across 7 languages, including low-resource languages not commonly seen in NLP benchmarks. |