One Country, 700+ Languages: NLP Challenges for Underrepresented Languages and Dialects in Indonesia (2022.acl-long)
Copied to clipboard
Alham Fikri Aji, Genta Indra Winata, Fajri Koto, Samuel Cahyawijaya, Ade Romadhony, Rahmad Mahendra, Kemal Kurniawan, David Moeljadi, Radityo Eko Prasojo, Timothy Baldwin, Jey Han Lau, Sebastian Ruder
| Challenge: | There are more than 700 languages spoken in Indonesia, equal to 10% of the world's languages, second only to Papua New Guinea. |
| Approach: | They focus on the languages spoken in Indonesia, the world's second most linguistically diverse nation, and the fourth most populous nation of the world. |
| Outcome: | The proposed model is based on the languages spoken in Indonesia, the world's second-most linguistically diverse nation, with 273 million people spread over 17,508 islands. |
Similar Papers
What Do Indonesians Really Need from Language Technology? A Nationwide Survey (2025.emnlp-main)
Copied to clipboard
| Challenge: | Despite efforts to develop NLP for Indonesia’s 700+ local languages, progress remains costly due to the need for direct engagement with native speakers. |
| Approach: | They conduct a nationwide survey to assess the actual needs of native Indonesian speakers. |
| Outcome: | The findings indicate that addressing language barriers is the most critical priority . concerns around privacy, bias, and the use of public data highlight the need for greater transparency and clear communication to support broader AI adoption. |
NusaX: Multilingual Parallel Sentiment Dataset for 10 Indonesian Local Languages (2023.eacl-main)
Copied to clipboard
Genta Indra Winata, Alham Fikri Aji, Samuel Cahyawijaya, Rahmad Mahendra, Fajri Koto, Ade Romadhony, Kemal Kurniawan, David Moeljadi, Radityo Eko Prasojo, Pascale Fung, Timothy Baldwin, Jey Han Lau, Rico Sennrich, Sebastian Ruder
| Challenge: | In Indonesia, many languages are endangered and some are even extinct due to the unavailability of data resources and benchmarks. |
| Approach: | They propose a high-quality multilingual parallel corpus that covers 10 local languages from Indonesia. |
| Outcome: | The proposed resource includes sentiment and machine translation datasets, and bilingual lexicons. |
Challenges and Strategies in Cross-Cultural NLP (2022.acl-long)
Copied to clipboard
Daniel Hershcovich, Stella Frank, Heather Lent, Miryam de Lhoneux, Mostafa Abdou, Stephanie Brandl, Emanuele Bugliarello, Laura Cabello Piqueras, Ilias Chalkidis, Ruixiang Cui, Constanza Fierro, Katerina Margatina, Phillip Rust, Anders Søgaard
| Challenge: | Various efforts have been made to accommodate linguistic diversity and serve speakers of many different languages. |
| Approach: | They propose a framework to examine cultural differences in NLP to better serve users . they argue that cultural knowledge, preferences and values can affect NLP practices . |
| Outcome: | The proposed framework examines how cultural knowledge, preferences and values can affect NLP practices. |
Some Languages are More Equal than Others: Probing Deeper into the Linguistic Disparity in the NLP World (2022.aacl-main)
Copied to clipboard
| Challenge: | Linguistic disparity in the NLP world is widely acknowledged, but the reasons behind it are rarely discussed within the field. |
| Approach: | They propose to categorise languages based on speaker population and vitality . they also analyse the distribution of language data resources and amount of NLP/CL research . |
| Outcome: | The proposed model identifies the reasons for the disparity and suggests ways to overcome it. |
Challenges of language technologies for the indigenous languages of the Americas (C18-1)
Copied to clipboard
| Challenge: | Indigenous languages of the American continent are highly diverse, but have received little attention from the technological perspective. |
| Approach: | They review the research, the digital resources and the available NLP systems for indigenous languages of the American continent . they stress the need of developing language resources and NLP tools for these languages . |
| Outcome: | The authors review the research and the available NLP systems on indigenous languages of the Americas . they argue that the lack of resources and tools can have a negative impact on the communities which depend on these languages . |
Bhaasha, Bhāṣā, Zaban: A Survey for Low-Resourced Languages in South Asia – Current Stage and Challenges (2025.findings-emnlp)
Copied to clipboard
| Challenge: | a survey examines the current efforts and challenges of NLP models for South Asian languages . there are more than 650 languages in South Asia, but many have very limited computational resources or are missing from existing models. |
| Approach: | a survey examines efforts and challenges of NLP for South Asian languages . they focus on transformer-based models such as BERT, T5, & GPT . findings highlight substantial issues, including missing data in critical domains . |
| Outcome: | The findings highlight significant issues, including missing data in critical domains . the survey aims to raise awareness within the NLP community for more targeted data curation . |
Scoping natural language processing in Indonesian and Malay for education applications (2022.acl-srw)
Copied to clipboard
| Challenge: | Limited natural language processing resources are available for Indonesian and Malay varieties and are difficult to locate. |
| Approach: | They propose to encourage collaboration and efficiency within NLP in Indonesian and Malay by identifying most published authors and research hubs. |
| Outcome: | The findings suggest that the field is dominated by exploratory corpus work, machine reading of text gathered from the Internet, and sentiment analysis. |
Building Representative Corpora from Illiterate Communities: A Reviewof Challenges and Mitigation Strategies for Developing Countries (2021.eacl-main)
Copied to clipboard
| Challenge: | Existing methods for collecting data from high-income countries (HICs) make implicit assumptions about literacy and internet access, but in low-income and sub-Saharan Africa (SSA) such assumptions may not hold for LICs where the bulk of the population lives. |
| Approach: | They propose a set of practical mitigation strategies to address the under-representation of illiterate communities in NLP corpora. |
| Outcome: | The proposed methods address the under-representation of illiterate communities in NLP corpora and propose mitigation strategies to help future work. |
Charting the Landscape of African NLP: Mapping Progress and Shaping the Road Ahead (2025.emnlp-main)
Copied to clipboard
| Challenge: | African languages are often left behind in state-of-the-art natural language processing systems and large language models. |
| Approach: | They analyze 884 research papers on NLP for African languages published over past five years . they identify key trends shaping the field and outline promising directions . |
| Outcome: | The findings identify key trends shaping the field and outline promising directions . the authors analyze 884 research papers on NLP for African languages published over the past five years . |
The State and Fate of Linguistic Diversity and Inclusion in the NLP World (2020.acl-main)
Copied to clipboard
| Challenge: | a small number of the over 7000 languages of the world are represented in the rapidly evolving language technologies and applications. |
| Approach: | They examine the relationship between types of languages, resources, and their representation in NLP conferences to understand the trajectory that different languages have followed over time. |
| Outcome: | The proposed model will help to bridge the gap between languages and their resources and convince the ACL community to prioritise the resolution of the predicaments highlighted. |