Papers with Indonesian
MERLIN: A Testbed for Multilingual Multimodal Entity Recognition and Linking (2026.tacl-1)
Copied to clipboard
Sathyanarayanan Ramamoorthy, Vishwa Shah, Simran Khanuja, Zaid Sheikh, Shan Jie, Ann Chia, Shearman Chua, Graham Neubig
| Challenge: | Existing methods for multilingual entity linking are limited by textual contexts and limited resources. |
| Approach: | They propose a testbed system for multilingual multimodal entity linking using BBC news articles paired with corresponding images in five languages. |
| Outcome: | The proposed system improves accuracy for entities with ambiguous textual contexts and models with weak multilingual abilities. |
NusaX: Multilingual Parallel Sentiment Dataset for 10 Indonesian Local Languages (2023.eacl-main)
Copied to clipboard
Genta Indra Winata, Alham Fikri Aji, Samuel Cahyawijaya, Rahmad Mahendra, Fajri Koto, Ade Romadhony, Kemal Kurniawan, David Moeljadi, Radityo Eko Prasojo, Pascale Fung, Timothy Baldwin, Jey Han Lau, Rico Sennrich, Sebastian Ruder
| Challenge: | In Indonesia, many languages are endangered and some are even extinct due to the unavailability of data resources and benchmarks. |
| Approach: | They propose a high-quality multilingual parallel corpus that covers 10 local languages from Indonesia. |
| Outcome: | The proposed resource includes sentiment and machine translation datasets, and bilingual lexicons. |
IndoLEM and IndoBERT: A Benchmark Dataset and Pre-trained Language Model for Indonesian NLP (2020.coling-main)
Copied to clipboard
| Challenge: | despite being spoken by 200 million people, the Indonesian language is underrepresented in NLP research. |
| Approach: | They propose a dataset for Indonesian that includes seven NLP tasks . they also propose 'indonesian language evaluation Montage' tasks that are based on previous work . |
| Outcome: | The proposed dataset shows that IndoBERT outperforms IndoLEM over most of the tasks. |
An Information-Theoretic Approach and Dataset for Probing Gender Stereotypes in Multilingual Masked Language Models (2022.findings-naacl)
Copied to clipboard
| Challenge: | Pretrained language models (PLMs) have been shown to encapsulate social biases, including those relating to gender and race. |
| Approach: | They propose a new bias measure based on Jensen–Shannon divergence that retains more information from the model output probabilities than other previously proposed bias measures. |
| Outcome: | The proposed measure outperforms CrowS-Pairs and other similar measures for non-English datasets. |
COPAL-ID: Indonesian Language Reasoning with Local Culture and Nuances (2024.naacl-long)
Copied to clipboard
| Challenge: | Existing multilingual language models struggle to capture local nuances and contexts that vary from culture to culture. |
| Approach: | They propose a public Indonesian language common sense reasoning dataset COPAL-ID . it incorporates Indonesian local and cultural nuances and provides a more natural portrayal of causal reasoning . |
| Outcome: | The proposed dataset is fluent and free from awkward phrases, unlike the previous dataset. |
Building Open Javanese and Sundanese Corpora for Multilingual Text-to-Speech (L18-1)
Copied to clipboard
Jaka Aris Eko Wibawa, Supheakmungkol Sarin, Chenfang Li, Knot Pipatsrisawat, Keshan Sodimana, Oddur Kjartansson, Alexander Gutkin, Martin Jansche, Linne Ha
| Challenge: | Using multi-speaker text-to-speech systems, we build systems for Javanese and Sundanese . progress in this direction is difficult because languages in the long tail of the distribution of the majority of the world's languages lack adequate linguistic resources . |
| Approach: | They present multi-speaker text-to-speech corpora for Javanese and Sundanese . they use mixed-gender recordings to build multi-language text-based systems . |
| Outcome: | The proposed multi-speaker text-to-speech systems outperform the systems constructed from a single language. |
Constructing Indonesian-English Travelogue Dataset (2024.lrec-main)
Copied to clipboard
| Challenge: | low-resource language research often hampered due to under-representation of how it is being used in reality. |
| Approach: | They propose to use a dataset comprising both Indonesian and English from personal travelogue articles . they used named and nominal expressions of four entity types related to travel . |
| Outcome: | The proposed dataset is more representative of how Indonesian language is being used in reality. |
IDK-MRC: Unanswerable Questions for Indonesian Machine Reading Comprehension (2022.emnlp-main)
Copied to clipboard
| Challenge: | Existing MRC datasets in Indonesian are inadequate because of the small size and limited question types. |
| Approach: | They propose to combine automatic and manual unanswerable question generation to minimize the cost of manual dataset construction while maintaining the dataset quality. |
| Outcome: | The proposed dataset significantly improves the performance of Indonesian MRC models, showing a large improvement for unanswerable questions. |
Improving Low-Resource Named Entity Recognition using Joint Sentence and Token Labeling (2020.acl-main)
Copied to clipboard
| Challenge: | Existing models for named entity recognition (NER) use sentence-level labels, which are expensive to obtain, to improve NER. |
| Approach: | They propose a sentence-level named entity recognition model that uses sentence-based labels that are easy to obtain. |
| Outcome: | The proposed model produces 3.78%, 4.20%, 2.08% improvements in F1 over the baseline on e-commerce product titles in Vietnamese, Thai, and Indonesian, respectively. |
Beyond Film Subtitles: Is YouTube the Best Approximation of Spoken Vocabulary? (2025.coling-main)
Copied to clipboard
Adam Nohejl, Frederikus Hudi, Eunike Andriani Kardinata, Shintaro Ozaki, Maria Angelica Riera Machin, Hongyu Sun, Justin Vasselli, Taro Watanabe
| Challenge: | Word frequency is a key variable in psycholinguistics, useful for modeling human familiarity with words . a recent study shows that frequency from YouTube subtitles is comparable to and often better than the best available resources. |
| Approach: | They use YouTube subtitles to construct frequency norms for five languages . they find they are comparable to and often better than the best currently available resources . |
| Outcome: | The proposed method improves on the best currently available resources for Chinese, English, Indonesian, Japanese, and Spanish. |
SLABERT Talk Pretty One Day: Modeling Second Language Acquisition with BERT (2023.acl-long)
Copied to clipboard
| Challenge: | NLP literature has not given enough attention to the phenomenon of negative transfer . positive transfer refers to the facilitating effects of one language in acquiring another and negative transfer refer to the negative effects between the learner's native [L1] and target [L2] languages. |
| Approach: | They build a Mutlilingual Age Ordered CHILDES dataset to understand the degree to which native Child-Directed Speech (CDS) can help or conflict with English language acquisition. |
| Outcome: | The proposed model enables us to understand the degree to which native Child-Directed Speech (CDS) can help or conflict with English language acquisition. |
IndoNLG: Benchmark and Resources for Evaluating Indonesian Natural Language Generation (2021.emnlp-main)
Copied to clipboard
Samuel Cahyawijaya, Genta Indra Winata, Bryan Wilie, Karissa Vincentio, Xiaohong Li, Adhiguna Kuncoro, Sebastian Ruder, Zhi Yuan Lim, Syafri Bahar, Masayu Khodra, Ayu Purwarianti, Pascale Fung
| Challenge: | Lack of publicly available NLG benchmarks for low-resource languages poses a challenge . authors show that IndoBART and IndoGPT achieve competitive performance on all tasks . |
| Approach: | They propose a benchmark to measure natural language generation progress in three low-resource languages of Indonesia . they use a corpus of pretraining datasets to build their models . |
| Outcome: | The proposed benchmark measures progress in Indonesian, Javanese, and Sundanese . the results highlight the importance of pretraining on closely related, localized languages . |
Visually Grounded Reasoning across Languages and Cultures (2021.emnlp-main)
Copied to clipboard
| Challenge: | a new protocol allows for a multilingual hierarchy of concepts and images based on native speakers . the results suggest that the current models are not robust enough to handle multilingual data . |
| Approach: | They propose a protocol to construct an ImageNet-style hierarchy representative of more languages and cultures. |
| Outcome: | The proposed protocol lets the selection of concepts and images be entirely driven by native speakers, rather than scraping them automatically. |
Rethinking Annotation: Can Language Learners Contribute? (2023.acl-long)
Copied to clipboard
| Challenge: | Researchers have traditionally recruited native speakers to provide annotations for benchmark datasets, but there are languages for which recruiting native speakers is difficult. |
| Approach: | They recruit 36 language learners and provide two types of additional resources and perform mini-tests to measure their language proficiency. |
| Outcome: | The proposed method improves learners' language proficiency in terms of vocabulary and grammar. |
NusaCrowd: Open Source Initiative for Indonesian NLP Resources (2023.findings-acl)
Copied to clipboard
Samuel Cahyawijaya, Holy Lovenia, Alham Fikri Aji, Genta Winata, Bryan Wilie, Fajri Koto, Rahmad Mahendra, Christian Wibisono, Ade Romadhony, Karissa Vincentio, Jennifer Santoso, David Moeljadi, Cahya Wirawan, Frederikus Hudi, Muhammad Satrio Wicaksono, Ivan Parmonangan, Ika Alfina, Ilham Firdausi Putra, Samsul Rahmadani, Yulianti Oenang, Ali Septiandri, James Jaya, Kaustubh Dhole, Arie Suryani, Rifki Afina Putri, Dan Su, Keith Stevens, Made Nindyatama Nityasya, Muhammad Adilazuarda, Ryan Hadiwijaya, Ryandito Diandaru, Tiezheng Yu, Vito Ghifari, Wenliang Dai, Yan Xu, Dyah Damapuspita, Haryo Wibowo, Cuk Tho, Ichwanul Karo Karo, Tirana Fatyanosa, Ziwei Ji, Graham Neubig, Timothy Baldwin, Sebastian Ruder, Pascale Fung, Herry Sujaini, Sakriani Sakti, Ayu Purwarianti
| Challenge: | Existing NLP research in Indonesian languages has been held back by factors such as language diversity, orthographic variation, resource limitation and other societal challenges. |
| Approach: | They present a collaborative initiative to collect and unify existing resources for Indonesian languages and open access to previously non-public resources. |
| Outcome: | The results show that the datasets are highly reliable and can be used to generate the first zero-shot benchmarks for natural language understanding and generation in Indonesian and the local languages of Indonesia. |
LORAXBENCH: A Multitask, Multilingual Benchmark Suite for 20 Indonesian Languages (2025.emnlp-main)
Copied to clipboard
| Challenge: | LORAXBENCH is a benchmark for low-resource languages of Indonesia . it covers reading comprehension, open domain QA, language inference, causal reasoning, translation, and cultural question answering across 20 languages. |
| Approach: | They propose a benchmark that focuses on low-resource languages of Indonesia and covers 6 diverse tasks: reading comprehension, open-domain QA, language inference, causal reasoning, translation, and cultural question answering. |
| Outcome: | The proposed benchmark covers reading comprehension, open-domain QA, language inference, causal reasoning, translation, and cultural question answering across 20 Indonesian languages. |
Monolingual Paraphrase Detection Corpus for Low Resource Pashto Language at Sentence Level (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing research on sentence-level paraphrase detection in Pashto has focused on English, but no work has been done on low-resource Pashtone. |
| Approach: | They propose to annotate sentences in Pashto to detect paraphrases . they will publicize a subset of 1,800 instances from their corpus, free from licensing issues. |
| Outcome: | The proposed corpus contains 6,727 sentences, encompassing 3,687 paraphrased and 3,040 non-paraphrased sentences. |
Re-Evaluating Evaluation for Multilingual Summarization (2024.emnlp-main)
Copied to clipboard
Jessica Forde, Ruochen Zhang, Lintang Sutawika, Alham Aji, Samuel Cahyawijaya, Genta Winata, Minghao Wu, Carsten Eickhoff, Stella Biderman, Ellie Pavlick
| Challenge: | Existing studies have shown that automated evaluation approaches correlate with human ratings in English, but this is unclear for other languages. |
| Approach: | They construct a small-scale pilot dataset containing article-summary pairs and human ratings in English, Chinese and Indonesian to measure the strength of summaries. |
| Outcome: | The results show that standard metrics are unreliable measures of quality in Chinese and Indonesian. |
Can LLM Generate Culturally Relevant Commonsense QA Data? Case Study in Indonesian and Sundanese (2024.emnlp-main)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are increasingly being used to generate synthetic data for training and evaluating models. |
| Approach: | They investigate the effectiveness of using Large Language Models to generate culturally relevant commonsense QA datasets for Indonesian and Sundanese languages using both LLMs and human annotators. |
| Outcome: | The proposed model generates 4.5K questions per language, compared with 4.5k for Indonesian and 4.5km for Sundanese. |
Do Language Models Understand Honorific Systems in Javanese? (2025.acl-long)
Copied to clipboard
Mohammad Rifqi Farhansyah, Iwan Darmawan, Adryan Kusumawardhana, Genta Indra Winata, Alham Fikri Aji, Derry Tanti Wijaya
| Challenge: | Despite its cultural and linguistic significance, there has been limited progress in developing a comprehensive corpus to capture these variations for natural language processing (NLP) tasks. |
| Approach: | They propose to use a dataset to capture the nuances of Unggah-Ungguh Basa, the Javanese speech etiquette framework, to assess the ability of language models to process various levels of Javanesi honorifics. |
| Outcome: | The proposed dataset encapsulates the nuances of Unggah-Ungguh Basa, the Javanese speech etiquette framework. |