Papers with Indonesian

20 papers
MERLIN: A Testbed for Multilingual Multimodal Entity Recognition and Linking (2026.tacl-1)

Copied to clipboard

Challenge: Existing methods for multilingual entity linking are limited by textual contexts and limited resources.
Approach: They propose a testbed system for multilingual multimodal entity linking using BBC news articles paired with corresponding images in five languages.
Outcome: The proposed system improves accuracy for entities with ambiguous textual contexts and models with weak multilingual abilities.
NusaX: Multilingual Parallel Sentiment Dataset for 10 Indonesian Local Languages (2023.eacl-main)

Copied to clipboard

Challenge: In Indonesia, many languages are endangered and some are even extinct due to the unavailability of data resources and benchmarks.
Approach: They propose a high-quality multilingual parallel corpus that covers 10 local languages from Indonesia.
Outcome: The proposed resource includes sentiment and machine translation datasets, and bilingual lexicons.
IndoLEM and IndoBERT: A Benchmark Dataset and Pre-trained Language Model for Indonesian NLP (2020.coling-main)

Copied to clipboard

Challenge: despite being spoken by 200 million people, the Indonesian language is underrepresented in NLP research.
Approach: They propose a dataset for Indonesian that includes seven NLP tasks . they also propose 'indonesian language evaluation Montage' tasks that are based on previous work .
Outcome: The proposed dataset shows that IndoBERT outperforms IndoLEM over most of the tasks.
An Information-Theoretic Approach and Dataset for Probing Gender Stereotypes in Multilingual Masked Language Models (2022.findings-naacl)

Copied to clipboard

Challenge: Pretrained language models (PLMs) have been shown to encapsulate social biases, including those relating to gender and race.
Approach: They propose a new bias measure based on Jensen–Shannon divergence that retains more information from the model output probabilities than other previously proposed bias measures.
Outcome: The proposed measure outperforms CrowS-Pairs and other similar measures for non-English datasets.
COPAL-ID: Indonesian Language Reasoning with Local Culture and Nuances (2024.naacl-long)

Copied to clipboard

Challenge: Existing multilingual language models struggle to capture local nuances and contexts that vary from culture to culture.
Approach: They propose a public Indonesian language common sense reasoning dataset COPAL-ID . it incorporates Indonesian local and cultural nuances and provides a more natural portrayal of causal reasoning .
Outcome: The proposed dataset is fluent and free from awkward phrases, unlike the previous dataset.
Building Open Javanese and Sundanese Corpora for Multilingual Text-to-Speech (L18-1)

Copied to clipboard

Challenge: Using multi-speaker text-to-speech systems, we build systems for Javanese and Sundanese . progress in this direction is difficult because languages in the long tail of the distribution of the majority of the world's languages lack adequate linguistic resources .
Approach: They present multi-speaker text-to-speech corpora for Javanese and Sundanese . they use mixed-gender recordings to build multi-language text-based systems .
Outcome: The proposed multi-speaker text-to-speech systems outperform the systems constructed from a single language.
Constructing Indonesian-English Travelogue Dataset (2024.lrec-main)

Copied to clipboard

Challenge: low-resource language research often hampered due to under-representation of how it is being used in reality.
Approach: They propose to use a dataset comprising both Indonesian and English from personal travelogue articles . they used named and nominal expressions of four entity types related to travel .
Outcome: The proposed dataset is more representative of how Indonesian language is being used in reality.
IDK-MRC: Unanswerable Questions for Indonesian Machine Reading Comprehension (2022.emnlp-main)

Copied to clipboard

Challenge: Existing MRC datasets in Indonesian are inadequate because of the small size and limited question types.
Approach: They propose to combine automatic and manual unanswerable question generation to minimize the cost of manual dataset construction while maintaining the dataset quality.
Outcome: The proposed dataset significantly improves the performance of Indonesian MRC models, showing a large improvement for unanswerable questions.
Improving Low-Resource Named Entity Recognition using Joint Sentence and Token Labeling (2020.acl-main)

Copied to clipboard

Challenge: Existing models for named entity recognition (NER) use sentence-level labels, which are expensive to obtain, to improve NER.
Approach: They propose a sentence-level named entity recognition model that uses sentence-based labels that are easy to obtain.
Outcome: The proposed model produces 3.78%, 4.20%, 2.08% improvements in F1 over the baseline on e-commerce product titles in Vietnamese, Thai, and Indonesian, respectively.
Beyond Film Subtitles: Is YouTube the Best Approximation of Spoken Vocabulary? (2025.coling-main)

Copied to clipboard

Challenge: Word frequency is a key variable in psycholinguistics, useful for modeling human familiarity with words . a recent study shows that frequency from YouTube subtitles is comparable to and often better than the best available resources.
Approach: They use YouTube subtitles to construct frequency norms for five languages . they find they are comparable to and often better than the best currently available resources .
Outcome: The proposed method improves on the best currently available resources for Chinese, English, Indonesian, Japanese, and Spanish.
SLABERT Talk Pretty One Day: Modeling Second Language Acquisition with BERT (2023.acl-long)

Copied to clipboard

Challenge: NLP literature has not given enough attention to the phenomenon of negative transfer . positive transfer refers to the facilitating effects of one language in acquiring another and negative transfer refer to the negative effects between the learner's native [L1] and target [L2] languages.
Approach: They build a Mutlilingual Age Ordered CHILDES dataset to understand the degree to which native Child-Directed Speech (CDS) can help or conflict with English language acquisition.
Outcome: The proposed model enables us to understand the degree to which native Child-Directed Speech (CDS) can help or conflict with English language acquisition.
IndoNLG: Benchmark and Resources for Evaluating Indonesian Natural Language Generation (2021.emnlp-main)

Copied to clipboard

Challenge: Lack of publicly available NLG benchmarks for low-resource languages poses a challenge . authors show that IndoBART and IndoGPT achieve competitive performance on all tasks .
Approach: They propose a benchmark to measure natural language generation progress in three low-resource languages of Indonesia . they use a corpus of pretraining datasets to build their models .
Outcome: The proposed benchmark measures progress in Indonesian, Javanese, and Sundanese . the results highlight the importance of pretraining on closely related, localized languages .
Visually Grounded Reasoning across Languages and Cultures (2021.emnlp-main)

Copied to clipboard

Challenge: a new protocol allows for a multilingual hierarchy of concepts and images based on native speakers . the results suggest that the current models are not robust enough to handle multilingual data .
Approach: They propose a protocol to construct an ImageNet-style hierarchy representative of more languages and cultures.
Outcome: The proposed protocol lets the selection of concepts and images be entirely driven by native speakers, rather than scraping them automatically.
Rethinking Annotation: Can Language Learners Contribute? (2023.acl-long)

Copied to clipboard

Challenge: Researchers have traditionally recruited native speakers to provide annotations for benchmark datasets, but there are languages for which recruiting native speakers is difficult.
Approach: They recruit 36 language learners and provide two types of additional resources and perform mini-tests to measure their language proficiency.
Outcome: The proposed method improves learners' language proficiency in terms of vocabulary and grammar.
NusaCrowd: Open Source Initiative for Indonesian NLP Resources (2023.findings-acl)

Copied to clipboard

Challenge: Existing NLP research in Indonesian languages has been held back by factors such as language diversity, orthographic variation, resource limitation and other societal challenges.
Approach: They present a collaborative initiative to collect and unify existing resources for Indonesian languages and open access to previously non-public resources.
Outcome: The results show that the datasets are highly reliable and can be used to generate the first zero-shot benchmarks for natural language understanding and generation in Indonesian and the local languages of Indonesia.
LORAXBENCH: A Multitask, Multilingual Benchmark Suite for 20 Indonesian Languages (2025.emnlp-main)

Copied to clipboard

Challenge: LORAXBENCH is a benchmark for low-resource languages of Indonesia . it covers reading comprehension, open domain QA, language inference, causal reasoning, translation, and cultural question answering across 20 languages.
Approach: They propose a benchmark that focuses on low-resource languages of Indonesia and covers 6 diverse tasks: reading comprehension, open-domain QA, language inference, causal reasoning, translation, and cultural question answering.
Outcome: The proposed benchmark covers reading comprehension, open-domain QA, language inference, causal reasoning, translation, and cultural question answering across 20 Indonesian languages.
Monolingual Paraphrase Detection Corpus for Low Resource Pashto Language at Sentence Level (2024.lrec-main)

Copied to clipboard

Challenge: Existing research on sentence-level paraphrase detection in Pashto has focused on English, but no work has been done on low-resource Pashtone.
Approach: They propose to annotate sentences in Pashto to detect paraphrases . they will publicize a subset of 1,800 instances from their corpus, free from licensing issues.
Outcome: The proposed corpus contains 6,727 sentences, encompassing 3,687 paraphrased and 3,040 non-paraphrased sentences.
Re-Evaluating Evaluation for Multilingual Summarization (2024.emnlp-main)

Copied to clipboard

Challenge: Existing studies have shown that automated evaluation approaches correlate with human ratings in English, but this is unclear for other languages.
Approach: They construct a small-scale pilot dataset containing article-summary pairs and human ratings in English, Chinese and Indonesian to measure the strength of summaries.
Outcome: The results show that standard metrics are unreliable measures of quality in Chinese and Indonesian.
Can LLM Generate Culturally Relevant Commonsense QA Data? Case Study in Indonesian and Sundanese (2024.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) are increasingly being used to generate synthetic data for training and evaluating models.
Approach: They investigate the effectiveness of using Large Language Models to generate culturally relevant commonsense QA datasets for Indonesian and Sundanese languages using both LLMs and human annotators.
Outcome: The proposed model generates 4.5K questions per language, compared with 4.5k for Indonesian and 4.5km for Sundanese.
Do Language Models Understand Honorific Systems in Javanese? (2025.acl-long)

Copied to clipboard

Challenge: Despite its cultural and linguistic significance, there has been limited progress in developing a comprehensive corpus to capture these variations for natural language processing (NLP) tasks.
Approach: They propose to use a dataset to capture the nuances of Unggah-Ungguh Basa, the Javanese speech etiquette framework, to assess the ability of language models to process various levels of Javanesi honorifics.
Outcome: The proposed dataset encapsulates the nuances of Unggah-Ungguh Basa, the Javanese speech etiquette framework.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations