| Challenge: | low-resource language research often hampered due to under-representation of how it is being used in reality. |
| Approach: | They propose to use a dataset comprising both Indonesian and English from personal travelogue articles . they used named and nominal expressions of four entity types related to travel . |
| Outcome: | The proposed dataset is more representative of how Indonesian language is being used in reality. |
Similar Papers
Arukikata Travelogue Dataset with Geographic Entity Mention, Coreference, and Link Annotation (2024.findings-eacl)
Copied to clipboard
Shohei Higashiyama, Hiroki Ouchi, Hiroki Teranishi, Hiroyuki Otomo, Yusuke Ide, Aitaro Yamamoto, Hiroyuki Shindo, Yuki Matsuda, Shoko Wakamiya, Naoya Inoue, Ikuya Yamada, Taro Watanabe
| Challenge: | et al., 2006) considers geographic relatedness among geo-entity mentions in document-level geoparsing. |
| Approach: | They present a Japanese travelogue dataset that considers geographic relatedness among geo-entity mentions. |
| Outcome: | The proposed dataset includes 200 travelogue documents with rich geo-entity information . it shows that human activities, mobility, and events are often described with natural language expressions of locations or geographic entities (geo-entities) |
NusaX: Multilingual Parallel Sentiment Dataset for 10 Indonesian Local Languages (2023.eacl-main)
Copied to clipboard
Genta Indra Winata, Alham Fikri Aji, Samuel Cahyawijaya, Rahmad Mahendra, Fajri Koto, Ade Romadhony, Kemal Kurniawan, David Moeljadi, Radityo Eko Prasojo, Pascale Fung, Timothy Baldwin, Jey Han Lau, Rico Sennrich, Sebastian Ruder
| Challenge: | In Indonesia, many languages are endangered and some are even extinct due to the unavailability of data resources and benchmarks. |
| Approach: | They propose a high-quality multilingual parallel corpus that covers 10 local languages from Indonesia. |
| Outcome: | The proposed resource includes sentiment and machine translation datasets, and bilingual lexicons. |
IndoNLI: A Natural Language Inference Dataset for Indonesian (2021.emnlp-main)
Copied to clipboard
| Challenge: | XLM-R model outperforms other pre-trained models in annotated data. |
| Approach: | They adapt the data collection protocol for MNLI and collect 18K sentence pairs annotated by crowd workers and experts. |
| Outcome: | The proposed dataset outperforms other pre-trained models on the expert-annotated data. |
Liputan6: A Large-scale Indonesian Dataset for Text Summarization (2020.aacl-main)
Copied to clipboard
| Challenge: | Despite having the fourth largest speaker population in the world, 1 Indonesian is under-represented in NLP. |
| Approach: | They propose to use a large-scale Indonesian summarization dataset to test extractive and abstractive summarizing methods. |
| Outcome: | The proposed methods are compared with multilingual and monolingual BERT-based models. |
IndoNLU: Benchmark and Resources for Evaluating Indonesian Natural Language Understanding (2020.aacl-main)
Copied to clipboard
Bryan Wilie, Karissa Vincentio, Genta Indra Winata, Samuel Cahyawijaya, Xiaohong Li, Zhi Yuan Lim, Sidik Soleman, Rahmad Mahendra, Pascale Fung, Syafri Bahar, Ayu Purwarianti
| Challenge: | Despite the availability of data on Indonesian, progress on this language is slow . available datasets are scattered, with a lack of documentation and minimal community engagement. |
| Approach: | They propose a resource for training, evaluation, and benchmarking on Indonesian natural language understanding tasks. |
| Outcome: | The proposed resource includes 12 tasks ranging from single sentence classification to pair-sentences sequence labeling with different levels of complexity. |
DaNewsroom: A Large-scale Danish Summarisation Dataset (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing datasets for automatic summarisation are English-oriented . however, only very limited datasets exist in languages other than English . |
| Approach: | They present the first large-scale non-English dataset specifically curated for automatic summarisation. |
| Outcome: | The proposed dataset is the first for the Danish language and is compared with existing datasets. |
Richer Output for Richer Countries: Uncovering Geographical Disparities in Generated Stories and Travel Recommendations (2025.findings-naacl)
Copied to clipboard
| Challenge: | a large body of work examines language models for biases concerning gender, race, occupation and religion . however, the impact of the encoded geographical knowledge on real-world applications has not been documented . |
| Approach: | They examine large language models for two common scenarios that require geographical knowledge: travel recommendations and geo-anchored story generation. |
| Outcome: | The results show that the language models are biased against poorer countries and poorer socioeconomic conditions. |
NusaCrowd: Open Source Initiative for Indonesian NLP Resources (2023.findings-acl)
Copied to clipboard
Samuel Cahyawijaya, Holy Lovenia, Alham Fikri Aji, Genta Winata, Bryan Wilie, Fajri Koto, Rahmad Mahendra, Christian Wibisono, Ade Romadhony, Karissa Vincentio, Jennifer Santoso, David Moeljadi, Cahya Wirawan, Frederikus Hudi, Muhammad Satrio Wicaksono, Ivan Parmonangan, Ika Alfina, Ilham Firdausi Putra, Samsul Rahmadani, Yulianti Oenang, Ali Septiandri, James Jaya, Kaustubh Dhole, Arie Suryani, Rifki Afina Putri, Dan Su, Keith Stevens, Made Nindyatama Nityasya, Muhammad Adilazuarda, Ryan Hadiwijaya, Ryandito Diandaru, Tiezheng Yu, Vito Ghifari, Wenliang Dai, Yan Xu, Dyah Damapuspita, Haryo Wibowo, Cuk Tho, Ichwanul Karo Karo, Tirana Fatyanosa, Ziwei Ji, Graham Neubig, Timothy Baldwin, Sebastian Ruder, Pascale Fung, Herry Sujaini, Sakriani Sakti, Ayu Purwarianti
| Challenge: | Existing NLP research in Indonesian languages has been held back by factors such as language diversity, orthographic variation, resource limitation and other societal challenges. |
| Approach: | They present a collaborative initiative to collect and unify existing resources for Indonesian languages and open access to previously non-public resources. |
| Outcome: | The results show that the datasets are highly reliable and can be used to generate the first zero-shot benchmarks for natural language understanding and generation in Indonesian and the local languages of Indonesia. |
Scoping natural language processing in Indonesian and Malay for education applications (2022.acl-srw)
Copied to clipboard
| Challenge: | Limited natural language processing resources are available for Indonesian and Malay varieties and are difficult to locate. |
| Approach: | They propose to encourage collaboration and efficiency within NLP in Indonesian and Malay by identifying most published authors and research hubs. |
| Outcome: | The findings suggest that the field is dominated by exploratory corpus work, machine reading of text gathered from the Internet, and sentiment analysis. |
Beyond Counting Datasets: A Survey of Multilingual Dataset Construction and Necessary Resources (2022.findings-emnlp)
Copied to clipboard
| Challenge: | Existing studies have examined the quality of labeled data in non-English languages. |
| Approach: | They annotate how datasets are created, input text and label sources, tools used to build them and what they study. |
| Outcome: | The results show that language-proficient NLP researchers' estimated availability correlates with dataset availability. |