Constructing Indonesian-English Travelogue Dataset (2024.lrec-main)

Copied to clipboard

Challenge: low-resource language research often hampered due to under-representation of how it is being used in reality.
Approach: They propose to use a dataset comprising both Indonesian and English from personal travelogue articles . they used named and nominal expressions of four entity types related to travel .
Outcome: The proposed dataset is more representative of how Indonesian language is being used in reality.

Similar Papers

Arukikata Travelogue Dataset with Geographic Entity Mention, Coreference, and Link Annotation (2024.findings-eacl)

Copied to clipboard

Challenge: et al., 2006) considers geographic relatedness among geo-entity mentions in document-level geoparsing.
Approach: They present a Japanese travelogue dataset that considers geographic relatedness among geo-entity mentions.
Outcome: The proposed dataset includes 200 travelogue documents with rich geo-entity information . it shows that human activities, mobility, and events are often described with natural language expressions of locations or geographic entities (geo-entities)
NusaX: Multilingual Parallel Sentiment Dataset for 10 Indonesian Local Languages (2023.eacl-main)

Copied to clipboard

Challenge: In Indonesia, many languages are endangered and some are even extinct due to the unavailability of data resources and benchmarks.
Approach: They propose a high-quality multilingual parallel corpus that covers 10 local languages from Indonesia.
Outcome: The proposed resource includes sentiment and machine translation datasets, and bilingual lexicons.
IndoNLI: A Natural Language Inference Dataset for Indonesian (2021.emnlp-main)

Copied to clipboard

Challenge: XLM-R model outperforms other pre-trained models in annotated data.
Approach: They adapt the data collection protocol for MNLI and collect 18K sentence pairs annotated by crowd workers and experts.
Outcome: The proposed dataset outperforms other pre-trained models on the expert-annotated data.
Liputan6: A Large-scale Indonesian Dataset for Text Summarization (2020.aacl-main)

Copied to clipboard

Challenge: Despite having the fourth largest speaker population in the world, 1 Indonesian is under-represented in NLP.
Approach: They propose to use a large-scale Indonesian summarization dataset to test extractive and abstractive summarizing methods.
Outcome: The proposed methods are compared with multilingual and monolingual BERT-based models.
IndoNLU: Benchmark and Resources for Evaluating Indonesian Natural Language Understanding (2020.aacl-main)

Copied to clipboard

Challenge: Despite the availability of data on Indonesian, progress on this language is slow . available datasets are scattered, with a lack of documentation and minimal community engagement.
Approach: They propose a resource for training, evaluation, and benchmarking on Indonesian natural language understanding tasks.
Outcome: The proposed resource includes 12 tasks ranging from single sentence classification to pair-sentences sequence labeling with different levels of complexity.
DaNewsroom: A Large-scale Danish Summarisation Dataset (2020.lrec-1)

Copied to clipboard

Challenge: Existing datasets for automatic summarisation are English-oriented . however, only very limited datasets exist in languages other than English .
Approach: They present the first large-scale non-English dataset specifically curated for automatic summarisation.
Outcome: The proposed dataset is the first for the Danish language and is compared with existing datasets.
Richer Output for Richer Countries: Uncovering Geographical Disparities in Generated Stories and Travel Recommendations (2025.findings-naacl)

Copied to clipboard

Challenge: a large body of work examines language models for biases concerning gender, race, occupation and religion . however, the impact of the encoded geographical knowledge on real-world applications has not been documented .
Approach: They examine large language models for two common scenarios that require geographical knowledge: travel recommendations and geo-anchored story generation.
Outcome: The results show that the language models are biased against poorer countries and poorer socioeconomic conditions.
NusaCrowd: Open Source Initiative for Indonesian NLP Resources (2023.findings-acl)

Copied to clipboard

Challenge: Existing NLP research in Indonesian languages has been held back by factors such as language diversity, orthographic variation, resource limitation and other societal challenges.
Approach: They present a collaborative initiative to collect and unify existing resources for Indonesian languages and open access to previously non-public resources.
Outcome: The results show that the datasets are highly reliable and can be used to generate the first zero-shot benchmarks for natural language understanding and generation in Indonesian and the local languages of Indonesia.
Scoping natural language processing in Indonesian and Malay for education applications (2022.acl-srw)

Copied to clipboard

Challenge: Limited natural language processing resources are available for Indonesian and Malay varieties and are difficult to locate.
Approach: They propose to encourage collaboration and efficiency within NLP in Indonesian and Malay by identifying most published authors and research hubs.
Outcome: The findings suggest that the field is dominated by exploratory corpus work, machine reading of text gathered from the Internet, and sentiment analysis.
Beyond Counting Datasets: A Survey of Multilingual Dataset Construction and Necessary Resources (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing studies have examined the quality of labeled data in non-English languages.
Approach: They annotate how datasets are created, input text and label sources, tools used to build them and what they study.
Outcome: The results show that language-proficient NLP researchers' estimated availability correlates with dataset availability.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations