Challenge: Existing research on word normalization in Indonesian language relies on static dictionaries and machine translation.
Approach: They propose to use Twitter to annotate Indonesian colloquial words with their standard forms and their word formation types/tags to perform morphological word normalization.
Outcome: The proposed dataset analyzes morphological word normalization on Indonesian colloquial Lexicons and provides a baseline for future work.

Similar Papers

IndoNLI: A Natural Language Inference Dataset for Indonesian (2021.emnlp-main)

Copied to clipboard

Challenge: XLM-R model outperforms other pre-trained models in annotated data.
Approach: They adapt the data collection protocol for MNLI and collect 18K sentence pairs annotated by crowd workers and experts.
Outcome: The proposed dataset outperforms other pre-trained models on the expert-annotated data.
Normalization of Indonesian-English Code-Mixed Twitter Data (D19-55)

Copied to clipboard

Challenge: Twitter is an excellent source of textual data for NLP researches, but it is noisy and often contains typos, slang terms, and non-standard abbreviations.
Approach: They propose a standardization system for Indonesian-English code-mixed Twitter data that includes tokenization, language identification, lexical normalization, and translation.
Outcome: The proposed standardization system is based on four modules for tokenization, language identification, lexical normalization, and translation.
Hierarchical Mapping for Crosslingual Word Embedding Alignment (2020.tacl-1)

Copied to clipboard

Challenge: Existing strategies that map word embeddings into a crosslingual space are biased towards the choice of the pivot language.
Approach: They propose to map any two languages into a different middle space by learning mappings across languages in a hierarchical way.
Outcome: The proposed strategy significantly improves vocabulary induction scores in all existing benchmarks and in a new non-English–centered benchmark.
Are Girls Neko or Shōjo? Cross-Lingual Alignment of Non-Isomorphic Embeddings with Iterative Normalization (P19-1)

Copied to clipboard

Challenge: Cross-lingual word embeddings (CLWE) are used to perform multilingual natural language processing tasks.
Approach: They propose a method that transforms monolingual embeddings to make orthogonal alignment easier by simultaneously enforcing that (1) individual word vectors are unit length, and (2) each language’s average vector is zero.
Outcome: The proposed method improves translation accuracy of three CLWE methods, with the largest improvement observed on English-Japanese (2% to 44% test accuracy).
Morphology-Aware Meta-Embeddings for Tamil (2021.naacl-srw)

Copied to clipboard

Challenge: In this work, we focus on producing morphologically enhanced word embeddings for Tamil, a highly agglutinative South Indian language with rich morphology that remains low-resource with regards to NLP tasks.
Approach: They present a first-ever word analogy dataset for Tamil using a rules-based segmenter and meta-embedding techniques.
Outcome: The proposed embeddings outperform baselines on the word analogy task by 16% and appear to mitigate a trade-off between semantic and morphological accuracy.
IndoNLU: Benchmark and Resources for Evaluating Indonesian Natural Language Understanding (2020.aacl-main)

Copied to clipboard

Challenge: Despite the availability of data on Indonesian, progress on this language is slow . available datasets are scattered, with a lack of documentation and minimal community engagement.
Approach: They propose a resource for training, evaluation, and benchmarking on Indonesian natural language understanding tasks.
Outcome: The proposed resource includes 12 tasks ranging from single sentence classification to pair-sentences sequence labeling with different levels of complexity.
NusaX: Multilingual Parallel Sentiment Dataset for 10 Indonesian Local Languages (2023.eacl-main)

Copied to clipboard

Challenge: In Indonesia, many languages are endangered and some are even extinct due to the unavailability of data resources and benchmarks.
Approach: They propose a high-quality multilingual parallel corpus that covers 10 local languages from Indonesia.
Outcome: The proposed resource includes sentiment and machine translation datasets, and bilingual lexicons.
Wiktionary Normalization of Translations and Morphological Information (2020.coling-main)

Copied to clipboard

Challenge: We extend the Yawipa Wiktionary Parser to extract and normalize translations from etymology glosses and morphological form-of relations.
Approach: They extend Yawipa to extract and normalize translations from etymology glosses . they propose a method to identify typos in translation annotations based on extracted morphological data .
Outcome: The proposed method improves on a standard attention baseline by using copy attention.
Scoping natural language processing in Indonesian and Malay for education applications (2022.acl-srw)

Copied to clipboard

Challenge: Limited natural language processing resources are available for Indonesian and Malay varieties and are difficult to locate.
Approach: They propose to encourage collaboration and efficiency within NLP in Indonesian and Malay by identifying most published authors and research hubs.
Outcome: The findings suggest that the field is dominated by exploratory corpus work, machine reading of text gathered from the Internet, and sentiment analysis.
IndoCL: Benchmarking Indonesian Language Development Assessment (2024.findings-emnlp)

Copied to clipboard

Challenge: Recent interest has surged in applying natural language processing (NLP) and machine learning (ML) to evaluate language development in both first (L1) and second (L2) language acquisition.
Approach: They propose to use an Indonesian corpus as a benchmark for LDA tasks and to use existing large-scale language models to improve performance.
Outcome: The proposed model extracts language-independent features, relieving laborious computation and reliance on specific language.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations