IndoCollex: A Testbed for Morphological Transformation of Indonesian Colloquial Words (2021.findings-acl)
Copied to clipboard
Haryo Akbarianto Wibowo, Made Nindyatama Nityasya, Afra Feyza Akyürek, Suci Fitriany, Alham Fikri Aji, Radityo Eko Prasojo, Derry Tanti Wijaya
| Challenge: | Existing research on word normalization in Indonesian language relies on static dictionaries and machine translation. |
| Approach: | They propose to use Twitter to annotate Indonesian colloquial words with their standard forms and their word formation types/tags to perform morphological word normalization. |
| Outcome: | The proposed dataset analyzes morphological word normalization on Indonesian colloquial Lexicons and provides a baseline for future work. |
Similar Papers
IndoNLI: A Natural Language Inference Dataset for Indonesian (2021.emnlp-main)
Copied to clipboard
| Challenge: | XLM-R model outperforms other pre-trained models in annotated data. |
| Approach: | They adapt the data collection protocol for MNLI and collect 18K sentence pairs annotated by crowd workers and experts. |
| Outcome: | The proposed dataset outperforms other pre-trained models on the expert-annotated data. |
Normalization of Indonesian-English Code-Mixed Twitter Data (D19-55)
Copied to clipboard
| Challenge: | Twitter is an excellent source of textual data for NLP researches, but it is noisy and often contains typos, slang terms, and non-standard abbreviations. |
| Approach: | They propose a standardization system for Indonesian-English code-mixed Twitter data that includes tokenization, language identification, lexical normalization, and translation. |
| Outcome: | The proposed standardization system is based on four modules for tokenization, language identification, lexical normalization, and translation. |
Hierarchical Mapping for Crosslingual Word Embedding Alignment (2020.tacl-1)
Copied to clipboard
| Challenge: | Existing strategies that map word embeddings into a crosslingual space are biased towards the choice of the pivot language. |
| Approach: | They propose to map any two languages into a different middle space by learning mappings across languages in a hierarchical way. |
| Outcome: | The proposed strategy significantly improves vocabulary induction scores in all existing benchmarks and in a new non-English–centered benchmark. |
Are Girls Neko or Shōjo? Cross-Lingual Alignment of Non-Isomorphic Embeddings with Iterative Normalization (P19-1)
Copied to clipboard
| Challenge: | Cross-lingual word embeddings (CLWE) are used to perform multilingual natural language processing tasks. |
| Approach: | They propose a method that transforms monolingual embeddings to make orthogonal alignment easier by simultaneously enforcing that (1) individual word vectors are unit length, and (2) each language’s average vector is zero. |
| Outcome: | The proposed method improves translation accuracy of three CLWE methods, with the largest improvement observed on English-Japanese (2% to 44% test accuracy). |
Morphology-Aware Meta-Embeddings for Tamil (2021.naacl-srw)
Copied to clipboard
| Challenge: | In this work, we focus on producing morphologically enhanced word embeddings for Tamil, a highly agglutinative South Indian language with rich morphology that remains low-resource with regards to NLP tasks. |
| Approach: | They present a first-ever word analogy dataset for Tamil using a rules-based segmenter and meta-embedding techniques. |
| Outcome: | The proposed embeddings outperform baselines on the word analogy task by 16% and appear to mitigate a trade-off between semantic and morphological accuracy. |
IndoNLU: Benchmark and Resources for Evaluating Indonesian Natural Language Understanding (2020.aacl-main)
Copied to clipboard
Bryan Wilie, Karissa Vincentio, Genta Indra Winata, Samuel Cahyawijaya, Xiaohong Li, Zhi Yuan Lim, Sidik Soleman, Rahmad Mahendra, Pascale Fung, Syafri Bahar, Ayu Purwarianti
| Challenge: | Despite the availability of data on Indonesian, progress on this language is slow . available datasets are scattered, with a lack of documentation and minimal community engagement. |
| Approach: | They propose a resource for training, evaluation, and benchmarking on Indonesian natural language understanding tasks. |
| Outcome: | The proposed resource includes 12 tasks ranging from single sentence classification to pair-sentences sequence labeling with different levels of complexity. |
NusaX: Multilingual Parallel Sentiment Dataset for 10 Indonesian Local Languages (2023.eacl-main)
Copied to clipboard
Genta Indra Winata, Alham Fikri Aji, Samuel Cahyawijaya, Rahmad Mahendra, Fajri Koto, Ade Romadhony, Kemal Kurniawan, David Moeljadi, Radityo Eko Prasojo, Pascale Fung, Timothy Baldwin, Jey Han Lau, Rico Sennrich, Sebastian Ruder
| Challenge: | In Indonesia, many languages are endangered and some are even extinct due to the unavailability of data resources and benchmarks. |
| Approach: | They propose a high-quality multilingual parallel corpus that covers 10 local languages from Indonesia. |
| Outcome: | The proposed resource includes sentiment and machine translation datasets, and bilingual lexicons. |
Wiktionary Normalization of Translations and Morphological Information (2020.coling-main)
Copied to clipboard
| Challenge: | We extend the Yawipa Wiktionary Parser to extract and normalize translations from etymology glosses and morphological form-of relations. |
| Approach: | They extend Yawipa to extract and normalize translations from etymology glosses . they propose a method to identify typos in translation annotations based on extracted morphological data . |
| Outcome: | The proposed method improves on a standard attention baseline by using copy attention. |
Scoping natural language processing in Indonesian and Malay for education applications (2022.acl-srw)
Copied to clipboard
| Challenge: | Limited natural language processing resources are available for Indonesian and Malay varieties and are difficult to locate. |
| Approach: | They propose to encourage collaboration and efficiency within NLP in Indonesian and Malay by identifying most published authors and research hubs. |
| Outcome: | The findings suggest that the field is dominated by exploratory corpus work, machine reading of text gathered from the Internet, and sentiment analysis. |
IndoCL: Benchmarking Indonesian Language Development Assessment (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Recent interest has surged in applying natural language processing (NLP) and machine learning (ML) to evaluate language development in both first (L1) and second (L2) language acquisition. |
| Approach: | They propose to use an Indonesian corpus as a benchmark for LDA tasks and to use existing large-scale language models to improve performance. |
| Outcome: | The proposed model extracts language-independent features, relieving laborious computation and reliance on specific language. |