Papers by Kemal Kurniawan
NusaX: Multilingual Parallel Sentiment Dataset for 10 Indonesian Local Languages (2023.eacl-main)
Copied to clipboard
Genta Indra Winata, Alham Fikri Aji, Samuel Cahyawijaya, Rahmad Mahendra, Fajri Koto, Ade Romadhony, Kemal Kurniawan, David Moeljadi, Radityo Eko Prasojo, Pascale Fung, Timothy Baldwin, Jey Han Lau, Rico Sennrich, Sebastian Ruder
| Challenge: | In Indonesia, many languages are endangered and some are even extinct due to the unavailability of data resources and benchmarks. |
| Approach: | They propose a high-quality multilingual parallel corpus that covers 10 local languages from Indonesia. |
| Outcome: | The proposed resource includes sentiment and machine translation datasets, and bilingual lexicons. |
One Country, 700+ Languages: NLP Challenges for Underrepresented Languages and Dialects in Indonesia (2022.acl-long)
Copied to clipboard
Alham Fikri Aji, Genta Indra Winata, Fajri Koto, Samuel Cahyawijaya, Ade Romadhony, Rahmad Mahendra, Kemal Kurniawan, David Moeljadi, Radityo Eko Prasojo, Timothy Baldwin, Jey Han Lau, Sebastian Ruder
| Challenge: | There are more than 700 languages spoken in Indonesia, equal to 10% of the world's languages, second only to Papua New Guinea. |
| Approach: | They focus on the languages spoken in Indonesia, the world's second most linguistically diverse nation, and the fourth most populous nation of the world. |
| Outcome: | The proposed model is based on the languages spoken in Indonesia, the world's second-most linguistically diverse nation, with 273 million people spread over 17,508 islands. |
On the Interplay between Human Label Variation and Model Fairness (2026.findings-eacl)
Copied to clipboard
| Challenge: | Existing studies on the impact of human label variation on model fairness have not explored the interaction between HLV and performance. |
| Approach: | They compare human label variation (HLV) training methods with four other methods . they find that HLV methods improve performance without harming fairness . |
| Outcome: | The proposed methods improve fairness without explicit debiasing under certain configurations. |
Unsupervised Cross-Lingual Transfer of Structured Predictors without Source Data (2022.naacl-main)
Copied to clipboard
| Challenge: | Recent successes of NLP systems require large amounts of labelled data for structured prediction tasks. |
| Approach: | They propose a method for unsupervised transfer from multiple input models for structured prediction using a cross-lingual setup. |
| Outcome: | The proposed method produces less noisy labels for the distant supervision. |
PPT: Parsimonious Parser Transfer for Unsupervised Cross-Lingual Adaptation (2021.eacl-main)
Copied to clipboard
| Challenge: | Existing methods for cross-lingual transfer use implicit supervision to parse low-resource languages without explicit supervision. |
| Approach: | They propose a method for unsupervised cross-lingual transfer that uses their output as implicit supervision as part of self-training on unlabelled text in the target language. |
| Outcome: | The proposed method improves over state-of-the-art models on both distant and nearby languages, despite being conceptually simpler. |