Papers by Hour Kaing

4 papers
PrahokBART: A Pre-trained Sequence-to-Sequence Model for Khmer Natural Language Generation (2025.coling-main)

Copied to clipboard

Challenge: Pre-trained sequence-to-sequence models are typically pretrained on extensive raw text corpora and fine-tuned on task-specific data.
Approach: They introduce a pre-trained sequence-to-sequence model trained from scratch for Khmer using carefully curated Khmer and English corpora.
Outcome: The proposed model outperforms existing models on three generative tasks and is data-efficient and effective in enhancing performance across various natural language generation tasks.
Supervised and Unsupervised Machine Translation for Myanmar-English and Khmer-English (D19-52)

Copied to clipboard

Challenge: Using cleaned and normalized noisy monolingual data, supervised neural and statistical machine translation systems performed among the best for the four translation directions.
Approach: They present supervised and unsupervised machine translation systems for the WAT2019 Myanmar-English and Khmer-English translation tasks.
Outcome: The proposed systems performed among the best for the four translation directions.
A Myanmar (Burmese)-English Named Entity Transliteration Dictionary (2020.lrec-1)

Copied to clipboard

Challenge: Currently, there are no data available for the transcription of borrowed English words in Myanmar . lack of resources is a problem for many understudied languages .
Approach: They construct a dictionary of Myanmar-English transliteration instances using a CC BY-NC-SA license.
Outcome: The proposed model outperforms the statistical model significantly on the character level.
Robust Neural Machine Translation for Abugidas by Glyph Perturbation (2024.eacl-short)

Copied to clipboard

Challenge: Neural machine translation systems are vulnerable when trained on limited data.
Approach: They propose to add noise to the training phase to increase robustness of NMT systems trained on limited data.
Outcome: The proposed training strategy overcomes noise and improves robustness for low-resource tasks for abugida glyphs.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations