Papers with Tetun
Tulun: Transparent and Adaptable Low-resource Machine Translation (2025.acl-demo)
Copied to clipboard
| Challenge: | a low-resource language that is the lingua franca in Timor-Leste lacks available corpora in the health domain. |
| Approach: | They propose a solution that combines neural MT with large language model-based post-editing guided by existing glossaries and translation memories. |
| Outcome: | The proposed system outperforms both standalone MT and LLM approaches across six low-resource languages on the FLORES dataset. |
Data Collection Pipeline for Low-Resource Languages: A Case Study on Constructing a Tetun Text Corpus (2024.lrec-main)
Copied to clipboard
| Challenge: | Labadain Crawler is a data collection pipeline designed to automate and optimize the process of constructing textual corpora from the web, with a specific target to low-resource languages. |
| Approach: | They propose a data collection pipeline built on top of Nutch, an open-source web crawler and data extraction framework, and a tokenizer and identifier for Tetun. |
| Outcome: | The proposed pipeline is based on Nutch, an open-source web crawler and data extraction framework, and is tested with Tetun, one of Timor-Leste’s official languages. |