KreolMorisienMT: A Dataset for Mauritian Creole Machine Translation (2022.findings-aacl)
Copied to clipboard
| Challenge: | Mauritian Creole is a French-based creole and a lingua franca of the Republic of Mauritius. |
| Approach: | They describe a dataset for benchmarking machine translation quality of Mauritian Creole. |
| Outcome: | The proposed dataset compares KreolMorisienMT with existing models and human evaluation reveals the systems’ high translation quality. |
Similar Papers
Kreyòl-MT: Building MT for Latin American, Caribbean and Colonial African Creole Languages (2024.naacl-long)
Copied to clipboard
Nathaniel Robinson, Raj Dabre, Ammon Shurtz, Rasul Dent, Onenamiyi Onesi, Claire Monroc, Loïc Grobol, Hasan Muhammad, Ashi Garg, Naome Etori, Vijay Murari Tiyyala, Olanrewaju Samuel, Matthew Stutzman, Bismarck Odoom, Sanjeev Khudanpur, Stephen Richardson, Kenton Murray
| Challenge: | Creole languages are used in much of Latin America, Africa and the Caribbean . a large multilingual bitext like ours has potential to build the best yet or first ever MT models for many languages . |
| Approach: | They present the largest cumulative dataset to date for Creole language MT . they provide MT models supporting all 41 Creoles in 172 translation directions . |
| Outcome: | The proposed model outperforms a genre-specific Creole MT model on its own benchmark for 23 of 34 translation directions. |
Automatic Speech Recognition and Query By Example for Creole Languages Documentation (2022.findings-acl)
Copied to clipboard
| Challenge: | CREAM project aims to provide linguists with new methods for language documentation based on automatic speech recognition and keyword-spotting. |
| Approach: | They propose to use one hour of annotated data to design an automatic speech recognition system for two Creole languages. |
| Outcome: | The proposed model is based on an hour of annotated data and is usable by linguists. |
IndicMT Eval: A Dataset to Meta-Evaluate Machine Translation Metrics for Indian Languages (2023.acl-long)
Copied to clipboard
Ananya Sai B, Tanay Dixit, Vignesh Nagarajan, Anoop Kunchukuttan, Pratyush Kumar, Mitesh M. Khapra, Raj Dabre
| Challenge: | Recent studies on machine translation systems focus on high-resource languages, but focus has shifted to low-resourced languages. |
| Approach: | They evaluate 16 metrics from a multidimensional quality metric dataset . they show pre-trained metrics have higher correlations with annotator scores . |
| Outcome: | The proposed evaluations show that pre-trained metrics outperform COMET on Indian languages. |
GuyLingo: The Republic of Guyana Creole Corpora (2024.naacl-short)
Copied to clipboard
| Challenge: | linguistic diversity across the globe encompasses a multitude of smaller, indigenous, and regional languages that lack the same level of computational support. |
| Approach: | They propose a corpus for advancing NLP research in the domain of Creolese in Guyana . they outline a framework for gathering and digitizing this corpus, including colloquial expressions, idioms, and regional variations in a low-resource language . |
| Outcome: | The proposed corpus includes colloquial expressions, idioms, and regional variations in a low-resource language. |
A fine-grained error analysis of NMT, SMT and RBMT output for English-to-Dutch (L18-1)
Copied to clipboard
| Challenge: | Since 2016, the landscape of automated translation has substantially changed with the arrival of neural machine translation (NMT). |
| Approach: | They propose to use an annotated SCATE corpus of MT errors to enrich the SCATE error taxonomy to fit the neural MT output. |
| Outcome: | The proposed system outperforms phrase-based and rule-based systems except for lexical issues. |
The SADID Evaluation Datasets for Low-Resource Spoken Language Machine Translation of Arabic Dialects (2020.coling-main)
Copied to clipboard
| Challenge: | Low-resource Machine Translation (LRT) models are still lagging behind on low-resourced language pairs due to the scarcity of parallel training data. |
| Approach: | They introduce benchmark datasets for Arabic and its dialects to examine their properties . they bootstrap existing parallel sentences and complement this with multilingual training . |
| Outcome: | The proposed method bootstraps existing parallel sentences and complements multilingual training to achieve strong baselines. |
SSA-COMET: Do LLMs Outperform Learned Metrics in Evaluating MT for Under-Resourced African Languages? (2025.emnlp-main)
Copied to clipboard
Senyu Li, Jiayi Wang, Felermino D. M. A. Ali, Colin Cherry, Daniel Deutsch, Eleftheria Briakou, Rui Sousa-Silva, Henrique Lopes Cardoso, Pontus Stenetorp, David Ifeoluwa Adelani
| Challenge: | Existing metrics for machine translation quality for under-resourced African languages suffer from limited language coverage and poor performance in low-resource settings. |
| Approach: | They propose a large-scale human-annotated machine translation evaluation dataset . they use a reference-based and reference-free evaluation model to compare MT quality . |
| Outcome: | The proposed models outperform AfriCOMET and the strongest LLM on low-resource languages. |
KC4MT: A High-Quality Corpus for Multilingual Machine Translation (2022.lrec-1)
Copied to clipboard
Vinh Van Nguyen, Ha Nguyen, Huong Thanh Le, Thai Phuong Nguyen, Tan Van Bui, Luan Nghia Pham, Anh Tuan Phan, Cong Hoang-Minh Nguyen, Viet Hong Tran, Anh Huu Tran
| Challenge: | In machine translation, Vietnamese is a low-resource language, and the quality of the training corpus is very low. |
| Approach: | They propose a method for building high-quality multilingual parallel corpus in news domain . they also publicize a corpus that includes 500.000 Vietnamese-Chinese bilingual sentence pairs . |
| Outcome: | The proposed method improves the quality of multilingual machine translation in Vietnamese, Laos, and Khmer . the public version includes 500.000 Vietnamese-Chinese bilingual sentence pairs . |
AfroMT: Pretraining Strategies and Reproducible Benchmarks for Translation of 8 African Languages (2021.emnlp-main)
Copied to clipboard
| Challenge: | Existing reproducible benchmarks for machine translation are limited to high-resource or well-represented languages. |
| Approach: | They propose to use AfroMT to develop a reproducible machine translation benchmark for eight widely spoken African languages and a suite of analysis tools to take into account their unique properties. |
| Outcome: | The proposed benchmarks show significant improvements when pretraining on 11 languages, with gains of up to 2 BLEU points over strong baselines. |
Error Analysis of Multilingual Language Models in Machine Translation: A Case Study of English-Amharic Translation (2024.emnlp-main)
Copied to clipboard
| Challenge: | Multilingual large language models have significantly advanced machine translation, yet challenges remain for low-resource languages like Amharic. |
| Approach: | They evaluated the performance of NLLB-200 and M2M in English-Amharic bidirectional translation using the Lesan AI dataset. |
| Outcome: | The proposed models outperformed the existing models in English-Amharic bidirectional translation using the Lesan AI dataset. |