The FLORES Evaluation Datasets for Low-Resource Machine Translation: Nepali–English and Sinhala–English (D19-1)
Copied to clipboard
Francisco Guzmán, Peng-Jen Chen, Myle Ott, Juan Pino, Guillaume Lample, Philipp Koehn, Vishrav Chaudhary, Marc’Aurelio Ranzato
| Challenge: | a vast majority of language pairs in the world are considered low-resource because they have little parallel data available. |
| Approach: | They propose to use a dataset to evaluate methods trained on low-resource language pairs . they report baseline performance using supervised, weakly supervised and semi-supervised settings . |
| Outcome: | The proposed evaluation datasets show that current state-of-the-art methods perform poorly on this benchmark, posing a challenge to the research community working on low-resource MT. |
Similar Papers
The Flores-101 Evaluation Benchmark for Low-Resource and Multilingual Machine Translation (2022.tacl-1)
Copied to clipboard
Naman Goyal, Cynthia Gao, Vishrav Chaudhary, Peng-Jen Chen, Guillaume Wenzek, Da Ju, Sanjana Krishnan, Marc’Aurelio Ranzato, Francisco Guzmán, Angela Fan
| Challenge: | a lack of good evaluation benchmarks hinders progress in low-resource and multilingual machine translation . despite advances in translation quality for a handful of languages, many low-source languages are not even supported by most popular translation engines. |
| Approach: | They propose a high-quality evaluation benchmark for machine translation using 3001 sentences from Wikipedia . they aim to improve evaluation of models on long tail of low-resource languages . |
| Outcome: | The proposed evaluation benchmarks are based on 3001 sentences extracted from Wikipedia . the results show that the models can be used to evaluate multilingual systems . |
Languages Still Left Behind: Toward a Better Multilingual Machine Translation Benchmark (2025.emnlp-main)
Copied to clipboard
Chihiro Taguchi, Seng Mai, Keita Kurabe, Yusuke Sakai, Georgina Agyei, Soudabeh Eslami, David Chiang
| Challenge: | Multilingual machine translation (MT) benchmarks are widely used to evaluate the capabilities of modern MT systems. |
| Approach: | They propose to use a multilingual machine translation benchmark to assess the capabilities of modern machine translation systems. |
| Outcome: | The FLORES+ benchmark claims to maintain a translation quality score of over 90% . however, the data in four languages falls short of the 90% quality standard . |
The SADID Evaluation Datasets for Low-Resource Spoken Language Machine Translation of Arabic Dialects (2020.coling-main)
Copied to clipboard
| Challenge: | Low-resource Machine Translation (LRT) models are still lagging behind on low-resourced language pairs due to the scarcity of parallel training data. |
| Approach: | They introduce benchmark datasets for Arabic and its dialects to examine their properties . they bootstrap existing parallel sentences and complement this with multilingual training . |
| Outcome: | The proposed method bootstraps existing parallel sentences and complements multilingual training to achieve strong baselines. |
Translation or Recitation? Calibrating Evaluation Scores for Machine Translation of Extremely Low-Resource Languages (2026.acl-short)
Copied to clipboard
| Challenge: | Existing studies show that performance across low-resource settings is variable, resulting in a significant barrier for the MT community. |
| Approach: | They propose to use FRED Difficulty Metrics to contextualize reported performance across different language pairs to determine whether breakthroughs reported in other contexts are artifacts of benchmark collection. |
| Outcome: | The proposed metrics explain a significant portion of result variability rather than model capability. |
Harnessing Multilinguality in Unsupervised Machine Translation for Rare Languages (2021.naacl-main)
Copied to clipboard
| Challenge: | Unsupervised translation systems have impressive performance on resource-rich language pairs . however, in more realistic settings, unsupervised systems perform poorly . |
| Approach: | They propose a model for 5 low-resource languages that leverages monolingual and auxiliary parallel data from other high-resourced languages. |
| Outcome: | The proposed model outperforms state-of-the-art models on low-resource languages . it also matches the current state- of-the art model for Nepali-English . |
Why should only High-Resource-Languages have all the fun? Pivot Based Evaluation in Low Resource Setting (2025.coling-main)
Copied to clipboard
| Challenge: | a limited number of evaluation metrics and resources are available for low-resource languages . a pivot-based evaluation framework is proposed to address these limitations . |
| Approach: | They propose a pivot-based evaluation framework that leverages advanced metrics for more meaningful evaluation. |
| Outcome: | The proposed framework extends the coverage of both lexical-based and embedding-based metrics even for languages not directly supported by advanced metrics. |
A Tulu Resource for Machine Translation (2024.lrec-main)
Copied to clipboard
| Challenge: | Using parallel datasets, we train a machine translation system in English–Tulu . |
| Approach: | They present a parallel dataset for English–Tulu translation using human translations into the multilingual machine translation resource FLORES-200. |
| Outcome: | The proposed model outperforms Google Translate by 19 BLEU points (in September 2023). |
SIB-200: A Simple, Inclusive, and Big Evaluation Dataset for Topic Classification in 200+ Languages and Dialects (2024.eacl-long)
Copied to clipboard
David Adelani, Hannah Liu, Xiaoyu Shen, Nikita Vassilyev, Jesujoba Alabi, Yanke Mao, Haonan Gao, En-Shiun Lee
| Challenge: | despite progress in building multilingual language models evaluation is limited to a few languages with available datasets . despite this, we create a large-scale open-sourced benchmark dataset for topic classification in 205 languages and dialects to address the lack of evaluation dataset for Natural Language Understanding (NLU). |
| Approach: | They create a large-scale open-sourced benchmark dataset for topic classification in 205 languages and dialects to address the lack of evaluation dataset for Natural Language Understanding (NLU). |
| Outcome: | The proposed dataset addresses the lack of evaluation dataset for Natural Language Understanding (NLU) for many languages, it is the first publicly available evaluation dataset. |
Improving Low-Resource Machine Translation for Formosan Languages Using Bilingual Lexical Resources (2024.findings-acl)
Copied to clipboard
| Challenge: | Using bilingual lexicons for low-resource languages can improve machine translation for low resource languages. |
| Approach: | They propose to use bilingual lexicons to improve machine translation for low-resource languages . they use parallel data and bilingual dictionaries to generate pseudo-parallel sentences . |
| Outcome: | The proposed techniques improve translation between Mandarin and Formosan languages and Spanish and Nahuatl, a language pair consisting of languages from completely different language families. |
Evaluating Machine Translation Datasets for Low-Web Data Languages: A Gendered Lens (2026.findings-acl)
Copied to clipboard
Hellina Hailu Nigatu, Bethelhem Yemane Mamo, Bontu Fufa Balcha, Debora Taye Tesfaye, Elbethel Daniel Zewdie, Ikram Behiru Nesiru, Jitu Ewnetu Hailu, Senait Mengesha Yayo
| Challenge: | afan oromo, amharic, and tigrinya are low-resourced languages . they are used for training, benchmarks, news, health, and sports . afono o'mara: quantity does not guarantee quality of MT datasets . |
| Approach: | They investigate the quality of machine translation datasets for three low-resourced languages . they found a large skew towards the male gender in the datasets . |
| Outcome: | The results show that training data has large representation of political and religious text, but benchmark datasets focus on news, health, and sports. |