Challenge: Current machine translation systems for low-resource languages have a particular failure mode: they tend to confuse words within a domain.
Approach: They propose a recall-based metric to measure the failure mode of machine translation systems for low-resource languages.
Outcome: The proposed model outperforms a lexicon-based translator in 122 low-resource languages.

Similar Papers

Translation or Recitation? Calibrating Evaluation Scores for Machine Translation of Extremely Low-Resource Languages (2026.acl-short)

Copied to clipboard

Challenge: Existing studies show that performance across low-resource settings is variable, resulting in a significant barrier for the MT community.
Approach: They propose to use FRED Difficulty Metrics to contextualize reported performance across different language pairs to determine whether breakthroughs reported in other contexts are artifacts of benchmark collection.
Outcome: The proposed metrics explain a significant portion of result variability rather than model capability.
Improving Low-Resource Machine Translation for Formosan Languages Using Bilingual Lexical Resources (2024.findings-acl)

Copied to clipboard

Challenge: Using bilingual lexicons for low-resource languages can improve machine translation for low resource languages.
Approach: They propose to use bilingual lexicons to improve machine translation for low-resource languages . they use parallel data and bilingual dictionaries to generate pseudo-parallel sentences .
Outcome: The proposed techniques improve translation between Mandarin and Formosan languages and Spanish and Nahuatl, a language pair consisting of languages from completely different language families.
Understanding In-Context Machine Translation for Low-Resource Languages: A Case Study on Manchu (2025.acl-long)

Copied to clipboard

Challenge: In-context machine translation (MT) with large language models can take advantage of linguistic resources such as grammar books and dictionaries.
Approach: They propose to use in-context machine translation (MT) with large language models to take advantage of linguistic resources such as grammar books and dictionaries.
Outcome: The proposed approach can take advantage of dictionaries and grammar books, but its performance is poor for many lowresource languages.
Fixing Rogue Memorization in Many-to-One Multilingual Translators of Extremely-Low-Resource Languages by Rephrasing Training Samples (2024.naacl-long)

Copied to clipboard

Challenge: Existing fine-tuning of large high-resource language models into multilingual machine translators is difficult for extremely lowresource languages.
Approach: They propose to fine-tune large high-resource language models into multilingual machine translators for extremely-lowresource languages such as endangered Indigenous languages.
Outcome: The proposed model halls are reformulated to improve translation accuracy and improve translation quality.
Machine Translation into Low-resource Language Varieties (2021.acl-short)

Copied to clipboard

Challenge: Current machine translation systems generate a "standard" target language, but many languages have multiple varieties that are different from the standard language.
Approach: They propose a framework to rapidly adapt machine translation systems to generate different target varieties . they propose to use no parallel data to generate languages close to, but different from, the standard target language .
Outcome: The proposed model improves on a system that generates Ukrainian and Belarusian in two languages with no parallel data.
Contextual Refinement of Translations: Large Language Models for Sentence and Document-Level Post-Editing (2024.naacl-long)

Copied to clipboard

Challenge: Large language models have demonstrated considerable success in various natural language processing tasks, but their performance in NMT tasks is still underexplored.
Approach: They propose to use LLMs as automatic post-editors rather than direct translators to improve BLEU and COMET performance.
Outcome: The proposed approach improves BLEU but COMET performance compared to in-context learning.
GATITOS: Using a New Multilingual Lexicon for Low-resource Machine Translation (2023.emnlp-main)

Copied to clipboard

Challenge: a new study explores the effectiveness of bilingual lexica in machine translation models . cross-lingual vocabulary alignment is still highly imperfect in these models, despite the success of supervised and self-supervised training.
Approach: They use a resource to improve translation performance on 200-language models . they show that lexica is more reliable than human-translated data .
Outcome: The proposed approach improves on 200-language translation models with lexical data augmentation . the proposed approach is open-source and has 168 tail languages .
Steering Large Language Models for Machine Translation with Finetuning and In-Context Learning (2023.findings-emnlp)

Copied to clipboard

Challenge: Large language models (LLMs) are a promising avenue for machine translation (MT) however, their effectiveness depends on the choice of few-shot examples and they often require extra post-processing due to overgeneration.
Approach: They propose a method that incorporates few-shot examples during finetuning to improve performance on MT tasks.
Outcome: The proposed method outperforms few-shot prompting while eliminating the need for in-context examples.
Very Large-Scale Lexical Resources to Enhance Chinese and Japanese Machine Translation (L18-1)

Copied to clipboard

Challenge: A major issue in machine translation applications is the recognition and translation of named entities.
Approach: They propose to integrate Very Large-Scale Lexical Resources (VLSLR) with lexicons to improve machine translation accuracy.
Outcome: The proposed lexical resources can enhance the quality of MT in general and NMT systems, which currently don't use lexicons.
Grammatical Error Correction through Round-Trip Machine Translation (2023.findings-eacl)

Copied to clipboard

Challenge: A decade ago the idea of using round-trip MT to guide grammatical error correction was not feasible due to the low quality of MT systems of the day.
Approach: They propose to use round-trip machine translation to guide grammatical error correction to preserve meaning while mapping its surface form from one language into another.
Outcome: The proposed system is re-examined across five languages and models of various sizes and yields consistent improvements.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations