Sampling and Filtering of Neural Machine Translation Distillation Data (2021.naacl-srw)
Copied to clipboard
| Challenge: | In most of neural machine translation distillation or stealing scenarios, the highest-scoring hypothesis of the target model is used to train a new model. |
| Approach: | They propose to use the highest-scoring hypothesis of the target model (teacher) to train a new model (student). |
| Outcome: | The proposed method improves the performance of MT models in English to Czech and with reference translations. |
Similar Papers
Selective Knowledge Distillation for Neural Machine Translation (2021.acl-long)
Copied to clipboard
| Challenge: | Neural Machine Translation models achieve state-of-the-art performance on many translation benchmarks. |
| Approach: | They propose a protocol that analyzes different impacts of samples by comparing various samples’ partitions. |
| Outcome: | The proposed methods yield up to +1.28 and +0.89 BLEU points improvements over the Transformer baseline, respectively. |
Data Sampling and (In)stability in Machine Translation Evaluation (2023.findings-acl)
Copied to clipboard
| Challenge: | a recent data sampling method skews the annotated data toward shorter documents, not necessarily representative of the full test set. |
| Approach: | They examine different approaches to human evaluation and ranking of machine translation systems at the conference on machine translation . they propose a method that uses available labour budget to sample data in a more representative manner . |
| Outcome: | The proposed method improves representation of document lengths and produces stable rankings of translation quality. |
A Self-Distillation Recipe for Neural Machine Translation (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing methods for Neural Machine Translation (NMT) have been proven effective in improving the performance of computer vision tasks without pre-training a teacher. |
| Approach: | They propose a rank-order augmented Pearson correlation loss and an iterative distillation method to prevent the discrepancy of predictions between the student and a stronger teacher from disturbing the training. |
| Outcome: | The proposed method can lead to significant improvements over the strong Transformer baseline on low/middle/high-resource tasks, obtaining comparable or better performance with fewer layers. |
MT-PATCHER: Selective and Extendable Knowledge Distillation from Large Language Models for Machine Translation (2024.naacl-long)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have shown their strong ability in the field of machine translation, yet they suffer from high computational cost and latency. |
| Approach: | They propose a framework which transfers knowledge from LLMs to existing MT models in a selective, comprehensive and proactive manner. |
| Outcome: | The proposed framework transfers knowledge from LLMs to existing MT models in a selective, comprehensive and proactive manner. |
Beyond the Mode: Sequence-Level Distillation of Multilingual Translation Models for Low-Resource Language Pairs (2025.findings-naacl)
Copied to clipboard
Aarón Galiano-Jiménez, Juan Antonio Pérez-Ortiz, Felipe Sánchez-Martínez, Víctor M. Sánchez-Cartagena
| Challenge: | Existing multilingual pre-trained models for low-resource languages have outperformed those trained from scratch for low resources due to high hardware requirements. |
| Approach: | They propose to use beam search to decode the whole output distribution of the teacher to improve student learning. |
| Outcome: | The proposed methods improve student model performance and reduce gender bias amplification common to beam search based methods. |
Selecting Backtranslated Data from Multiple Sources for Improved Neural Machine Translation (2020.acl-main)
Copied to clipboard
| Challenge: | incorporating backtranslated data from different sources has led to improved results in machine translation (MT) |
| Approach: | They use a low-resource use-case and a high-resourced language pair to test different backtranslation scenarios and employ data selection to optimise the synthetic corpora. |
| Outcome: | The proposed method reduces the amount of data used while maintaining high-quality MT systems. |
Towards Understanding and Improving Knowledge Distillation for Neural Machine Translation (2023.acl-long)
Copied to clipboard
| Challenge: | Existing knowledge distillation techniques for neural machine translation lack special treatment on the top-1 information, which is limiting the potential of KD. |
| Approach: | They propose a method to distill knowledge from top-1 predictions of teachers and a technique to infuse more additional knowledge by distilling on the data without ground-truth targets. |
| Outcome: | The proposed method outperforms the vanilla word-level KD and outperfies the existing methods on three different students with different capacity gaps. |
Why Skip If You Can Combine: A Simple Knowledge Distillation Technique for Intermediate Layers (2020.emnlp-main)
Copied to clipboard
| Challenge: | Existing knowledge distillation techniques are not suitable for deep learning tasks due to memory constraints. |
| Approach: | They propose to combine knowledge from a large teacher network into a student network (S) they propose to use a combinatorial mechanism to inject layer-level supervision from T to S . |
| Outcome: | The proposed model outperforms existing models in PortugueseEnglish, TurkishEnglish and EnglishGerman directions and students trained using it have 50% fewer parameters and can deliver comparable results to 12-layer teachers. |
Online Distilling from Checkpoints for Neural Machine Translation (N19-1)
Copied to clipboard
| Challenge: | Existing neural machine translation models have a deep structure with large amounts of parameters, making them hard to train. |
| Approach: | They propose an online method to generate a teacher model from checkpoints . they show steady improvement over a strong self-attention-based baseline system . |
| Outcome: | The proposed method improves on-the-fly on several datasets and language pairs. |
Collective Wisdom: Improving Low-resource Neural Machine Translation using Adaptive Knowledge Distillation (2020.coling-main)
Copied to clipboard
| Challenge: | Existing approaches to train high-quality NMT models in bilingually low-resource scenarios are limited by the scarcity of parallel sentence-pairs. |
| Approach: | They propose to distill the knowledge of teacher models to a single student model by using knowledge distillation. |
| Outcome: | The proposed approach achieves up to +0.9 BLEU score improvements compared to strong baselines. |