Accelerating Asynchronous Stochastic Gradient Descent for Neural Machine Translation (D18-1)
Copied to clipboard
| Challenge: | In order to achieve faster training we increase the mini-batch size and scale the learning rate accordingly. |
| Approach: | They propose a technique that delays gradient updates by increasing the mini-batch size to improve the model's convergence. |
| Outcome: | The proposed technique can train a shallow machine translation system 27% faster than an optimized baseline with negligible penalty in BLEU. |
Similar Papers
Combining Global Sparse Gradients with Local Gradients in Distributed Neural Network Training (D19-1)
Copied to clipboard
| Challenge: | In recent years, neural network models have grown dramatically in terms of number of parameters, so exchanging gradients during data-parallel training is costly in terms both of bandwidth and time. |
| Approach: | They propose to combine the compressed global gradient with the local gradient to restore Transformer convergence while RNNs converge faster. |
| Outcome: | The proposed method restores transformer convergence while RNNs converge faster. |
Hybrid-Regressive Paradigm for Accurate and Speed-Robust Neural Machine Translation (2023.findings-acl)
Copied to clipboard
| Challenge: | Autoregressive translation (NAT) is less robust in decoding batch size and hardware settings than NAT. |
| Approach: | They propose a two-stage translation prototype that prompts a small number of AT predictions and fills in previously skipped tokens at once. |
| Outcome: | The proposed translation prototype achieves comparable translation quality with AT while having 1.5x faster inference speed regardless of batch size and device. |
Fully Non-autoregressive Neural Machine Translation: Tricks of the Trade (2021.findings-acl)
Copied to clipboard
| Challenge: | Existing non-autoregressive neural machine translation models are slow to learn the dependency between output tokens. |
| Approach: | They propose to use fully non-autoregressive neural machine translation (NAT) to predict tokens with single forward of neural networks. |
| Outcome: | The proposed model achieves state-of-the-art results on three translation benchmarks with comparable performance to autoregressive and iterative NAT systems. |
Improving Non-Autoregressive Neural Machine Translation via Modeling Localness (2022.coling-1)
Copied to clipboard
| Challenge: | Existing non-autoregressive neural machine translation models suffer from poor localization quality due to sequential dependencies within the target sentence. |
| Approach: | They propose to introduce local information into NAT models by explicitly introducing local information about surrounding words into the encoder and decoder sides to achieve localness-aware representations. |
| Outcome: | The proposed method can achieve significant improvements over strong NAT baselines. |
Non-Autoregressive Neural Machine Translation: A Call for Clarity (2022.emnlp-main)
Copied to clipboard
| Challenge: | Non-autoregressive translation models require a single forward pass to generate the output sequence instead of iteratively producing each predicted token. |
| Approach: | They propose to use a single forward pass to generate the output sequence instead of iteratively producing each predicted token. |
| Outcome: | The proposed models improve translation quality and speed under third-party testing environments. |
Tricks for Training Sparse Translation Models (2022.naacl-main)
Copied to clipboard
| Challenge: | Multitask learning with an unbalanced data distribution skews model learning towards high resource tasks. |
| Approach: | They propose to use a temperature heating mechanism and dense pre-training to mitigate this by training models with a fixed model capacity. |
| Outcome: | The proposed techniques improve performance on two multilingual translation benchmarks compared to BASELayers and Dense scaling baselines and in combination, more than 2x model convergence speed. |
Incremental Decoding and Training Methods for Simultaneous Translation in Neural Machine Translation (N18-2)
Copied to clipboard
| Challenge: | a tunable agent decides the best segmentation strategy for a user-defined BLEU loss and Average Proportion (AP) constraint. |
| Approach: | They propose a tunable agent which decides the best segmentation strategy for a user-defined BLEU loss and average proportion (AP) constraint. |
| Outcome: | The proposed agent outperforms existing Wait-if-diff and Wait-If-worse agents on BLEU with a lower latency. |
Bridging the Gap between Training and Inference for Neural Machine Translation (P19-1)
Copied to clipboard
| Challenge: | Neural Machine Translation generates target words sequentially while at inference it has to generate the entire sequence from scratch. |
| Approach: | They propose to use ground truth and inference to generate target words sequentially while at inference it has to generate the entire sequence from scratch. |
| Outcome: | Experiments on Chinese->English and WMT’14 English->German translation tasks show that the proposed model can achieve significant improvements on multiple datasets. |
Compact Personalized Models for Neural Machine Translation (D18-1)
Copied to clipboard
| Challenge: | a large proportion of model parameters can be frozen during adaptation with minimal or no reduction in translation quality. |
| Approach: | They propose gradient-based domain adaptation methods for self-attentive machine translation models . they encourage structured sparsity in the set of offset tensors during learning . |
| Outcome: | The proposed method achieves high space and time efficiency using sparse models . the results compare the proposed method with incremental adaptation . |
Revisiting the Weaknesses of Reinforcement Learning for Neural Machine Translation (2021.naacl-main)
Copied to clipboard
| Challenge: | In neural sequence-to-sequence learning, Reinforcement Learning (RL) has gained popularity due to the suitability of Policy Gradient (PG) methods for the end-to end training paradigm. |
| Approach: | They propose to let the model explore the output space beyond the reference output that is used for standard cross-entropy minimization by reinforcing model outputs according to their quality, effectively increasing the likelihood of higher-quality samples. |
| Outcome: | The proposed model explores the output space beyond the reference output that is used for cross-entropy minimization, increasing the likelihood of higher-quality samples. |