Self-generated Replay Memories for Continual Neural Machine Translation (2024.naacl-long)
Copied to clipboard
| Challenge: | Neural Machine Translation systems exhibit strong performance in several different languages, but their ability to learn continuously is limited by catastrophic forgetting. |
| Approach: | They propose a method that leverages a key property of encoder-decoder Transformers, i.e. their generative ability, to continuously learn Neural Machine Translation systems. |
| Outcome: | The proposed approach can counteract catastrophic forgetting without explicit memorization of training data. |
Similar Papers
Continual Learning of Neural Machine Translation within Low Forgetting Risk Regions (2022.emnlp-main)
Copied to clipboard
| Challenge: | Currently, continuous learning methods suffer from catastrophic forgetting problem, causing model to forget previous knowledge while learning new knowledge. |
| Approach: | They propose a two-stage continuous learning method based on local features of the real loss to avoid catastrophic forgetting problem. |
| Outcome: | The proposed method achieves significant improvements on domain adaptation and more challenging language adaptation tasks. |
Generative Replay Inspired by Hippocampal Memory Indexing for Continual Language Learning (2023.eacl-main)
Copied to clipboard
| Challenge: | Continual learning (CL) is a fundamental requirement for human-like general intelligence (Parisi et al., 2019). |
| Approach: | They propose to control sample generation using compressed features of previous training samples by using hippocampal memory indexing to enhance the generative replay. |
| Outcome: | The proposed method outperforms current generative replay methods and generates training samples from previous tasks. |
Learn To Remember: Transformer with Recurrent Memory for Document-Level Machine Translation (2022.findings-naacl)
Copied to clipboard
| Challenge: | Recent studies have shown that the effective use of contextual information between sentences can achieve better performance in document-level machine translation. |
| Approach: | They propose a recurrent memory unit to the Transformer to support the information exchange between the sentence and previous context. |
| Outcome: | The proposed model outperforms the previous work on TED and News by 0.91 s-BLEU and 1.49 d-BLUE on average. |
Is Encoder-Decoder Redundant for Neural Machine Translation? (2022.aacl-main)
Copied to clipboard
| Challenge: | Encoder-decoder architecture is widely adopted for sequence-to-sequence modeling tasks. |
| Approach: | They propose to combine bilingual and multilingual translations to train a language model to do translation. |
| Outcome: | The proposed approach performs on par with the baseline encoder-decoder Transformer . the proposed approach is compared with the translation model in the target language . |
Multi-split Reversible Transformers Can Enhance Neural Machine Translation (2021.eacl-main)
Copied to clipboard
| Challenge: | Large-scale transformers have been shown to improve neural machine translation performance but training these wider and deeper networks could be extremely memory intensive. |
| Approach: | They propose a multi-split based reversible transformer and a backpropagation algorithm that does not need to store activations for most layers. |
| Outcome: | The proposed model outperforms the vanilla transformer by at least 1.4 BLEU points in three datasets. |
Self-Attentive Residual Decoder for Neural Machine Translation (N18-1)
Copied to clipboard
| Challenge: | Neural sequence-to-sequence networks with attention have been used for machine translation . however, the target-side context is limited and the model lacks the ability to capture non-syntactic dependencies among words. |
| Approach: | They propose a sequence-to-sequence network with attention that captures contextual information at each time-step prediction through an attention mechanism. |
| Outcome: | The proposed model outperforms a neural MT baseline and memory and self-attention network on three language pairs. |
Continual-learning for Modelling Low-Resource Languages from Large Language Models (2026.eacl-long)
Copied to clipboard
| Challenge: | Existing models for low-resource languages with catastrophic forgetting pose several challenges, including learning to model multi-lingual scenarios. |
| Approach: | They propose to employ a continual learning strategy using parts-of-speech code-switching and replay adapter strategies to mitigate catastrophic forgetting gap while training LLM from LLM. |
| Outcome: | The proposed architecture is able to train LLMs from LLM and mitigate catastrophic forgetting gap on vision language tasks. |
Syntactically Supervised Transformers for Faster Neural Machine Translation (P19-1)
Copied to clipboard
| Challenge: | Standard decoders for neural machine translation generate a single token per timestep, which slows inference . a series of controlled experiments demonstrates that SynST decodes sentences 5x faster than the baseline autoregressive Transformer. |
| Approach: | They propose a syntactically supervised Transformer that generates all target tokens in one shot . synST is a variant of the Transformer architecture that autoregressively predicts a chunked parse tree . |
| Outcome: | The proposed method decodes sentences 5x faster than the baseline method on En-De and En-Fr datasets while achieving higher BLEU scores. |
Towards Incremental Transformers: An Empirical Analysis of Transformer Models for Incremental NLU (2021.emnlp-main)
Copied to clipboard
| Challenge: | Recent work attempts to apply incremental processing to NLUs but this is computationally expensive and does not scale efficiently for long sequences. |
| Approach: | They propose to apply Transformers incrementally via restart-incrementality by repeatedly feeding, to an unchanged model, increasingly longer input prefixes to produce partial outputs. |
| Outcome: | The proposed model has better incremental performance and faster inference speed compared to the standard Transformer and LT with restart-incrementality, at the cost of part of the non-incremental quality. |
Leitner-Guided Memory Replay for Cross-lingual Continual Learning (2024.naacl-long)
Copied to clipboard
| Challenge: | Various continual learning approaches have proposed to mitigate catastrophic forgetting by restricting the data buffer or limiting the data size of a model. |
| Approach: | They propose to use a human-inspired spaced-repetition technique to prioritize examples for cross-lingual continual learning. |
| Outcome: | The proposed approach significantly and consistently decreases forgetting while maintaining accuracy across natural language understanding tasks, language orders, and languages. |