Challenge: Neural Machine Translation systems exhibit strong performance in several different languages, but their ability to learn continuously is limited by catastrophic forgetting.
Approach: They propose a method that leverages a key property of encoder-decoder Transformers, i.e. their generative ability, to continuously learn Neural Machine Translation systems.
Outcome: The proposed approach can counteract catastrophic forgetting without explicit memorization of training data.

Similar Papers

Continual Learning of Neural Machine Translation within Low Forgetting Risk Regions (2022.emnlp-main)

Copied to clipboard

Challenge: Currently, continuous learning methods suffer from catastrophic forgetting problem, causing model to forget previous knowledge while learning new knowledge.
Approach: They propose a two-stage continuous learning method based on local features of the real loss to avoid catastrophic forgetting problem.
Outcome: The proposed method achieves significant improvements on domain adaptation and more challenging language adaptation tasks.
Generative Replay Inspired by Hippocampal Memory Indexing for Continual Language Learning (2023.eacl-main)

Copied to clipboard

Challenge: Continual learning (CL) is a fundamental requirement for human-like general intelligence (Parisi et al., 2019).
Approach: They propose to control sample generation using compressed features of previous training samples by using hippocampal memory indexing to enhance the generative replay.
Outcome: The proposed method outperforms current generative replay methods and generates training samples from previous tasks.
Learn To Remember: Transformer with Recurrent Memory for Document-Level Machine Translation (2022.findings-naacl)

Copied to clipboard

Challenge: Recent studies have shown that the effective use of contextual information between sentences can achieve better performance in document-level machine translation.
Approach: They propose a recurrent memory unit to the Transformer to support the information exchange between the sentence and previous context.
Outcome: The proposed model outperforms the previous work on TED and News by 0.91 s-BLEU and 1.49 d-BLUE on average.
Is Encoder-Decoder Redundant for Neural Machine Translation? (2022.aacl-main)

Copied to clipboard

Challenge: Encoder-decoder architecture is widely adopted for sequence-to-sequence modeling tasks.
Approach: They propose to combine bilingual and multilingual translations to train a language model to do translation.
Outcome: The proposed approach performs on par with the baseline encoder-decoder Transformer . the proposed approach is compared with the translation model in the target language .
Multi-split Reversible Transformers Can Enhance Neural Machine Translation (2021.eacl-main)

Copied to clipboard

Challenge: Large-scale transformers have been shown to improve neural machine translation performance but training these wider and deeper networks could be extremely memory intensive.
Approach: They propose a multi-split based reversible transformer and a backpropagation algorithm that does not need to store activations for most layers.
Outcome: The proposed model outperforms the vanilla transformer by at least 1.4 BLEU points in three datasets.
Self-Attentive Residual Decoder for Neural Machine Translation (N18-1)

Copied to clipboard

Challenge: Neural sequence-to-sequence networks with attention have been used for machine translation . however, the target-side context is limited and the model lacks the ability to capture non-syntactic dependencies among words.
Approach: They propose a sequence-to-sequence network with attention that captures contextual information at each time-step prediction through an attention mechanism.
Outcome: The proposed model outperforms a neural MT baseline and memory and self-attention network on three language pairs.
Continual-learning for Modelling Low-Resource Languages from Large Language Models (2026.eacl-long)

Copied to clipboard

Challenge: Existing models for low-resource languages with catastrophic forgetting pose several challenges, including learning to model multi-lingual scenarios.
Approach: They propose to employ a continual learning strategy using parts-of-speech code-switching and replay adapter strategies to mitigate catastrophic forgetting gap while training LLM from LLM.
Outcome: The proposed architecture is able to train LLMs from LLM and mitigate catastrophic forgetting gap on vision language tasks.
Syntactically Supervised Transformers for Faster Neural Machine Translation (P19-1)

Copied to clipboard

Challenge: Standard decoders for neural machine translation generate a single token per timestep, which slows inference . a series of controlled experiments demonstrates that SynST decodes sentences 5x faster than the baseline autoregressive Transformer.
Approach: They propose a syntactically supervised Transformer that generates all target tokens in one shot . synST is a variant of the Transformer architecture that autoregressively predicts a chunked parse tree .
Outcome: The proposed method decodes sentences 5x faster than the baseline method on En-De and En-Fr datasets while achieving higher BLEU scores.
Towards Incremental Transformers: An Empirical Analysis of Transformer Models for Incremental NLU (2021.emnlp-main)

Copied to clipboard

Challenge: Recent work attempts to apply incremental processing to NLUs but this is computationally expensive and does not scale efficiently for long sequences.
Approach: They propose to apply Transformers incrementally via restart-incrementality by repeatedly feeding, to an unchanged model, increasingly longer input prefixes to produce partial outputs.
Outcome: The proposed model has better incremental performance and faster inference speed compared to the standard Transformer and LT with restart-incrementality, at the cost of part of the non-incremental quality.
Leitner-Guided Memory Replay for Cross-lingual Continual Learning (2024.naacl-long)

Copied to clipboard

Challenge: Various continual learning approaches have proposed to mitigate catastrophic forgetting by restricting the data buffer or limiting the data size of a model.
Approach: They propose to use a human-inspired spaced-repetition technique to prioritize examples for cross-lingual continual learning.
Outcome: The proposed approach significantly and consistently decreases forgetting while maintaining accuracy across natural language understanding tasks, language orders, and languages.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations