Overcoming Catastrophic Forgetting During Domain Adaptation of Neural Machine Translation (N19-1)
Copied to clipboard
| Challenge: | Neural Machine Translation (NMT) performs poorly without large training corpora. |
| Approach: | They propose a machine learning method that retains the majority of general-domain performance lost in continued training without degrading in-domain. |
| Outcome: | The proposed method retains the majority of general-domain performance lost in continued training without degrading in-domain performances. |
Similar Papers
Domain adapted machine translation: What does catastrophic forgetting forget and why? (2024.emnlp-main)
Copied to clipboard
| Challenge: | Neural Machine Translation (NMT) models can be specialized by domain adaptation, often fine-tuning on a dataset of interest. |
| Approach: | They propose a novel approach to understanding catastrophic forgetting during NMT adaptation by investigating the relationship between the data and the in-domain vocabulary coverage. |
| Outcome: | The proposed model can be specialized by fine-tuning on a domain of interest, but can fail to achieve the predicted quality of the target domain. |
Investigating Catastrophic Forgetting During Continual Training for Neural Machine Translation (2020.coling-main)
Copied to clipboard
| Challenge: | Neural machine translation models suffer from catastrophic forgetting during continual training . models tend to overfit to frequent observations in the in-domain data but forget previously learned knowledge. |
| Approach: | They investigated the causes of catastrophic forgetting in NMT models by examining their parameters and modules. |
| Outcome: | The proposed model forgets previously learned knowledge and swings to fit new data . the results show that some parameters are important for both the general-domain and in-domain translation and the great change of them during continual training brings about the performance decline in general- domain. |
Mitigating the Diminishing Effect of Elastic Weight Consolidation (2022.coling-1)
Copied to clipboard
| Challenge: | Existing work addresses catastrophic forgetting in sequential training by fine-tuning pre-trained language models on different datasets. |
| Approach: | They propose to rescale the components of EWC to mitigate catastrophic forgetting by mixing new and old training data and retraining the model from scratch. |
| Outcome: | The proposed method requires smaller values for the trade-off parameters to achieve comparable results to EWC on natural language inference and fact-checking tasks. |
Improving the Quality Trade-Off for Neural Machine Translation Multi-Domain Adaptation (2021.emnlp-main)
Copied to clipboard
| Challenge: | Building neural machine translation systems to perform well on a specific target domain remains a challenge. |
| Approach: | They propose to train a single NMT system per language pair that performs well across multiple domains. |
| Outcome: | The proposed approach improves the Pareto frontier on this task. |
Domain Adaptive Inference for Neural Machine Translation (P19-1)
Copied to clipboard
| Challenge: | Neural Machine Translation models are effective when trained on broad domains with large datasets, such as news translation. |
| Approach: | They propose a novel approach for adaptive ensemble weighting for Neural Machine Translation by extending Bayesian Interpolation with source information. |
| Outcome: | The proposed approach improves performance on Spanish-English and English-German tasks without the need for the domain label. |
An Empirical Investigation Towards Efficient Multi-Domain Language Model Pre-training (2020.emnlp-main)
Copied to clipboard
| Challenge: | Pre-training large language models is a standard practice in the natural language processing community. |
| Approach: | They propose to use elastic weight consolidation to mitigate catastrophic forgetting when pre-trained large language models are evaluated on generic benchmarks. |
| Outcome: | The proposed model achieves state-of-the-art on out-of domain tasks with minimal pre-training . elastic weight consolidation provides best overall scores yielding only a 0.33% drop in performance across seven generic tasks while remaining competitive in bio-medical tasks. |
Continual Learning of Neural Machine Translation within Low Forgetting Risk Regions (2022.emnlp-main)
Copied to clipboard
| Challenge: | Currently, continuous learning methods suffer from catastrophic forgetting problem, causing model to forget previous knowledge while learning new knowledge. |
| Approach: | They propose a two-stage continuous learning method based on local features of the real loss to avoid catastrophic forgetting problem. |
| Outcome: | The proposed method achieves significant improvements on domain adaptation and more challenging language adaptation tasks. |
Pruning-then-Expanding Model for Domain Adaptation of Neural Machine Translation (2021.naacl-main)
Copied to clipboard
| Challenge: | Existing methods for domain adaptation suffer from catastrophic forgetting, large domain divergence, and model explosion. |
| Approach: | They propose a method which prunes the model and keeps the important neurons or parameters responsible for both general-domain and in-domain translation. |
| Outcome: | The proposed method improves on different language pairs and domains compared with strong baselines. |
Overcoming Catastrophic Forgetting beyond Continual Learning: Balanced Training for Neural Machine Translation (2022.acl-long)
Copied to clipboard
| Challenge: | Neural networks tend to gradually forget the previously learned knowledge when learning multiple tasks sequentially from dynamic data distributions. |
| Approach: | They propose a method that iteratively provides complementary knowledge to student models by dynamically updating teacher models trained on specific data orders. |
| Outcome: | The proposed method improves on multiple machine translation tasks and improves performance over baseline systems. |
More Than Catastrophic Forgetting: Integrating General Capabilities For Domain-Specific LLMs (2024.emnlp-main)
Copied to clipboard
Chengyuan Liu, Yangyang Kang, Shihang Wang, Lizhi Qing, Fubang Zhao, Chao Wu, Changlong Sun, Kun Kuang, Fei Wu
| Challenge: | a recent study shows that performance on general tasks decreases after Large Language Models are fine-tuned on domain-specific tasks. |
| Approach: | They propose a general capability integration approach to integrate general capabilities and domain knowledge within a single instance. |
| Outcome: | The proposed method improves performance on domain-specific tasks by integrating general capabilities and domain knowledge. |