Self-Attentive Residual Decoder for Neural Machine Translation (N18-1)

Copied to clipboard

Challenge: Neural sequence-to-sequence networks with attention have been used for machine translation . however, the target-side context is limited and the model lacks the ability to capture non-syntactic dependencies among words.
Approach: They propose a sequence-to-sequence network with attention that captures contextual information at each time-step prediction through an attention mechanism.
Outcome: The proposed model outperforms a neural MT baseline and memory and self-attention network on three language pairs.

Similar Papers

Refining Source Representations with Relation Networks for Neural Machine Translation (C18-1)

Copied to clipboard

Challenge: Existing neural machine translation frameworks that forget distant information and disregard relationship between source and target words are not effective.
Approach: They propose to use relation networks to learn better representations of the source . they propose to associate source words with each other to help retain their relationships .
Outcome: Experiments show that the proposed approach outperforms the encoder-decoder framework on several datasets.
Recurrent Attention for Neural Machine Translation (2021.emnlp-main)

Copied to clipboard

Challenge: Recent research questions the importance of dot-product self-attention in Transformer models and shows that most attention heads learn simple positional patterns.
Approach: They propose a novel mechanism to replace dot-product self-attention with a recurrent atteNtion mechanism that directly learns attention weights without token-to-token interaction.
Outcome: The proposed model outperforms the Transformer model on translation tasks with fewer parameters and inference time.
Modeling Recurrence for Transformer (N19-1)

Copied to clipboard

Challenge: Existing studies show that the lack of recurrence modeling hinders the development of a translation model.
Approach: They propose to model recurrence for Transformer with an additional recurrent encoder.
Outcome: The proposed model outperforms the deep model on EnglishGerman and ChineseEnglish translation tasks.
Self-generated Replay Memories for Continual Neural Machine Translation (2024.naacl-long)

Copied to clipboard

Challenge: Neural Machine Translation systems exhibit strong performance in several different languages, but their ability to learn continuously is limited by catastrophic forgetting.
Approach: They propose a method that leverages a key property of encoder-decoder Transformers, i.e. their generative ability, to continuously learn Neural Machine Translation systems.
Outcome: The proposed approach can counteract catastrophic forgetting without explicit memorization of training data.
Improving Non-Autoregressive Neural Machine Translation via Modeling Localness (2022.coling-1)

Copied to clipboard

Challenge: Existing non-autoregressive neural machine translation models suffer from poor localization quality due to sequential dependencies within the target sentence.
Approach: They propose to introduce local information into NAT models by explicitly introducing local information about surrounding words into the encoder and decoder sides to achieve localness-aware representations.
Outcome: The proposed method can achieve significant improvements over strong NAT baselines.
Why Self-Attention? A Targeted Evaluation of Neural Machine Translation Architectures (D18-1)

Copied to clipboard

Challenge: Recent studies show that non-recurrent architectures outperform RNNs in neural machine translation.
Approach: They hypothesize that CNNs and self-attentional networks could extract semantic features from source text.
Outcome: The proposed architectures outperform RNNs on two tasks: subject-verb agreement and word sense disambiguation.
Deconvolution-Based Global Decoding for Neural Machine Translation (C18-1)

Copied to clipboard

Challenge: Existing models for Neural Machine Translation (NMT) use Recurrent Neural Network (RNN) to generate translation word by word following a sequential order.
Approach: They propose a Neural Machine Translation (NMT) model that decodes the sequence with the guidance of its structural prediction of the target-side context.
Outcome: The proposed model is more competitive compared with the state-of-the-art methods and reduces repetition with the instruction from the target-side context for decoding.
Context-aware Decoder for Neural Machine Translation using a Target-side Document-Level Language Model (2021.naacl-main)

Copied to clipboard

Challenge: Neural machine translation models that incorporate inter-sentential contexts can be trained only in document-level parallel data with sentential alignments.
Approach: They propose a method to perform context-aware decoding with any pre-trained translation model . their method uses sentence-level parallel data and target-side document-level monolingual data .
Outcome: The proposed method performs context-aware decoding on English to Russian translation using BLEU and contrastive tests.
Document Context Neural Machine Translation with Memory Networks (P18-1)

Copied to clipboard

Challenge: Experimental results show that our model exploits both source and target document context.
Approach: They propose a document-level neural machine translation model which takes both source and target document context into account using memory networks.
Outcome: The proposed model outperforms previous work in terms of BLEU and METEOR in English translations.
Syntax-Based Attention Masking for Neural Machine Translation (2021.naacl-srw)

Copied to clipboard

Challenge: Existing approaches to extend transformers to source-side trees are linearized into sequences, but they are limited by positional encodings.
Approach: They propose a method for extending transformers to source-side trees by using masks based on tree positions . they define a number of masks that limit self-attention based upon relationships among tree nodes .
Outcome: The proposed method improves on translations from English to germany and English to english and germany by +2.1 BLEU.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations