Challenge: Attention models are often used to justify the model’s decision in generating a token but it has not been rigorously established to what extent attention is a reliable source of information in NMT.
Approach: They propose to use attention models to modify crucial aspects of the trained attention model to produce function and content words in the translation process.
Outcome: The proposed models preserve function and content words in the translation process compared to state-of-the-art models.

Similar Papers

Towards Understanding Neural Machine Translation with Word Importance (D19-1)

Copied to clipboard

Challenge: Neural machine translation (NMT) has advanced the state-of-the-art on various language pairs, but the interpretability of NMT remains unsatisfactory.
Approach: They propose to attribute NMT output to every input word using a gradient-based method to measure word importance.
Outcome: The proposed method is superior on identifying input words with higher influence on translation performance.
Look Harder: A Neural Machine Translation Model with Hard Attention (P19-1)

Copied to clipboard

Challenge: Soft-attention based Neural Machine Translation models attend all the words in the source sequence for each target token, which makes them ineffective for long sequence translation.
Approach: They propose a hard-attention based NMT model which selects a subset of source tokens for each target token to effectively handle long sequence translation.
Outcome: The proposed model performs better on long sequences and achieves significant improvement on English-German and English-French translation tasks compared to soft-attention based models.
Do Multilingual Neural Machine Translation Models Contain Language Pair Specific Attention Heads? (2021.findings-acl)

Copied to clipboard

Challenge: Recent studies on multilingual representations focus on whether there is an emergence of language-independent representations or whether multilingual models partition their weights among different languages.
Approach: They analyze encoder self-attention and encoder-decoder attention heads in a multilingual neural translation model.
Outcome: The proposed model is based on a multilingual neural translation model with a language-independent representation.
Attention Weights in Transformer NMT Fail Aligning Words Between Sequences but Largely Explain Model Predictions (2021.findings-emnlp)

Copied to clipboard

Challenge: Using attention weights, we show that NMT models make alignment errors by relying on uninformative tokens from the source sequence.
Approach: They propose to use attention weights to regulate alignment errors in NMT models . they propose methods that largely reduce the word alignment error rate compared to standard induced alignments from attention weighted tokens.
Outcome: The proposed methods reduce the word alignment error rate compared to standard induced alignments from attention weights.
Training with Adversaries to Improve Faithfulness of Attention in Neural Machine Translation (2020.aacl-srw)

Copied to clipboard

Challenge: Existing approaches to measure faithfulness of neural machine translation models are based on stress tests and a novel objective that rewards faithful behaviour by the model through probability divergence.
Approach: They propose a measure of faithfulness for neural machine translation models based on stress tests and measuring faithfulness based upon how often the model output changes.
Outcome: The proposed objective increases faithfulness without reducing translation quality and can even improve translation quality in some cases.
Is Attention Explanation? An Introduction to the Debate (2022.acl-long)

Copied to clipboard

Challenge: Attention has been used in various tasks of NLP and other fields of machine learning to increase performance and provide some explanations.
Approach: They propose to use attention as an explanation for deep learning models to increase performance . they propose to apply attention weights to queries and queries based on scalar scores .
Outcome: The proposed model can be used to increase performance while providing some explanations.
Neural Hidden Markov Model for Machine Translation (P18-2)

Copied to clipboard

Challenge: Attention-based neural machine translation models selectively focus on specific source positions to produce a translation.
Approach: They propose to replace the attention component with a neural hidden Markov model that selectively focuss on specific source positions to produce a translation.
Outcome: The proposed model performs better than the state-of-the-art attention-based models on the GermanEnglish and ChineseEnglish translation tasks.
Training Deeper Neural Machine Translation Models with Transparent Attention (D18-1)

Copied to clipboard

Challenge: Existing NMT models are shallow in comparison to convolutional models used for both text and vision tasks.
Approach: They propose to modify the attention mechanism to ease the optimization of deeper models by a simple modification to the seq2seq with attention paradigm.
Outcome: The proposed model achieves consistent gains of 0.7-1.1 BLEU on the benchmark WMT’14 English-German and WMT'15 Czech-English tasks.
Attention is not Explanation (N19-1)

Copied to clipboard

Challenge: Attention mechanisms have seen wide adoption in neural NLP models.
Approach: They perform extensive experiments to assess the degree to which attention weights provide meaningful "explanations" they find that attention weighted inputs are often uncorrelated with gradient-based measures of feature importance .
Outcome: The proposed model is based on a distribution over attended-to input units . the findings show that attention weights are often uncorrelated with features .
Measuring and Improving Faithfulness of Attention in Neural Machine Translation (2021.eacl-main)

Copied to clipboard

Challenge: Existing evidence for faithfulness of neural machine translation models is lacking.
Approach: They propose a novel objective that rewards faithful behaviour by the model through probability divergence and a differentiable objective that can increase faithfulness without reducing the translation quality.
Outcome: The proposed objective increases faithfulness without reducing translation quality and can even improve translation quality in some cases.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations