Challenge: Recent studies show that non-recurrent architectures outperform RNNs in neural machine translation.
Approach: They hypothesize that CNNs and self-attentional networks could extract semantic features from source text.
Outcome: The proposed architectures outperform RNNs on two tasks: subject-verb agreement and word sense disambiguation.

Similar Papers

How Much Attention Do You Need? A Granular Analysis of Neural Machine Translation Architectures (P18-1)

Copied to clipboard

Challenge: Neural Machine Translation (NMT) has been replaced by convolutional or self-attentional approaches.
Approach: They propose an architecture definition language that allows for a flexible combination of common building blocks.
Outcome: The proposed architectures can bring recurrent and convolutional models close to the Transformer architecture, but not using self-attention.
Assessing the Ability of Self-Attention Networks to Learn Word Order (P19-1)

Copied to clipboard

Challenge: Existing studies have attributed SAN to being weak at learning positional information for sequence modeling due to lack of recurrence structure.
Approach: They propose a word reordering detection task to quantify how well word order information is learned by SAN and RNN.
Outcome: The proposed task quantifies how well word order information learned by SAN and RNN is learned.
How Can Self-Attention Networks Recognize Dyck-n Languages? (2020.findings-emnlp)

Copied to clipboard

Challenge: Recent work has explored the generalized Dyck-n (Dn) languages .
Approach: They compare the performance of two variants of self-attention networks for Dyck-n (Dn) languages with a starting symbol.
Outcome: The proposed model can generalize to longer sequences and deeper dependencies.
Character-Level Translation with Self-attention (2020.acl-main)

Copied to clipboard

Challenge: Existing models for character-level neural machine translation operate on word-level, which makes them memory inefficient because of large vocabulary sizes.
Approach: They propose a transformer-based model and a novel variant that uses convolutions to combine information from nearby characters to facilitate character interactions.
Outcome: The proposed model outperforms the standard transformer model and learns more robust character alignments on bilingual and multilingual translation datasets.
Recurrent Attention for Neural Machine Translation (2021.emnlp-main)

Copied to clipboard

Challenge: Recent research questions the importance of dot-product self-attention in Transformer models and shows that most attention heads learn simple positional patterns.
Approach: They propose a novel mechanism to replace dot-product self-attention with a recurrent atteNtion mechanism that directly learns attention weights without token-to-token interaction.
Outcome: The proposed model outperforms the Transformer model on translation tasks with fewer parameters and inference time.
Recurrent Attention Networks for Long-text Modeling (2023.findings-acl)

Copied to clipboard

Challenge: Existing approaches to encoding long documents using self-attention have been limited by quadratic computational complexities and limited application in long text processing.
Approach: They propose a long-document encoding model that allows the recurrent operation of self-attention.
Outcome: The proposed model extracts global semantics in token-level and document-level representations, making it inherently compatible with both sequential and sequential tasks.
Towards Better Modeling Hierarchical Structure for Self-Attention with Ordered Neurons (D19-1)

Copied to clipboard

Challenge: Recent studies have shown that a hybrid of self-attention networks (SANs) and recurrent neural networks (RNNs) outperforms both individual architectures, while not much is known about why the hybrid models work.
Approach: They propose to use an advanced variant of self-attention networks (SANs) to enhance the strength of hybrid models by introducing a syntax-oriented inductive bias to perform tree-like composition.
Outcome: The proposed model outperforms both individual models and a standard hybrid model on a machine translation task.
Convolutional Self-Attention Networks (N19-1)

Copied to clipboard

Challenge: Existing models of self-attention networks lack the ability to capture dependencies regardless of distance and can be enhanced with multi-head attention.
Approach: They propose a convolutional self-attention network which can be enhanced by multi-head attention by allowing the model to attend to information from different representation subspaces.
Outcome: The proposed model outperforms existing models on improving locality of SANs on different language pairs and model settings.
Phrase-level Self-Attention Networks for Universal Sentence Encoding (D18-1)

Copied to clipboard

Challenge: Phrase-level self-attention networks (PSAN) can capture context dependencies at the phrase level instead of the sentence level.
Approach: They propose to perform self-attention across words inside a phrase to capture context dependencies at the phrase level and use the gated memory updating mechanism to refine each word’s representation hierarchically with longer-term context dependency captured in a larger phrase.
Outcome: The proposed model can achieve state-of-the-art performance across a plethora of NLP tasks including binary and multi-class classification, natural language inference and sentence similarity.
Enhancing Machine Translation with Dependency-Aware Self-Attention (2020.acl-main)

Copied to clipboard

Challenge: Currently, most neural machine translation models rely on pairs of parallel sentences, assuming syntactic information is automatically learned by an attention mechanism.
Approach: They propose a parameter-free, dependency-aware self-attention mechanism that integrates syntactic knowledge into a Transformer model and propose 'a parameter free approach' they also propose - a novel mechanism that improves translation quality for long sentences and in low-resource scenarios.
Outcome: The proposed approach improves translation quality on English-German and English-Turkish translation tasks and in low-resource scenarios.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations