Papers with SANs
Towards Better Modeling Hierarchical Structure for Self-Attention with Ordered Neurons (D19-1)
Copied to clipboard
| Challenge: | Recent studies have shown that a hybrid of self-attention networks (SANs) and recurrent neural networks (RNNs) outperforms both individual architectures, while not much is known about why the hybrid models work. |
| Approach: | They propose to use an advanced variant of self-attention networks (SANs) to enhance the strength of hybrid models by introducing a syntax-oriented inductive bias to perform tree-like composition. |
| Outcome: | The proposed model outperforms both individual models and a standard hybrid model on a machine translation task. |
Self-Attention with Structural Position Representations (D19-1)
Copied to clipboard
| Challenge: | Experimental results show that SANs can't encode positions of input words . SAN's are currently lacking in encoding positions of words based on position-unaware "bagof-words" theory . |
| Approach: | They propose to augment SANs with structural position representations to capture latent structure of input sentence. |
| Outcome: | The proposed approach consistently outperforms the sequential representations on translation tasks. |
Self-Attention with Cross-Lingual Position Representation (2020.acl-main)
Copied to clipboard
| Challenge: | Position encoding (PE) is used to preserve word order information for natural language processing tasks, generating fixed position indices for input sequences. |
| Approach: | They propose to augment SANs with cross-lingual position representations to model bilingually aware latent structure for the input sentence. |
| Outcome: | The proposed model significantly improves translation quality over baselines on EnglishGerman, JapaneseEnglish, and ChineseEnglish translation tasks. |
How Does Selective Mechanism Improve Self-Attention Networks? (2020.acl-main)
Copied to clipboard
| Challenge: | Experimental results show that selective SANs outperform the standard SAN by paying more attention to content words that contribute to the meaning of the sentence. |
| Approach: | They propose to implement selective SANs with a flexible Gumbel-Softmax to improve word order encoding and structure modeling. |
| Outcome: | The proposed system outperforms the standard SANs on several representative NLP tasks including natural language inference, semantic role labelling, and machine translation. |
Increasing Learning Efficiency of Self-Attention Networks through Direct Position Interactions, Learnable Temperature, and Convoluted Attention (2020.coling-main)
Copied to clipboard
| Challenge: | SANs are an integral part of successful neural networks such as Transformer . training SAN on a task or pretraining them on language modeling requires large amounts of data and compute resources. |
| Approach: | They propose to modify SANs to enable faster learning, i.e., higher accuracies after fewer update steps. |
| Outcome: | The proposed modifications enable faster learning, i.e., higher accuracies after fewer update steps. |
Convolutional Self-Attention Networks (N19-1)
Copied to clipboard
| Challenge: | Existing models of self-attention networks lack the ability to capture dependencies regardless of distance and can be enhanced with multi-head attention. |
| Approach: | They propose a convolutional self-attention network which can be enhanced by multi-head attention by allowing the model to attend to information from different representation subspaces. |
| Outcome: | The proposed model outperforms existing models on improving locality of SANs on different language pairs and model settings. |
Corpora for Document-Level Neural Machine Translation (2020.lrec-1)
Copied to clipboard
| Challenge: | Document-level machine translation models translate sentences in isolation, but there are three main problems for document-level models. |
| Approach: | They propose to use document-level machine translation to capture discourse dependencies across sentences by considering a document as a whole. |
| Outcome: | The proposed method captures discourse dependencies across sentences by considering a document as a whole. |