Papers with SANs

7 papers
Towards Better Modeling Hierarchical Structure for Self-Attention with Ordered Neurons (D19-1)

Copied to clipboard

Challenge: Recent studies have shown that a hybrid of self-attention networks (SANs) and recurrent neural networks (RNNs) outperforms both individual architectures, while not much is known about why the hybrid models work.
Approach: They propose to use an advanced variant of self-attention networks (SANs) to enhance the strength of hybrid models by introducing a syntax-oriented inductive bias to perform tree-like composition.
Outcome: The proposed model outperforms both individual models and a standard hybrid model on a machine translation task.
Self-Attention with Structural Position Representations (D19-1)

Copied to clipboard

Challenge: Experimental results show that SANs can't encode positions of input words . SAN's are currently lacking in encoding positions of words based on position-unaware "bagof-words" theory .
Approach: They propose to augment SANs with structural position representations to capture latent structure of input sentence.
Outcome: The proposed approach consistently outperforms the sequential representations on translation tasks.
Self-Attention with Cross-Lingual Position Representation (2020.acl-main)

Copied to clipboard

Challenge: Position encoding (PE) is used to preserve word order information for natural language processing tasks, generating fixed position indices for input sequences.
Approach: They propose to augment SANs with cross-lingual position representations to model bilingually aware latent structure for the input sentence.
Outcome: The proposed model significantly improves translation quality over baselines on EnglishGerman, JapaneseEnglish, and ChineseEnglish translation tasks.
How Does Selective Mechanism Improve Self-Attention Networks? (2020.acl-main)

Copied to clipboard

Challenge: Experimental results show that selective SANs outperform the standard SAN by paying more attention to content words that contribute to the meaning of the sentence.
Approach: They propose to implement selective SANs with a flexible Gumbel-Softmax to improve word order encoding and structure modeling.
Outcome: The proposed system outperforms the standard SANs on several representative NLP tasks including natural language inference, semantic role labelling, and machine translation.
Increasing Learning Efficiency of Self-Attention Networks through Direct Position Interactions, Learnable Temperature, and Convoluted Attention (2020.coling-main)

Copied to clipboard

Challenge: SANs are an integral part of successful neural networks such as Transformer . training SAN on a task or pretraining them on language modeling requires large amounts of data and compute resources.
Approach: They propose to modify SANs to enable faster learning, i.e., higher accuracies after fewer update steps.
Outcome: The proposed modifications enable faster learning, i.e., higher accuracies after fewer update steps.
Convolutional Self-Attention Networks (N19-1)

Copied to clipboard

Challenge: Existing models of self-attention networks lack the ability to capture dependencies regardless of distance and can be enhanced with multi-head attention.
Approach: They propose a convolutional self-attention network which can be enhanced by multi-head attention by allowing the model to attend to information from different representation subspaces.
Outcome: The proposed model outperforms existing models on improving locality of SANs on different language pairs and model settings.
Corpora for Document-Level Neural Machine Translation (2020.lrec-1)

Copied to clipboard

Challenge: Document-level machine translation models translate sentences in isolation, but there are three main problems for document-level models.
Approach: They propose to use document-level machine translation to capture discourse dependencies across sentences by considering a document as a whole.
Outcome: The proposed method captures discourse dependencies across sentences by considering a document as a whole.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations