Challenge: Existing studies have investigated the multi-head self-attention mechanism of transformers.
Approach: They propose to use a human-in-the-loop pipeline to discover task-specific attention patterns and inject them into transformer models to improve their accuracy.
Outcome: The proposed methods improve the performance of transformer models by incorporating predefined patterns into their attention matrices.

Similar Papers

Telling BERT’s Full Story: from Local Attention to Global Aggregation (2021.eacl-main)

Copied to clipboard

Challenge: Recent work discouraging the use of attention distributions for explaining a model’s behaviour suggests that attention distribution can provide insights into local behaviour of attention heads.
Approach: They propose a distinction between local patterns revealed by attention and global patterns that refer back to the input and analyze BERT from both angles.
Outcome: The proposed model can explain local behaviour of attention heads by comparing local and global patterns from both angles.
A Multiscale Visualization of Attention in the Transformer Model (P19-3)

Copied to clipboard

Challenge: Various tools have been developed to visualize attention in NLP models, ranging from attention-matrix heatmaps to bipartite graph representations.
Approach: They propose an open-source tool that visualizes attention at multiple scales and provides a unique perspective on the attention mechanism.
Outcome: The proposed model outperforms OpenAI GPT-2 and BERT on several language modeling benchmarks.
Roles and Utilization of Attention Heads in Transformer-based Neural Language Models (2020.acl-main)

Copied to clipboard

Challenge: Sentence encoders based on transformer architectures have shown promising results on various natural language understanding tasks.
Approach: They propose a sentence representation method that takes advantage of most influential attention heads.
Outcome: The proposed method improves performance on the downstream tasks.
How Far Does BERT Look At: Distance-based Clustering and Analysis of BERT’s Attention (2020.coling-main)

Copied to clipboard

Challenge: Recent work on multi-head attention mechanism shows heuristics and clues in analyzing various aspects of the mechanism.
Approach: They propose to cluster attention heatmaps into significantly different patterns through unsupervised clustering on top of a set of proposed features.
Outcome: The proposed features can explain and calibrate different attention heads in Transformer models.
Fixed Encoder Self-Attention Patterns in Transformer-Based Machine Translation (2020.findings-emnlp)

Copied to clipboard

Challenge: Recent studies have shown that attention heads learn simple positional patterns .
Approach: They propose to replace all but one attention head of each encoder layer with simple fixed – non-learnable – attentive patterns that are solely based on position and do not require external knowledge.
Outcome: The proposed model improves translation quality and improves BLEU scores by up to 3 points in low-resource scenarios.
Attention Head Masking for Inference Time Content Selection in Abstractive Summarization (2021.naacl-main)

Copied to clipboard

Challenge: Existing studies show that multi-heads attentions at the same layer collectively guide the summarization.
Approach: They propose an inference-time attention head masking mechanism that works on encoder-decoder attentions to pinpoint salient content at inference time.
Outcome: The proposed technique outperforms state-of-the-art models on CNN/DailyMail and New York Times datasets and is data-efficient.
Self-Attention Guided Copy Mechanism for Abstractive Summarization (2020.acl-main)

Copied to clipboard

Challenge: Abstractive summarization models have been widely used to extract words from source into summary, but how to ensure that important words in source are copied remains a challenge.
Approach: They propose a Transformer-based model to enhance copy mechanism by identifying the importance of each source word based on the degree centrality.
Outcome: The proposed model outperforms baseline methods on CNN/Daily Mail and Gigaword datasets.
Guiding Attention for Self-Supervised Learning with Transformers (2020.findings-emnlp)

Copied to clipboard

Challenge: Recent studies show that self-attention patterns in trained models contain a majority of non-linguistic regularities.
Approach: They propose a technique to allow efficient self-supervised learning with bi-directional Transformers by using an auxiliary loss function to guide attention heads to conform to such patterns.
Outcome: The proposed method achieves state-of-the-art in low-resource settings and is agnostic to pre-training objectives.
Adaptive Transformers for Learning Multimodal Representations (2020.acl-srw)

Copied to clipboard

Challenge: Existing approaches for learning visiolinguistic representations with transformers are over-parametrized and require extensive training.
Approach: They propose to extend attention spans, sparse, and structured dropout methods to learn more about how the network perceives the complexity of input sequences.
Outcome: The proposed approaches improve on language semantics and visiolinguistic representations, but are often over-parametrized and require large amounts of computation.
Highlight-Transformer: Leveraging Key Phrase Aware Attention to Improve Abstractive Multi-Document Summarization (2021.findings-acl)

Copied to clipboard

Challenge: Existing models do not consider key phrases in determining attention weights of self-attention . Existing work does not consider the importance of key phrases when determining weights .
Approach: They propose a model with highlighting mechanism to assign greater attention weights to key phrases . they propose two structures of highlighting attention for each head and the multihead highlighting . experimental results show that their proposed model significantly outperforms the baseline model .
Outcome: The proposed model outperforms the baseline models on a multi-news dataset.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations