Challenge: Transformer architecture is composed of multi-head attention, which has been extensively analyzed.
Approach: They extended the scope of the analysis of Transformers from solely the attention patterns to the whole attention block, i.e., multi-head attention, residual connection, and layer normalization.
Outcome: The proposed method incorporates the whole attention block, i.e., multi-head attention, residual connection, and layer normalization into the analysis.

Similar Papers

RealFormer: Transformer Likes Residual Attention (2021.findings-acl)

Copied to clipboard

Challenge: Existing techniques to create Residual Attention Layer Transformer networks outperform the canonical Transformer on a wide spectrum of tasks.
Approach: They propose a technique to create Residual Attention Layer Transformer networks that outperform the canonical Transformer on a wide spectrum of tasks.
Outcome: The proposed technique outperforms the canonical Transformer on a wide spectrum of tasks including Masked Language Modeling, GLUE, SQUAD, Neural Machine Translation, WikiHop, HotpotQA, Natural Questions, and OpenKP.
VISIT: Visualizing and Interpreting the Semantic Information Flow of Transformers (2023.findings-emnlp)

Copied to clipboard

Challenge: Recent work in interpretability suggests we can project weights and hidden states of transformer-based language models (LMs) to their vocabulary space, a transformation that makes them more human interpretable.
Approach: They propose a tool to visualize a forward pass of Generative Pre-trained Transformers as an interactive flow graph with nodes representing neurons or hidden states and edges representing interactions between them.
Outcome: The proposed visualization simplifies huge amounts of data into easy-to-read graphs that can reflect the models’ internal processing, uncovering the contribution of each component to the models' final prediction.
Probing for Bridging Inference in Transformer Language Models (2021.naacl-main)

Copied to clipboard

Challenge: Pre-trained transformer language models are capable of bridging inference, but they lack the commonsense knowledge to capture syntactic information.
Approach: They investigate whether pre-trained transformer language models capture bridging inference . they use a masked token prediction task to investigate attention heads in BERT .
Outcome: The proposed model significantly captures bridging inference, the authors show . the distance between anaphor-antecedent and context plays an important role in the inference .
TLM: Token-Level Masking for Transformers (2023.emnlp-main)

Copied to clipboard

Challenge: Structured dropout approaches have been investigated to regularize the multi-head attention mechanism in Transformers.
Approach: They propose a new regularization scheme based on token-level rather than structure-level to reduce overfitting by manipulating the connections between tokens in the multi-head attention via masking.
Outcome: The proposed regularization scheme outperforms attention dropout and DropHead on 18 datasets and can establish a new record on the data-to-text benchmark Rotowire (18.93 BLEU).
Roles and Utilization of Attention Heads in Transformer-based Neural Language Models (2020.acl-main)

Copied to clipboard

Challenge: Sentence encoders based on transformer architectures have shown promising results on various natural language understanding tasks.
Approach: They propose a sentence representation method that takes advantage of most influential attention heads.
Outcome: The proposed method improves performance on the downstream tasks.
Transformer Grammars: Augmenting Transformer Language Models with Syntactic Inductive Biases at Scale (2022.tacl-1)

Copied to clipboard

Challenge: a novel class of Transformer language models that combine expressive power, scalability, and strong performance of Transformers and recursive syntactic compositions.
Approach: They introduce Transformer Grammars, a class of Transformer language models that combine expressive power and recursive syntactic compositions.
Outcome: The proposed model outperforms strong baselines on sentence-level language modeling perplexity and syntax-sensitive language evaluation metrics.
Mask Attention Networks: Rethinking and Strengthen Transformer (2021.naacl-main)

Copied to clipboard

Challenge: Existing research explores to enhance the two sublayers separately to improve the capability of Transformer for text representation.
Approach: They propose to combine SAN and Feed-Forward Networks to create a dynamic mask attention network with a learnable mask matrix which can model localness adaptively.
Outcome: The proposed model outperforms the original Transformer on translation and text summarization tasks.
Multiformer: A Head-Configurable Transformer-Based Model for Direct Speech Translation (2022.naacl-srw)

Copied to clipboard

Challenge: Existing approaches to address speech tasks with a self-attention mechanism are expensive and lead to information loss.
Approach: They propose a Transformer-based model which uses different attention mechanisms on each head to bias the self-attention towards the extraction of more diverse token interactions.
Outcome: The proposed model outperforms baseline models by 0.7 BLEU in the speech task.
A Meta-Learning Perspective on Transformers for Causal Language Modeling (2024.findings-acl)

Copied to clipboard

Challenge: Mechanisms of the Transformer architecture for causal language modeling are not well understood.
Approach: They propose a meta-learning view of the Transformer architecture when trained for a causal language modeling task by explicating an inner optimization process that may happen within the Transformer.
Outcome: The proposed model is based on a self-attention mechanism and has been widely used in natural language processing, computer vision, and scientific discovery.
Context Analysis for Pre-trained Masked Language Models (2020.findings-emnlp)

Copied to clipboard

Challenge: Pre-trained language models that learn contextualized word representations from a large un-annotated corpus have become a standard component for many downstream NLP tasks.
Approach: They propose to use a masking and gradient approach to evaluate the impact of context on the word representation.
Outcome: The proposed model architectures are architecture agnostic and gradient based.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations