Papers with Self-attention

6 papers
Efficient Content-Based Sparse Attention with Routing Transformers (2021.tacl-1)

Copied to clipboard

Challenge: Self-attention suffers from quadratic computation and memory requirements with respect to sequence length . despite its effectiveness, self-attention models suffer from quadratic computation and a limited set of locations .
Approach: They propose to learn dynamic sparse attention patterns that avoid allocating computation and memory to attend to content unrelated to the query of interest.
Outcome: The proposed model outperforms similar sparse attention models on language modeling and image generation on Wikitext-103 .
Does Self-Attention Need Separate Weights in Transformers? (2025.naacl-industry)

Copied to clipboard

Challenge: Experimental results show a 66.53% reduction in parameter size within the attention block and competitive accuracy improvements of 3.55% and 0.89% over symmetric and pairwise attention-based models, respectively.
Approach: They propose a simplified approach where a single weight matrix is used for Keys, Queries, and Values instead of separate matrices for each.
Outcome: The proposed approach outperforms the BERT baseline on GLUE tasks even outperforming the standard BERT model in handling noisy and out-of-domain data.
Rethinking Self-Attention: Towards Interpretability in Neural Parsing (2020.findings-emnlp)

Copied to clipboard

Challenge: Recent work shows that attention mechanisms provide arguably explainable attention distributions that can help to interpret predictions.
Approach: They propose a new self-attention layer where attention heads represent labels.
Outcome: The proposed model obtains state-of-the-art results on the Penn Treebank and Chinese Treebank.
Self-Attentional Models for Lattice Inputs (P19-1)

Copied to clipboard

Challenge: Existing work has extended recurrent neural networks to model lattice inputs but these models suffer from slow computation speeds.
Approach: They propose to extend the paradigm of self-attention to handle lattice inputs by adding probabilistic reachability masks that incorporate latticae structure into the model and support lattics if available.
Outcome: The proposed model outperforms baseline models while being much faster to compute than previous models.
CoCA: Fusing Position Embedding with Collinear Constrained Attention in Transformers for Long Context Window Extending (2024.acl-long)

Copied to clipboard

Challenge: Existing models that use self-attention and position embedding have anomalous behavior that hinder long context window extrapolation.
Approach: They propose a collinear constraint between Q and K to integrate RoPE and self-attention.
Outcome: The proposed model integrates self-attention and position embedding into LLMs without fine-tuning.
ChunkAttention: Efficient Self-Attention with Prefix-Aware KV Cache and Two-Phase Partition (2024.acl-long)

Copied to clipboard

Challenge: Experiments show that ChunkAttention can speed up the self-attention kernel by 3.2-4.8 compared to the start-of-the-art implementation.
Approach: They propose a prefix-aware self-attention module that can detect matching prompt prefixes across multiple requests and share their key/value tensors in memory at runtime.
Outcome: The proposed module can speed up the self-attention kernel by 3.2-4.8 compared to the start-of-the-art implementation, with the length of the system prompt ranging from 1024 to 4096.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations