Challenge: Recent studies show that Transformers can process longer sequences because of their complexity and time scales quadratic to the sequence length.
Approach: They propose an efficient Transformer model with adaptive attention that can select useful tokens automatically in sparse attention by learnable position vectors.
Outcome: The proposed model can select useful tokens automatically in sparse attention by learnable position vectors.

Similar Papers

Efficient Long-Range Transformers: You Need to Attend More, but Not Necessarily at Every Layer (2023.findings-emnlp)

Copied to clipboard

Challenge: Pretrained transformer models have demonstrated remarkable performance across various natural language processing tasks.
Approach: They propose a transformer variant with mixed attention spans that leverages the attention mechanism to capture long- and short-range dependencies in the sequence.
Outcome: The proposed model can achieve competitive performance to models with full attention while reducing computational cost (75%)
Long-range Sequence Modeling with Predictable Sparse Attention (2022.acl-long)

Copied to clipboard

Challenge: Existing approaches to capture global context dependencies in sequence modeling suffer from quadratic complexity in time and memory usage.
Approach: They propose an efficient Transformer architecture for fast long-range sequence modeling with a sparse attention matrix and a hidden state cross module.
Outcome: The proposed architecture outperforms the standard multi-head attention and its variants in various long-sequence tasks with low computational costs.
Learning Adaptive Axis Attentions in Fine-tuning: Beyond Fixed Sparse Attention Patterns (2022.findings-acl)

Copied to clipboard

Challenge: Adaptive Axis Attention learns different attention patterns for each task and model layer . sparse attention patterns do not improve the run time of the models but they reduce model memory requirements .
Approach: They propose a method that learns different attention patterns for each Transformer layer . they propose 'adaptive axis attention' method that identifies important tokens .
Outcome: The proposed method does not require pre-training to accommodate sparse attention patterns.
LongT5: Efficient Text-To-Text Transformer for Long Sequences (2022.findings-naacl)

Copied to clipboard

Challenge: Recent work has shown that increasing the input length or increasing model size can improve the performance of Transformer-based neural models.
Approach: They propose a model that integrates attention ideas from long-input transformers and adopts pre-training strategies from summarization pre-train into the scalable T5 architecture.
Outcome: The proposed model outperforms the original T5 models on several summarization and question answering tasks and achieves state-of-the-art results.
Efficient Content-Based Sparse Attention with Routing Transformers (2021.tacl-1)

Copied to clipboard

Challenge: Self-attention suffers from quadratic computation and memory requirements with respect to sequence length . despite its effectiveness, self-attention models suffer from quadratic computation and a limited set of locations .
Approach: They propose to learn dynamic sparse attention patterns that avoid allocating computation and memory to attend to content unrelated to the query of interest.
Outcome: The proposed model outperforms similar sparse attention models on language modeling and image generation on Wikitext-103 .
The Sparse Frontier: Sparse Attention Trade-offs in Transformer LLMs (2026.findings-acl)

Copied to clipboard

Challenge: Sparse attention is a promising strategy to extend long-context capabilities in LLMs . but its efficiency–accuracy trade-offs remain unclear due to the lack of comprehensive evaluation .
Approach: They evaluate sparse attention methods across multiple model families and sizes . they find larger sparser models outperform smaller dense ones at equivalent cost .
Outcome: The proposed methods outperform smaller sparse models at equivalent cost and improve the Pareto frontier.
Adaptive Attention Span in Transformers (P19-1)

Copied to clipboard

Challenge: We extend the maximum context size of a neural network called Transformer to 8k characters.
Approach: They propose a self-attention mechanism that can learn its optimal attention span . this allows for models with longer context and the capability to catch longer dependencies.
Outcome: The proposed model achieves state-of-the-art performance on text8 and enwiki8 using 8k characters with no loss of performance, and maintains control over memory footprint and computational time.
The NLP Task Effectiveness of Long-Range Transformers (2023.eacl-main)

Copied to clipboard

Challenge: Existing benchmarks on long-range attention models have not been sufficient to develop efficient Transformers and their practical application on complex NLP tasks.
Approach: They propose to benchmark 7 Transformer variants on 5 difficult NLP tasks and 7 datasets to examine their capacity for long-range attention.
Outcome: The proposed models have advantages on content selection and query-guided decoding, but they come with previously unrecognized drawbacks such as insufficient attention to distant tokens and accumulated approximation error.
Sparsifying Transformer Models with Trainable Representation Pooling (2022.acl-long)

Copied to clipboard

Challenge: Existing approaches to sparsify attention in the Transformer model are based on quadratic memory complexity and a lack of information for each word.
Approach: They propose a method to sparsify attention in a Transformer model by learning to select the most-informative token representations during the training process.
Outcome: The proposed model performs better than the current SOTA model while being 1.8 faster during training, 4.5 faster inference and 13 more efficient in the decoder.
ClusterFormer: Neural Clustering Attention for Efficient and Effective Transformer (2022.acl-long)

Copied to clipboard

Challenge: Existing sparse attention methods use fixed patterns to select words without considering similarities between words.
Approach: They propose a neural clustering method which integrates into the Self-Attention Mechanism in Transformer and integrates it into the target task.
Outcome: The proposed method outperforms two typical sparse attention methods on translation, text classification, and text matching tasks while having a comparable or even better time and memory efficiency.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations