Challenge: Self-attention mechanisms have made striking state-of-the-art (SOTA) progress in various sequence learning tasks, attending to all the global contexts at different locations.
Approach: They propose a gated component self-dependency units (SDU) that incorporates LSTM-styled gating units to replenish internal semantic importance within the multi-dimensional latent space of individual representations.
Outcome: The proposed system could boost the performance of the Transformer modules by allowing them to skip connections and step towards suboptimal points during the optimization process.

Similar Papers

Token-Wise Kernels (TWiKers) for Vicinity-Aware Attention in Transformers (2026.findings-eacl)

Copied to clipboard

Challenge: Token-Wise Kernels (TWiKers) are a novel enhancement to transformers that learn token-specific convolutional kernels applied to the keys or values.
Approach: They propose a transformer enhancement that learns token-specific convolutional kernels applied to the keys or values.
Outcome: The proposed transformers learn token-specific convolutional kernels applied to the keys or values . the results show that content words retain self-focus while function words shift attention toward their neighbors .
Recurrent Attention for Neural Machine Translation (2021.emnlp-main)

Copied to clipboard

Challenge: Recent research questions the importance of dot-product self-attention in Transformer models and shows that most attention heads learn simple positional patterns.
Approach: They propose a novel mechanism to replace dot-product self-attention with a recurrent atteNtion mechanism that directly learns attention weights without token-to-token interaction.
Outcome: The proposed model outperforms the Transformer model on translation tasks with fewer parameters and inference time.
Fixed Encoder Self-Attention Patterns in Transformer-Based Machine Translation (2020.findings-emnlp)

Copied to clipboard

Challenge: Recent studies have shown that attention heads learn simple positional patterns .
Approach: They propose to replace all but one attention head of each encoder layer with simple fixed – non-learnable – attentive patterns that are solely based on position and do not require external knowledge.
Outcome: The proposed model improves translation quality and improves BLEU scores by up to 3 points in low-resource scenarios.
LAIT: Efficient Multi-Segment Encoding in Transformers with Layer-Adjustable Interaction (2023.acl-long)

Copied to clipboard

Challenge: In many NLP tasks, the input text can be seen as a sequence of related segments.
Approach: They propose a layer-adjustable interactions framework that contextualizes token representations by attending to all other tokens at each layer, leading to quadratic increase in compute effort with the input length.
Outcome: The proposed model reduces 30-50% of attention FLOPs while maintaining high accuracy.
Efficient Content-Based Sparse Attention with Routing Transformers (2021.tacl-1)

Copied to clipboard

Challenge: Self-attention suffers from quadratic computation and memory requirements with respect to sequence length . despite its effectiveness, self-attention models suffer from quadratic computation and a limited set of locations .
Approach: They propose to learn dynamic sparse attention patterns that avoid allocating computation and memory to attend to content unrelated to the query of interest.
Outcome: The proposed model outperforms similar sparse attention models on language modeling and image generation on Wikitext-103 .
Inceptive Transformers: Enhancing Contextual Representations through Multi-Scale Feature Learning Across Domains and Languages (2025.emnlp-main)

Copied to clipboard

Challenge: Encoder transformer models compress information from all tokens into a single [CLS] token to represent global context.
Approach: They propose a 1-D convolution module that augments token representations with multi-scale local features to improve performance.
Outcome: Experiments on five diverse tasks show that the proposed framework outperforms baseline models by 1% to 14% while maintaining efficiency.
Generalized Attention Flow: Feature Attribution for Transformer Models via Maximum Flow (2025.acl-long)

Copied to clipboard

Challenge: Existing approaches to feature attributions rely on attention weights and attention weightings.
Approach: They propose a feature attribution method that replaces attention weights with the generalized Information Tensor to enhance the performance of Transformer-based models.
Outcome: The proposed method outperforms state-of-the-art feature attribution methods on sequence classification tasks and provides a more reliable interpretation of Transformer model outputs.
FBS: Modeling Native Parallel Reading inside a Transformer (2026.findings-acl)

Copied to clipboard

Challenge: Existing acceleration methods largely patch the autoregressive pipeline and miss core human-reading ingredients.
Approach: They propose a trainable loop that injects a causal loop into Transformers via a 'parafoveal' approach.
Outcome: The proposed model improves quality-efficiency trade-off without increasing parameters . ablations show the three modules are complementary .
Fine-Tuning Pre-trained Transformers into Decaying Fast Weights (2022.emnlp-main)

Copied to clipboard

Challenge: Autoregressive Transformers incur O(T) complexity during per-token generation due to the self-attention mechanism.
Approach: They propose a kernel-based method to approximate causal self-attention by replacing it with recurrent formulations with various update rules and feature maps to achieve O(1) time and memory complexity.
Outcome: The proposed method outperforms prior methods and retains 99% of attention’s performance on WikiText-103 against more complex attention substitutes.
Enhancing Machine Translation with Dependency-Aware Self-Attention (2020.acl-main)

Copied to clipboard

Challenge: Currently, most neural machine translation models rely on pairs of parallel sentences, assuming syntactic information is automatically learned by an attention mechanism.
Approach: They propose a parameter-free, dependency-aware self-attention mechanism that integrates syntactic knowledge into a Transformer model and propose 'a parameter free approach' they also propose - a novel mechanism that improves translation quality for long sentences and in low-resource scenarios.
Outcome: The proposed approach improves translation quality on English-German and English-Turkish translation tasks and in low-resource scenarios.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations