Highway Transformer: Self-Gating Enhanced Self-Attentive Networks (2020.acl-main)
Copied to clipboard
| Challenge: | Self-attention mechanisms have made striking state-of-the-art (SOTA) progress in various sequence learning tasks, attending to all the global contexts at different locations. |
| Approach: | They propose a gated component self-dependency units (SDU) that incorporates LSTM-styled gating units to replenish internal semantic importance within the multi-dimensional latent space of individual representations. |
| Outcome: | The proposed system could boost the performance of the Transformer modules by allowing them to skip connections and step towards suboptimal points during the optimization process. |
Similar Papers
Token-Wise Kernels (TWiKers) for Vicinity-Aware Attention in Transformers (2026.findings-eacl)
Copied to clipboard
| Challenge: | Token-Wise Kernels (TWiKers) are a novel enhancement to transformers that learn token-specific convolutional kernels applied to the keys or values. |
| Approach: | They propose a transformer enhancement that learns token-specific convolutional kernels applied to the keys or values. |
| Outcome: | The proposed transformers learn token-specific convolutional kernels applied to the keys or values . the results show that content words retain self-focus while function words shift attention toward their neighbors . |
Recurrent Attention for Neural Machine Translation (2021.emnlp-main)
Copied to clipboard
| Challenge: | Recent research questions the importance of dot-product self-attention in Transformer models and shows that most attention heads learn simple positional patterns. |
| Approach: | They propose a novel mechanism to replace dot-product self-attention with a recurrent atteNtion mechanism that directly learns attention weights without token-to-token interaction. |
| Outcome: | The proposed model outperforms the Transformer model on translation tasks with fewer parameters and inference time. |
Fixed Encoder Self-Attention Patterns in Transformer-Based Machine Translation (2020.findings-emnlp)
Copied to clipboard
| Challenge: | Recent studies have shown that attention heads learn simple positional patterns . |
| Approach: | They propose to replace all but one attention head of each encoder layer with simple fixed – non-learnable – attentive patterns that are solely based on position and do not require external knowledge. |
| Outcome: | The proposed model improves translation quality and improves BLEU scores by up to 3 points in low-resource scenarios. |
LAIT: Efficient Multi-Segment Encoding in Transformers with Layer-Adjustable Interaction (2023.acl-long)
Copied to clipboard
Jeremiah Milbauer, Annie Louis, Mohammad Javad Hosseini, Alex Fabrikant, Donald Metzler, Tal Schuster
| Challenge: | In many NLP tasks, the input text can be seen as a sequence of related segments. |
| Approach: | They propose a layer-adjustable interactions framework that contextualizes token representations by attending to all other tokens at each layer, leading to quadratic increase in compute effort with the input length. |
| Outcome: | The proposed model reduces 30-50% of attention FLOPs while maintaining high accuracy. |
Efficient Content-Based Sparse Attention with Routing Transformers (2021.tacl-1)
Copied to clipboard
| Challenge: | Self-attention suffers from quadratic computation and memory requirements with respect to sequence length . despite its effectiveness, self-attention models suffer from quadratic computation and a limited set of locations . |
| Approach: | They propose to learn dynamic sparse attention patterns that avoid allocating computation and memory to attend to content unrelated to the query of interest. |
| Outcome: | The proposed model outperforms similar sparse attention models on language modeling and image generation on Wikitext-103 . |
Inceptive Transformers: Enhancing Contextual Representations through Multi-Scale Feature Learning Across Domains and Languages (2025.emnlp-main)
Copied to clipboard
| Challenge: | Encoder transformer models compress information from all tokens into a single [CLS] token to represent global context. |
| Approach: | They propose a 1-D convolution module that augments token representations with multi-scale local features to improve performance. |
| Outcome: | Experiments on five diverse tasks show that the proposed framework outperforms baseline models by 1% to 14% while maintaining efficiency. |
Generalized Attention Flow: Feature Attribution for Transformer Models via Maximum Flow (2025.acl-long)
Copied to clipboard
| Challenge: | Existing approaches to feature attributions rely on attention weights and attention weightings. |
| Approach: | They propose a feature attribution method that replaces attention weights with the generalized Information Tensor to enhance the performance of Transformer-based models. |
| Outcome: | The proposed method outperforms state-of-the-art feature attribution methods on sequence classification tasks and provides a more reliable interpretation of Transformer model outputs. |
FBS: Modeling Native Parallel Reading inside a Transformer (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing acceleration methods largely patch the autoregressive pipeline and miss core human-reading ingredients. |
| Approach: | They propose a trainable loop that injects a causal loop into Transformers via a 'parafoveal' approach. |
| Outcome: | The proposed model improves quality-efficiency trade-off without increasing parameters . ablations show the three modules are complementary . |
Fine-Tuning Pre-trained Transformers into Decaying Fast Weights (2022.emnlp-main)
Copied to clipboard
| Challenge: | Autoregressive Transformers incur O(T) complexity during per-token generation due to the self-attention mechanism. |
| Approach: | They propose a kernel-based method to approximate causal self-attention by replacing it with recurrent formulations with various update rules and feature maps to achieve O(1) time and memory complexity. |
| Outcome: | The proposed method outperforms prior methods and retains 99% of attention’s performance on WikiText-103 against more complex attention substitutes. |
Enhancing Machine Translation with Dependency-Aware Self-Attention (2020.acl-main)
Copied to clipboard
| Challenge: | Currently, most neural machine translation models rely on pairs of parallel sentences, assuming syntactic information is automatically learned by an attention mechanism. |
| Approach: | They propose a parameter-free, dependency-aware self-attention mechanism that integrates syntactic knowledge into a Transformer model and propose 'a parameter free approach' they also propose - a novel mechanism that improves translation quality for long sentences and in low-resource scenarios. |
| Outcome: | The proposed approach improves translation quality on English-German and English-Turkish translation tasks and in low-resource scenarios. |