Challenge: Existing studies focus on compressing the Key-Value cache or grouping attention heads, while overlooking redundancy between layers.
Approach: They propose a lightweight substitute for self-attention in well-trained LLMs that uses feed-forward networks to align attention heads between adjacent layers and low-rank matrices to approximate differences in layer-wise attention weights.
Outcome: The proposed model reduces redundancy by sharing weights across layers while maintaining high response quality while reducing redundant calculations within 53% 84% of the total layers.

Similar Papers

Head-wise Shareable Attention for Large Language Models (2024.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs) suffer from huge number of parameters, which restricts their deployment on edge devices.
Approach: They propose two methods that share parameters across attention heads to reduce memory usage and reduce performance drop by using coarse-grained weight sharing rules.
Outcome: The proposed methods reuse pre-trained weights without retraining and then share, denoted as PostShare.
Pit One Against Many: Leveraging Attention-head Embeddings for Parameter-efficient Multi-head Attention (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing pre-trained language models have produced performance gains in various tasks but come with large computational requirements.
Approach: They propose an alternative module that uses only a single shared projection matrix and multiple head embeddings (MHE) they demonstrate that MHE attention is substantially more memory efficient compared to alternative attention mechanisms.
Outcome: The proposed model is more memory efficient compared to the current model while achieving high retention ratio on several downstream tasks.
IAM: Efficient Inference through Attention Mapping between Different-scale LLMs (2025.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) are a challenge due to their internal reasoning processes.
Approach: They propose an algorithm that can optimize attention matrices by performing attention mapping between small and large LLMs.
Outcome: The proposed framework can reduce KV cache usage by 22.1% and accelerate prefill by 15% without sacrificing performance.
Data-Informed Global Sparseness in Attention Mechanisms for Deep Neural Networks (2024.lrec-main)

Copied to clipboard

Challenge: Attention pruning techniques have been developed to identify and exploit sparseness . previous work has taken pioneering steps to discover and explain the sparsity in attention patterns .
Approach: They propose a framework that observes attention patterns in a fixed dataset and generates a global sparseness mask.
Outcome: The proposed approach saves 90% of computations and maintains quality of results.
Efficient Sparse Attention needs Adaptive Token Release (2024.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) have demonstrated remarkable capabilities across a wide array of text-centric tasks, however, their ‘large’ scale introduces significant computational and storage challenges, particularly in managing the key-value states of the transformer, which limits their wider applicability.
Approach: They propose to release resources from caches and rebuild key-value states by a lightweight controller module to approximate an ideal top-K sparse attention.
Outcome: The proposed method achieves a significant throughput improvement of 221.8% over full attention and a model with 7 billion tokens.
Deputy: Accelerating Large Language Model Inference with Dynamic Low-Rank Substitution (2026.findings-acl)

Copied to clipboard

Challenge: Existing dynamic schemes such as early-exit and layer-drop reduce FLOPs but break batch processing or introduce KV-cache inconsistency.
Approach: They propose a dynamic low-rank substitution framework that employs a lightweight decision module at each layer to dynamically determine the execution branch for different tokens.
Outcome: The proposed model reduces computation by approximately 40% compared to the original dense model while outperforming existing baseline methods.
Does Self-Attention Need Separate Weights in Transformers? (2025.naacl-industry)

Copied to clipboard

Challenge: Experimental results show a 66.53% reduction in parameter size within the attention block and competitive accuracy improvements of 3.55% and 0.89% over symmetric and pairwise attention-based models, respectively.
Approach: They propose a simplified approach where a single weight matrix is used for Keys, Queries, and Values instead of separate matrices for each.
Outcome: The proposed approach outperforms the BERT baseline on GLUE tasks even outperforming the standard BERT model in handling noisy and out-of-domain data.
Layer Duplication in LLMs (2025.findings-emnlp)

Copied to clipboard

Challenge: Duplicating multihead self-attention layers once or twice yields significant performance boosts.
Approach: They investigate the effect of duplicating multihead self-attention layers in large language models .
Outcome: The proposed method improves performance across a range of language tasks with or without fine-tuning.
A Systematic Study of Cross-Layer KV Sharing for Efficient LLM Inference (2025.naacl-short)

Copied to clipboard

Challenge: Recent studies have shown that sharing key-value (KV) cache across layers is effective in efficient inference of large language models.
Approach: They propose a unified framework that covers several recent methods and their novel variants to investigate cross-layer KV sharing.
Outcome: The proposed framework achieves higher throughput and better performance when reducing the size of the key-value cache by 2 while maintaining competitive performance.
Convolutions and Self-Attention: Re-interpreting Relative Positions in Pre-trained Language Models (2021.acl-long)

Copied to clipboard

Challenge: Recent work has shown that convolutions have been successful in natural language learning.
Approach: They propose a convolutional approach to construct relative position embeddings in self-attention layers and propose 'compact attention' they propose multiple ways to integrate convolutions into Transformer self- attention.
Outcome: The proposed composite attention improves performance on multiple downstream tasks, replacing absolute position embeddings, and is more expressive than convolutions in NLP.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations