Challenge: Existing methods for fine-tuning require caching of intermediate activations to update weights during the backward pass.
Approach: They develop a method to reduce memory usage in fine-tuning of transformers by backpropagating through just a subset of input tokens.
Outcome: The proposed method reduces memory usage and memory footprint on large transformer models . it can be easily combined with existing methods like LoRA, reducing memory cost .

Similar Papers

TARE: Lightweight Token-Aware Representation Editing for Fine-tuning Transformer-like Models (2026.acl-long)

Copied to clipboard

Challenge: Existing PEFT methods can be costly and underfit token-level contexts.
Approach: They propose a PEFT method that performs fine-grained, token-specific edits with a small additional inference overhead and minimal tuning.
Outcome: The proposed method outperforms state-of-the-art methods in 8 tasks and GLUE with a minimal tuning overhead and inference overhead.
Quantized Side Tuning: Fast and Memory-Efficient Tuning of Quantized Large Language Models (2024.acl-long)

Copied to clipboard

Challenge: Existing methods to finetun large language models (LLMs) only update a small number of trainable parameters, or attempt to reduce the memory footprint during the training phase of the finetune process.
Approach: They propose quantized side tuing (QST) which quantizes an LLM’s model weights into 4-bit to reduce the memory footprint of the original weights.
Outcome: The proposed method reduces the memory footprint of the model weights, optimizer states, and intermediate activations while reducing the memory requirements.
TRAMS: Training-free Memory Selection for Long-range Language Modeling (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing methods like Transformer-XL are plagued by ineffective memory selections due to the high number of tokens involved in attention calculation.
Approach: They propose a plug-and-play strategy that selects tokens participating in attention calculation based on one simple metric and ignores the other ones.
Outcome: The proposed strategy keeps tokens with high attention scores and ignores the other ones on word-level and character-level benchmarks without additional training or adding additional parameters.
Forward Knows Efficient Backward Path: Saliency-Guided Memory-Efficient Fine-tuning of Large Language Models (2025.acl-long)

Copied to clipboard

Challenge: a number of fine-tuning approaches are available to improve performance of large language models.
Approach: They propose a memory-efficient method to minimize memory associated with cached intermediate activations.
Outcome: The proposed method minimizes memory associated with cached intermediate activations while preserving accuracy.
Finetuning Pretrained Transformers into RNNs (2021.emnlp-main)

Copied to clipboard

Challenge: Efficient transformers outperform recurrent neural networks in natural language generation, but this comes with significant computational cost and memory footprint during generation.
Approach: They propose to convert a pretrained transformer into its efficient recurrent counterpart, improving efficiency while maintaining accuracy.
Outcome: The proposed transformers outperform recurrent neural networks in natural language generation but come with significant computational and memory footprint during generation.
Fine-Tuning Pre-trained Transformers into Decaying Fast Weights (2022.emnlp-main)

Copied to clipboard

Challenge: Autoregressive Transformers incur O(T) complexity during per-token generation due to the self-attention mechanism.
Approach: They propose a kernel-based method to approximate causal self-attention by replacing it with recurrent formulations with various update rules and feature maps to achieve O(1) time and memory complexity.
Outcome: The proposed method outperforms prior methods and retains 99% of attention’s performance on WikiText-103 against more complex attention substitutes.
Understanding and Overcoming the Challenges of Efficient Transformer Quantization (2021.emnlp-main)

Copied to clipboard

Challenge: Recent advances in transformer quantization have shown remarkable improvement in many Natural Language Processing tasks and beyond.
Approach: They propose a novel quantization scheme for transformers that can be quantized to ultra-low bit-widths, leading to significant memory savings with a minimum accuracy loss.
Outcome: The proposed methods achieve state-of-the-art results on the GLUE benchmark using BERT, while preserving memory and accuracy.
PARA: Parameter-Efficient Fine-tuning with Prompt-Aware Representation Adjustment (2024.emnlp-industry)

Copied to clipboard

Challenge: Existing methods for parameter-efficient fine-tuning excel in the context of single-backbone multi-tenant applications.
Approach: They propose to integrate a lightweight vector generator within each Transformer layer to improve prompt-aware representation adjustment.
Outcome: The proposed method surpasses current benchmarks in terms of performance despite having a similar number of adjustable parameters.
State-offset Tuning: State-based Parameter-Efficient Fine-Tuning for State Space Models (2025.acl-short)

Copied to clipboard

Challenge: State Space Models (SSMs) have emerged as efficient alternatives to Transformers, but their application to SSMs remains unexplored.
Approach: They propose a state-based PEFT method that adjusts state directly instead of using external prompts.
Outcome: The proposed method is based on state-offset tuning, which directly affects state at every timestep.
MEFT: Memory-Efficient Fine-Tuning through Sparse Adapter (2024.acl-long)

Copied to clipboard

Challenge: Parameter-Efficient Fine-tuning (PEFT) methods are limited on knowledge-intensive tasks due to the limited number of trainable parameters.
Approach: They propose a mechanism that fine-tunes Large Language Models with larger adapters . they store and update the parameters of larger adapter adapters on the CPU .
Outcome: The proposed method achieves comparable results to those obtained with larger memory capacities over the limited bandwidth of PCI Express (PCIe).

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations