Papers with Lillama

1 papers
Efficient One-shot Compression via Low-Rank Local Feature Distillation (2025.naacl-long)

Copied to clipboard

Challenge: Existing structured pruning approaches for large language models require calibration data and costly continued pretraining on billions of tokens to recover lost performance.
Approach: They propose a method that locally distills activations with low-rank weights . they compress Mixtral-8x7B on a single GPU and Phi-2 3B by 40% .
Outcome: The proposed method compresses Mixtral-8x7B on a single A100 GPU, removing 10 billion parameters while retaining over 95% of its original performance.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations