Papers by Junda Su
In Defense of Structural Sparse Adapters for Concurrent LLM Serving (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Large language models (LLMs) require adapters to fine tune performance without extensive retraining. |
| Approach: | They propose a system that uses structurally sparse adapters to serve LLMs with multiple structurally-sparse axons. |
| Outcome: | The proposed system achieves 2.12 speedup over low-rank adapters on 96 adapters with a single GPU. |