Papers by Harshith Goka
zFLoRA: Zero-Latency Fused Low-Rank Adapters (2025.emnlp-main)
Copied to clipboard
| Challenge: | Large language models (LLMs) are increasingly deployed with task-specific adapters catering to multiple downstream applications. |
| Approach: | They propose a low-latency fused low-rank adapter that introduces zero latency overhead on top of the base model. |
| Outcome: | The proposed adapter reduces the inference time of the model by 2.5x . the proposed adapters are tested on 18 different tasks on different platforms . |