Papers with L4Q
L4Q: Parameter Efficient Quantization-Aware Fine-Tuning on Large Language Models (2025.acl-long)
Copied to clipboard
| Challenge: | Quantization-aware PEFT methods have been developed to reduce memory and computational costs associated with large language models. |
| Approach: | They propose a method that integrates Quantization-Aware Training (QAT) with LoRA to reduce memory overhead and improve model accuracy. |
| Outcome: | The proposed method significantly reduces QAT’s memory overhead while preserving the advantage of QAT in producing fully quantized LLMs with high accuracy. |