Papers by Bingzheng Liu
pQuant: Towards Effective Low-Bit Language Models via Decoupled Linear Quantization-Aware Training (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing methods for building efficient large language models with sub 2-bit weights are lacking in accuracy and scalability. |
| Approach: | They propose a method that decouples parameters by splitting linear layers into two specialized branches. |
| Outcome: | The proposed method achieves state-of-the-art performance in extremely low-bit quantization. |