Papers by George Constantinides
Revisiting Block-based Quantisation: What is Important for Sub-8-bit LLM Inference? (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing quantisation methods mainly focus on 8-bit LLMs . a lack of scaling offsets in the quantisation process limits the use of LLM inference. |
| Approach: | They propose to use block quantisations to reduce scaling offsets in Large language models . they find that the block quantizations reduce scaling only from an arithmetic perspective . |
| Outcome: | The proposed methods reduce scaling offsets solely from an arithmetic perspective without additional treatments in the computational path. |