Papers by Jianyi Cheng
Revisiting Block-based Quantisation: What is Important for Sub-8-bit LLM Inference? (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing quantisation methods mainly focus on 8-bit LLMs . a lack of scaling offsets in the quantisation process limits the use of LLM inference. |
| Approach: | They propose to use block quantisations to reduce scaling offsets in Large language models . they find that the block quantizations reduce scaling only from an arithmetic perspective . |
| Outcome: | The proposed methods reduce scaling offsets solely from an arithmetic perspective without additional treatments in the computational path. |
Refining Salience-Aware Sparse Fine-Tuning Strategies for Language Models (2025.acl-long)
Copied to clipboard
| Challenge: | Existing methods for fine-tuning large language models require expensive training on consumer-grade hardwares. |
| Approach: | They propose a sparsity-based approach that introduces trainable sparse adaptations to the weight matrices in the model and offers greater flexibility in selecting fine-tuned parameters. |
| Outcome: | The proposed method outperforms other methods for a simple yet effective baseline for nLP tasks while sacrificing performance. |