Papers by Weixiao Zhan
Distillation Traps and Guards: A Calibration Knob for LLM Distillability (2026.acl-long)
Copied to clipboard
| Challenge: | Knowledge distillation (KD) transfers capabilities from large language models (LLMs) to smaller students, yet it can fail unpredictably and also underpins model leakage risks. |
| Approach: | They propose a method that allows teachers to control their distillability via reinforcement fine-tuning (RFT) they propose to use tail noise, off-policy instability, and the teacher–student gap to improve KD. |
| Outcome: | The proposed method outperforms SFT and KD baselines and can be used to protect teachers and students from bottlenecks. |