Papers by Xiaohui Zhou
BiLD: Bi-directional Logits Difference Loss for Large Language Model Distillation (2025.coling-main)
Copied to clipboard
| Challenge: | Knowledge distillation (KD) is a method for reducing model size while preserving performance. |
| Approach: | They propose a method to distill large language models at the logit level by transferring knowledge from a large teacher model to a smaller student model. |
| Outcome: | The proposed method outperforms supervised fine-tuning, vanilla KL loss and five other distillation methods on 13 datasets. |
Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Large language models (LLMs) have demonstrated exceptional performance across diverse natural language processing tasks. |
| Approach: | They propose a method for converting multi-head attention into grouped-query attention with any compression ratio of KV heads. |
| Outcome: | The proposed method can compress up to 87.5% KV heads of LLaMA2-7B model and 75% Kv heads of Sheared-LLa MA-1.3B with acceptable performance degradation. |
Joint Geometrical and Statistical Domain Adaptation for Cross-domain Code Vulnerability Detection (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing approaches to detect code vulnerability are limited by labeled training data on target domains. |
| Approach: | They propose a cross-domain code vulnerability detection framework called MNCRI . they propose mutual nearest neighbor contrastive learning to align the source and target domains . |
| Outcome: | The proposed framework outperforms state-of-the-art methods in cross-domain code vulnerability detection tasks. |
Ambiguity Awareness Optimization: Towards Semantic Disambiguation for Direct Preference Optimization (2025.emnlp-main)
Copied to clipboard
| Challenge: | Direct Preference Optimization (DPO) is a widely used reinforcement learning from human feedback (RLHF) method across various domains. |
| Approach: | They propose an approach that automatically re-weights ambiguous content to reduce ambiguities by calculating semantic similarity from preference pairs. |
| Outcome: | The proposed approach outperforms state-of-the-art approaches in performance across multiple model scales and widely adopted benchmark datasets. |
AIMMerging: Adaptive Iterative Model Merging Using Training Trajectories for Language Model Continual Learning (2025.emnlp-main)
Copied to clipboard
Yujie Feng, Jian Li, Xiaoyu Dong, Pengfei Xu, Xiaohui Zhou, Yujia Zhang, Zexin Lu, Yasha Wang, Alan Zhao, Xu Chu, Xiao-Ming Wu
| Challenge: | Recent model merging-based methods struggle to effectively manage the trade-off between learning new knowledge and preventing catastrophic forgetting. |
| Approach: | They propose a model merging framework that utilizes learning and forgetting signals from the training trajectory to dynamically monitor the model’s training status. |
| Outcome: | The proposed framework achieves significant performance improvements over existing state-of-the-art methods on three CL benchmarks with various model sizes (from 770M to 13B). |