Meta-KD: A Meta Knowledge Distillation Framework for Language Model Compression across Domains (2021.acl-long)
Copied to clipboard
| Challenge: | Pre-trained language models have been successful in NLP tasks, but their large size and long inference time limit their deployment in real-time applications. |
| Approach: | They propose a meta-teacher model that captures transferable knowledge across domains and passes it to students. |
| Outcome: | The proposed model can distill large teacher models into small student models with guidance from the meta-teacher. |
Similar Papers
Knowledge Distillation with Reptile Meta-Learning for Pretrained Language Model Compression (2022.coling-1)
Copied to clipboard
| Challenge: | Knowledge distillation (KD) can transfer knowledge from the original model into a compact model to achieve model compression. |
| Approach: | They propose a knowledge distillation method with reptile meta-learning to facilitate the transfer of knowledge from the teacher to the student. |
| Outcome: | Extensive experiments on the GLUE benchmark show the proposed method performs better than previous methods. |
BERT Learns to Teach: Knowledge Distillation with Meta Learning (2022.acl-long)
Copied to clipboard
| Challenge: | Existing knowledge distillation methods are based on teacher model, but have drawbacks . a teacher model is fixed during training, but meta learning can improve student performance . |
| Approach: | They propose a meta learning framework that allows the teacher network to learn to better transfer knowledge to the student network. |
| Outcome: | Experiments show that MetaDistil can improve on existing methods and is less sensitive to student capacity and hyperparameters. |
Knowledge Distillation for Language Models (2025.naacl-tutorial)
Copied to clipboard
| Challenge: | Knowledge distillation (KD) aims to transfer knowledge from a teacher to a student . this tutorial will cover topics ranging from LLM sequence compression to LLM self-distillation . |
| Approach: | They propose to introduce intermediate-layer matching and prediction matching . they will then present advanced techniques such as reinforcement learning-based KD and multi-teacher distillation . |
| Outcome: | This tutorial aims to provide participants with a comprehensive understanding of the techniques and applications of knowledge distillation for language models. |
AD-KD: Attribution-Driven Knowledge Distillation for Language Model Compression (2023.acl-long)
Copied to clipboard
| Challenge: | Existing knowledge distillation methods focus on the transfer of model-specific knowledge but overlook data-specific information. |
| Approach: | They propose an attribution-driven knowledge distillation approach which explores the token-level rationale behind the teacher model and transfers attribution knowledge to the student model. |
| Outcome: | The proposed method outperforms state-of-the-art methods on the GLUE benchmark and shows that it is more efficient than existing methods. |
HRKD: Hierarchical Relational Knowledge Distillation for Cross-domain Language Model Compression (2021.emnlp-main)
Copied to clipboard
| Challenge: | Large pre-trained language models (PLMs) have shown overwhelming performances on many tasks, but their large size and slow inference speed have hindered practical deployments. |
| Approach: | They propose a hierarchical relational knowledge distillation method to capture hierarchic and domain relational information. |
| Outcome: | The proposed method outperforms existing methods on multi-domain datasets and is highly reproducible. |
Tutoring Helps Students Learn Better: Improving Knowledge Distillation for BERT with Tutor Network (2022.emnlp-main)
Copied to clipboard
| Challenge: | Existing knowledge distillation approaches for language models have overlooked the difficulty of training examples. |
| Approach: | They propose a framework that controls difficulty of training examples during pre-training by a tutor network. |
| Outcome: | The proposed framework outperforms state-of-the-art KD methods with student models on the GLUE benchmark. |
ReAugKD: Retrieval-Augmented Knowledge Distillation For Pre-trained Language Models (2023.acl-short)
Copied to clipboard
Jianyi Zhang, Aashiq Muhamed, Aditya Anantharaman, Guoyin Wang, Changyou Chen, Kai Zhong, Qingjun Cui, Yi Xu, Belinda Zeng, Trishul Chilimbi, Yiran Chen
| Challenge: | Knowledge distillation (KD) is an effective compression technique to derive a smaller student model from a larger teacher model by transferring the knowledge embedded in the teacher's network. |
| Approach: | They propose a framework and loss function that preserves the semantic similarities of teacher and student training examples to enable the student to retrieve from the knowledge base effectively. |
| Outcome: | The proposed framework preserves the semantic similarities of teacher and student training examples to achieve state-of-the-art performance on the GLUE benchmark. |
Towards Zero-Shot Knowledge Distillation for Natural Language Processing (2021.emnlp-main)
Copied to clipboard
| Challenge: | Knowledge distillation (KD) is a common knowledge transfer algorithm used for model compression across a variety of deep learning based natural language processing (NLP) solutions. |
| Approach: | They propose to use teacher training data for model compression . they investigate six tasks and find they can achieve between 75% and 92% of the teacher’s classification score while compressing the model 30 times. |
| Outcome: | The proposed solution achieves between 75% and 92% of the teacher’s classification score while compressing the model 30 times. |
Meta-Learning Adaptive Knowledge Distillation for Efficient Biomedical Natural Language Processing (2022.findings-aacl)
Copied to clipboard
| Challenge: | Existing knowledge distillation methods have been proposed to reduce the size of large models for biomedical natural language processing tasks. |
| Approach: | They propose a meta-learning approach which adaptively learns parameters that enable optimal rate of knowledge exchange between teacher and student models from the distillation data during knowledge distillation. |
| Outcome: | The proposed method improves the performance of knowledge distillation methods on two biomedical natural language processing tasks. |
Beyond One-Step Distillation: Bridging the Capacity Gap in Small Language Models via Multi-Step Knowledge Transfer (2026.eacl-srw)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) excel across diverse tasks but remain too large for efficient on-device deployment. |
| Approach: | They revisit multi-step knowledge distillation as an effective remedy . they demonstrate that MSKD improves ROUGE-L and perplexity over single-step approaches . |
| Outcome: | The proposed approach improves ROUGE-L and perplexity over single-step approaches . large language models are too large for efficient on-device deployment, the authors show . |