Challenge: knowledge distillation (KD) targeting attention should selectively accelerate syntax acquisition, a study finds . logit-based KD dramatically improves data-efficiency, attention-based one provides minimal benefit even for syntactic tasks.
Approach: a study predicts that knowledge distillation targeting attention should selectively accelerate syntax acquisition . a systolic analysis of student models compared to logit-based knowledge distillations .
Outcome: a new study shows that knowledge distillation (KD) targeting attention accelerates syntax acquisition . the hypothesis is tested on syntactic benchmarks and perplexity.

Similar Papers

Efficient Transformer Knowledge Distillation: A Performance Review (2023.emnlp-industry)

Copied to clipboard

Challenge: Pretrained transformer language models have been gaining popularity in the field of natural language processing . however, there is no study into the intersection of these two fields .
Approach: They propose a method to extract knowledge from transformers to produce high-performing efficient attention models with low costs.
Outcome: The proposed model compression method preserves up to 98.6% of original model performance across short-context tasks and up to 95.8% on long-concept Named Entity Recognition tasks while decreasing inference times by up to 57%.
Align-to-Distill: Trainable Attention Alignment for Knowledge Distillation in Neural Machine Translation (2024.lrec-main)

Copied to clipboard

Challenge: Existing knowledge distillation approaches to NMT often rely on heuristics when deciding which teacher layers to distill from.
Approach: They propose an approach to align student attention heads with their teacher counterparts by heuristics to solve a feature mapping problem.
Outcome: The proposed strategy shows gains of +3.61 and +0.63 BLEU points for WMT-2022 DeDsb and WMT-2014 EnDe compared to baselines.
Sparse Logit Sampling: Accelerating Knowledge Distillation in LLMs (2025.acl-long)

Copied to clipboard

Challenge: Knowledge distillation is a cost-effective technique to distill knowledge in Large Language Models, if the teacher output logits can be pre-computed and cached.
Approach: They propose an importance-sampling-based method which provides unbiased estimates, preserves the gradient in expectation, and requires storing significantly sparser logits.
Outcome: The proposed method enables faster training of student models with marginal overhead (10%) compared to cross-entropy based training, while maintaining competitive performance compared with full distillation.
Scalable Syntax-Aware Language Models Using Knowledge Distillation (P19-1)

Copied to clipboard

Challenge: Prior work has shown that syntactic neural language models learn from small amounts of training data more effectively than sequential models.
Approach: They propose a knowledge distillation technique that transfers knowledge from a syntactic language model trained on a small corpus to an LSTM language model and enables it to develop a more structurally sensitive representation of the larger training data.
Outcome: The proposed method improves on baseline syntactic evaluations on LSTMs with a higher level of accuracy than previous methods.
Unveiling the Magic: Investigating Attention Distillation in Retrieval-Augmented Generation (2024.naacl-short)

Copied to clipboard

Challenge: Retrieval-augmented generation framework addresses the limitations of large language models by enabling real-time knowledge updates for more accurate answers.
Approach: They propose to use attention distillation to improve retrieval-augmented language models' learning performance by identifying key factors influencing their workflow and proposing indicators for optimizing models’ training methods and avoiding ineffective training.
Outcome: The proposed framework improves the learning performance of large language models in the training phase but also reduces the impact of ineffective training.
On the Generalization vs Fidelity Paradox in Knowledge Distillation (2025.findings-acl)

Copied to clipboard

Challenge: Knowledge distillation (KD) is a key technique for compressing large language models into smaller ones while preserving performance.
Approach: They propose to use knowledge distillation to compress large language models into smaller ones while preserving performance.
Outcome: The proposed technique improves the performance of smaller models by 10% while providing only marginal benefits for larger models.
Understanding and Improving Knowledge Distillation for Quantization Aware Training of Large Transformer Encoders (2022.emnlp-main)

Copied to clipboard

Challenge: Knowledge distillation (KD) has been used for quantization-aware training to improve the ability of a lightweight model with the transferred knowledge from the teacher.
Approach: They propose two methods to improve attention recovery of quantized large Transformers by combining attention-map and attention-output losses.
Outcome: The proposed methods achieve state-of-the-art accuracy for quantized large Transformer encoder models with sub-2-bit weight quantization.
Generation-Distillation for Efficient Natural Language Understanding in Low-Data Settings (D19-61)

Copied to clipboard

Challenge: Recent research points to knowledge distillation as a potential solution for NLU tasks.
Approach: They propose a training approach that distills large finetuned LMs into a small network using unlabeled training examples.
Outcome: The proposed approach outperforms BERT training approaches while using 300 times fewer parameters.
Dynamic Knowledge Distillation for Pre-trained Language Models (2021.emnlp-main)

Copied to clipboard

Challenge: Existing methods conduct knowledge distillation statically, e.g., student model aligns output distribution to teacher model on pre-defined training dataset.
Approach: They propose a dynamic knowledge distillation that empowers the student to adjust the learning procedure according to its competency . they find it is promising and provide discussions on potential future directions towards more efficient methods .
Outcome: The proposed method can boost student model performance while accelerating training . the proposed method reduces memory usage and accelerates model inference .
Dataset Distillation with Attention Labels for Fine-tuning BERT (2023.acl-short)

Copied to clipboard

Challenge: Specifically, we propose to introduce attention labels, which can efficiently distill the knowledge from the original dataset and transfer it to the transformer models via attention probabilities.
Approach: They propose to introduce attention labels which can efficiently distill the knowledge from the original dataset and transfer it to the transformer models via attention probabilities.
Outcome: The proposed methods perform impressively in four different NLP tasks and achieve 93.2% accuracy in AGNews, which is 98.5% of the original dataset even with only one sample per class and only one gradient step.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations