Papers by Jongpil Kim
A Study on Knowledge Distillation from Weak Teacher for Scaling Up Pre-trained Language Models (2023.findings-acl)
Copied to clipboard
| Challenge: | a study shows that DWT can be effective in the vision domain and natural language processing pre-training stages. |
| Approach: | They examine three key factors to optimize Distillation from Weak Teacher (DWT) DWT is a method of transferring knowledge from a weaker teacher model to a larger student model to improve its performance. |
| Outcome: | a new study examines three key factors to optimize DWT in NLP pre-training scenarios . the impact of teacher model quality and guidelines for adjusting the weighting value for DW T loss are examined . |
Co-training and Co-distillation for Quality Improvement and Compression of Language Models (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Knowledge Distillation (KD) compresses expensive pre-trained language models . however, most smaller models fail to surpass performance of larger model . |
| Approach: | They propose a framework that co-trains two models while mutually distilling knowledge to improve performance and inference speed together. |
| Outcome: | The proposed framework outperforms the original larger model by 1.66 on the GLUE benchmark. |