Papers by De-Chuan Zhan
Maximizing the Effectiveness of Larger BERT Models for Compression (2025.acl-long)
Copied to clipboard
| Challenge: | Existing methods for capturing large BERT models as teachers do not fully exploit the potential advantages of larger teachers. |
| Approach: | They propose a method that leverages a pretrained teacher model to guide the training of a lightweight student model to enhance knowledge transfer. |
| Outcome: | The proposed method enhances knowledge transfer by leveraging a pretrained teacher model to guide the training of a lightweight student model. |
Logits-Based Block Pruning with Affine Transformations for Large Language Models (2026.findings-eacl)
Copied to clipboard
| Challenge: | Existing methods for pruning models rely on calibration data and neglect cumulative effects of pruning on subsequent blocks. |
| Approach: | They propose to use the Logit Disruption Score (LDS) to measure the impact of pruning by comparing the cosine similarity between the logits of the original and pruned models. |
| Outcome: | Experiments on multiple datasets show that the proposed pruning technique reduces reliance on calibration data and improves generalization, achieving competitive results with existing methods. |
DART: Distilling Autoregressive Reasoning to Silent Thought (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing models that use Chain-of-Thought (CoT) have been slow to deploy in real-time applications due to its autoregressive nature. |
| Approach: | They propose a framework that replaces autoregressive CoT with non-autoregressive Silent Thought (ST) the framework uses a lightweight Reasoning Evolvement Module to align hidden states with the CoT pathway and a Reasoning Embedment Module (REM) during inference, only the ST pathway is activated, enabling the ST tokens to evolve into informative embeddings. |
| Outcome: | The proposed framework replaces autoregressive CoT with non-autoregressive Silent Thought (ST) it enables LLMs to generate answers directly from ST tokens without additional computational cost . |