Papers by Wangshu Zhang
AdapterDistillation: Non-Destructive Task Composition with Knowledge Distillation (2023.emnlp-industry)
Copied to clipboard
| Challenge: | Recent work on learning from multiple tasks has shown that adding an extra fusion layer to implement knowledge composition is non-scalable for some applications. |
| Approach: | They propose a two-stage knowledge distillation algorithm to extract task specific knowledge by using local data to train a student adapter. |
| Outcome: | Experiments on frequently asked question retrieval in task-oriented dialog systems validate the efficiency of AdapterDistillation. |
Query Distillation: BERT-based Distillation for Ensemble Ranking (2020.coling-industry)
Copied to clipboard
| Challenge: | Recent years have witnessed substantial progress in the development of neural ranking networks, but an increasingly heavy computational burden due to growing numbers of parameters and the adoption of model ensembles. |
| Approach: | They propose a two-stage distillation method that allows a smaller student model to be trained while benefiting from the better performance of the teacher model. |
| Outcome: | The proposed method shows higher-quality rankings compared to the teacher model. |
Improving Knowledge Production Efficiency With Question Answering on Conversation (2023.acl-industry)
Copied to clipboard
| Challenge: | Existing researches on conversation-based QA focus on document-based tasks . current researche focuses on document based tasks, but there is a lack of researche on conversation based qa . |
| Approach: | They propose a multi-span extraction model on conversation-based QA and introduce continual pre-training and multi-task learning schemes to further improve model performance. |
| Outcome: | The proposed model outperforms baseline on two Chinese datasets and will be released for research purposes. |