Papers by Taiqiang Wu

8 papers
TencentPretrain: A Scalable and Flexible Toolkit for Pre-training Models of Different Modalities (2023.acl-demo)

Copied to clipboard

Challenge: Several pre-training models of different modalities are showing a rising trend of homogeneity in their model structures.
Approach: They propose a toolkit that supports pre-training models of different modalities.
Outcome: The proposed toolkit can match the performance of the original implementations on text, vision, and audio benchmarks.
Multi-stage Distillation Framework for Cross-Lingual Semantic Similarity Matching (2022.findings-naacl)

Copied to clipboard

Challenge: Existing studies have shown that cross-lingual knowledge distillation can improve the performance of pre-trained models for cross-linguistic similarity matching tasks.
Approach: They propose a multi-stage distillation framework for constructing a small-size but high-performance cross-lingual model using contrastive learning, bottleneck, and parameter recurrent strategies.
Outcome: The proposed model can compress the size of XLM-R and MiniLM by more than 50% while the performance is only reduced by about 1%.
Mixture-of-Subspaces in Low-Rank Adaptation (2024.emnlp-main)

Copied to clipboard

Challenge: Using a subspace-inspired Low-Rank Adaptation method, large language models can be optimized for downstream tasks using parameter-efficient finetuning.
Approach: They propose a subspace-inspired Low-Rank Adaptation method that decomposes LoRA weights into two subspaces and merges them into the frozen original weight.
Outcome: The proposed method outperforms LoRA on commonsense reasoning, visual instruction tuning, and subject-driven text-to-image generation tasks.
Revisiting Model Interpolation for Efficient Reasoning (2026.acl-long)

Copied to clipboard

Challenge: Existing models that interpolate weights of two specialized models can be abused for efficient reasoning.
Approach: They propose to merge two specialized models and create a model that combines efficiency and efficiency.
Outcome: The proposed method outperforms existing models on efficiency and effectiveness.
Rethinking Kullback-Leibler Divergence in Knowledge Distillation for Large Language Models (2025.coling-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) have been used in Knowledge Distillation (KD) to compress large models.
Approach: They propose a Kullback-Leiber divergence method which adaptively allocates weights to combine RKL and FKL to reduce the size of Large Language Models (LLMs).
Outcome: The proposed method outperforms baselines and improves diversity and quality of generated responses.
ProFit: Leveraging High-Value Signals in SFT via Probability-Guided Token Selection (2026.findings-acl)

Copied to clipboard

Challenge: Traditional fine-tuning ignores one-to-many nature of language, leading to overfitting . authors propose a method to fine- tune LLMs by leveraging tokens.
Approach: They propose a method to fine-tune Large Language Models by leveraging tokens to mask low-probability tokens.
Outcome: The proposed method outperforms baselines on general reasoning and mathematical benchmarks.
Edge-free but Structure-aware: Prototype-Guided Knowledge Distillation from GNNs to MLPs (2025.coling-main)

Copied to clipboard

Challenge: Existing methods to train low-latency multilayer perceptrons (MLPs) on graph tasks are based on graph nodes and lack graph structural information.
Approach: They propose to distill graph structural information from Graph Neural Networks (GNNs) to low-latency multilayer perceptrons (MLPs) on graph tasks.
Outcome: The proposed method does not require graph edges (edge-free setting) yet learns structure-aware MLPs.
Weight-Inherited Distillation for Task-Agnostic BERT Compression (2024.findings-naacl)

Copied to clipboard

Challenge: Knowledge Distillation (KD) is a predominant approach for BERT compression.
Approach: They propose a weight-inherited distillation method which directly transfers knowledge from the teacher to a compact student model by inheriting the weights.
Outcome: The proposed method outperforms state-of-the-art KD-based methods on GLUE and SQUAD benchmarks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations