Papers by Shuai Shao
Two-Stage Regularization-Based Structured Pruning for LLMs (2026.acl-long)
Copied to clipboard
Mingkuan Feng, Jinyang Wu, Siyuan Liu, Shuai Zhang, Hongjian Fang, Ruihan Jin, Feihu Che, Pengpeng Shao, Zhengqi Wen, Jianhua Tao
| Challenge: | Structural pruning is a promising solution for large language models . prior structured pruning methods remove unimportant parameters based on certain metrics . |
| Approach: | They propose a structural pruning method that iteratively learns the weights of transformer layers by adding their l1-norm to the loss function. |
| Outcome: | The proposed pruning method outperforms strong layer-wise pruning methods without requiring retraining. |
RadialRouter: Structured Representation for Efficient and Robust Large Language Models Routing (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Current routing methods are limited in exploring the connection between query and LLM characteristics. |
| Approach: | They propose a framework for LLM routing that uses a transformer-based backbone and a radial structure to articulate the query-LLMs relationship. |
| Outcome: | The proposed framework outperforms existing routing methods by 9.2% and 5.8% on RouterBench. |
On Orthogonality Constraints for Transformers (2021.acl-short)
Copied to clipboard
Aston Zhang, Alvin Chan, Yi Tay, Jie Fu, Shuohang Wang, Shuai Zhang, Huajie Shao, Shuochao Yao, Roy Ka-Wei Lee
| Challenge: | a dedicated study on orthogonality constraints for transformers has been lacking . plug-and-play constraints increase the BLEU of transformers . |
| Approach: | They propose to use plug-and-play constraints to encourage matrices to be orthogonal for numerical stability. |
| Outcome: | The proposed constraint increases the BLEU on the large-scale WMT’16 EnDe benchmark by a factor of 28.4 to 29.6. |
Luna: A Lightweight Evaluation Model to Catch Language Model Hallucinations with High Accuracy and Low Cost (2025.coling-industry)
Copied to clipboard
| Challenge: | Retrieval-augmented generation (RAG) systems are crucial for enhancing the capabilities of large language models (LLMs) in industry applications. |
| Approach: | They propose a DeBERTA-large encoder for hallucination detection in RAG settings that is fine-tuned for halluination detection. |
| Outcome: | The proposed model outperforms GPT-3.5 and commercial evaluation frameworks on the hallucination detection task, with 97% and 91% reduction in cost and latency, respectively. |
Pandora’s Box or Aladdin’s Lamp: A Comprehensive Analysis Revealing the Role of RAG Noise in Large Language Models (2025.acl-long)
Copied to clipboard
| Challenge: | Retrieval-Augmented Generation (RAG) has emerged as a promising approach to address hallucinations in large language models (LLMs). |
| Approach: | They define seven distinct noise types from a linguistic perspective and establish a Noise RAG Benchmark (NoiserBench) they propose to evaluate noise that is beneficial to LLMs and noise that's harmful to LRMs. |
| Outcome: | The proposed framework consists of seven distinct noise types from a linguistic perspective and includes multiple datasets and reasoning tasks. |
Distribution Shift Alignment Helps LLMs Simulate Survey Response Distributions (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing methods to simulate survey responses are based on zero-shot methods, but they are sensitive to prompt changes and deviate from the real-world distributions. |
| Approach: | They propose a distribution shift alignment method that aligns both the output distributions and the distribution shifts across different backgrounds to provide results closer to the true distribution than the training data. |
| Outcome: | The proposed method outperforms zero-shot methods on five public survey datasets and reduces the required real data by 53.48-69.12%. |