Papers by Shuai LI

2 papers
Reinforcement Learning on Pre-Training Data (2026.acl-long)

Copied to clipboard

Challenge: Recent progress in large language models is driven by scaling of training compute through pre-training with nexttoken prediction (NTP) or post-training (RL) Pre-training using NTP enables models to acquire extensive knowledge and skills from general data, but it suffers from data inefficiency and catastrophic forgetting in continual learning settings.
Approach: They propose to scale training compute through pre-training with next-token prediction (NTP) or post-training by scaling reinforcement learning (RL) to improve learning from general data.
Outcome: Experiments on multiple benchmarks and models show that the proposed approach improves continual pre-training and provides a strong foundation for post-training on Qwen3-8B-Base.
Distribution Shift Alignment Helps LLMs Simulate Survey Response Distributions (2026.findings-acl)

Copied to clipboard

Challenge: Existing methods to simulate survey responses are based on zero-shot methods, but they are sensitive to prompt changes and deviate from the real-world distributions.
Approach: They propose a distribution shift alignment method that aligns both the output distributions and the distribution shifts across different backgrounds to provide results closer to the true distribution than the training data.
Outcome: The proposed method outperforms zero-shot methods on five public survey datasets and reduces the required real data by 53.48-69.12%.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations