Papers by Xingjin Wang

2 papers
Entropy Scheduling in Reinforcement Learning for Large Language Models (2026.findings-acl)

Copied to clipboard

Challenge: entropy in reinforcement learning functions analogously to the learning rate in LLMs.
Approach: They propose an entropy scheduling system that optimizes different pre-set goals by controlling and scheduling entropicy at each step of the RL process.
Outcome: The proposed method improves AIME2024 from 50.9 to 54.9 within 40 training steps.
LDM2: A Large Decision Model Imitating Human Cognition with Dynamic Memory Enhancement (2023.findings-emnlp)

Copied to clipboard

Challenge: Extensive experiments conducted in two interactive environments have shown that our LDM2 outperforms the baselines in terms of both score and success rate.
Approach: They propose a large decision model with memory that leverages a dynamic memory mechanism to construct dynamic prompts, guiding the LLMs in making proper decisions according to the faced state.
Outcome: The proposed model outperforms baseline models in two interactive environments in terms of score and success rate.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations