Papers by Yangen Hu
Reward Alignment Optimization: A Direct Point-wise Alignment Approach (2026.acl-long)
Copied to clipboard
| Challenge: | Existing Direct Alignment Algorithms (DAAs) are limiting in generalizaiton to implicit rewards. |
| Approach: | They propose a point-wise direct alignment method that uses an explicit reward model to specify exact target generation probabilities and align the policy offline towards them. |
| Outcome: | The proposed method outperforms existing direct alignment algorithms while enabling controllable target probability distributions. |
Enhancing Efficiency and Exploration in Reinforcement Learning for LLMs (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing approaches allocate an equal number of rollouts to all questions during the RL process, which is inefficient. |
| Approach: | They propose a mechanism for dynamically allocating rollout budgets based on the difficulty of the problems, enabling more efficient RL training. |
| Outcome: | The proposed model improves response precision while preserving exploratory ability to uncover potential correct pathways. |