Papers by Kehao Chen
Beyond Majority Voting: Towards Fine-grained and More Reliable Reward Signal for Test-Time Reinforcement Learning (2026.acl-long)
Copied to clipboard
| Challenge: | RLVR is a paradigm for improving reasoning ability of large language models . but voting results often induce confirmation bias and suffer from sparse rewards . |
| Approach: | They propose a framework integrating model confidence and dynamic subgroup partitioning to address these issues. |
| Outcome: | The proposed framework outperforms recent baselines on multiple models and benchmarks. |