Papers by Xuexiang Wen
Verifier-Free RL for LLMs via Intrinsic Gradient-Norm Reward (2026.findings-acl)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have demonstrated remarkable capabilities in reasoning tasks and general tasks. |
| Approach: | They propose a "Verifier-free Intrinsic Gradient-Norm Reward" that uses only the policy model itself. |
| Outcome: | The proposed reward outperforms the state-of-the-art RLIF baseline INTUITOR on math benchmarks and shows cross-domain transfer to code benchmarks when trained only on math data. |