Beyond Reasoning Gains: Mitigating General-Capability Forgetting in Large Reasoning Models (2026.findings-acl)
Copied to clipboard
Hoang Phan, Xianjun Yang, Yuanshun Yao, Jingyu Zhang, Shengjie Bi, Xiaocheng Tang, Madian Khabsa, Lijuan Liu, Deren Lei
| Challenge: | Reinforcement learning with verifiable rewards (RLVR) has delivered impressive gains in mathematical and multimodal reasoning . however, the recipe introduces a significant risk of capability regression, where models forget foundational skills after prolonged training without employing regularization strategies. |
| Approach: | They propose a replay strategy with dynamic objective reweighting for general knowledge preservation using short-horizon signals of convergence and instability. |
| Outcome: | The proposed method preserves general capabilities and improves reasoning . it can be applied to existing RLVR pipelines without training additional models or tuning . |
Similar Papers
Turning Failures into Value: Negative Experience Replay for RLVR via Confidence Gating and Boundary Failure Sampling (2026.acl-long)
Copied to clipboard
| Challenge: | Existing experience replay methods for RLVR ignore sample inefficiency . expensive reasoning trajectories are discarded immediately after a single gradient update . |
| Approach: | They propose a method to replay failure trajectories to improve model refinement . they propose 'nexGRPO' which employs mid-confidence gating to filter invalid noise and saturated errors. |
| Outcome: | The proposed model outperforms strong baaselines and achieves improved out-of-distribution generalization. |
Knowledge-to-Verification: Exploring RLVR for LLMs in Knowledge-Intensive Domains (2026.acl-long)
Copied to clipboard
Zhonghang Yuan, Zhefan Wang, Fang Hu, Zihong Chen, Jinzhe Li, Gang Li, Jie Ying, Huanjun Kong, Songyang Zhang, Nanqing Dong
| Challenge: | Recent large language models (LLMs) have demonstrated remarkable progress in reasoning, but their applications on knowledge-intensive domains have not been explored due to the scarcity of high-quality verifiable data. |
| Approach: | They propose a framework that extends reinforcement learning with verifiable rewards (RLVR) to knowledge-intensive domains through automated verififiability data synthesis while enabling verification of the LLM's reasoning process. |
| Outcome: | Extensive experiments show that the proposed framework enhances the reasoning of large language models in knowledge-intensive domains without significantly compromising the model’s general capabilities. |
Crossing the Reward Bridge: Expanding Reinforcement Learning with Verifiable Rewards Across Diverse Domains (2026.acl-long)
Copied to clipboard
| Challenge: | Reinforcement learning with verifiable rewards (RLVR) has been effective on structured tasks, but its reliance on simple, rule-based verifiers creates a bottleneck. |
| Approach: | They propose a framework that uses a generative verifier to provide soft, probabilistic rewards. |
| Outcome: | The proposed framework outperforms existing models up to 10x their size and can be scalable and effective. |
Revisiting Entropy in Reinforcement Learning for Large Reasoning Models (2026.findings-acl)
Copied to clipboard
Renren Jin, Pengzhi Gao, Yuqi Ren, Zhuowen Han, Tongxuan Zhang, Wuwei Huang, Wei Liu, Jian Luan, Deyi Xiong
| Challenge: | Reinforcement learning with verifiable rewards (RLVR) has emerged as a paradigm for enhancing the reasoning capabilities of large language models. |
| Approach: | They propose a positive-advantage reweighting approach that regulates model entropy by adjusting the loss weights assigned to tokens with positive advantages during RLVR training. |
| Outcome: | The proposed approach regulates model entropy by adjusting loss weights assigned to tokens with positive advantages during RLVR training while maintaining competitive performance. |
DARL: Encouraging Diverse Answers for General Reasoning without Verifiers (2026.findings-acl)
Copied to clipboard
| Challenge: | Recent efforts such as RLPR have extended RLVR to general domains, enabling training on broader datasets and achieving improvements over RL PR. |
| Approach: | They propose a framework that encourages the generation of diverse answers within a controlled deviation range from the reference while preserving alignment with it. |
| Outcome: | Extensive experiments on 13 benchmarks show that DARL surpasses RLPR in both reasoning accuracy and output diversity. |
RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization (2026.acl-long)
Copied to clipboard
Yihong Dong, Xue Jiang, Yongding Tao, Huanyu Liu, Kechi Zhang, Lili Mou, Rongyu Cao, Yingwei MA, Jue Chen, Binhua Li, Zhi Jin, Fei Huang, Yongbin Li, Ge Li
| Challenge: | Reinforcement Learning with Verifiable Reward (RLVR) has significantly advanced the complex reasoning abilities of Large Language Models (LLMs). |
| Approach: | They propose a hybrid-policy optimization approach that synergizes internal exploitation with external data to achieve stronger reasoning capabilities. |
| Outcome: | The proposed approach achieves state-of-the-art performance on six math reasoning benchmarks and superior performance on out-of distribution reasoning tasks. |
Entropy-Aware Reshaping of Reinforcement Signals for Multi-Answer Reasoning (2026.findings-acl)
Copied to clipboard
| Challenge: | Reinforcement learning with verifiable rewards (RLVR) is a standard post-training paradigm for large language models. |
| Approach: | They propose a framework that reshapes how learning signals are normalized and aggregated. |
| Outcome: | Experiments on MCTACO and MMLU-Multi show that the proposed framework improves accuracy, training stability and cross-dataset transfer performance. |
Rethinking Entropy Interventions in RLVR: An Entropy Change Perspective (2026.acl-long)
Copied to clipboard
Zhezheng Hao, Hong Wang, Haoyang Liu, Jian Luo, Jiarui Yu, Hande Dong, Qiang Lin, Can Wang, Jiawei Chen
| Challenge: | Large Language Models (LLMs) have remarkable reasoning capabilities in complex tasks such as mathematics and coding. |
| Approach: | They propose an entropy-modulation method that adaptively reweighs tokens based on theoretically-estimated entropic variations. |
| Outcome: | The proposed method outperforms state-of-the-art methods in six mathematical reasoning and three coding benchmarks. |
Does RLVR Extend Reasoning Boundaries? Investigating Capability Expansion in Vision-Language Models (2026.acl-long)
Copied to clipboard
| Challenge: | Recent studies suggest that RLVR amplifies behaviors inherent to the pre-training distribution rather than inducing new capabilities. |
| Approach: | They propose a framework for RLVR that extends the spatial reasoning boundary . they use a mapping framework where the difficulty is precisely regulated by path length and number of turns . |
| Outcome: | The proposed framework extends the spatial reasoning boundary on two real-world navigation benchmarks. |
Balancing Classification and Calibration Performance in Decision-Making LLMs via Calibration Aware Reinforcement Learning (2026.findings-acl)
Copied to clipboard
| Challenge: | Large language models (LLMs) are increasingly deployed in decision-making tasks where accuracy and reliable confidence estimates are essential. |
| Approach: | They propose a calibration-aware reinforcement learning formulation that directly adjusts decision-token probabilities. |
| Outcome: | The proposed model preserves RLVR’s accuracy level while mitigating overconfidence, reducing ECE scores up to 9 points. |