Papers with ToRL
GeometryZero: Advancing Geometry Solving via Group Contrastive Policy Optimization (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing methods for auxiliary construction training are expensive and underperform . Existing Corresponding Author training methods lack self-correction capabilities in reasoning chains. |
| Approach: | They propose a reinforcement learning framework that rewards auxiliary construction with geometric reasoning by grouping construction rewards with a Length Reward. |
| Outcome: | Experiments on Geometry3K and MathVista show that GeometryZero outperforms baselines on auxiliary constructions. |