Papers by Hangliang Ren
LSRL: Process-Supervised GRPO on Latent Recurrent States Improves Mathematical Reasoning (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Latent-recurrent language models solve tasks by iteratively refining hidden states rather than emitting chain-of-thought tokens. |
| Approach: | They propose a process-supervised variant of Guided Reward Policy Optimization that rewards latent steps at every latent step. |
| Outcome: | The proposed model improves absolute accuracy by +4.27 points on GSM-8K and +2.06 points on MathQA. |