Papers by Huaijie Wang
Offline Reinforcement Learning for LLM Multi-step Reasoning (2025.findings-acl)
Copied to clipboard
| Challenge: | Large language models (LLMs) are increasingly applied to complex tasks requiring multi-step reasoning. |
| Approach: | They propose an offline method for enhancing multi-step reasoning by optimizing the soft Bellman Equation by combining a policy model and a value function. |
| Outcome: | The proposed method surpasses existing methods on multi-step reasoning benchmarks and can be extended to multi-iteration frameworks when additional resources are available. |