Challenge: Recent advances in code generation focus on optimizing the thought process, but lack effective process supervision, making it difficult to optimize the thoughts.
Approach: They propose a method that leverages the code execution feedback to build a code PRM by collecting a large dataset of thought traces and then training it to take both the reasoning process and code execution as input.
Outcome: The proposed approach outperforms baselines and strong LLMs in the inference stage.

Similar Papers

R-PRM: Reasoning-Driven Process Reward Modeling (2025.emnlp-main)

Copied to clipboard

Challenge: Existing Process Reward Models (PRMs) output evaluation scores directly, limiting both learning efficiency and evaluation accuracy.
Approach: They propose a Reasoning-Driven Process Reward Modeling (R-PRM) which activates inherent reasoning to enhance process-level evaluation.
Outcome: The proposed model outperforms baseline models on ProcessBench and PRMBench by 13.9 and 8.5 F1 scores.
A Comprehensive Survey of Process Reward Models: Data Generation, Model Construction, and Usage (2026.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) have advanced reasoning ability, yet conventional alignment remains dominated by outcome reward models that judge only final answers.
Approach: They summarize applications across math, code, text, multimodal reasoning, robotics, and agents . goal is to clarify design spaces, reveal open challenges, and guide future research toward fine-grained, robust reasoning alignment.
Outcome: The proposed model enables finer credit assignment, richer diagnostics, and improved robustness.
Process Reward Models Meet Planning: Generating Precise and Scalable Datasets for Step-Level Rewards (2026.acl-long)

Copied to clipboard

Challenge: Existing PRM datasets are expensive to construct and limited to the mathematical domain.
Approach: They propose a method to generate a corpus of one million reasoning steps using the Planning Domain Definition Language.
Outcome: The proposed model generates a corpus of approximately one million reasoning steps across various PDDL domains and trains them.
Process-Supervised Reward Models for Verifying Clinical Note Generation: A Scalable Approach Guided by Domain Expertise (2025.emnlp-main)

Copied to clipboard

Challenge: Currently, no automated, scalable method exists to evaluate the quality of LLM-generated clinical notes, leaving manual evaluation the gold standard.
Approach: They propose a framework for training PRMs to deliver step-level reward signals for LLM-generated clinical notes.
Outcome: The proposed framework outperforms reasoning and non-reasoning models on key evaluations and selects physician-preferred clinical notes with 56.2% accuracy.
PhysPRM: A Generative Process Reward Model with Fine-grained Diagnosis for Physics Problem Solving (2026.findings-acl)

Copied to clipboard

Challenge: Existing Large Language Models (LLMs) struggle with physics problem solving due to difficulties in decoding implicit constraints and maintaining physical consistency.
Approach: They propose a Generative PRM that treats evaluation as a generative task . it produces fine-grained diagnoses comprising critiques, final judgments, and specific error types .
Outcome: The proposed model improves performance across seven benchmarks in Best-of-N and critique refinement strategies.
Out of Distribution, Out of Luck: Process Rewards Misguide Reasoning Models (2026.eacl-short)

Copied to clipboard

Challenge: 80% of reasoning model outputs respond to formatting artifacts rather than mathematical content.
Approach: They evaluate process reward models that provide step-level feedback during inference . they identify distinct reward prediction patterns that differentiate reasoning from non-reasoning model outputs .
Outcome: The proposed model fails to enhance and sometimes degrade reasoning model performance.
The Lessons of Developing Process Reward Models in Mathematical Reasoning (2025.findings-acl)

Copied to clipboard

Challenge: a recent study shows that process reward models can make mistakes, leading to wrong conclusions.
Approach: They propose a consensus filtering mechanism that integrates MC estimation with LLM-as-a-judge to improve model performance and data efficiency.
Outcome: The proposed model outperforms existing open-source alternatives and provides practical guidelines for future research.
Retrieval-Augmented Process Reward Model for Generalizable Mathematical Reasoning (2025.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) have advanced mathematical reasoning, but they still struggle with out-of-distribution (OOD) issues.
Approach: They propose a framework to evaluate the logical validity of reasoning steps . they retrieves semantically similar questions and steps for PRM as a warmup .
Outcome: The proposed framework outperforms baseline models on multiple real-world datasets.
An Efficient and Precise Training Data Construction Framework for Process-supervised Reward Model in Mathematical Reasoning (2025.acl-long)

Copied to clipboard

Challenge: Existing methods for constructing process supervision training data are costly or suffer from poor quality.
Approach: They propose a framework called EpicPRM which annotates each intermediate reasoning step based on its quantified contribution and uses an adaptive binary search algorithm to enhance annotation precision and efficiency.
Outcome: The proposed framework improves annotation precision and efficiency and can be used to train a high-quality training dataset with 50k annotated intermediate steps.
Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision (2025.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs) are prone to hallucination, especially during multihop tasks.
Approach: They propose a hierarchical, erroraware discriminative PRM that classifies math errors at each step and combines finegrained signals to estimate step correctness.
Outcome: The proposed model outperforms the prior best in a new stateof-theart PRMScore of 67.7 on a 400Ksample dataset .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations