Home Knowledge Base Process Reward Model (PRM)

Process Reward Model (PRM) is a reward model that assigns scores to each intermediate reasoning step rather than only the final answer — enabling fine-grained training signal for multi-step reasoning tasks where step-level correctness matters more than final outcome.

ORM vs. PRM

PRM Training

PRM Applications

Math Reasoning Results

Challenges

Process reward models are the key to closing the gap between raw reasoning capability and reliable problem-solving — by rewarding correct thinking processes rather than just correct answers, PRMs enable the kind of robust multi-step reasoning that characterizes mathematical expertise.

process reward modelprmreasoning rewardoutcome reward modelormreward hacking

Explore 500+ Semiconductor & AI Topics

From EUV lithography to CUDA optimization — search the full knowledge base or chat with our AI assistant.