A PRM assigns a score to each reasoning step; these scores are used for training, reranking, or inference-time search.
Rewarding only the final outcome is sparse and doesn't indicate which reasoning step was wrong.