Compute RPE = received reward − predicted reward; the signal updates predictions and the policy.
An agent must know how much the actual outcome deviates from expectations to update its behavior.