Compute δ = r + γV(s′) − V(s) and update V(s) toward reducing this error.
We need to learn state values without waiting for episode end and without a full model of the environment.