V(s)/Q(s,a) estimate expected discounted return; they are learned via TD or Monte Carlo and used to choose actions.
An agent needs an estimate of the long-term value of states/actions, not just the immediate reward.