Reward Function
How it works
1) The environment is modelled as an MDP with state and action spaces and transition dynamics. 2) At each step t the agent takes action a_t in state s_t, transitions to s_{t+1}, and receives a scalar reward r_{t+1} = R(s_t, a_t, s_{t+1}). 3) The return is the discounted sum G_t = Σ_{k≥0} γ^k r_{t+k+1}, where γ down-weights distant rewards. 4) The agent seeks a policy π maximising the expected return E_π[G_t]. 5) Value functions V_π(s) and Q_π(s,a) estimate the expected return and drive policy improvement (e.g. in PPO, GRPO). 6) When the reward is sparse, reward shaping is used; potential-based shaping F = γΦ(s') − Φ(s) leaves the optimal policy unchanged. 7) In RLHF the reward function is replaced by a reward model trained on human preference pairs, and the policy is optimised with an RL algorithm (e.g. PPO) with a KL penalty against the base model.
Problem solved
Reinforcement learning needs an unambiguous, measurable definition of an agent's goal. The reward function solves this by reducing any goal to a scalar signal that can be optimised with RL methods — without having to supply the agent with correct actions in advance (as in supervised learning).
Components
A single real number returned by the environment after each transition; higher means a better transition with respect to the goal.
The function assigning a reward to a (state, action) pair or (state, action, next state) triple. It defines the task goal.
The cumulative, usually discounted, sum of future rewards that the agent maximises. The discount factor γ weights temporally distant rewards.
An auxiliary signal added to the reward to accelerate learning. Potential-based shaping preserves the optimal policy (Ng et al., 1999).
Official
A neural network approximating the reward function, trained on human preferences when an explicit reward cannot be defined (e.g. text quality).
Official
Implementation
The agent maximises the reward signal in ways unintended by the designer, exploiting loopholes in the reward definition.
When reward appears only at the goal, the agent rarely encounters a signal and learning is very slow or fails.
An arbitrary additional reward signal can shift the optimum and teach the agent unintended behaviour.
Overly large or variable reward values destabilise gradient estimation and training.
In RLHF the policy can over-optimise an imperfect reward model, degrading true quality (Goodhart's law).
Evolution
Bellman formalises MDPs and the optimality equation, in which reward is part of the problem specification.
The textbook "Reinforcement Learning: An Introduction" establishes the reward function as the central element of RL and states the reward hypothesis.
Ng, Harada and Russell prove that shaping F = γΦ(s')−Φ(s) preserves the optimal policy (policy invariance).
"Concrete Problems in AI Safety" (Amodei et al.) frames reward hacking as a key problem of poorly specified rewards.
Christiano et al. show a reward model can be trained from human preferences instead of hand-specifying a reward — the basis of RLHF.
Ouyang et al. use a reward model from human rankings and PPO to align language models, popularising RLHF.
Hyperparameters (configurable axes)
Weight of future vs. immediate rewards (γ ∈ [0,1)). Lower γ = short horizon, higher γ = long planning horizon.
How often a non-zero reward appears. Sparse rewards hinder exploration; dense rewards ease learning but risk reward hacking.
The range of reward values. Normalisation/clipping stabilise training and prevent single terms from dominating.
The potential used in potential-based shaping. A well-chosen Φ speeds learning without changing the optimal policy.
Computational complexity
Time complexity: O(1) na przejście (analityczna); O(forward pass) dla modelu nagrody.
Hardware requirements
A reward function is itself a specification/mapping with no hardware preference; its compute cost depends on the implementation.
When the reward is a learned model (reward model in RLHF), evaluating it is a neural-network forward pass that benefits from GPUs.