Robots Atlas>ROBOTS ATLAS
Training

Reward Function

1998ActivePublished: 28 September 2026Updated: 28 September 2026Published
Key innovation
Formalises an agent's goal as a single scalar reward signal, turning open-ended "what we want" into a measurable, optimisable objective — the foundation of all reinforcement learning.
Category
Training
Abstraction level
Primitive
Operation level
TrainingPost-trainingRobot control
Use cases
Reinforcement learningRobotics: control, manipulation, locomotionReward shaping and sim-to-realRLHF: language-model fine-tuningGames and decision-making agentsReinforcement Learning with Verifiable Rewards (RLVR)Recommender systems and sequential optimisation

How it works

1) The environment is modelled as an MDP with state and action spaces and transition dynamics. 2) At each step t the agent takes action a_t in state s_t, transitions to s_{t+1}, and receives a scalar reward r_{t+1} = R(s_t, a_t, s_{t+1}). 3) The return is the discounted sum G_t = Σ_{k≥0} γ^k r_{t+k+1}, where γ down-weights distant rewards. 4) The agent seeks a policy π maximising the expected return E_π[G_t]. 5) Value functions V_π(s) and Q_π(s,a) estimate the expected return and drive policy improvement (e.g. in PPO, GRPO). 6) When the reward is sparse, reward shaping is used; potential-based shaping F = γΦ(s') − Φ(s) leaves the optimal policy unchanged. 7) In RLHF the reward function is replaced by a reward model trained on human preference pairs, and the policy is optimised with an RL algorithm (e.g. PPO) with a KL penalty against the base model.

Problem solved

Reinforcement learning needs an unambiguous, measurable definition of an agent's goal. The reward function solves this by reducing any goal to a scalar signal that can be optimised with RL methods — without having to supply the agent with correct actions in advance (as in supervised learning).

Components

Scalar reward signalImmediate feedback signal

A single real number returned by the environment after each transition; higher means a better transition with respect to the goal.

Reward mapping RTask goal specification

The function assigning a reward to a (state, action) pair or (state, action, next state) triple. It defines the task goal.

ReturnQuantity being optimised

The cumulative, usually discounted, sum of future rewards that the agent maximises. The discount factor γ weights temporally distant rewards.

Shaping termSpeeds learning without changing the optimum

An auxiliary signal added to the reward to accelerate learning. Potential-based shaping preserves the optimal policy (Ng et al., 1999).

Official

Reward modelReplaces a hand-specified reward

A neural network approximating the reward function, trained on human preferences when an explicit reward cannot be defined (e.g. text quality).

Official

Implementation

Implementation pitfalls
Reward hacking / misspecificationCritical

The agent maximises the reward signal in ways unintended by the designer, exploiting loopholes in the reward definition.

Fix:Careful specification, adversarial testing, preference-based reward models, penalties for undesired behaviour, monitoring.
Sparse reward and poor explorationHigh

When reward appears only at the goal, the agent rarely encounters a signal and learning is very slow or fails.

Fix:Reward shaping, curriculum learning, intrinsic motivation, hindsight experience replay.
Shaping that changes the optimal policyHigh

An arbitrary additional reward signal can shift the optimum and teach the agent unintended behaviour.

Fix:Use potential-based shaping (Ng et al., 1999), which guarantees invariance of the optimal policy.
Unstable reward scaleMedium

Overly large or variable reward values destabilise gradient estimation and training.

Fix:Reward normalisation and clipping, return standardisation, appropriate γ.
Reward-model over-optimisation (Goodhart)High

In RLHF the policy can over-optimise an imperfect reward model, degrading true quality (Goodhart's law).

Fix:KL penalty against the base model, early stopping, periodic reward-model refresh.

Evolution

Original paper · 1998 · MIT Press · Richard S. Sutton
Reinforcement Learning: An Introduction
Richard S. Sutton, Andrew G. Barto
1957
Markov Decision Processes and dynamic programming

Bellman formalises MDPs and the optimality equation, in which reward is part of the problem specification.

1998
Reward hypothesis and canonical RL formulation (Sutton & Barto)
Inflection point

The textbook "Reinforcement Learning: An Introduction" establishes the reward function as the central element of RL and states the reward hypothesis.

1999
Potential-based reward shaping
Inflection point

Ng, Harada and Russell prove that shaping F = γΦ(s')−Φ(s) preserves the optimal policy (policy invariance).

2016
Reward hacking / misspecification in AI safety

"Concrete Problems in AI Safety" (Amodei et al.) frames reward hacking as a key problem of poorly specified rewards.

2017
Learning reward functions from human preferences
Inflection point

Christiano et al. show a reward model can be trained from human preferences instead of hand-specifying a reward — the basis of RLHF.

2022
Reward model in RLHF at scale (InstructGPT)
Inflection point

Ouyang et al. use a reward model from human rankings and PPO to align language models, popularising RLHF.

Hyperparameters (configurable axes)

Discount factor γCritical

Weight of future vs. immediate rewards (γ ∈ [0,1)). Lower γ = short horizon, higher γ = long planning horizon.

Reward density (sparse vs dense)High

How often a non-zero reward appears. Sparse rewards hinder exploration; dense rewards ease learning but risk reward hacking.

Reward scale and normalisationMedium

The range of reward values. Normalisation/clipping stabilise training and prevent single terms from dominating.

Shaping potential ΦMedium

The potential used in potential-based shaping. A well-chosen Φ speeds learning without changing the optimal policy.

Computational complexity

Time complexity: O(1) na przejście (analityczna); O(forward pass) dla modelu nagrody.

Hardware requirements

Primary

A reward function is itself a specification/mapping with no hardware preference; its compute cost depends on the implementation.

Good fit

When the reward is a learned model (reward model in RLHF), evaluating it is a neural-network forward pass that benefits from GPUs.