Robots Atlas>ROBOTS ATLAS
Training

Post-Training

2022ActivePublished: 29 September 2026Updated: 29 September 2026Published
Key innovation
The umbrella stage of adapting a pretrained base model after pretraining — aligning it to tasks, instructions and human preferences rather than learning more raw knowledge.
Category
Training
Abstraction level
Paradigm
Operation level
Post-trainingTrainingModel
Use cases
Aligning chatbots (ChatGPT, Claude, Gemini)Teaching instruction followingBoosting reasoning (RLVR)Reducing harmful outputs (alignment)Fine-tuning VLA policies in robotics

How it works

After pretraining the model goes through one or more phases: (1) SFT on instruction–response pairs; (2) preference optimisation — a reward model is trained on human comparisons and the policy is optimised (RLHF/PPO) or direct preference optimisation (DPO) is used; (3) optionally RLVR with an automatic, verifiable reward (e.g. code/math correctness) and distillation from a teacher model. Phases can be combined and iterated.

Problem solved

A pretrained base model predicts text but does not follow instructions, can be unhelpful or unsafe, and does not keep a desired format. Post-training closes the gap between "can model language" and "is a useful assistant".

Components

Supervised fine-tuning (SFT)Bootstraps assistant behaviour

Fine-tuning on high-quality instruction–response pairs.

Preference optimisation (RLHF/DPO)Shapes quality and safety

Aligns to human preferences via a reward model + RL or directly (DPO).

RLVRStrengthens reasoning

Reinforcement learning with an automatic, verifiable reward (code, math).

Evolution

Original paper · 2022 · Long Ouyang
Training language models to follow instructions with human feedback
Long Ouyang, Jeff Wu, Xu Jiang, et al.
2022
InstructGPT establishes SFT+RLHF as the post-training standard
Inflection point
2023
DPO simplifies preference alignment without a separate reward model
2025
RLVR becomes a key part of post-training for reasoning models