Post-Training
How it works
After pretraining the model goes through one or more phases: (1) SFT on instruction–response pairs; (2) preference optimisation — a reward model is trained on human comparisons and the policy is optimised (RLHF/PPO) or direct preference optimisation (DPO) is used; (3) optionally RLVR with an automatic, verifiable reward (e.g. code/math correctness) and distillation from a teacher model. Phases can be combined and iterated.
Problem solved
A pretrained base model predicts text but does not follow instructions, can be unhelpful or unsafe, and does not keep a desired format. Post-training closes the gap between "can model language" and "is a useful assistant".
Components
Fine-tuning on high-quality instruction–response pairs.
Aligns to human preferences via a reward model + RL or directly (DPO).
Reinforcement learning with an automatic, verifiable reward (code, math).