Post-training¶
Post-training turns a base model (a raw next-token predictor) into an assistant (a model that follows instructions, answers, refuses, calls tools, and behaves as a coherent agent).
The classic pipeline, popularized by InstructGPT:
- Supervised fine-tuning (SFT) — train on high-quality instruction/response pairs.
- Preference data — collect comparisons between responses.
- Reward modeling — train a model that scores responses.
- RLHF with PPO — optimize the policy against the reward model.
Modern open recipes often replace step 4 with offline preference optimization:
- DPO and friends — directly optimize on preference pairs without a reward model. Includes IPO, KTO, ORPO, SimPO, CPO, RLVR, GRPO.
Where each method shines¶
| Goal | Best fit |
|---|---|
| Teach format, style, refusal patterns | SFT |
| Teach taste over plausible alternatives | DPO or PPO |
| Optimize against scalable AI feedback | RLAIF / Constitutional AI |
| Optimize against verifiable outcomes (math, code, JSON) | RLVR (see DPO and friends) |
| Optimize chain-of-thought reasoning | GRPO-style methods (see Reasoning) |
Common failure pattern¶
The single most common post-training mistake is applying preference optimization too aggressively. A KL-budget that's too loose, a reward model that's been over-trained, or a DPO β that's too low produces models that:
- Are verbose and meandering (length-biased reward)
- Refuse benign requests (over-cautious reward)
- Sound confident on things they don't know (sycophancy)
- Lose niche capabilities the base model had (mode collapse)
The cure is anchoring: keep the policy close to the SFT model (KL penalty for PPO, β in DPO), evaluate broadly (not just on the preference set), and stop early.