Skip to content

Post-training

Post-training turns a base model (a raw next-token predictor) into an assistant (a model that follows instructions, answers, refuses, calls tools, and behaves as a coherent agent).

The classic pipeline, popularized by InstructGPT:

  1. Supervised fine-tuning (SFT) — train on high-quality instruction/response pairs.
  2. Preference data — collect comparisons between responses.
  3. Reward modeling — train a model that scores responses.
  4. RLHF with PPO — optimize the policy against the reward model.

Modern open recipes often replace step 4 with offline preference optimization:

  1. DPO and friends — directly optimize on preference pairs without a reward model. Includes IPO, KTO, ORPO, SimPO, CPO, RLVR, GRPO.

Where each method shines

Goal Best fit
Teach format, style, refusal patterns SFT
Teach taste over plausible alternatives DPO or PPO
Optimize against scalable AI feedback RLAIF / Constitutional AI
Optimize against verifiable outcomes (math, code, JSON) RLVR (see DPO and friends)
Optimize chain-of-thought reasoning GRPO-style methods (see Reasoning)

Common failure pattern

The single most common post-training mistake is applying preference optimization too aggressively. A KL-budget that's too loose, a reward model that's been over-trained, or a DPO β that's too low produces models that:

  • Are verbose and meandering (length-biased reward)
  • Refuse benign requests (over-cautious reward)
  • Sound confident on things they don't know (sycophancy)
  • Lose niche capabilities the base model had (mode collapse)

The cure is anchoring: keep the policy close to the SFT model (KL penalty for PPO, β in DPO), evaluate broadly (not just on the preference set), and stop early.