Skip to content

Direct Preference Optimization and related methods

Goal

Use preference pairs directly without training a separate reward model or running online RL.

DPO (Rafailov et al., 2023) reframes preference optimization so the language model is directly optimized to prefer chosen responses over rejected responses. The DPO paper describes it as simpler and computationally lighter than RLHF, avoiding explicit reward-model fitting and sampling from the policy during fine-tuning.

The DPO loss is:

\[ \mathcal{L}_{\text{DPO}} = -\log\sigma\left(\beta \log\frac{\pi_\theta(y_w|x)}{\pi_{\text{ref}}(y_w|x)} - \beta \log\frac{\pi_\theta(y_l|x)}{\pi_{\text{ref}}(y_l|x)}\right) \]

where y_w is the chosen ("winning") response and y_l is the rejected one. The reference model is the SFT model, frozen.

DPO data format

Same as preference data:

{
  "prompt": "...",
  "chosen": "...",
  "rejected": "..."
}

Benefits

  • Simpler than PPO
  • More stable
  • Easier to implement
  • No separate reward model required
  • Works well for many chat-alignment tasks
  • ~5–10× cheaper to run than full PPO

Limitations

  • Depends heavily on quality of preference pairs
  • Offline — only sees the data you have, can't explore.
  • Less natural for interactive environments
  • May not explore beyond the dataset
  • Can still over-optimize style biases (length, verbosity)
  • Sensitive to β — too low and the policy drifts; too high and nothing changes.

The post-DPO landscape is rich. Each variant addresses specific weaknesses:

Method Key idea When to use
IPO (Azar et al., 2023) Replaces sigmoid with squared loss; less prone to overfitting Noisy preference data
KTO (Ethayarajh et al., 2024) Uses thumbs-up/down (single labels), not pairs When you only have pointwise feedback
ORPO (Hong et al., 2024) Combines SFT + odds-ratio preference loss in one step Skip the SFT-then-DPO two-stage pipeline
SimPO (Meng et al., 2024) Length-normalized DPO without reference model Reference-free; addresses length bias
CPO (Xu et al., 2024) Contrastive preference for translation Translation/structured output
RLAIF (Lee et al., 2023) RLHF with AI labels in place of human labels Safety, when you have a strong judge
GRPO (Shao et al., 2024) Group-relative PPO; no value model Reasoning models — used by DeepSeek-R1
RLVR (Lambert et al., 2024) RL with verifiable rewards (no RM needed) Math, code, instruction-following with checkers

Which one is best depends on data type, task verifiability, compute, and stability needs.

How to choose

  • Default for chat alignment: DPO. It is simple, well-supported, and well-tuned defaults exist.
  • If you have only thumbs-up/down: KTO.
  • If you want to skip the SFT stage entirely: ORPO.
  • If length bias is killing you: SimPO.
  • For math / code / structured tasks with verifiers: RLVR.
  • For reasoning models trained against verifiable rewards: GRPO.
  • For interactive agents that need exploration: PPO/RLHF still has a place.

Iterative DPO

A common pattern in modern open recipes:

  1. SFT model M0.
  2. Generate responses with M0, label preferences (human or AI), DPO → M1.
  3. Generate responses with M1, label preferences (the new candidates are better, so the labels are sharper), DPO → M2.
  4. Repeat for 2–4 rounds.

Iterative DPO ("self-rewarding" or "online DPO" in the literature) consistently outperforms a single DPO pass because the preference data stays on-policy.

Practical tips

  • Run SFT first. Most variants need a reasonable SFT base. ORPO is the exception.
  • Start with β around 0.1. Lower → more drift, more learning, more risk. Higher → safer, less change.
  • Hold out a chat eval set like AlpacaEval 2 (length-controlled) or MT-Bench. Watch for regressions, not just preference-set wins.
  • Always check length. DPO has a strong length-bias failure mode. Plot mean response length over training.
  • Don't over-train. 1–3 epochs at most. DPO loss can keep going down while quality regresses.
  • Co-evaluate with capability evals. MMLU, HumanEval, IFEval, GSM8K — DPO can silently regress capabilities the SFT model had.

Further reading