Skip to content

Preference data

Goal

Collect comparisons that tell the model which response is better.

A preference example usually looks like:

{
  "prompt": "Explain why the sky is blue.",
  "chosen": "The sky looks blue because molecules in the air scatter shorter blue wavelengths...",
  "rejected": "The sky is blue because it reflects the ocean."
}

Preference data teaches:

  • Helpfulness
  • Honesty
  • Harmlessness
  • Concision
  • Style
  • Refusal quality
  • Reasoning quality
  • Tool-use appropriateness
  • Factual calibration

Sources of preference data

Human labelers

Humans rank model outputs.

Pros:

  • Captures real user preferences
  • Good for style, helpfulness, safety

Cons:

  • Expensive
  • Noisy
  • Labeler disagreement
  • Hard for expert domains

Expert annotators

Used for medicine, law, coding, science, safety.

Pros: higher correctness. Cons: very expensive, hard to scale.

AI feedback

A stronger model judges outputs.

Constitutional AI uses a written set of principles and model-generated feedback to train for harmlessness with less direct human labeling.

Pros:

  • Scales cheaply
  • Can cover many edge cases

Cons:

  • Judge bias
  • Weakness inherited from teacher
  • Can reward plausible but wrong answers

Verifiable rewards

Used when correctness can be checked automatically.

Examples:

  • Math answer matches
  • Code passes unit tests
  • JSON validates
  • Function call executes
  • SQL returns expected result
  • Theorem proof checks
  • Browser task succeeds

Tülu 3 introduced Reinforcement Learning with Verifiable Rewards (RLVR) for tasks such as math and instruction following where outcomes can be automatically checked. RLVR is a major reason recent open models close the gap to frontier proprietary ones on math and code.

How to construct preference pairs

You need meaningfully different candidates. Common patterns:

  1. Same model, multiple samples. Sample 4–8 responses from the SFT model with high temperature. Have a judge (human or AI) rank them. Use top-vs-bottom as chosen/rejected.
  2. Two models. Sample one response from a stronger model and one from a weaker model. Often used for distillation-flavored preference data.
  3. Edited pairs. Take a chosen response and produce a rejected one by deleting a key fact, breaking format, or making it sycophantic.
  4. Per-dimension pairs. Build separate sets for "helpful but unsafe vs safe", "long vs concise", "reasoned vs guessed".
  5. Synthetic via constitutional critique. Have the model self-critique against principles, then revise. The original is rejected, the revision is chosen.

Pair diversity beats pair volume

A small, well-curated set covering a range of failure modes (sycophancy, hallucination, verbosity, format breakage, refusal calibration) trains better preference policies than a giant dataset of low-stakes style preferences.

Length bias and other gotchas

Preference data has notorious systematic biases:

  • Length bias — labelers (and AI judges) prefer longer answers, even when shorter is better.
  • Style bias — labelers prefer particular tones (confident, structured, with bullet points).
  • First-position bias — when comparing two answers, labelers favor whichever they read first.
  • Recency bias — labelers favor whichever model they used last.
  • Verbose-correctness illusion — wordy hedging reads as careful, even when content is wrong.

Mitigations:

  • Length-balance your chosen/rejected pairs explicitly.
  • Randomize order shown to labelers.
  • Use length-controlled metrics (LC-Win-Rate, AlpacaEval 2 length-controlled).
  • Add explicit "shorter is better" preference examples.
  • Calibrate AI judges against held-out human labels.

Practical tips

  • Define the rubric explicitly. A one-sentence "which is better?" produces noisy labels. A two-page rubric with worked examples produces useful labels.
  • Measure inter-labeler agreement. Below ~70% agreement on a held-out set, your data is noisy enough that you should fix the rubric or task before scaling.
  • Don't reuse SFT prompts as preference prompts unchanged. The model will be near-optimal on them, so all candidates look similar. Use harder, longer-tail prompts.
  • Save metadata. For each pair: prompt source, who/what labeled it, when, which models generated each candidate, what rubric version. You'll want this for ablation.

Further reading