Preference data¶
Goal¶
Collect comparisons that tell the model which response is better.
A preference example usually looks like:
{
"prompt": "Explain why the sky is blue.",
"chosen": "The sky looks blue because molecules in the air scatter shorter blue wavelengths...",
"rejected": "The sky is blue because it reflects the ocean."
}
Preference data teaches:
- Helpfulness
- Honesty
- Harmlessness
- Concision
- Style
- Refusal quality
- Reasoning quality
- Tool-use appropriateness
- Factual calibration
Sources of preference data¶
Human labelers¶
Humans rank model outputs.
Pros:
- Captures real user preferences
- Good for style, helpfulness, safety
Cons:
- Expensive
- Noisy
- Labeler disagreement
- Hard for expert domains
Expert annotators¶
Used for medicine, law, coding, science, safety.
Pros: higher correctness. Cons: very expensive, hard to scale.
AI feedback¶
A stronger model judges outputs.
Constitutional AI uses a written set of principles and model-generated feedback to train for harmlessness with less direct human labeling.
Pros:
- Scales cheaply
- Can cover many edge cases
Cons:
- Judge bias
- Weakness inherited from teacher
- Can reward plausible but wrong answers
Verifiable rewards¶
Used when correctness can be checked automatically.
Examples:
- Math answer matches
- Code passes unit tests
- JSON validates
- Function call executes
- SQL returns expected result
- Theorem proof checks
- Browser task succeeds
Tülu 3 introduced Reinforcement Learning with Verifiable Rewards (RLVR) for tasks such as math and instruction following where outcomes can be automatically checked. RLVR is a major reason recent open models close the gap to frontier proprietary ones on math and code.
How to construct preference pairs¶
You need meaningfully different candidates. Common patterns:
- Same model, multiple samples. Sample 4–8 responses from the SFT model with high temperature. Have a judge (human or AI) rank them. Use top-vs-bottom as chosen/rejected.
- Two models. Sample one response from a stronger model and one from a weaker model. Often used for distillation-flavored preference data.
- Edited pairs. Take a chosen response and produce a rejected one by deleting a key fact, breaking format, or making it sycophantic.
- Per-dimension pairs. Build separate sets for "helpful but unsafe vs safe", "long vs concise", "reasoned vs guessed".
- Synthetic via constitutional critique. Have the model self-critique against principles, then revise. The original is rejected, the revision is chosen.
Pair diversity beats pair volume
A small, well-curated set covering a range of failure modes (sycophancy, hallucination, verbosity, format breakage, refusal calibration) trains better preference policies than a giant dataset of low-stakes style preferences.
Length bias and other gotchas¶
Preference data has notorious systematic biases:
- Length bias — labelers (and AI judges) prefer longer answers, even when shorter is better.
- Style bias — labelers prefer particular tones (confident, structured, with bullet points).
- First-position bias — when comparing two answers, labelers favor whichever they read first.
- Recency bias — labelers favor whichever model they used last.
- Verbose-correctness illusion — wordy hedging reads as careful, even when content is wrong.
Mitigations:
- Length-balance your chosen/rejected pairs explicitly.
- Randomize order shown to labelers.
- Use length-controlled metrics (LC-Win-Rate, AlpacaEval 2 length-controlled).
- Add explicit "shorter is better" preference examples.
- Calibrate AI judges against held-out human labels.
Practical tips¶
- Define the rubric explicitly. A one-sentence "which is better?" produces noisy labels. A two-page rubric with worked examples produces useful labels.
- Measure inter-labeler agreement. Below ~70% agreement on a held-out set, your data is noisy enough that you should fix the rubric or task before scaling.
- Don't reuse SFT prompts as preference prompts unchanged. The model will be near-optimal on them, so all candidates look similar. Use harder, longer-tail prompts.
- Save metadata. For each pair: prompt source, who/what labeled it, when, which models generated each candidate, what rubric version. You'll want this for ablation.
Further reading¶
- Bai et al., "Training a Helpful and Harmless Assistant with RLHF", 2022 — Anthropic's classic preference-collection writeup. arxiv.org/abs/2204.05862
- Constitutional AI — Bai et al., 2022. arxiv.org/abs/2212.08073
- Tülu 3 — RLVR — Lambert et al., 2024. arxiv.org/abs/2411.15124
- UltraFeedback — Cui et al., 2023. Large open synthetic preference corpus. arxiv.org/abs/2310.01377
- HelpSteer / HelpSteer2 — NVIDIA. Human-rated preference data. arxiv.org/abs/2406.08673
- Singhal et al., "A Long Way to Go: Investigating Length Correlations in RLHF", 2023. The length-bias paper. arxiv.org/abs/2310.03716