Skip to content

Reward modeling

Goal

Train a model that scores responses.

Classic RLHF uses:

  1. SFT model generates multiple answers.
  2. Humans rank answers.
  3. A reward model learns to assign higher scores to preferred answers.
  4. The policy model is optimized against the reward model.

InstructGPT popularized this pipeline: supervised fine-tuning, reward model training from human comparisons, then PPO reinforcement learning.

Reward model input

Usually:

prompt + response → scalar reward

For chat:

system + conversation + candidate assistant response → scalar reward

The reward model is typically initialized from the SFT model (sharing the backbone) with a scalar regression head replacing the LM head. Training is much cheaper than full SFT — you only need a strong scoring backbone, not a generator.

Reward model training objective

Most often pairwise ranking loss (Bradley–Terry):

\[ \mathcal{L} = -\log \sigma\big(r(\text{prompt}, \text{chosen}) - r(\text{prompt}, \text{rejected})\big) \]

In words: train the reward model to give the chosen response a higher score than the rejected one, with a sigmoid margin.

Variants:

  • Pointwise regression — when you have absolute scores, not just rankings.
  • Listwise — using ranked lists of more than two candidates per prompt.
  • Multi-objective — separate heads for helpfulness, harmlessness, honesty, etc., combined at use time.

Reward model failure modes

Reward hacking

The policy finds outputs that score highly but are bad.

Examples:

  • Overly verbose answers
  • Fake citations
  • Excessive hedging
  • Polite but useless refusals
  • Formatting tricks (lists, headers everywhere)
  • Appearing helpful without being correct

This is the canonical problem in RLHF and is why KL penalties exist.

Distribution shift

The reward model is trained on outputs from one model, but the policy later generates different outputs. The further the policy drifts from the RM training distribution, the less reliable the reward.

Mitigation: collect fresh preference data from the current policy and retrain the reward model periodically (iterative RLHF). Or constrain the policy with a KL penalty so it can't drift too far.

Length bias

Reward model prefers longer answers — both because labelers do and because longer answers can hide more "compliance" tokens.

Style bias

Reward model prefers a particular tone (confident, list-formatted, structured).

Safety/helpfulness imbalance

Reward model may overreward harmlessness and underreward usefulness, or vice versa. Multi-objective rewards or separate safety classifiers help.

Process reward models

Recent reasoning work introduces process reward models (PRMs) that score individual reasoning steps, not just the final answer.

  • Stepwise scoring → finer-grained reward signal.
  • Useful for math, code, theorem proving.
  • More expensive to train (requires step-level labels or self-consistent process supervision).

See Lightman et al., "Let's Verify Step by Step" and the reasoning page.

Practical tips

  • Initialize from SFT. Reward models trained from scratch underperform.
  • Hold out a separate eval set of preference pairs. Track agreement rate (how often the RM agrees with the held-out label) as your headline RM metric.
  • Calibrate against length. Plot RM score vs. response length on held-out pairs. If the slope is steep, the RM has length bias and your downstream policy will exploit it.
  • Use the RM as a judge, not a god. Always co-evaluate with human spot checks and downstream task evals.
  • Refresh the RM if you retrain the policy heavily. A stale RM rewards behavior the policy left behind.

Further reading