Skip to content

Safety training

Goal

Make the model helpful while reducing harmful behavior.

Safety is not one stage. It appears in:

A safe deployed system is the product of all of these layers, not any single one.

Safety data categories

  • Violence
  • Self-harm
  • Medical advice
  • Legal advice
  • Financial advice
  • Cybersecurity
  • Biosecurity
  • Weapons
  • Hate and harassment
  • Sexual content
  • Privacy
  • Child safety
  • Fraud
  • Extremism
  • Political persuasion
  • Misinformation
  • Prompt injection
  • Data exfiltration

Each category has its own policy decisions: where the helpful/harmful line is, when to redirect, when to refuse, when to provide harm-reduction information.

Refusal quality

Bad refusal:

I cannot help with that.

Better refusal:

I can't help with instructions for breaking into an account. I can help you recover access to your own account or set up stronger authentication.

A good refusal is:

  • Brief
  • Clear
  • Non-judgmental
  • Specific enough to be useful
  • Redirects to safe alternatives
  • Does not leak harmful details (don't list which steps you're refusing to help with)

Soft-refusal training is a force multiplier

Most safety regressions in deployed models come from over-refusal — refusing benign requests because they superficially resemble unsafe ones. Including hundreds of "looks unsafe but is actually fine" examples in SFT (e.g., "how do I kill a process in Linux") prevents this.

Constitutional AI

Constitutional AI (Bai et al., 2022) uses a set of principles — a "constitution" — to guide model critique and revision of its own outputs, followed by supervised and reinforcement-learning phases. It is a way to scale harmlessness training with less direct human labeling.

The pattern:

  1. Sample a response.
  2. Ask the model to critique its own response against principles.
  3. Ask the model to revise.
  4. Use the revision (chosen) vs. original (rejected) as preference data.

The "constitution" is just a set of natural-language principles — short, readable, auditable. This is much more scalable than per-example human labeling and produces models that explain their refusals consistently.

Safety failure modes

  • Over-refusal — refusing benign requests.
  • Under-refusal — answering things the policy says no to.
  • Policy inconsistency — refusing one phrasing, helping with another equivalent one.
  • Jailbreak susceptibility — adversarial prompts bypass the policy.
  • Indirect prompt injection — retrieved content overrides system instructions. See Red teaming.
  • Tool misuse — model uses tools to exfiltrate data or trigger destructive actions.
  • Sensitive data leakage — model reveals memorized training data, system prompts, or user history.
  • Sycophantic unsafe agreement — user pushes, model caves and helps.
  • Unsafe completion after benign prefix — first 80% looks fine, then drifts unsafe.
  • Multilingual safety gaps — refusal works in English but fails in lower-resource languages.

Multi-layer policy

Treat safety as a defense in depth problem:

  1. Pretraining filter — remove the worst data classes.
  2. SFT refusal data — teach refusal style and policy.
  3. Preference / RLHF — refine refusal calibration and helpfulness tradeoff.
  4. Inference-time classifier — catch unsafe inputs before they reach the model.
  5. Output classifier — catch unsafe outputs before they reach the user.
  6. Tool sandboxing — limit blast radius of agentic actions.
  7. Monitoring & audit logs — detect emerging patterns and incident-respond.

Skipping any layer is a tempting shortcut. Don't.

Practical tips

  • Write down the policy. A vague "be helpful but safe" mandate produces vague labels. Write a 5–20 page policy with worked examples per category.
  • Calibrate refusal vs helpfulness explicitly. Use a benign-but-edgy eval set (e.g., XSTest) to track over-refusal as carefully as you track under-refusal.
  • Translate policies into multiple languages. Multilingual safety is the most common gap in open-weights models.
  • Test adversarially before launch. Internal red team + external red team (academics, security pros) before any deployment past internal testing.
  • Plan for rollbacks. Safety bugs are emergencies. You need to be able to roll back the model in under an hour.

Further reading