Safety training¶
Goal¶
Make the model helpful while reducing harmful behavior.
Safety is not one stage. It appears in:
- Data filtering
- SFT
- Preference data
- RLHF / DPO
- Constitutional AI / RLAIF
- Red teaming
- Inference-time policies (system prompts, classifiers, output filters)
- Monitoring
A safe deployed system is the product of all of these layers, not any single one.
Safety data categories¶
- Violence
- Self-harm
- Medical advice
- Legal advice
- Financial advice
- Cybersecurity
- Biosecurity
- Weapons
- Hate and harassment
- Sexual content
- Privacy
- Child safety
- Fraud
- Extremism
- Political persuasion
- Misinformation
- Prompt injection
- Data exfiltration
Each category has its own policy decisions: where the helpful/harmful line is, when to redirect, when to refuse, when to provide harm-reduction information.
Refusal quality¶
Bad refusal:
I cannot help with that.
Better refusal:
I can't help with instructions for breaking into an account. I can help you recover access to your own account or set up stronger authentication.
A good refusal is:
- Brief
- Clear
- Non-judgmental
- Specific enough to be useful
- Redirects to safe alternatives
- Does not leak harmful details (don't list which steps you're refusing to help with)
Soft-refusal training is a force multiplier
Most safety regressions in deployed models come from over-refusal — refusing benign requests because they superficially resemble unsafe ones. Including hundreds of "looks unsafe but is actually fine" examples in SFT (e.g., "how do I kill a process in Linux") prevents this.
Constitutional AI¶
Constitutional AI (Bai et al., 2022) uses a set of principles — a "constitution" — to guide model critique and revision of its own outputs, followed by supervised and reinforcement-learning phases. It is a way to scale harmlessness training with less direct human labeling.
The pattern:
- Sample a response.
- Ask the model to critique its own response against principles.
- Ask the model to revise.
- Use the revision (chosen) vs. original (rejected) as preference data.
The "constitution" is just a set of natural-language principles — short, readable, auditable. This is much more scalable than per-example human labeling and produces models that explain their refusals consistently.
Safety failure modes¶
- Over-refusal — refusing benign requests.
- Under-refusal — answering things the policy says no to.
- Policy inconsistency — refusing one phrasing, helping with another equivalent one.
- Jailbreak susceptibility — adversarial prompts bypass the policy.
- Indirect prompt injection — retrieved content overrides system instructions. See Red teaming.
- Tool misuse — model uses tools to exfiltrate data or trigger destructive actions.
- Sensitive data leakage — model reveals memorized training data, system prompts, or user history.
- Sycophantic unsafe agreement — user pushes, model caves and helps.
- Unsafe completion after benign prefix — first 80% looks fine, then drifts unsafe.
- Multilingual safety gaps — refusal works in English but fails in lower-resource languages.
Multi-layer policy¶
Treat safety as a defense in depth problem:
- Pretraining filter — remove the worst data classes.
- SFT refusal data — teach refusal style and policy.
- Preference / RLHF — refine refusal calibration and helpfulness tradeoff.
- Inference-time classifier — catch unsafe inputs before they reach the model.
- Output classifier — catch unsafe outputs before they reach the user.
- Tool sandboxing — limit blast radius of agentic actions.
- Monitoring & audit logs — detect emerging patterns and incident-respond.
Skipping any layer is a tempting shortcut. Don't.
Practical tips¶
- Write down the policy. A vague "be helpful but safe" mandate produces vague labels. Write a 5–20 page policy with worked examples per category.
- Calibrate refusal vs helpfulness explicitly. Use a benign-but-edgy eval set (e.g., XSTest) to track over-refusal as carefully as you track under-refusal.
- Translate policies into multiple languages. Multilingual safety is the most common gap in open-weights models.
- Test adversarially before launch. Internal red team + external red team (academics, security pros) before any deployment past internal testing.
- Plan for rollbacks. Safety bugs are emergencies. You need to be able to roll back the model in under an hour.
Further reading¶
- Bai et al., "Constitutional AI", 2022. arxiv.org/abs/2212.08073
- Bai et al., "Training a Helpful and Harmless Assistant", 2022. arxiv.org/abs/2204.05862
- Llama Guard — Meta. Open-weights safety classifier. arxiv.org/abs/2312.06674
- WildGuard — AI2. Comprehensive safety classifier and benchmark. arxiv.org/abs/2406.18495
- XSTest — Röttger et al. Tests over-refusal of safe prompts. arxiv.org/abs/2308.01263
- Anthropic's Acceptable Use Policy — published policy as a reference. anthropic.com/legal/aup
- OpenAI Safety policies — openai.com/policies
- NIST AI Risk Management Framework — nist.gov/itl/ai-risk-management-framework