Compact checklists¶
Use these as gate-checks before declaring a stage done.
Pretraining checklist¶
- Data provenance documented
- PII and secrets filtered
- Deduplication complete
- Benchmark decontamination run
- Tokenizer validated
- Data mixture tested on small models
- Scaling curves estimated
- Training stability tested
- Eval harness ready
- Checkpoint recovery tested
See Base pretraining.
SFT checklist¶
- Chat template stable
- Assistant-token masking correct (with a unit test that asserts which positions contribute)
- Multi-turn examples included
- Tool examples included
- Safety examples included
- Structured-output examples included
- Domain examples reviewed
- Over-refusal measured (e.g., XSTest)
- Format adherence measured (IFEval, schema validators)
- Held-out SFT eval set in place
- Per-category eval reported (general, code, math, reasoning, multi-turn, tool)
See SFT.
Preference checklist¶
- Preference rubric clear and written down
- Labeler agreement measured
- Chosen / rejected pairs high quality
- Length bias controlled (length-balanced or length-controlled metric)
- Safety / helpfulness balanced
- Reward hacking monitored
- KL or reference constraint tuned
- Human eval confirms gains (not just preference-set wins)
- Capability evals (MMLU, HumanEval, IFEval) hold up
- AlpacaEval 2 LC or Arena-Hard improvement is real
See Preference data and DPO and friends.
Tool-use checklist¶
- No-tool examples included
- Single-tool examples covering full schema
- Multi-tool / chained-tool trajectories included
- Invalid tool calls tested
- Tool errors tested (timeouts, permission, empty results)
- Prompt-injection attempts in retrieved content tested
- Schema validation measured (JSON parse rate, schema match rate)
- Tool-result faithfulness measured (no fabricated tool outputs)
- Permission boundaries enforced (sandbox, audit log)
- BFCL / τ-bench scores collected
- End-to-end task success on private agent eval
See Tool use.
Long-context checklist¶
- Position scaling chosen (default: YaRN)
- Continued pretraining on long sequences run
- Mixed short + long sequences in training
- Needle-in-a-haystack passes at full context
- RULER multi-skill scores collected
- LongBench / InfiniteBench scores collected
- Lost-in-the-middle tested
- Serving stack tuned for long KV cache (paged attention, eviction, cache quantization)
See Long-context.
Safety checklist¶
- Written policy with worked examples per category
- Refusal style data in SFT
- Soft-refusal / "looks unsafe but is fine" data included
- Multilingual safety data included
- Safety classifier on input (e.g., Llama Guard / WildGuard)
- Safety classifier on output
- Tool sandboxing in place
- Internal red team complete
- External red team complete (when stakes warrant)
- XSTest / over-refusal scores collected
- HarmBench / AdvBench scores collected
- Disclosure & rollback plan documented
See Safety and Red teaming.
Deployment checklist¶
- Latency measured (p50/p95/p99 prefill and decode)
- Cost measured (cost per request, cost per million tokens)
- Refusal rate monitored
- Tool-call failure rate monitored
- JSON validity rate monitored
- Safety incidents tracked
- User feedback loop active
- Rollback possible in under an hour
- Model / version metadata logged with every request
- Evaluation set updated after failures
- Canary rollout policy defined
- Privacy / retention policy documented
See Serving and Continuous improvement.