Skip to content

Compact checklists

Use these as gate-checks before declaring a stage done.

Pretraining checklist

  • Data provenance documented
  • PII and secrets filtered
  • Deduplication complete
  • Benchmark decontamination run
  • Tokenizer validated
  • Data mixture tested on small models
  • Scaling curves estimated
  • Training stability tested
  • Eval harness ready
  • Checkpoint recovery tested

See Base pretraining.

SFT checklist

  • Chat template stable
  • Assistant-token masking correct (with a unit test that asserts which positions contribute)
  • Multi-turn examples included
  • Tool examples included
  • Safety examples included
  • Structured-output examples included
  • Domain examples reviewed
  • Over-refusal measured (e.g., XSTest)
  • Format adherence measured (IFEval, schema validators)
  • Held-out SFT eval set in place
  • Per-category eval reported (general, code, math, reasoning, multi-turn, tool)

See SFT.

Preference checklist

  • Preference rubric clear and written down
  • Labeler agreement measured
  • Chosen / rejected pairs high quality
  • Length bias controlled (length-balanced or length-controlled metric)
  • Safety / helpfulness balanced
  • Reward hacking monitored
  • KL or reference constraint tuned
  • Human eval confirms gains (not just preference-set wins)
  • Capability evals (MMLU, HumanEval, IFEval) hold up
  • AlpacaEval 2 LC or Arena-Hard improvement is real

See Preference data and DPO and friends.

Tool-use checklist

  • No-tool examples included
  • Single-tool examples covering full schema
  • Multi-tool / chained-tool trajectories included
  • Invalid tool calls tested
  • Tool errors tested (timeouts, permission, empty results)
  • Prompt-injection attempts in retrieved content tested
  • Schema validation measured (JSON parse rate, schema match rate)
  • Tool-result faithfulness measured (no fabricated tool outputs)
  • Permission boundaries enforced (sandbox, audit log)
  • BFCL / τ-bench scores collected
  • End-to-end task success on private agent eval

See Tool use.

Long-context checklist

  • Position scaling chosen (default: YaRN)
  • Continued pretraining on long sequences run
  • Mixed short + long sequences in training
  • Needle-in-a-haystack passes at full context
  • RULER multi-skill scores collected
  • LongBench / InfiniteBench scores collected
  • Lost-in-the-middle tested
  • Serving stack tuned for long KV cache (paged attention, eviction, cache quantization)

See Long-context.

Safety checklist

  • Written policy with worked examples per category
  • Refusal style data in SFT
  • Soft-refusal / "looks unsafe but is fine" data included
  • Multilingual safety data included
  • Safety classifier on input (e.g., Llama Guard / WildGuard)
  • Safety classifier on output
  • Tool sandboxing in place
  • Internal red team complete
  • External red team complete (when stakes warrant)
  • XSTest / over-refusal scores collected
  • HarmBench / AdvBench scores collected
  • Disclosure & rollback plan documented

See Safety and Red teaming.

Deployment checklist

  • Latency measured (p50/p95/p99 prefill and decode)
  • Cost measured (cost per request, cost per million tokens)
  • Refusal rate monitored
  • Tool-call failure rate monitored
  • JSON validity rate monitored
  • Safety incidents tracked
  • User feedback loop active
  • Rollback possible in under an hour
  • Model / version metadata logged with every request
  • Evaluation set updated after failures
  • Canary rollout policy defined
  • Privacy / retention policy documented

See Serving and Continuous improvement.