Skip to content

Deployment

Pretraining and post-training make the model. Deployment makes it useful.

This section covers the full operational lifecycle:

The most common production failures

In rough order of frequency:

  1. Latency / cost regressions after a model update — quantization or context-length changes break serving assumptions.
  2. Refusal regressions — over-refusal of benign prompts users used to get help with.
  3. Format regressions — JSON parse rate drops, tool calls become unreliable.
  4. Hallucination — model confidently wrong, especially after preference optimization.
  5. Jailbreak / prompt injection in production inputs that didn't appear in test sets.
  6. Drift — user behavior changes faster than the eval set captures it.

The remedies — robust eval suites, gradual rollouts, observability, fast rollback — are operational, not algorithmic.