Deployment¶
Pretraining and post-training make the model. Deployment makes it useful.
This section covers the full operational lifecycle:
- Compression & efficiency — quantization, distillation, speculative decoding, KV-cache, LoRA.
- Evaluation — benchmarks, human eval, LLM-as-judge, calibration.
- Red teaming — find failures before users do.
- Serving — KV cache, batching, routing, observability.
- Continuous improvement — logs → failures → data → retrain.
The most common production failures¶
In rough order of frequency:
- Latency / cost regressions after a model update — quantization or context-length changes break serving assumptions.
- Refusal regressions — over-refusal of benign prompts users used to get help with.
- Format regressions — JSON parse rate drops, tool calls become unreliable.
- Hallucination — model confidently wrong, especially after preference optimization.
- Jailbreak / prompt injection in production inputs that didn't appear in test sets.
- Drift — user behavior changes faster than the eval set captures it.
The remedies — robust eval suites, gradual rollouts, observability, fast rollback — are operational, not algorithmic.