Deployment / serving¶
Serving stack¶
Common components:
- Model server — vLLM, SGLang, TensorRT-LLM, Triton, TGI.
- Tokenizer service — usually colocated with the model server.
- KV-cache manager — paged attention, prefix caching, eviction.
- Batching engine — continuous batching to mix prefill and decode.
- Router — pick which model serves which request.
- Safety filters — input classifier, output classifier, content moderation.
- Tool orchestrator — function-calling middleware.
- Retrieval system — vector database + reranker.
- Logging / monitoring — request logs, traces, metrics.
- Evaluation harness — periodic offline eval on production samples.
- Rollout system — canary, A/B, gradual rollout, fast rollback.
Inference controls¶
Important parameters:
| Parameter | What it does |
|---|---|
| Temperature | Randomness of sampling. 0 = greedy. Typical 0.0–1.0. |
| Top-p (nucleus) | Truncate to smallest set of tokens with cumulative prob ≥ p. |
| Top-k | Truncate to top-k tokens. |
| Max tokens | Hard cap on output length. |
| Stop sequences | Strings that terminate generation. |
| Repetition penalty | Discourage repeating tokens. |
| Presence/frequency penalty | OpenAI-style alternatives to repetition penalty. |
| Tool-choice policy | auto, none, or force a specific tool. |
| JSON / schema constraints | Constrained decoding. See structured output. |
| Safety policy routing | Optionally route to a more conservative model for risky inputs. |
Routing¶
Production systems may route between:
- Small fast model (low-stakes, simple)
- Large high-quality model (complex queries)
- Code model (programming queries)
- Reasoning model (math, planning)
- Vision model (image inputs)
- Safety classifier (refusal decisions)
- Retrieval-augmented mode (factual queries)
- Tool-using agent mode (tasks requiring action)
Routing decisions can be made by:
- A small classifier
- The model itself ("do you need to think hard about this?")
- Heuristics on input length, image presence, tool calls available
The economic upside of routing is large: most queries are easy and cheap; reserving the expensive model for the hard ones saves serious money at scale.
Monitoring¶
Track:
- Latency — p50/p95/p99 for prefill, decode, end-to-end.
- Cost — tokens in/out per request, $/request.
- Error rate — 4xx/5xx, timeouts, OOMs.
- Refusal rate — overall, by category, by user segment.
- Tool-call failure rate — invalid args, missing tool, errored tool.
- JSON validity — for structured-output endpoints.
- User satisfaction — thumbs up/down, stars, retention.
- Safety incidents — flagged outputs, escalations, abuse reports.
- Hallucination reports — user-flagged factual errors.
- Drift over time — distribution shifts in inputs and outputs.
- Data retention / privacy compliance — auditable trail of what was logged and why.
Observability practices¶
- Sample, don't store everything. Full request logging is expensive and risky. Sample 1–5% with metadata, full sample for errors and flagged outputs.
- Hash sensitive inputs. PII in user prompts is a liability. Either redact or hash.
- Trace across the stack. Distributed tracing across retrieval → model → tool calls → response makes debugging multi-hop failures tractable.
- Replay capability. When a bug surfaces, you should be able to replay the exact request against a different model or configuration.
Practical tips¶
- Continuous batching is non-negotiable. Static batching wastes 10–100× throughput at typical workloads.
- Prefix caching is free money. A few percent of system-prompt overlap unlocks large savings; common in agent workflows.
- Plan for fast rollback. A model bug needs to be reversible in under an hour. Maintain previous-version capacity.
- Canary new models. Roll out to 1% → 5% → 25% → 100% with eval gates between each step.
- Monitor refusal rate at the same priority as latency. A spike in refusals after a model update is a real product problem.
- Run shadow traffic — a copy of production traffic against the new model, with comparison metrics, before any user-visible rollout.
Reference serving frameworks¶
- vLLM — high-throughput open serving with paged attention. github.com/vllm-project/vllm
- SGLang — fast LLM and VLM serving with structured outputs. github.com/sgl-project/sglang
- TensorRT-LLM — NVIDIA. Highest throughput, more setup. github.com/NVIDIA/TensorRT-LLM
- TGI (Text Generation Inference) — HuggingFace.
llama.cpp— for CPU / edge / Apple Silicon. github.com/ggerganov/llama.cpp- MLC LLM — cross-platform local inference. github.com/mlc-ai/mlc-llm
Further reading¶
- Kwon et al., "vLLM / PagedAttention", 2023. arxiv.org/abs/2309.06180
- Yu et al., "Orca: A Distributed Serving System", 2022. The continuous-batching paper. usenix.org
- Pope et al., "Efficiently Scaling Transformer Inference", 2022. arxiv.org/abs/2211.05102
- Yao et al., "Distributed Inference and Fine-tuning of Large Language Models Over The Internet", 2023.
- HuggingFace Inference Endpoints docs — huggingface.co/docs/inference-endpoints