Skip to content

Deployment / serving

Serving stack

Common components:

  • Model server — vLLM, SGLang, TensorRT-LLM, Triton, TGI.
  • Tokenizer service — usually colocated with the model server.
  • KV-cache manager — paged attention, prefix caching, eviction.
  • Batching engine — continuous batching to mix prefill and decode.
  • Router — pick which model serves which request.
  • Safety filters — input classifier, output classifier, content moderation.
  • Tool orchestrator — function-calling middleware.
  • Retrieval system — vector database + reranker.
  • Logging / monitoring — request logs, traces, metrics.
  • Evaluation harness — periodic offline eval on production samples.
  • Rollout system — canary, A/B, gradual rollout, fast rollback.

Inference controls

Important parameters:

Parameter What it does
Temperature Randomness of sampling. 0 = greedy. Typical 0.0–1.0.
Top-p (nucleus) Truncate to smallest set of tokens with cumulative prob ≥ p.
Top-k Truncate to top-k tokens.
Max tokens Hard cap on output length.
Stop sequences Strings that terminate generation.
Repetition penalty Discourage repeating tokens.
Presence/frequency penalty OpenAI-style alternatives to repetition penalty.
Tool-choice policy auto, none, or force a specific tool.
JSON / schema constraints Constrained decoding. See structured output.
Safety policy routing Optionally route to a more conservative model for risky inputs.

Routing

Production systems may route between:

  • Small fast model (low-stakes, simple)
  • Large high-quality model (complex queries)
  • Code model (programming queries)
  • Reasoning model (math, planning)
  • Vision model (image inputs)
  • Safety classifier (refusal decisions)
  • Retrieval-augmented mode (factual queries)
  • Tool-using agent mode (tasks requiring action)

Routing decisions can be made by:

  • A small classifier
  • The model itself ("do you need to think hard about this?")
  • Heuristics on input length, image presence, tool calls available

The economic upside of routing is large: most queries are easy and cheap; reserving the expensive model for the hard ones saves serious money at scale.

Monitoring

Track:

  • Latency — p50/p95/p99 for prefill, decode, end-to-end.
  • Cost — tokens in/out per request, $/request.
  • Error rate — 4xx/5xx, timeouts, OOMs.
  • Refusal rate — overall, by category, by user segment.
  • Tool-call failure rate — invalid args, missing tool, errored tool.
  • JSON validity — for structured-output endpoints.
  • User satisfaction — thumbs up/down, stars, retention.
  • Safety incidents — flagged outputs, escalations, abuse reports.
  • Hallucination reports — user-flagged factual errors.
  • Drift over time — distribution shifts in inputs and outputs.
  • Data retention / privacy compliance — auditable trail of what was logged and why.

Observability practices

  • Sample, don't store everything. Full request logging is expensive and risky. Sample 1–5% with metadata, full sample for errors and flagged outputs.
  • Hash sensitive inputs. PII in user prompts is a liability. Either redact or hash.
  • Trace across the stack. Distributed tracing across retrieval → model → tool calls → response makes debugging multi-hop failures tractable.
  • Replay capability. When a bug surfaces, you should be able to replay the exact request against a different model or configuration.

Practical tips

  • Continuous batching is non-negotiable. Static batching wastes 10–100× throughput at typical workloads.
  • Prefix caching is free money. A few percent of system-prompt overlap unlocks large savings; common in agent workflows.
  • Plan for fast rollback. A model bug needs to be reversible in under an hour. Maintain previous-version capacity.
  • Canary new models. Roll out to 1% → 5% → 25% → 100% with eval gates between each step.
  • Monitor refusal rate at the same priority as latency. A spike in refusals after a model update is a real product problem.
  • Run shadow traffic — a copy of production traffic against the new model, with comparison metrics, before any user-visible rollout.

Reference serving frameworks

Further reading