Skip to content

Continuous improvement

A modern LLM product is never done. The most reliable competitive moat is a tight feedback loop: real usage → failure analysis → targeted training data → updated evals → new model → deploy → repeat.

Feedback loop

  1. Collect user interactions — request, response, tool calls, latency, user feedback signals.
  2. Filter and privacy-process logs — redact PII, sample, retain on documented schedule.
  3. Identify failures — automatic (errors, low confidence, low ratings) and manual (random sampling, user reports).
  4. Create targeted evals — every recurring failure becomes a permanent eval case.
  5. Generate or label improved data — write or synthesize SFT/preference pairs that fix the failure.
  6. Fine-tune or preference-train — typically a small targeted update, not a from-scratch retrain.
  7. Red team — verify the fix doesn't introduce new vulnerabilities.
  8. A/B test — run shadow or split traffic.
  9. Deploy gradually — canary → percentage rollout.
  10. Monitor — watch for regressions on existing capabilities.

This loop typically runs on a 1-to-4-week cadence at mature shops.

Common improvement data

  • Bad answers corrected by experts — the highest-leverage SFT data you can get.
  • Failed tool calls — invalid JSON, wrong tool, hallucinated arguments.
  • User thumbs down — noisy but plentiful.
  • Escalated safety cases — anything a human reviewer flagged as harmful.
  • Hallucination reports — user-flagged or RAG-faithfulness-detected factual errors.
  • Format failures — schema validation logs.
  • Long-context failures — questions where the answer was in context but the model missed it.
  • Domain-specific mistakes — caught by domain SMEs reviewing samples.
  • Ambiguous prompts where clarification was needed — teach the model to ask.

Beware

Do not blindly train on user feedback.

Reasons:

  • Users may be wrong. A thumbs-down on a correct refusal is bad training signal.
  • Feedback is biased. Engaged users disproportionately complain about edge cases.
  • Logs contain private data. PII leakage through training is a serious risk.
  • Malicious users can poison data. If "user prefers this" influences training, attackers will craft inputs to manipulate the model.
  • Thumbs-up may reward style over correctness. Users like confident, structured, lengthy answers — even when shorter and uncertain would be more honest.

The remedy is a review pipeline: every piece of user-derived training data should pass through some combination of automatic filtering, expert review, and decontamination before it becomes a training signal.

Privacy and retention

  • Document your retention policy. Users (and regulators) need to know what is logged, for how long, and for what purposes.
  • Provide opt-out. Users should be able to opt their data out of training.
  • Run PII detection on logs. Redact before storage when possible; redact at retrieval otherwise.
  • Audit access. Who can read raw logs? Tracked in an audit log itself?
  • Plan for deletion requests. Deleting a user's data from a trained model is hard; the legally-cleanest answer is "we don't train on their data" rather than "we deleted it after training."

Metrics that matter for continuous improvement

  • Time-to-fix — from a reported failure to a deployed fix. The lower, the more nimble your team.
  • Regression rate — how often a new model release introduces a regression. The lower, the better your eval coverage.
  • Eval coverage — what fraction of historical failures are still caught by your current eval suite?
  • Drift — distributional change in inputs and outputs over time.
  • Cost per token, cost per request — must be tracked alongside quality, since improvements often add cost.

Practical tips

  • Make the failure → eval → fix loop explicit. Have a system (or a spreadsheet, at minimum) that tracks every reported failure through to its eval case and its fix.
  • Invest in synthetic data tools. Generating preference pairs that target specific failure modes is a powerful capability that compounds over time.
  • Don't retrain from scratch each cycle. Targeted SFT or DPO updates on top of the previous best model is faster and safer.
  • Save every model. When something regresses, you'll want to bisect.
  • Hold a "what surprised us" review monthly. What did we learn that changes our roadmap?

Further reading

  • Christiano et al., "Deep Reinforcement Learning from Human Preferences", 2017. The grandparent of all modern feedback loops. arxiv.org/abs/1706.03741
  • Carlini et al., "Poisoning Web-Scale Training Datasets is Practical", 2023. Why naive feedback loops are dangerous. arxiv.org/abs/2302.10149
  • Anthropic — "Many-Shot Jailbreaking", 2024. Shows how adversarial inputs evolve.
  • OpenAI Model Spec — explicit policy and behavior documentation. openai.com/index/openai-model-spec
  • Anthropic Usage Policies / Model Card archive — example of how a vendor documents iterative changes.