Skip to content

A realistic end-to-end recipe

Here is a practical recipe for building a strong open-style LLM. Each stage links back to the detailed page on this wiki.

Stage A: Define target

Decide:

  • Model size
  • Context length
  • Languages
  • Domains
  • License constraints
  • Safety bar
  • Serving budget
  • Hardware
  • Intended users
  • Tool requirements

The single biggest mistake at this stage is not deciding. A diffuse goal ("a generally good chat model") produces a diffuse model. Pick a specific intended user and use case, even if you broaden later.

Stage B: Build tokenizer

See Tokenizer training.

Train on a representative sample:

  • Web
  • Code
  • Math
  • Multilingual
  • Chat templates
  • Tool-call JSON
  • Domain data

Validate:

  • Compression ratio
  • Non-English efficiency
  • Code tokenization
  • Special-token behavior
  • No accidental splitting of control tokens

Reserve special-token slots for tool calls, role markers, vision placeholders, and safety/control tokens before pretraining.

Stage C: Build pretraining corpus

See Data collection and cleaning & filtering.

Create:

  • Raw data lake (versioned, immutable)
  • Cleaned corpus
  • Deduplicated corpus
  • Filtered high-quality corpus
  • Benchmark decontamination report
  • Data mixture configs
  • Data versioning

The data lake is the source of truth; everything downstream is reproducible from it.

Stage D: Train small models first

Before training a huge model, train:

  • 100M
  • 500M
  • 1B
  • 3B

Use these to test:

  • Data mixture
  • Loss curves
  • Tokenizer
  • Architecture
  • Optimizer
  • Scaling trends
  • Eval harness
  • Stability

The cost of these runs is tiny relative to the main run. The bugs they surface are not.

Stage E: Full pretraining

See Base pretraining.

Train base model.

Monitor:

  • Training loss
  • Validation loss by domain
  • Gradient norms
  • LR schedule
  • Throughput
  • Hardware failures
  • Benchmark snapshots
  • Toxicity snapshots
  • Memorization checks

Save checkpoints frequently. Keep at least every-1000-steps checkpoints for the first few percent of training (where instabilities show up) and every-10000-steps checkpoints later.

Stage F: Mid-training

See Continued pretraining / mid-training.

Optional but common.

Train on:

  • More code
  • More math
  • More high-quality documents
  • More multilingual
  • More domain data
  • Long-form data

Use lower LR and careful replay. This is where most of the headline-eval improvement happens for a fixed compute budget.

Stage G: Long-context extension

See Long-context extension.

If needed:

  • Modify positional scaling (default: YaRN)
  • Continue pretraining on long sequences
  • Mix short and long examples
  • Evaluate retrieval and reasoning across positions
  • Tune serving stack for long KV cache

Stage H: SFT

See Supervised fine-tuning.

Train on high-quality instruction data.

Mix:

  • General chat
  • Reasoning
  • Code
  • Multilingual
  • Tool use
  • Safety
  • Structured output
  • Domain tasks

Use assistant-token-only loss. Verify your loss mask with a unit test. 1–3 epochs at low LR.

Stage I: Preference optimization

See DPO and friends and RLHF/PPO.

Choose one:

  • DPO for simplicity (most chat-alignment cases)
  • PPO/RLHF for more flexible reward optimization
  • RLVR for verifiable tasks (math, code, IFEval)
  • GRPO for reasoning models
  • RLAIF / Constitutional feedback for scalable safety

Iterate (2–4 rounds) when possible: regenerate preferences from the new policy, retrain.

Stage J: Specialized post-training

See the Specialization section.

Add targeted capability training:

  • Function calling
  • JSON schema adherence
  • RAG grounding
  • Agentic workflows
  • Code execution
  • Math verification
  • Domain-specific behaviors

Stage K: Safety and red team

See Safety and Red teaming.

Run:

  • Static safety evals
  • Human red team
  • Automated jailbreaks
  • Tool security tests
  • Multilingual safety tests
  • Domain expert safety review

Translate every successful attack into a permanent eval case.

Stage L: Compression and deployment

See Compression and Serving.

Optimize:

  • Quantization
  • KV cache
  • Batching
  • Routing
  • Speculative decoding
  • Latency / cost tradeoffs

Re-evaluate the compressed model on the full eval suite — quantization quality regressions are real.

Stage M: Monitor and iterate

See Continuous improvement.

Create an ongoing loop:

Logs → failure mining → data → evals → training → deployment

This is the moat. Most of the gap between a good open-weights launch and a great deployed product is in the speed and quality of this cycle.

How long does this all take?

Rough order of magnitudes for a small (1B–7B) frontier-style open model with a small team:

Stage Time
Stages A–C: target, tokenizer, corpus 2–6 months
Stage D: small-model recipe validation 2–4 weeks
Stage E: full pretraining 1–4 weeks of compute
Stage F: mid-training 1–2 weeks
Stage G: long-context 1–2 weeks
Stages H–I: SFT + preference 4–8 weeks (data is the bottleneck)
Stage J: specialization 4–12 weeks (parallel)
Stage K: safety + red team 2–4 weeks (continuous)
Stage L: compression + deployment 2–4 weeks
Stage M: continuous Forever

For a frontier-grade run, multiply most of these by 2–10×. Data collection, in particular, can dominate.