A realistic end-to-end recipe¶
Here is a practical recipe for building a strong open-style LLM. Each stage links back to the detailed page on this wiki.
Stage A: Define target¶
Decide:
- Model size
- Context length
- Languages
- Domains
- License constraints
- Safety bar
- Serving budget
- Hardware
- Intended users
- Tool requirements
The single biggest mistake at this stage is not deciding. A diffuse goal ("a generally good chat model") produces a diffuse model. Pick a specific intended user and use case, even if you broaden later.
Stage B: Build tokenizer¶
See Tokenizer training.
Train on a representative sample:
- Web
- Code
- Math
- Multilingual
- Chat templates
- Tool-call JSON
- Domain data
Validate:
- Compression ratio
- Non-English efficiency
- Code tokenization
- Special-token behavior
- No accidental splitting of control tokens
Reserve special-token slots for tool calls, role markers, vision placeholders, and safety/control tokens before pretraining.
Stage C: Build pretraining corpus¶
See Data collection and cleaning & filtering.
Create:
- Raw data lake (versioned, immutable)
- Cleaned corpus
- Deduplicated corpus
- Filtered high-quality corpus
- Benchmark decontamination report
- Data mixture configs
- Data versioning
The data lake is the source of truth; everything downstream is reproducible from it.
Stage D: Train small models first¶
Before training a huge model, train:
- 100M
- 500M
- 1B
- 3B
Use these to test:
- Data mixture
- Loss curves
- Tokenizer
- Architecture
- Optimizer
- Scaling trends
- Eval harness
- Stability
The cost of these runs is tiny relative to the main run. The bugs they surface are not.
Stage E: Full pretraining¶
See Base pretraining.
Train base model.
Monitor:
- Training loss
- Validation loss by domain
- Gradient norms
- LR schedule
- Throughput
- Hardware failures
- Benchmark snapshots
- Toxicity snapshots
- Memorization checks
Save checkpoints frequently. Keep at least every-1000-steps checkpoints for the first few percent of training (where instabilities show up) and every-10000-steps checkpoints later.
Stage F: Mid-training¶
See Continued pretraining / mid-training.
Optional but common.
Train on:
- More code
- More math
- More high-quality documents
- More multilingual
- More domain data
- Long-form data
Use lower LR and careful replay. This is where most of the headline-eval improvement happens for a fixed compute budget.
Stage G: Long-context extension¶
If needed:
- Modify positional scaling (default: YaRN)
- Continue pretraining on long sequences
- Mix short and long examples
- Evaluate retrieval and reasoning across positions
- Tune serving stack for long KV cache
Stage H: SFT¶
Train on high-quality instruction data.
Mix:
- General chat
- Reasoning
- Code
- Multilingual
- Tool use
- Safety
- Structured output
- Domain tasks
Use assistant-token-only loss. Verify your loss mask with a unit test. 1–3 epochs at low LR.
Stage I: Preference optimization¶
See DPO and friends and RLHF/PPO.
Choose one:
- DPO for simplicity (most chat-alignment cases)
- PPO/RLHF for more flexible reward optimization
- RLVR for verifiable tasks (math, code, IFEval)
- GRPO for reasoning models
- RLAIF / Constitutional feedback for scalable safety
Iterate (2–4 rounds) when possible: regenerate preferences from the new policy, retrain.
Stage J: Specialized post-training¶
See the Specialization section.
Add targeted capability training:
- Function calling
- JSON schema adherence
- RAG grounding
- Agentic workflows
- Code execution
- Math verification
- Domain-specific behaviors
Stage K: Safety and red team¶
See Safety and Red teaming.
Run:
- Static safety evals
- Human red team
- Automated jailbreaks
- Tool security tests
- Multilingual safety tests
- Domain expert safety review
Translate every successful attack into a permanent eval case.
Stage L: Compression and deployment¶
See Compression and Serving.
Optimize:
- Quantization
- KV cache
- Batching
- Routing
- Speculative decoding
- Latency / cost tradeoffs
Re-evaluate the compressed model on the full eval suite — quantization quality regressions are real.
Stage M: Monitor and iterate¶
Create an ongoing loop:
Logs → failure mining → data → evals → training → deployment
This is the moat. Most of the gap between a good open-weights launch and a great deployed product is in the speed and quality of this cycle.
How long does this all take?¶
Rough order of magnitudes for a small (1B–7B) frontier-style open model with a small team:
| Stage | Time |
|---|---|
| Stages A–C: target, tokenizer, corpus | 2–6 months |
| Stage D: small-model recipe validation | 2–4 weeks |
| Stage E: full pretraining | 1–4 weeks of compute |
| Stage F: mid-training | 1–2 weeks |
| Stage G: long-context | 1–2 weeks |
| Stages H–I: SFT + preference | 4–8 weeks (data is the bottleneck) |
| Stage J: specialization | 4–12 weeks (parallel) |
| Stage K: safety + red team | 2–4 weeks (continuous) |
| Stage L: compression + deployment | 2–4 weeks |
| Stage M: continuous | Forever |
For a frontier-grade run, multiply most of these by 2–10×. Data collection, in particular, can dominate.