The big picture¶
A modern LLM is not "trained once." It is built through a sequence of stages, each with its own data, objective, evaluation, and failure modes. Understanding the shape of the pipeline before diving into any single stage is the most useful thing you can do.
The full pipeline¶
- Data collection and filtering — Curate a massive, diverse, legal corpus.
- Tokenizer training — Learn an efficient vocabulary across languages, code, math, and markup.
- Base model pretraining — Predict the next token over trillions of tokens.
- Continued pretraining / mid-training — Targeted domain or quality adaptation.
- Long-context extension — From 4K to 128K+ tokens.
- Supervised fine-tuning — Turn the raw predictor into an assistant.
- Preference optimization — RLHF, DPO, or RLVR.
- Reasoning, tool use, function calling — Specialized post-training capabilities.
- Safety post-training — Refusal quality and multi-layer policy.
- Compression and serving — Quantization, distillation, KV-cache.
- Evaluation, red teaming, deployment, monitoring — Continuous quality control.
- Continuous improvement loop — Logs → failures → data → retrain.
The core mental model¶
Think of the base model as learning the world and language. Post-training teaches it how to behave.
This split matters because it predicts where to put effort:
- Need more knowledge about a topic? → pretraining or continued pretraining.
- Need different behavior (refusal style, format, tone)? → SFT and preference optimization.
- Need action (tool calling, agents)? → tool-trajectory data and post-training.
- Need calibration (knowing when not to know)? → preference data, RLHF, careful eval.
- Need safety? → multi-layer: data filtering + SFT + preference + inference policy + monitoring.
What changed in the last few years¶
The pipeline above is now widely understood, but the emphasis keeps shifting:
- Pretraining data quality beats raw quantity — see Chinchilla, then RefinedWeb, then C4-style filtering, then quality classifiers, then synthetic-data-augmented mixes.
- SFT data quality and diversity beats SFT data volume — see LIMA (1,000 carefully curated examples).
- Preference optimization is moving offline — DPO, IPO, KTO, ORPO, SimPO, CPO replaced or supplemented PPO at many shops because they are simpler and more stable.
- Verifiable rewards (math checkers, unit tests, exact-match) increasingly drive reasoning post-training — see Tülu 3 RLVR.
- Long context is now table stakes — RoPE scaling + YaRN + FlashAttention made 128K-context models routine.
- Tool use is a first-class capability — not a prompt-engineering trick.
- Safety is a multi-stage, multi-layer concern — not a single filter.
What hasn't changed¶
- Decoder-only Transformer with next-token prediction is still the dominant architecture.
- Cross-entropy loss is still the loss.
- Adam-family optimizers still win in practice.
- Scaling laws still bind: model size and tokens scale together.
- Evaluation is still hard, contamination is still everywhere, and human eval is still essential.
Reading the rest of the wiki¶
Each chapter is self-contained. You can read top to bottom, or jump to whichever stage you are working on. Cross-references are dense — follow the links.
If you are pressed for time, the canonical "I want the whole picture" path is:
- The big picture (this page)
- Data collection
- Base pretraining
- Supervised fine-tuning
- DPO and friends
- Tool use
- Evaluation
- What matters most
Further reading¶
- Llama 3 technical report — Meta, 2024. The most thorough single-document description of a modern frontier-grade open recipe. arxiv.org/abs/2407.21783
- OLMo 2 — AI2, 2024. Fully open weights, data, and recipe. arxiv.org/abs/2501.00656
- Tülu 3 — AI2, 2024. Open post-training recipe with RLVR. arxiv.org/abs/2411.15124
- InstructGPT — Ouyang et al., 2022. The original SFT → RM → PPO recipe. arxiv.org/abs/2203.02155
- Chinchilla scaling laws — Hoffmann et al., 2022. arxiv.org/abs/2203.15556
- Andrej Karpathy — Intro to LLMs (1-hour talk) — A free, exceptional overview. youtube.com/watch?v=zjkBMFhNj_g
- Andrej Karpathy — Let's build GPT, from scratch — Hands-on companion. youtube.com/watch?v=kCc8FmEb1nY