Open LLM Training Wiki¶
A practitioner-oriented knowledge base for training modern large language models — from raw text to deployed assistant.
The exact recipes used by frontier labs are proprietary, but the broad pipeline is now visible from open technical reports — Llama 3, OLMo 2, Tülu 3, InstructGPT, DPO, Toolformer, YaRN, Constitutional AI, FlashAttention, Chinchilla, and others. This wiki distills that pipeline into one navigable place.
Each section in the source guide becomes its own page here, with the full text of the guide plus curated further-reading pointers.
How a modern LLM is built¶
A modern LLM is not "trained once." It is built through a sequence of stages, each with its own data, objective, failure modes, and evaluations:
-
1. Data collection & filtering
Web, books, code, math, multilingual, synthetic. The single most important asset.
-
Compress text efficiently across languages, code, and markup.
-
Decoder-only Transformer with RoPE, GQA, RMSNorm, SwiGLU. Dense or MoE.
-
Next-token prediction over trillions of tokens. Teaches language and the world.
-
Targeted domain adaptation before instruction tuning.
-
RoPE scaling, YaRN, FlashAttention, training data that actually uses the window.
-
Turn a raw text predictor into an assistant. Teaches behavior and format.
-
Comparisons that tell the model which response is better.
-
Learn a scalar score over responses for downstream RL.
-
Optimize the SFT policy against a reward model with KL anchoring.
-
Direct preference optimization — simpler, more stable than full PPO.
-
Math, code, planning. Verifiable rewards. Process vs. outcome supervision.
-
13. Tool use & function calling
When to call, what to call, how to handle errors and recover.
-
14. Retrieval-augmented generation
Grounded, cited, abstaining when evidence is insufficient.
-
Reliable JSON / SQL / typed-object generation with schema constraints.
-
Refusal quality, multi-layer policy, Constitutional AI.
-
Vision encoder + projector + multimodal SFT.
-
Medicine, law, finance, code — without destroying general ability.
-
Quantization, distillation, speculative decoding, LoRA.
-
Benchmarks, human eval, LLM-as-judge, calibration.
-
Find failures before users do.
-
KV cache, batching, routing, observability.
-
Logs → failures → data → evals → retrain.
-
A realistic end-to-end recipe and per-stage checklists.
The simplest useful mental model¶
Pretraining builds the brain. Continued pretraining specializes the knowledge. SFT teaches the interface. Preference optimization teaches taste and judgment. Tool training teaches action. Safety training teaches boundaries. Evaluation and monitoring keep the system honest.
Where to start reading¶
If you have 15 minutes: read The big picture and What matters most.
If you have an afternoon: walk the pipeline top to bottom — start at Data collection, follow the nav.
If you are building right now: jump to The end-to-end recipe and the matching checklists.
If you want papers and repos: see the consolidated Resources & reading page.
Conventions¶
- Goal: what each stage is trying to achieve, in one paragraph.
- Data / Methods: concrete formats, mixtures, and recipes.
- Failure modes: what breaks in practice and how it shows up.
- Further reading: a short, curated list of papers, repos, and posts. Not exhaustive — opinionated.
Inspired by Andrej Karpathy's knowledge base, the Chinchilla scaling laws, the Llama 3 paper, and the open-weights ecosystem.