Pretraining¶
Pretraining is where the model learns language and the world. It is the most expensive phase, and the one where decisions about data and scale matter most.
This section is split into three pages:
- Base pretraining — next-token prediction at scale. Scaling laws, data mixtures, training mechanics, parallelism, stability.
- Continued pretraining / mid-training — targeted adaptation of an already-pretrained model. Domain, data refresh, quality upweighting.
- Long-context extension — RoPE scaling, YaRN, training data that actually uses the long window.
How to think about pretraining¶
Pretraining produces the base model. Base models are not chatty — they are next-token predictors. Without SFT, prompting them is awkward (you have to write completions, not questions).
Pretraining is what teaches:
- Grammar and vocabulary
- World knowledge
- Style and register
- Latent reasoning patterns
- Most of what the model "knows"
Post-training is what teaches:
- How to behave as an assistant
- Format adherence
- Refusal style
- Tool-calling protocols
- Calibration ("I don't know")
This split is the most important mental model in modern LLM development. Most quality problems trace back to which side of this split they belong on.