What matters most¶
The full pipeline has hundreds of decisions. These are the ones that disproportionately determine outcomes.
Data quality beats data quantity after a point¶
Massive low-quality data is not enough. Filtering, deduplication, and mixture design are core capabilities.
A modern frontier-grade pretraining recipe spends as much engineering effort on data as on model architecture and training infrastructure combined. See Data cleaning & filtering.
Pretraining teaches knowledge; post-training teaches behavior¶
SFT and preference optimization make the model usable, but they do not replace broad pretraining.
If your model doesn't know facts X, Y, Z, you cannot SFT them in (cheaply). If your model knows them but answers awkwardly, SFT is exactly the right tool. See The big picture.
Tool use is its own capability¶
You need explicit data for tool selection, argument generation, error handling, and final response synthesis.
Models trained without dedicated tool data plateau on agentic tasks no matter how strong the base. See Tool use.
Long context is not just longer RoPE¶
A model can accept 128K tokens and still fail to use them well. Long-context data and evals matter.
Especially: needle-in-a-haystack is necessary, not sufficient. RULER, LongBench, and "lost in the middle" tests reveal the actual usable context. See Long-context.
Preference optimization is powerful but easy to overdo¶
Over-optimization can produce models that are polite, verbose, evasive, or reward-hacking.
Watch length, watch refusal rate, watch capability regressions on independent eval sets. β / KL settings matter more than people expect. See DPO and friends.
Safety needs multiple layers¶
Do not rely on one filter, one SFT set, or one reward model.
Defense in depth: pretraining filter + SFT refusal data + preference / RLHF + inference-time classifier + tool sandbox + monitoring. See Safety.
Evaluation must be adversarial and product-specific¶
Benchmarks are useful, but production failures often appear in messy real workflows.
Build a private eval set written by your team, never published. Run automated regressions on every checkpoint. Pair quantitative scores with qualitative spot checks. See Evaluation.
The simplest useful mental model¶
Pretraining builds the brain. Continued pretraining specializes the knowledge. SFT teaches the interface. Preference optimization teaches taste and judgment. Tool training teaches action. Safety training teaches boundaries. Evaluation and monitoring keep the system honest.
Most production problems trace back to attempting to solve a problem on the wrong layer of this stack.