Skip to content

What matters most

The full pipeline has hundreds of decisions. These are the ones that disproportionately determine outcomes.

Data quality beats data quantity after a point

Massive low-quality data is not enough. Filtering, deduplication, and mixture design are core capabilities.

A modern frontier-grade pretraining recipe spends as much engineering effort on data as on model architecture and training infrastructure combined. See Data cleaning & filtering.

Pretraining teaches knowledge; post-training teaches behavior

SFT and preference optimization make the model usable, but they do not replace broad pretraining.

If your model doesn't know facts X, Y, Z, you cannot SFT them in (cheaply). If your model knows them but answers awkwardly, SFT is exactly the right tool. See The big picture.

Tool use is its own capability

You need explicit data for tool selection, argument generation, error handling, and final response synthesis.

Models trained without dedicated tool data plateau on agentic tasks no matter how strong the base. See Tool use.

Long context is not just longer RoPE

A model can accept 128K tokens and still fail to use them well. Long-context data and evals matter.

Especially: needle-in-a-haystack is necessary, not sufficient. RULER, LongBench, and "lost in the middle" tests reveal the actual usable context. See Long-context.

Preference optimization is powerful but easy to overdo

Over-optimization can produce models that are polite, verbose, evasive, or reward-hacking.

Watch length, watch refusal rate, watch capability regressions on independent eval sets. β / KL settings matter more than people expect. See DPO and friends.

Safety needs multiple layers

Do not rely on one filter, one SFT set, or one reward model.

Defense in depth: pretraining filter + SFT refusal data + preference / RLHF + inference-time classifier + tool sandbox + monitoring. See Safety.

Evaluation must be adversarial and product-specific

Benchmarks are useful, but production failures often appear in messy real workflows.

Build a private eval set written by your team, never published. Run automated regressions on every checkpoint. Pair quantitative scores with qualitative spot checks. See Evaluation.

The simplest useful mental model

Pretraining builds the brain. Continued pretraining specializes the knowledge. SFT teaches the interface. Preference optimization teaches taste and judgment. Tool training teaches action. Safety training teaches boundaries. Evaluation and monitoring keep the system honest.

Most production problems trace back to attempting to solve a problem on the wrong layer of this stack.