Skip to content

Continued pretraining / mid-training

Goal

Improve the pretrained base model in targeted domains before instruction tuning.

This is still usually next-token prediction, not chat training.

Use cases:

  • Code specialization
  • Medical/domain adaptation
  • Legal adaptation
  • Finance adaptation
  • Math/science improvement
  • Multilingual improvement
  • Long-context adaptation
  • Data refresh after base pretraining
  • Improving style or quality distribution

Why not jump straight to SFT?

SFT teaches behavior. Continued pretraining teaches distributional knowledge and fluency.

For example, if you want a biomedical model, SFT alone may teach it to answer politely, but continued pretraining on biomedical literature can improve its internal domain representations — vocabulary, factual associations, reasoning patterns specific to the field. SFT can then turn those improved representations into helpful behavior.

A useful rule of thumb:

  • If your model doesn't know the words, continued pretraining.
  • If your model knows the words but won't use them helpfully, SFT.
  • If your model uses them but with the wrong taste, preference optimization.

How it differs from base pretraining

Aspect Base pretraining Continued pretraining
Starting point Random init Trained base model
Data Broad, diverse Targeted, often narrower
Token budget Trillions Billions to tens of billions
Learning rate Peak ~1e-4 to 1e-3 Lower; ~10–30% of base peak
Schedule Long cosine + cooldown Short, often with replay
Risk Underfitting Catastrophic forgetting

Risks

Catastrophic forgetting

Too much narrow-domain training can degrade general ability.

Mitigation:

  • Mix in general data (commonly 5–30% of the original pretraining mixture as "replay").
  • Use low learning rate.
  • Train for limited tokens.
  • Monitor broad evals, not just domain evals.
  • Use replay data — actual samples from the original pretraining mix.

Overfitting to style

A model continued-pretrained on narrow text may imitate that style too strongly.

Example: a model continued-pretrained on legal contracts may then over-use legalese in casual chat. Mitigate with style-diverse replay and explicit SFT later.

Benchmark contamination

Domain corpora (especially synthetic ones) may contain benchmark answers. Re-decontaminate against your eval suite before mid-training, even if you decontaminated the original corpus.

Safety regression

Domain data may contain unsafe procedural instructions (medical, biosec, security). Domain experts often draft content assuming a controlled audience. After continued pretraining, rerun safety evals before declaring the model done.

Mid-training as a distinct stage

OLMo 2 and others now treat mid-training as a named stage between base pretraining and SFT. The pattern:

  1. Finish base pretraining on a broad mix.
  2. Define a "mid-training" mix with explicit upweighting of code, math, instruction-like text, long-form documents, and any domain corpora.
  3. Resume training with a fresh learning-rate decay schedule — typically a short cosine or linear decay over a few percent of the original token budget.
  4. Save the mid-trained checkpoint as the input to SFT.

This is essentially a structured, named version of the "annealing" / "cooldown" trick — see Base pretraining.

Practical tips

  • Always include replay. Even 5% of the original pretraining mix dramatically reduces forgetting.
  • Lower the learning rate. Continuing at the base peak LR is the most common cause of catastrophic forgetting.
  • Run general evals every few hundred steps, not just domain evals. Watch for regression.
  • Decontaminate again. Domain corpora often leak benchmarks.
  • Save the pre-mid-training checkpoint. If mid-training regresses something important, you'll want to roll back and try a different mix.

Further reading