Continued pretraining / mid-training¶
Goal¶
Improve the pretrained base model in targeted domains before instruction tuning.
This is still usually next-token prediction, not chat training.
Use cases:
- Code specialization
- Medical/domain adaptation
- Legal adaptation
- Finance adaptation
- Math/science improvement
- Multilingual improvement
- Long-context adaptation
- Data refresh after base pretraining
- Improving style or quality distribution
Why not jump straight to SFT?¶
SFT teaches behavior. Continued pretraining teaches distributional knowledge and fluency.
For example, if you want a biomedical model, SFT alone may teach it to answer politely, but continued pretraining on biomedical literature can improve its internal domain representations — vocabulary, factual associations, reasoning patterns specific to the field. SFT can then turn those improved representations into helpful behavior.
A useful rule of thumb:
- If your model doesn't know the words, continued pretraining.
- If your model knows the words but won't use them helpfully, SFT.
- If your model uses them but with the wrong taste, preference optimization.
How it differs from base pretraining¶
| Aspect | Base pretraining | Continued pretraining |
|---|---|---|
| Starting point | Random init | Trained base model |
| Data | Broad, diverse | Targeted, often narrower |
| Token budget | Trillions | Billions to tens of billions |
| Learning rate | Peak ~1e-4 to 1e-3 | Lower; ~10–30% of base peak |
| Schedule | Long cosine + cooldown | Short, often with replay |
| Risk | Underfitting | Catastrophic forgetting |
Risks¶
Catastrophic forgetting¶
Too much narrow-domain training can degrade general ability.
Mitigation:
- Mix in general data (commonly 5–30% of the original pretraining mixture as "replay").
- Use low learning rate.
- Train for limited tokens.
- Monitor broad evals, not just domain evals.
- Use replay data — actual samples from the original pretraining mix.
Overfitting to style¶
A model continued-pretrained on narrow text may imitate that style too strongly.
Example: a model continued-pretrained on legal contracts may then over-use legalese in casual chat. Mitigate with style-diverse replay and explicit SFT later.
Benchmark contamination¶
Domain corpora (especially synthetic ones) may contain benchmark answers. Re-decontaminate against your eval suite before mid-training, even if you decontaminated the original corpus.
Safety regression¶
Domain data may contain unsafe procedural instructions (medical, biosec, security). Domain experts often draft content assuming a controlled audience. After continued pretraining, rerun safety evals before declaring the model done.
Mid-training as a distinct stage¶
OLMo 2 and others now treat mid-training as a named stage between base pretraining and SFT. The pattern:
- Finish base pretraining on a broad mix.
- Define a "mid-training" mix with explicit upweighting of code, math, instruction-like text, long-form documents, and any domain corpora.
- Resume training with a fresh learning-rate decay schedule — typically a short cosine or linear decay over a few percent of the original token budget.
- Save the mid-trained checkpoint as the input to SFT.
This is essentially a structured, named version of the "annealing" / "cooldown" trick — see Base pretraining.
Practical tips¶
- Always include replay. Even 5% of the original pretraining mix dramatically reduces forgetting.
- Lower the learning rate. Continuing at the base peak LR is the most common cause of catastrophic forgetting.
- Run general evals every few hundred steps, not just domain evals. Watch for regression.
- Decontaminate again. Domain corpora often leak benchmarks.
- Save the pre-mid-training checkpoint. If mid-training regresses something important, you'll want to roll back and try a different mix.
Further reading¶
- Ke et al., "Continual Pre-training of Language Models", 2022. arxiv.org/abs/2302.03241
- Gupta et al., "Continual Pre-Training of Large Language Models: How to (re)warm your model?", 2023. arxiv.org/abs/2308.04014
- Ibrahim et al., "Simple and Scalable Strategies to Continually Pre-train Large Language Models", 2024. arxiv.org/abs/2403.08763
- Code Llama paper, 2023. Continued pretraining of Llama 2 on code. arxiv.org/abs/2308.12950
- Med-PaLM 2 — domain-adapted medical model. arxiv.org/abs/2305.09617
- OLMo 2 paper, 2024 — formalizes mid-training. arxiv.org/abs/2501.00656
- MiniCPM technical report, 2024 — explicit annealing/cooldown phase. arxiv.org/abs/2404.06395