Specialization¶
After SFT and preference optimization, most production-grade models go through one or more specialized post-training stages that target specific capabilities:
- Reasoning — math, code, planning, scientific reasoning. Increasingly trained with verifiable rewards and process supervision.
- Tool use & function calling — when to call, what to call, how to handle errors and combine results.
- Retrieval-augmented generation — using retrieved evidence, citing sources, abstaining.
- Structured output — reliable JSON, SQL, typed objects, schema adherence.
- Safety training — refusal quality, multi-layer policy, Constitutional AI.
- Multimodal extension — vision, audio, document understanding.
- Domain specialization — medicine, law, finance, code.
These stages can be done in parallel, sequentially, or interleaved. The choice depends on what tradeoffs you're willing to make and how much you can afford to evaluate.
How specializations interact¶
Specialization rarely stays clean. Common interactions:
- Reasoning training improves tool use. Models that learn to think step by step also call tools more correctly.
- Tool use can hurt pure reasoning. Models trained to reach for a calculator may over-rely on it for tasks they could do internally.
- RAG training can hurt closed-book QA. A model trained to "consult evidence" may abstain more than is helpful when no evidence is provided.
- Safety training can hurt helpfulness. Over-refusal is the canonical failure.
- Domain training can hurt general behavior. Continued pretraining on legal text can make casual chat sound like a contract.
The fix is always the same: broad evaluation + balanced data mixes + iterative refinement.