Domain specialization¶
Goal¶
Adapt a general model to a domain without destroying general ability.
Domains:
- Medicine
- Law
- Finance
- Education
- Customer support
- Coding
- Clinical documentation
- Scientific research
- Enterprise knowledge work
Techniques¶
Continued pretraining¶
Best for domain language and knowledge. See Continued pretraining. Run on cleaned, license-cleared domain corpora — clinical notes (de-identified), legal documents, financial filings, etc. — typically with replay from the original mix.
SFT¶
Best for domain tasks and style. Domain experts curate or review instruction-response pairs. The volume can be modest (LIMA-style) if quality is high.
Preference optimization¶
Best for domain-specific quality preferences — what counts as a good clinical note, what counts as a sound legal argument, etc. Expert annotators are essential here.
RAG¶
Best for changing private knowledge — internal documentation, frequently updated regulations, customer-specific context. See RAG. RAG plus a moderately specialized model often outperforms a heavily fine-tuned model with stale parametric knowledge.
Adapters / LoRA¶
Useful when compute is limited or many domain variants are needed. Train a small adapter (a few hundred MB) per domain or customer instead of full fine-tuning.
Tool integration¶
Useful when correctness depends on external systems — calculators, simulators, EHR APIs, document search, accounting software.
How to choose¶
| Need | Best technique |
|---|---|
| The model doesn't know the vocabulary | Continued pretraining |
| The model knows the words, doesn't follow the format | SFT |
| The model produces plausible but low-taste answers | Preference optimization |
| The knowledge changes weekly | RAG |
| You need 50 customer-specific variants | LoRA / adapters |
| Correctness needs an external system | Tool integration |
In practice, production systems combine most of these — a base model continued-pretrained on the domain, SFT'd on domain tasks, preference-tuned for taste, then deployed with RAG over fresh content and tools for system access.
Domain data concerns¶
For clinical or legal models, quality and provenance matter enormously.
Important:
- Licensed data — confirm you have rights for training. Clinical notes are usually covered by HIPAA / data-use agreements; legal documents by court records; financial filings by SEC publication. Consult counsel.
- Expert review — domain experts review SFT data, not just labelers.
- De-identification — automated PII removal + manual spot-check.
- Source dates — out-of-date guidance is dangerous. Track when each document was authored.
- Jurisdiction — legal and medical advice varies by jurisdiction; the model should know which it's giving.
- Guideline versions — which version of the standard / guideline / DSM / ICD was the data based on?
- Uncertainty handling — domain models that confidently answer when they shouldn't are worse than general models.
- Escalation behavior — when should the model hand off to a human professional?
- Clear non-replacement framing — "This is not medical advice; consult a clinician."
Practical tips¶
- Don't replace, augment. A medical model that wraps a clinician's workflow is far more useful (and safer) than one that "replaces the doctor."
- Build the eval first. Have domain experts hand-label 200–500 evaluation cases before you start training. You will use this set for years.
- Test general capabilities continuously. Every domain-specialized model checkpoint should be evaluated on broad benchmarks (MMLU, IFEval, safety) — not just the domain ones.
- Run domain red teams. Domain experts (oncologists, securities lawyers) find failure modes that generic red teamers miss.
- Document the data. A reproducible record of which documents trained the model is essential for legal review and incident response.
Further reading¶
- Med-PaLM 2 — Singhal et al., 2023. Reference for clinical LLM evaluation. arxiv.org/abs/2305.09617
- BloombergGPT, 2023. A large finance-domain LLM. arxiv.org/abs/2303.17564
- Code Llama, 2023. arxiv.org/abs/2308.12950
- DeepSeek-Coder-V2, 2024. State-of-the-art open code model with detailed report. arxiv.org/abs/2406.11931
- PMC-LLaMA — biomedical specialization. arxiv.org/abs/2304.14454
- Lawyer LLaMA, 2023. arxiv.org/abs/2305.15062
- LegalBench — Guha et al., 2023. Legal-reasoning evaluation. arxiv.org/abs/2308.11462
- MedHELM — comprehensive medical eval. crfm.stanford.edu/helm/medical