Supervised fine-tuning / instruction tuning¶
Goal¶
Teach the base model to behave like an assistant.
SFT converts a raw text predictor into a model that follows instructions, answers directly, respects roles, formats outputs, handles multi-turn chat, and refuses some unsafe requests.
InstructGPT used supervised fine-tuning as the first stage before reward modeling and PPO-based RLHF. Most modern recipes still start the same way.
Data format¶
Typical chat example:
{
"messages": [
{
"role": "system",
"content": "You are a helpful, honest, and concise assistant."
},
{
"role": "user",
"content": "Explain photosynthesis in simple terms."
},
{
"role": "assistant",
"content": "Photosynthesis is how plants use sunlight..."
}
]
}
Tool-use example:
{
"messages": [
{
"role": "system",
"content": "Use tools when needed."
},
{
"role": "user",
"content": "What is the weather in Paris?"
},
{
"role": "assistant",
"tool_calls": [
{
"name": "get_weather",
"arguments": { "location": "Paris, France" }
}
]
},
{
"role": "tool",
"name": "get_weather",
"content": "{\"temperature_c\": 12, \"condition\": \"rain\"}"
},
{
"role": "assistant",
"content": "It is 12°C and rainy in Paris."
}
]
}
Loss masking¶
Usually compute loss only on assistant tokens, not user/system tokens.
Why?
- The model should learn to produce assistant responses.
- It should condition on user messages, not imitate them.
- It should not learn to generate system prompts.
For tool-call training, you may compute loss on:
- Assistant natural-language tokens
- Tool-call JSON tokens
- Function name
- Arguments
But not usually on tool-result tokens — those come from the environment, not the model.
Loss-masking bugs are silent and devastating
A common bug: training loss looks fine, but the assistant-token mask is off-by-one or excludes the EOS, so the model never learns to stop. Inspect a few samples after tokenization and confirm exactly which positions contribute to the loss.
Instruction-tuning data categories¶
A strong SFT mix usually includes many categories.
General instruction following¶
Examples: explain, summarize, rewrite, compare, plan, classify, extract, brainstorm.
Purpose: basic helpfulness, natural assistant tone, task compliance.
Reasoning and math¶
Examples: GSM8K-like grade-school math, MATH-style competition problems, logic puzzles, word problems, scientific reasoning, multi-step planning.
Purpose: better decomposition, more robust intermediate reasoning, better arithmetic and symbolic manipulation.
Code¶
Examples: generate function, explain code, debug, write tests, refactor, translate between languages, use libraries, multi-file reasoning.
Purpose: programming skill, structured output, exact following of constraints.
Knowledge QA¶
Examples: factual QA, open-book QA, domain-specific QA, citation-aware QA, "I don't know" examples.
Purpose: direct answer style, calibrated factuality, reduced hallucination.
Multi-turn conversation¶
Examples: follow-up questions, corrections, user changes constraints, memory within conversation, clarification, refusal after earlier benign context, tool-result incorporation.
Purpose: dialogue coherence, state tracking, avoiding one-shot behavior.
Style control¶
Examples: "explain like I'm five", "use academic tone", "write as a checklist", "be concise", "use JSON only", "no markdown".
Purpose: controllability, format adherence.
Function calling and tools¶
Examples: API call selection, argument generation, no-tool-needed examples, multiple tools, tool errors, tool-result synthesis, schema adherence.
Purpose: reliable agent behavior. See Tool use for more.
Safety and refusal¶
Examples: disallowed requests, allowed alternatives, medical/legal/financial caution, self-harm support, privacy-preserving responses, dual-use boundary cases.
Purpose: safe behavior before RLHF/DPO. See Safety.
Example SFT data mix¶
There is no universal recipe, but a practical mix might look like:
| Category | Rough share |
|---|---|
| General instructions | 25–40% |
| Reasoning / math | 10–20% |
| Code | 10–25% |
| Knowledge QA | 10–20% |
| Multi-turn chat | 5–15% |
| Tool / function calling | 5–15% |
| Structured output | 5–10% |
| Safety / refusal | 5–15% |
| Domain-specific data | depends |
The exact balance depends on the target model:
- Coding assistant: much more code, tests, repo-level tasks.
- Medical assistant: more clinical QA, uncertainty, citations, safety.
- Agentic model: more tool trajectories and environment feedback.
- Consumer chatbot: more general chat, safety, style, multilingual.
Quality matters more than size¶
LIMA showed surprisingly strong behavior from only 1,000 carefully curated examples on top of a strong base model, suggesting that SFT quality and diversity can matter more than sheer volume.
That does not mean 1,000 examples is enough for production. It means SFT is often about teaching interaction style and format, not injecting all knowledge.
Read your SFT data
Sample 100 examples uniformly from your SFT dataset and read every one. You will find: bad answers labeled good, off-policy refusals, length-biased examples, tool calls with wrong schemas, verbose disclaimers, and accidental benchmark contamination. Fixing 5% of your data is usually worth more than 5× the volume.
SFT failure modes¶
Over-refusal¶
The model refuses benign requests. Often caused by an over-aggressive safety mix or noisy refusal labels.
Sycophancy¶
The model agrees with false user claims. Often caused by training data where the assistant was rewarded for agreement.
Verbosity drift¶
The model becomes long-winded. Often caused by length bias in synthetic data.
Format brittleness¶
The model fails strict JSON or schema constraints. See Structured output.
Capability regression¶
Too much narrow SFT can reduce base reasoning or coding skill. Mix in general data and run broad evals.
Assistant persona overfitting¶
The model imitates the tone of the dataset too strongly (e.g., signature openings, formulaic disclaimers).
Low diversity¶
The model handles common instruction phrasings but fails unusual ones. Diversity in prompts matters as much as diversity in answers.
Open instruction datasets worth knowing¶
- OpenAssistant / OASST1 — early open conversational dataset.
- Tülu 3 SFT mix — high-quality, well-documented mixture from AI2. huggingface.co/datasets/allenai/tulu-3-sft-mixture
- No Robots — Stanford. Small, hand-written, very high quality.
- UltraChat — large synthetic conversational corpus.
- Magpie — self-generated SFT data via instruction-tuned models. arxiv.org/abs/2406.08464
- ShareGPT-style logs — user/assistant logs (license varies; check before use).
- WildChat — real user-assistant conversations, opt-in. arxiv.org/abs/2405.01470
Practical tips¶
- Pick the chat template before generating data. Changing it later means re-rendering and possibly retraining.
- Mask correctly and verify with a unit test. A loss-mask test that asserts which positions contribute is worth its weight in gold.
- Track per-category eval, not just aggregate. A regression in code quality is invisible in average eval scores.
- Don't over-train. SFT is fast to overfit. Most recipes do 1–3 epochs at low LR (~1e-5 to 5e-6).
- Keep a held-out SFT eval that mirrors your data mix. Loss on this set is a better leading indicator than benchmark scores.
Further reading¶
- InstructGPT — Ouyang et al., 2022. arxiv.org/abs/2203.02155
- LIMA — Zhou et al., 2023. arxiv.org/abs/2305.11206
- Tülu 3 — AI2, 2024. The clearest open recipe for modern post-training. arxiv.org/abs/2411.15124
- Self-Instruct — Wang et al., 2022. Synthetic instruction generation. arxiv.org/abs/2212.10560
- Magpie — Xu et al., 2024. arxiv.org/abs/2406.08464
- WildChat — Zhao et al., 2024. arxiv.org/abs/2405.01470
- HuggingFace
trl—SFTTraineris the de-facto reference implementation. github.com/huggingface/trl axolotl— popular open SFT/DPO training framework. github.com/axolotl-ai-cloud/axolotlopen-instruct— AI2's reference Tülu trainer. github.com/allenai/open-instruct