Skip to content

Supervised fine-tuning / instruction tuning

Goal

Teach the base model to behave like an assistant.

SFT converts a raw text predictor into a model that follows instructions, answers directly, respects roles, formats outputs, handles multi-turn chat, and refuses some unsafe requests.

InstructGPT used supervised fine-tuning as the first stage before reward modeling and PPO-based RLHF. Most modern recipes still start the same way.

Data format

Typical chat example:

{
  "messages": [
    {
      "role": "system",
      "content": "You are a helpful, honest, and concise assistant."
    },
    {
      "role": "user",
      "content": "Explain photosynthesis in simple terms."
    },
    {
      "role": "assistant",
      "content": "Photosynthesis is how plants use sunlight..."
    }
  ]
}

Tool-use example:

{
  "messages": [
    {
      "role": "system",
      "content": "Use tools when needed."
    },
    {
      "role": "user",
      "content": "What is the weather in Paris?"
    },
    {
      "role": "assistant",
      "tool_calls": [
        {
          "name": "get_weather",
          "arguments": { "location": "Paris, France" }
        }
      ]
    },
    {
      "role": "tool",
      "name": "get_weather",
      "content": "{\"temperature_c\": 12, \"condition\": \"rain\"}"
    },
    {
      "role": "assistant",
      "content": "It is 12°C and rainy in Paris."
    }
  ]
}

Loss masking

Usually compute loss only on assistant tokens, not user/system tokens.

Why?

  • The model should learn to produce assistant responses.
  • It should condition on user messages, not imitate them.
  • It should not learn to generate system prompts.

For tool-call training, you may compute loss on:

  • Assistant natural-language tokens
  • Tool-call JSON tokens
  • Function name
  • Arguments

But not usually on tool-result tokens — those come from the environment, not the model.

Loss-masking bugs are silent and devastating

A common bug: training loss looks fine, but the assistant-token mask is off-by-one or excludes the EOS, so the model never learns to stop. Inspect a few samples after tokenization and confirm exactly which positions contribute to the loss.

Instruction-tuning data categories

A strong SFT mix usually includes many categories.

General instruction following

Examples: explain, summarize, rewrite, compare, plan, classify, extract, brainstorm.

Purpose: basic helpfulness, natural assistant tone, task compliance.

Reasoning and math

Examples: GSM8K-like grade-school math, MATH-style competition problems, logic puzzles, word problems, scientific reasoning, multi-step planning.

Purpose: better decomposition, more robust intermediate reasoning, better arithmetic and symbolic manipulation.

Code

Examples: generate function, explain code, debug, write tests, refactor, translate between languages, use libraries, multi-file reasoning.

Purpose: programming skill, structured output, exact following of constraints.

Knowledge QA

Examples: factual QA, open-book QA, domain-specific QA, citation-aware QA, "I don't know" examples.

Purpose: direct answer style, calibrated factuality, reduced hallucination.

Multi-turn conversation

Examples: follow-up questions, corrections, user changes constraints, memory within conversation, clarification, refusal after earlier benign context, tool-result incorporation.

Purpose: dialogue coherence, state tracking, avoiding one-shot behavior.

Style control

Examples: "explain like I'm five", "use academic tone", "write as a checklist", "be concise", "use JSON only", "no markdown".

Purpose: controllability, format adherence.

Function calling and tools

Examples: API call selection, argument generation, no-tool-needed examples, multiple tools, tool errors, tool-result synthesis, schema adherence.

Purpose: reliable agent behavior. See Tool use for more.

Safety and refusal

Examples: disallowed requests, allowed alternatives, medical/legal/financial caution, self-harm support, privacy-preserving responses, dual-use boundary cases.

Purpose: safe behavior before RLHF/DPO. See Safety.

Example SFT data mix

There is no universal recipe, but a practical mix might look like:

Category Rough share
General instructions 25–40%
Reasoning / math 10–20%
Code 10–25%
Knowledge QA 10–20%
Multi-turn chat 5–15%
Tool / function calling 5–15%
Structured output 5–10%
Safety / refusal 5–15%
Domain-specific data depends

The exact balance depends on the target model:

  • Coding assistant: much more code, tests, repo-level tasks.
  • Medical assistant: more clinical QA, uncertainty, citations, safety.
  • Agentic model: more tool trajectories and environment feedback.
  • Consumer chatbot: more general chat, safety, style, multilingual.

Quality matters more than size

LIMA showed surprisingly strong behavior from only 1,000 carefully curated examples on top of a strong base model, suggesting that SFT quality and diversity can matter more than sheer volume.

That does not mean 1,000 examples is enough for production. It means SFT is often about teaching interaction style and format, not injecting all knowledge.

Read your SFT data

Sample 100 examples uniformly from your SFT dataset and read every one. You will find: bad answers labeled good, off-policy refusals, length-biased examples, tool calls with wrong schemas, verbose disclaimers, and accidental benchmark contamination. Fixing 5% of your data is usually worth more than 5× the volume.

SFT failure modes

Over-refusal

The model refuses benign requests. Often caused by an over-aggressive safety mix or noisy refusal labels.

Sycophancy

The model agrees with false user claims. Often caused by training data where the assistant was rewarded for agreement.

Verbosity drift

The model becomes long-winded. Often caused by length bias in synthetic data.

Format brittleness

The model fails strict JSON or schema constraints. See Structured output.

Capability regression

Too much narrow SFT can reduce base reasoning or coding skill. Mix in general data and run broad evals.

Assistant persona overfitting

The model imitates the tone of the dataset too strongly (e.g., signature openings, formulaic disclaimers).

Low diversity

The model handles common instruction phrasings but fails unusual ones. Diversity in prompts matters as much as diversity in answers.

Open instruction datasets worth knowing

  • OpenAssistant / OASST1 — early open conversational dataset.
  • Tülu 3 SFT mix — high-quality, well-documented mixture from AI2. huggingface.co/datasets/allenai/tulu-3-sft-mixture
  • No Robots — Stanford. Small, hand-written, very high quality.
  • UltraChat — large synthetic conversational corpus.
  • Magpie — self-generated SFT data via instruction-tuned models. arxiv.org/abs/2406.08464
  • ShareGPT-style logs — user/assistant logs (license varies; check before use).
  • WildChat — real user-assistant conversations, opt-in. arxiv.org/abs/2405.01470

Practical tips

  • Pick the chat template before generating data. Changing it later means re-rendering and possibly retraining.
  • Mask correctly and verify with a unit test. A loss-mask test that asserts which positions contribute is worth its weight in gold.
  • Track per-category eval, not just aggregate. A regression in code quality is invisible in average eval scores.
  • Don't over-train. SFT is fast to overfit. Most recipes do 1–3 epochs at low LR (~1e-5 to 5e-6).
  • Keep a held-out SFT eval that mirrors your data mix. Loss on this set is a better leading indicator than benchmark scores.

Further reading