Reasoning specialization¶
Goal¶
Improve performance on math, code, logic, scientific reasoning, planning, and multi-step problem solving.
This is the area where the field has moved fastest in the last 18 months. Verifiable-reward RL (RLVR) and group-relative policy optimization (GRPO), demonstrated at scale in models like DeepSeek-R1, have closed much of the gap between open and frontier reasoning models.
Data types¶
Human-written reasoning¶
High quality but expensive. Domain experts walk through problems step by step. Used as a high-trust seed set.
Synthetic chain-of-thought¶
Generated by stronger models.
Risks:
- Incorrect reasoning traces (the answer is right but the steps are wrong)
- Overlong reasoning (verbose but circular)
- Teacher-model style imitation
- Leakage of benchmark problems
Rejection sampling¶
Generate many candidate solutions, keep those that pass a verifier.
Examples:
- Math answer checker
- Unit tests
- Symbolic solver
- Theorem checker
- Exact-match answer
Rejection-sampling fine-tuning is the workhorse of modern reasoning post-training. It converts unreliable synthetic generations into a high-quality, verified dataset.
Process supervision¶
Train on intermediate steps, not just final answer.
Useful when you can label or verify steps. The canonical work is Lightman et al., "Let's Verify Step by Step", which showed process reward models outperform outcome reward models on math.
Outcome supervision¶
Only final answer is scored.
Easier to scale, but may reward lucky guesses. In practice, outcome supervision dominates for code (unit tests are natural step-skipping outcome rewards) and process supervision dominates for math.
Verifiable RL (RLVR)¶
Use automatic checkers as rewards.
This is especially important for modern reasoning models. Tülu 3's RLVR is an example of post-training with verifiable rewards for tasks with checkable outcomes.
DeepSeek-R1 trained a strong reasoning model using GRPO with rule-based rewards (correct answer, correct format) and no human labels for the reasoning traces — a major proof point that long chain-of-thought can emerge from RL alone.
The recipe (as of 2025)¶
A representative modern reasoning post-training recipe:
- SFT seed: a small, high-quality set of human-written or strong-teacher-distilled reasoning traces.
- Cold-start data: thousands of problems with verifier-validated solutions (rejection sampling).
- RLVR / GRPO: scale up with rule-based rewards on math, code, instruction-following.
- Distillation: distill the resulting reasoning model into smaller models for serving.
DeepSeek-R1 used a slight variant: pure RL from the base model (no SFT seed), discovering long chain-of-thought emerges spontaneously, then SFT + RL with the resulting traces.
Reasoning failure modes¶
- Correct final answer for wrong reason — the model gets lucky.
- Brittle prompt sensitivity — small wording changes flip the answer.
- Overthinking simple tasks — long chain-of-thought for "what's 2+2".
- Verbose hidden scratchpads —
<thinking>blocks that ramble for thousands of tokens. - Poor calibration — confident wrong answers.
- Memorized benchmark patterns — strong on GSM8K, weak on novel arithmetic.
- Good math but poor real-world reasoning — formal vs. informal generalization gap.
- Unit-test overfitting in code — passes the given tests, fails edge cases.
Practical tips¶
- Verifiers are gold. Whatever you can check automatically, you can train on. Build the verifier infrastructure before scaling RL.
- Curate the SFT seed carefully. A small set of good traces matters far more than a large set of mediocre ones.
- Watch for thought-token explosion. Set a max-length budget and reward concision. Otherwise the model learns to ramble because longer traces correlate with success on hard problems.
- Track per-domain accuracy. Math, code, and logic all behave differently. Aggregate "reasoning" scores hide regressions.
- Test on novel problems. Benchmark contamination in math (GSM8K, MATH, AIME) is rampant. Use private eval sets.
- Distillation is real. A strong reasoning teacher can lift smaller students dramatically; this is how most "reasoning capability" propagates across model sizes today.
Further reading¶
- Wei et al., "Chain-of-Thought Prompting Elicits Reasoning", 2022. arxiv.org/abs/2201.11903
- Lightman et al., "Let's Verify Step by Step", 2023. arxiv.org/abs/2305.20050
- Cobbe et al., "Training Verifiers to Solve Math Word Problems" (GSM8K), 2021. arxiv.org/abs/2110.14168
- Hendrycks et al., "Measuring Mathematical Problem Solving With the MATH Dataset", 2021. arxiv.org/abs/2103.03874
- Shao et al., "DeepSeekMath / GRPO", 2024. arxiv.org/abs/2402.03300
- DeepSeek-R1, 2025. The clearest open writeup of large-scale RL-trained reasoning. arxiv.org/abs/2501.12948
- Tülu 3 — RLVR, 2024. arxiv.org/abs/2411.15124
- Open-source reproductions:
Open-R1(github.com/huggingface/open-r1),simple_GRPO,verl(github.com/volcengine/verl).