Evaluation¶
Goal¶
Measure capability, safety, robustness, and product behavior.
No single benchmark is enough. The model that wins on MMLU but produces invalid JSON 10% of the time is unshippable.
Evaluation categories¶
General knowledge¶
- MMLU / MMLU-Pro — multi-task knowledge.
- TriviaQA — open-domain QA.
- GPQA — graduate-level science. Hard, contamination-resistant.
Reasoning¶
- GSM8K — grade-school math.
- MATH — competition math.
- BBH — Big-Bench Hard.
- ARC-Challenge — grade-school science with reasoning.
- AIME / AMC — competition math (high-end reasoning).
Code¶
- HumanEval — function-completion. Saturated; useful as a sanity check.
- MBPP — basic Python problems.
- LiveCodeBench — continuously updated, harder to contaminate.
- SWE-bench / SWE-bench Verified — real-world software-engineering tasks.
- Unit-test pass rates on a private code test set.
Instruction following¶
- IFEval — verifiable instruction-following constraints. arxiv.org/abs/2311.07911
- Format adherence (custom tests for your output formats)
- Constraint following ("respond in fewer than 100 words", "only use words starting with 'a'")
Chat quality¶
- MT-Bench — multi-turn LLM-judge eval.
- AlpacaEval 2 (length-controlled) — automatic LLM-judge eval, length-corrected.
- Arena-Hard — adversarially-curated battles with strong judges.
- Human preference comparisons in your own product.
- Chatbot Arena — public Elo from blind battles.
Long context¶
- RULER — multi-skill long-context. arxiv.org/abs/2404.06654
- Needle-in-a-haystack — sanity check.
- LongBench, InfiniteBench — multi-task long-context.
- Repository-level QA (specific to your codebase).
Tool use¶
- BFCL (Berkeley Function-Calling Leaderboard) — single + multi-turn function calling.
- τ-bench — Sierra. Realistic multi-turn tool use.
- API-Bank — agentic tool selection.
- End-to-end task success on your own tool set.
Safety¶
- Llama Guard / WildGuard — classifier-based scoring of refusal behavior.
- XSTest — over-refusal of safe prompts. arxiv.org/abs/2308.01263
- HarmBench — adversarial robustness.
- AdvBench — jailbreak attempts.
- Custom red-team eval set.
Factuality¶
- TruthfulQA — adversarial common misconceptions.
- SimpleQA — short-form fact recall (OpenAI). arxiv.org/abs/2411.04368
- Closed-book vs open-book grounded QA on your own data.
- Citation correctness.
- Hallucination rate from human review.
Calibration¶
Ask:
- Does the model know when it does not know?
- Does confidence track correctness?
- Does it ask clarifying questions appropriately?
Calibration is rarely tested but increasingly important — over-confident wrong answers are a major reason RLHF'd models can be worse than base models for some tasks.
Human evaluation¶
Still essential.
Humans judge:
- Helpfulness
- Style
- Subtle correctness
- Safety nuance
- Domain expertise
- User satisfaction
But human evals are expensive and noisy, so they need careful rubrics.
Make human eval cheap to repeat
A small (200–500 example) human eval set you can rerun every model iteration is more valuable than a 10,000 example set you run once. Frequency beats scale.
LLM-as-judge¶
Useful but dangerous.
Pros:
- Fast
- Cheap
- Scalable
- Good for early iteration
Cons:
- Bias toward certain styles (verbose, structured)
- Can miss factual errors
- Can prefer verbose answers
- Can be gamed
- Self-preference (judges prefer outputs from models in their own family)
Best practice: calibrate judge models against human labels. If your judge agrees with humans 85% of the time on a held-out set, you can trust it for trend-tracking; if it agrees 60%, you can't.
Decontamination matters here too¶
If your eval set leaks into pretraining or SFT, your scores are fiction. Decontaminate every stage:
- Pretraining corpus → eval set
- Continued-pretraining corpus → eval set
- SFT data → eval set
- Preference data → eval set
n-gram (13-gram) overlap and embedding similarity should both be checked.
Practical tips¶
- Build your own private eval. Public benchmarks contaminate fast. A few hundred examples written by your team, never published, is the most trustworthy signal you'll have.
- Track regressions, not just absolute scores. A 1% drop on IFEval is a real signal that something broke.
- Run evals automatically on every checkpoint. A merge that improves chat quality but regresses code by 5% should be caught before deploy.
- Don't tune to your eval. If you start optimizing the model against your private eval, you lose its independence. Keep a holdout-of-the-holdout.
- Pair quantitative + qualitative. Read 50 model outputs after every major change. Numbers miss systematic style regressions.
Further reading¶
- Liang et al., "HELM (Holistic Evaluation of Language Models)", 2022. arxiv.org/abs/2211.09110
- Hendrycks et al., "MMLU", 2020. arxiv.org/abs/2009.03300
- Suzgun et al., "BIG-Bench Hard", 2022. arxiv.org/abs/2210.09261
- Zheng et al., "Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena", 2023. arxiv.org/abs/2306.05685
- Dubois et al., "Length-Controlled AlpacaEval", 2024. arxiv.org/abs/2404.04475
- Zhou et al., "IFEval", 2023. arxiv.org/abs/2311.07911
- EleutherAI
lm-evaluation-harness— standard eval framework. github.com/EleutherAI/lm-evaluation-harness - OpenAI
evals— github.com/openai/evals inspect_ai— UK AISI. Solid modern evals framework. github.com/UKGovernmentBEIS/inspect_ai- HELM Lite / HELM Instruct — Stanford CRFM. crfm.stanford.edu/helm
- Chatbot Arena — chat.lmsys.org