Skip to content

Evaluation

Goal

Measure capability, safety, robustness, and product behavior.

No single benchmark is enough. The model that wins on MMLU but produces invalid JSON 10% of the time is unshippable.

Evaluation categories

General knowledge

  • MMLU / MMLU-Pro — multi-task knowledge.
  • TriviaQA — open-domain QA.
  • GPQA — graduate-level science. Hard, contamination-resistant.

Reasoning

  • GSM8K — grade-school math.
  • MATH — competition math.
  • BBH — Big-Bench Hard.
  • ARC-Challenge — grade-school science with reasoning.
  • AIME / AMC — competition math (high-end reasoning).

Code

  • HumanEval — function-completion. Saturated; useful as a sanity check.
  • MBPP — basic Python problems.
  • LiveCodeBench — continuously updated, harder to contaminate.
  • SWE-bench / SWE-bench Verified — real-world software-engineering tasks.
  • Unit-test pass rates on a private code test set.

Instruction following

  • IFEval — verifiable instruction-following constraints. arxiv.org/abs/2311.07911
  • Format adherence (custom tests for your output formats)
  • Constraint following ("respond in fewer than 100 words", "only use words starting with 'a'")

Chat quality

  • MT-Bench — multi-turn LLM-judge eval.
  • AlpacaEval 2 (length-controlled) — automatic LLM-judge eval, length-corrected.
  • Arena-Hard — adversarially-curated battles with strong judges.
  • Human preference comparisons in your own product.
  • Chatbot Arena — public Elo from blind battles.

Long context

  • RULER — multi-skill long-context. arxiv.org/abs/2404.06654
  • Needle-in-a-haystack — sanity check.
  • LongBench, InfiniteBench — multi-task long-context.
  • Repository-level QA (specific to your codebase).

Tool use

  • BFCL (Berkeley Function-Calling Leaderboard) — single + multi-turn function calling.
  • τ-bench — Sierra. Realistic multi-turn tool use.
  • API-Bank — agentic tool selection.
  • End-to-end task success on your own tool set.

Safety

  • Llama Guard / WildGuard — classifier-based scoring of refusal behavior.
  • XSTest — over-refusal of safe prompts. arxiv.org/abs/2308.01263
  • HarmBench — adversarial robustness.
  • AdvBench — jailbreak attempts.
  • Custom red-team eval set.

Factuality

  • TruthfulQA — adversarial common misconceptions.
  • SimpleQA — short-form fact recall (OpenAI). arxiv.org/abs/2411.04368
  • Closed-book vs open-book grounded QA on your own data.
  • Citation correctness.
  • Hallucination rate from human review.

Calibration

Ask:

  • Does the model know when it does not know?
  • Does confidence track correctness?
  • Does it ask clarifying questions appropriately?

Calibration is rarely tested but increasingly important — over-confident wrong answers are a major reason RLHF'd models can be worse than base models for some tasks.

Human evaluation

Still essential.

Humans judge:

  • Helpfulness
  • Style
  • Subtle correctness
  • Safety nuance
  • Domain expertise
  • User satisfaction

But human evals are expensive and noisy, so they need careful rubrics.

Make human eval cheap to repeat

A small (200–500 example) human eval set you can rerun every model iteration is more valuable than a 10,000 example set you run once. Frequency beats scale.

LLM-as-judge

Useful but dangerous.

Pros:

  • Fast
  • Cheap
  • Scalable
  • Good for early iteration

Cons:

  • Bias toward certain styles (verbose, structured)
  • Can miss factual errors
  • Can prefer verbose answers
  • Can be gamed
  • Self-preference (judges prefer outputs from models in their own family)

Best practice: calibrate judge models against human labels. If your judge agrees with humans 85% of the time on a held-out set, you can trust it for trend-tracking; if it agrees 60%, you can't.

Decontamination matters here too

If your eval set leaks into pretraining or SFT, your scores are fiction. Decontaminate every stage:

  • Pretraining corpus → eval set
  • Continued-pretraining corpus → eval set
  • SFT data → eval set
  • Preference data → eval set

n-gram (13-gram) overlap and embedding similarity should both be checked.

Practical tips

  • Build your own private eval. Public benchmarks contaminate fast. A few hundred examples written by your team, never published, is the most trustworthy signal you'll have.
  • Track regressions, not just absolute scores. A 1% drop on IFEval is a real signal that something broke.
  • Run evals automatically on every checkpoint. A merge that improves chat quality but regresses code by 5% should be caught before deploy.
  • Don't tune to your eval. If you start optimizing the model against your private eval, you lose its independence. Keep a holdout-of-the-holdout.
  • Pair quantitative + qualitative. Read 50 model outputs after every major change. Numbers miss systematic style regressions.

Further reading