Core Principle
For each capability axis in Dimensions of LLM Quality, there is a small set of canonical benchmarks the field has converged on. Knowing which benchmark maps to which axis lets you read a model release post intelligently and spot the gaps the lab didn’t report.
Last refreshed 2026-07-17 with a fresh research pass (prior body was April 2026 and had gone stale — saturation levels move fast).
Why This Matters
Labs pick which benchmarks to put on their release page. If a model release lists MMLU-Pro and GPQA but not SWE-bench, that’s a signal — not just an omission. You can only read those signals if you know what each benchmark measures and what the saturation level is.
Benchmarks also rot. MMLU and HumanEval are saturated and contaminated. Citing them as evidence today is a tell that someone isn’t paying attention. By mid-2026 this rot has reached the previous frontier: SWE-bench Verified and GPQA Diamond are both near-saturated, and the differentiators have moved to a new tier of ultra-long-horizon and novel-reasoning benchmarks.
The single most important reading rule (2026)
Most agent/coding benchmark numbers measure a system, not a model. Bare-model, vendor-scaffolded, and full-system scores for the same benchmark (e.g. GAIA, May 2026) diverge by 30-50 points. The model number tells you about the LLM; the scaffolded number tells you about the vendor’s product; the system number tells you about the integrator’s stack. Always ask which is being quoted before comparing across labs.
Evidence/Examples
General reasoning & knowledge
The 2026 landscape settled on three headline benchmarks; the older ones are effectively dead for frontier comparison.
- ARC-AGI-2 — novel abstract reasoning (visual grid puzzles), Chollet’s “fluid intelligence” test. Launched March 2025 at 0% for all models; by mid-2026 the frontier climbed fast (GPT-5.5 ~85%, Gemini 3.1 Pro ~77%, Opus 4.6 ~69%; human avg 66%, grand-prize threshold >85%). Now the sharpest reasoning differentiator, but climbing toward saturation — ARC-AGI-3 already exists. Leaderboard · Paper
- HLE (Humanity’s Last Exam) — ~2,500 expert-authored, multi-domain, multimodal frontier questions (Scale AI + CAIS). Frontier only in the low-20s % as of May 2026 → the most headroom of any knowledge bench (~2-3 yrs). Leaderboard · Site
- GPQA Diamond — PhD-level science, “Google-proof”. Approaching saturation (~94%, Gemini 3.1 Pro, Feb 2026); discriminating power for top models is narrowing. Still quoted, but a cluster near ceiling. Paper
- MMLU-Pro — successor to MMLU; still useful mid-tier, but frontier models cluster high. Paper
- MMLU / BIG-Bench Hard (BBH) — saturated/contaminated; ignore for frontier comparisons.
Math
- AIME 2025 — current frontier differentiator (2024 set is aging).
- FrontierMath — research-level, mostly unsolved; the durable hard one. Epoch AI
- MATH — largely saturated at the frontier; use for smaller models. Paper
Coding — the tier that still differentiates (2026)
SWE-bench Verified saturated in 2026 (top models cluster ~88%), so the live ranking moved to harder successors. All of these have independent leaderboards — check those, not the lab’s release-day self-report (the GLM 5.2 / Kimi K3 headline numbers were self-quoted; the independent boards tell a soberer story, e.g. FrontierSWE is led by Claude Fable 5, not the open models).
- SWE-bench Verified — real GitHub issue fixes; was the gold standard, now saturating. Most-reproduced coding bench; still a useful floor. Site
- SWE-bench Pro — Scale AI (Aug 2025, ICLR 2026); harder/less-contaminated issue set, scores fall to ~55-70%. The direct Verified successor. Standardized Scale SEAL board is the apples-to-apples run. Public leaderboard
- FrontierSWE — Proximal Labs; ultra-long-horizon (20 h/task) implementation, perf-eng, and ML-research projects. Largely unsaturated — most models barely make progress. Independently tracked. LLM Stats · Epoch AI · Site
- SWE-Marathon — Abundant AI (June 2026); 20 ultra-long tasks (~27M tokens/attempt: compilers, kernels, prod services). Frontier <30% (Opus 4.8 topped ~26%). Notable: 13.8% of rollouts showed reward-hacking — a caution about verifier-gaming on long tasks. Leaderboard · Paper
- Terminal-Bench — agent driving a real shell/CLI end-to-end; stresses tool-use + environment interaction, not just diffs.
- DeepSWE — Datacurve; another frontier coding-agent measure. Site
- LiveCodeBench — refreshed from LeetCode, contamination-resistant; competitive-programming style. Site
- Aider polyglot — multi-file edits across languages. Leaderboard
- Legacy: HumanEval / MBPP — saturated, contaminated; ignore.
Long-horizon autonomy (the 2026 headline metric)
- METR Task-Completion Time Horizon — not a pass/fail suite but the frontier-tracking metric: the human-time length of task a model finishes with 50% reliability. Doubling ~every 7 months historically, accelerating to ~4-4.3 months since 2023; as of mid-2026 the frontier pushes past ~16 h (beyond what METR’s suite reliably measures). This is the number to cite for “how autonomous are agents really.” METR · Epoch AI
Tool use / agents
- τ²-bench (tau2-bench) — Sierra; successor to tau-bench. Multi-turn tool-agent-user with policy adherence across retail / airline / telecom (+ banking, voice). Current standard for realistic tool agents. Repo
- BFCL (Berkeley Function Calling Leaderboard) — still canonical for raw function-calling. Site
- GAIA — general assistant tasks; watch the system-vs-model caveat above. Paper
Computer / browser use
- OSWorld — computer-use on a real desktop (files, apps, GUI).
- WebArena — multi-step real-website browser tasks.
Long context
Multimodal
- MMMU — college-level multi-discipline vision. Site
- MathVista — visual math; ChartQA / DocVQA — charts and documents.
Contamination-resistant aggregates (start here)
- Artificial Analysis Intelligence Index — composite cross-capability score; the trusted 30-second snapshot (see LLM Comparison Sources). Harder to cherry-pick than a single bench — which is why GLM 5.2 and Kimi K3 both quoted their Index rank.
- LiveBench — rotated monthly; reasoning, math, coding, language, IF, data. Site
Implications
- Check the saturation date first. SWE-bench Verified (~88%) and GPQA Diamond (~94%) crossed into the dead zone in 2026; ARC-AGI-2 is climbing there. The live differentiators are FrontierSWE, SWE-Marathon, HLE, and METR time-horizon.
- Independent board > release-day self-report. The 2026 successors all have independent leaderboards (Scale SEAL, LLM Stats, Epoch AI). When a lab quotes its own coding numbers (GLM 5.2, Kimi K3), go find the model on the independent board before believing the ranking.
- System vs model vs scaffold (see reading rule above) is the single biggest source of misleading agent numbers.
- Arena/Elo “WebDev” and blind-preference scores are the distrusted class (gameable, selective disclosure — see LLM Comparison Sources). A 2-on-Arena claim is weaker than an independent SWE-bench Pro run.
- Contamination stays the silent killer — prefer benchmarks that refresh (LiveBench, LiveCodeBench) or hold out private items (SWE-bench Pro private split, FrontierSWE).
- For agentic coding specifically: SWE-bench Pro (Scale SEAL) for one-shot fixes + FrontierSWE / SWE-Marathon for long-horizon + METR time-horizon for the trend line beat any single release-day number. For tool use, τ²-bench + BFCL.
Related Ideas
- Dimensions of LLM Quality — what each benchmark is trying to measure
- LLM Comparison Sources — leaderboards that aggregate these benchmarks
- GLM 5.2 - Near-Frontier Open-Weight Coding Model / Kimi K3 - Largest Open-Weight Model, Frontier-Trading Benchmarks — mid-2026 releases whose self-reported numbers this note contextualizes
- Small Local LLMs as Judges
Questions
- What’s the half-life of a benchmark before contamination kills it? (LiveCodeBench refreshes monthly for a reason.)
- How well does benchmark performance actually transfer to private workloads? (Anecdotally: loose correlation, not tight.)
- Once METR’s suite tops out past ~16 h, what replaces it as the long-horizon trend line?
- How much of the open-weight (GLM/Kimi) “frontier parity” survives once numbers come from independent boards rather than release tables?