State of AI Benchmarks: Q3 2026

Frontier benchmarks in Q3 2026 are telling three different stories at once: knowledge and reasoning tests are converging near their ceilings, coding and agentic benchmarks are fragmenting into incompatible versions that produce wildly different rankings for the same nominal test, and independent trackers routinely disagree with each other by more than the gap separating the models they rank.

Knowledge and reasoning: the ceiling has arrived

GPQA Diamond scores now cluster in a tight band at the top. As of August 25, 2026, BenchLM.ai’s tracker had Gemini 3.1 Pro Preview and GPT-5.6 Sol tied at 94.1%, with GPT-5.5 close behind at 93.5%. A separate run on Vals.ai put Gemini 3.1 Pro fractionally ahead of Claude Opus 4.7, 94.3% to 94.2% — a gap smaller than the benchmark’s own run-to-run noise. That’s consistent with what we flagged in our GPQA explainer: score differences at the top of GPQA Diamond are no longer doing much work distinguishing frontier models.

MMLU-Pro shows the same pattern from the other direction. Frontier models sit functionally saturated above 88% across most 2026 tracking sites, a trajectory we’ve followed since our piece on why the original MMLU died and what replaced it.

Humanity’s Last Exam (HLE), built specifically to resist saturation, is the outlier — and the benchmark where trackers disagree most violently. A September 4, 2026 snapshot has Claude Fable 5.1 leading at 65%, with Opus 5 at 64.7% and Mythos 5 at 64.5%. Two days later, a different configuration of the same tracker showed Claude Fable 5 leading at just 55.5%, with Opus 5 at 54.9% and GPT-6 Astra at 54.7% — a nearly 10-point swing with no model change behind it. The stakes of that gap are real: Stanford’s 2026 AI Index reports frontier models gained roughly 30 percentage points on HLE in a single year, which makes a 10-point tracker disagreement equal to nearly a third of a year’s progress. Details on the benchmark’s design are in our Humanity’s Last Exam piece.

Coding benchmarks: SWE-bench compresses, LiveCodeBench and Aider Polyglot disagree

SWE-bench Verified is nearing the same compression GPQA hit earlier this year. BenchLM.ai’s September 4, 2026 leaderboard has Claude Opus 5 at 96%, Claude Mythos 5 at 95.5%, and Claude Fable 5 at 95% — a top three within one point of each other. A separate Vals.ai run of Opus 5 scored it higher still, at 97%. That compression is the subject of our SWE-bench explainer, and it isn’t happening in a vacuum: contamination audits circulating this year found training-data overlap on SWE-Bench across frontier model families, and separately flagged a majority of “hard” tasks as having flawed test cases — the same failure mode documented years earlier in our HumanEval contamination case study.

LiveCodeBench, designed specifically to resist that kind of contamination via a rolling problem-release cutoff, still produces conflicting rankings. BenchLM.ai’s August 2026 leaderboard put Qwen3.7 Max on top at 91.6%, ahead of Qwen3.7 Plus (89.6%) and GLM-4.7 (84.9%). A separate tracker had DeepSeek-V4-Pro-Max leading instead, at 93.5% — a different model and a nearly two-point-higher score from the same nominal benchmark. As we noted covering LiveCodeBench’s contamination-resistant design, a rolling cutoff controls for training overlap but does nothing to standardize harness, sampling, or scoring protocol across trackers.

Aider Polyglot, built to counter SWE-bench’s Python bias, shows a different failure mode entirely: staleness. The most complete public tracker still lists GPT-5 on top at 88.0% across 22 evaluated models — a ranking that predates GPT-5.5, GPT-5.6 Sol, Claude Fable 5, and GPT-6 Astra, all of which now lead other coding benchmarks. As our Aider Polyglot piece noted, being hard for vendors to game doesn’t help if the leaderboard is also slow to update with current models.

Agentic benchmarks are fragmenting into incompatible versions

Nowhere is version fragmentation more visible than Terminal-Bench, which now has at least four actively tracked versions producing incompatible leaderboards in the same quarter. The original Terminal-Bench has GPT-5.6 Sol leading at 65.9% (BenchLM.ai, Sep 1, 2026). Terminal-Bench 2.0 has GPT-5.5 on top at 82.7% (llm-stats.com). Terminal-Bench 2.1 has Grok 4.6 leading at 88.4%, ahead of GLM 5.3 (88.2%) and DeepSeek V4 Pro (87.9%), per CodingFleet’s tracker. And Terminal-Bench 3.0, the hardest variant, has Claude Opus 5 leading a much lower field at 42.7%, with GPT-5.6 Sol at 34.6% and Claude Fable 5 at 34.0%. Four citations of “Terminal-Bench” in the same quarter can describe scores 20 to 60 points apart depending on which version is meant — a versioning problem we flagged in our Terminal-bench explainer.

tau2-bench, the active successor to the original tau-bench (frozen at its 2024 model set, where Claude 3.5 Sonnet still tops the retail leaderboard at 69.2%), shows the same tracker split as HLE: one August 16, 2026 leaderboard has GLM-5.2 at 99.1%, while another puts GLM-5 at just 89.7% on the nominally same benchmark. As covered in our tau-bench explainer, multi-turn tool-calling scores are unusually sensitive to how strictly a tracker enforces policy adherence, which is the likely source of a 9-point gap.

OSWorld-Verified is comparatively well-behaved: BenchLM.ai’s September 4, 2026 snapshot has Qwen3.8 Max leading at 86.1%, with Claude Fable 5 and Claude Mythos 5 tied at 85%. Zoom out, though, and Stanford’s AI Index frames a much bigger story: aggregate agent success on OSWorld-style computer-use tasks rose from roughly 12% to roughly 66% over the past year — the largest year-over-year jump the Index tracked, and the subject of our OSWorld explainer on why computer-use scores stayed so low for so long. For the full map of which agentic benchmark measures what, see our AI agent benchmarks guide.

ARC-AGI-2: the rare benchmark still separating models by double digits

ARC-AGI-2 is the exception to the compression story. BenchLM.ai’s September 4, 2026 leaderboard has GPT-6 Astra leading at 95%, with GPT-5.6 Sol at 92.5% and Claude Opus 5 at 90.4% — a wider spread than GPQA, HLE, or SWE-bench show at the top. ARC Prize’s own materials put average human performance at 66% and the grand-prize threshold above 85%, meaning the top three models here have now cleared both the human baseline and the prize bar — a milestone our ARC-AGI explainer covers in more detail.

Q3 2026 benchmark snapshot

Benchmark Top model Score Tracker Date
GPQA Diamond Gemini 3.1 Pro Preview / GPT-5.6 Sol (tied) 94.1% BenchLM.ai Aug 25, 2026
SWE-bench Verified Claude Opus 5 96% (BenchLM.ai) / 97% (Vals.ai) BenchLM.ai / Vals.ai Sep 4, 2026
Humanity’s Last Exam Claude Fable 5.1 / Claude Fable 5 65% vs. 55.5% (two snapshots) BenchLM.ai Sep 4 & Sep 6, 2026
LiveCodeBench Qwen3.7 Max / DeepSeek-V4-Pro-Max 91.6% vs. 93.5% BenchLM.ai / other tracker Aug 2026
Terminal-Bench (v1/2.0/2.1/3.0) GPT-5.6 Sol / GPT-5.5 / Grok 4.6 / Claude Opus 5 65.9% / 82.7% / 88.4% / 42.7% BenchLM.ai, llm-stats.com, CodingFleet Aug–Sep 2026
tau2-bench GLM-5.2 / GLM-5 99.1% vs. 89.7% Two trackers Aug 16, 2026
OSWorld-Verified Qwen3.8 Max 86.1% BenchLM.ai Sep 4, 2026
ARC-AGI-2 GPT-6 Astra 95% BenchLM.ai Sep 4, 2026
Aider Polyglot GPT-5 (stale ranking) 88.0% llm-stats.com last major refresh predates GPT-5.5/5.6

Why the numbers keep moving

Stanford’s AI Index frames this as a “jagged frontier”: Gemini Deep Think won IMO gold in 2025, yet the best model on ClockBench reads analog clocks correctly only 50.1% of the time versus 90.1% for humans. Separately, Artificial Analysis reported that six labs now field a model scoring above 50 on its Intelligence Index, and the AI Index notes that as of March 2026, Anthropic, xAI, Google, OpenAI, Alibaba, and DeepSeek were clustered within 25 Elo points of each other on arena-style leaderboards — evidence that competition has shifted toward cost and reliability rather than raw capability at the very top.

Some of the disagreement documented above is protocol, not capability: harness choice, sampling temperature, retry budgets, and judge configuration all move scores by several points on the same underlying model, a dynamic we detail in our MT-Bench and LLM-as-judge piece. If you’re choosing which numbers to trust when evaluating a model for your own use case, our guide to evaluating LLMs walks through how to weight benchmark claims against your own task-specific tests.

FAQ

Why do different trackers report different scores for the same benchmark?
Public leaderboards vary in harness (the scaffolding and tools given to the model), sampling settings, retry/attempt budgets, and which benchmark version they’re running. Two snapshots taken days apart from BenchLM.ai on Humanity’s Last Exam, for example, showed a roughly 10-point swing with no model change — a configuration difference, not a capability difference.

Is SWE-bench Verified still a meaningful benchmark now that top scores are near 96–97%?
It’s meaningful for detecting regressions and for ranking non-frontier models, but for the top tier the score gaps (roughly one point separating first from third place) are smaller than the disagreement between trackers measuring the same model, and contamination concerns raised in 2026 audits mean a high score alone doesn’t rule out memorization.

Which agentic benchmark should I trust right now — Terminal-Bench, tau2-bench, or OSWorld?
None in isolation. Terminal-Bench has four incompatible versions currently in circulation with scores 20–60 points apart for the same model. tau2-bench shows a similar spread between trackers. OSWorld-Verified is the most stable of the three this quarter, but all three measure different task distributions, so the right choice depends on whether your use case looks more like terminal automation, customer-service tool calling, or general desktop computer use.

Last updated September 8, 2026. This page is refreshed as benchmarks and scores move.

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top