State of AI Benchmarks: Q3 2026
Q3 2026 benchmark data: knowledge tests saturating, coding scores compressing, and agentic benchmarks fragmenting into versions that disagree by 20-60 points.
Q3 2026 benchmark data: knowledge tests saturating, coding scores compressing, and agentic benchmarks fragmenting into versions that disagree by 20-60 points.
No single ‘AI visibility score’ is empirically defensible yet. Here’s the metric hierarchy we’re using to measure brand AI readiness — and where the research disagrees.
Profound, Ahrefs Brand Radar, Semrush, Peec AI, and Otterly.AI each define “AI visibility” differently. Here’s what each platform’s own docs say it measures.
Public ARC-AGI-2 leaderboards show models above 90%. The official, cost-capped ARC Prize contest tied to $2M shows 24%. Here’s why.
A sourced, seven-layer framework for auditing whether AI engines can crawl, parse, cite, and correctly attribute your brand.
OSWorld-Verified is saturating near 85-86%, but its harder successor, OSWorld 2.0, still tops out at 20.6% binary task completion. Here’s why.
64% of buyers open their first AI prompt with a category or competitor query. Here’s what G2’s own survey data shows buyers actually ask at each funnel stage.
MT-Bench proved LLM judges can match human agreement — but 2026 research shows position, verbosity, and self-enhancement bias run deeper than the original paper implied.
GPTBot, ClaudeBot, and PerplexityBot don’t wait for redirect chains or slow servers. A wrong status code can quietly erase a page from AI answers.
LiveCodeBench fights contamination by dating problems, not hiding them. Three trackers disagree on today’s leader — here’s what the scores mean.