FrontierMath Explained: Why Epoch AI Rewrote 42% of Its Own Problems
Epoch AI’s FrontierMath math benchmark got a 42%-of-problems rewrite in June 2026. Here’s what changed, why scores jumped, and where trackers disagree.
Epoch AI’s FrontierMath math benchmark got a 42%-of-problems rewrite in June 2026. Here’s what changed, why scores jumped, and where trackers disagree.
Ahrefs, Pew, Seer Interactive, and SparkToro data on how AI Overviews cut organic click-through rates in 2026 — and where the numbers disagree.
Grok’s 2026 record: rising hallucination rates, a UK ICO/Ofcom probe into image generation, and documented Grokipedia accuracy problems.
GDPval grades AI against real professional deliverables. The top score jumped from 38.8% to 74.1% in a year, but trackers disagree on who leads now.
Q3 2026 benchmark data: knowledge tests saturating, coding scores compressing, and agentic benchmarks fragmenting into versions that disagree by 20-60 points.
No single ‘AI visibility score’ is empirically defensible yet. Here’s the metric hierarchy we’re using to measure brand AI readiness — and where the research disagrees.
Profound, Ahrefs Brand Radar, Semrush, Peec AI, and Otterly.AI each define “AI visibility” differently. Here’s what each platform’s own docs say it measures.
Public ARC-AGI-2 leaderboards show models above 90%. The official, cost-capped ARC Prize contest tied to $2M shows 24%. Here’s why.
A sourced, seven-layer framework for auditing whether AI engines can crawl, parse, cite, and correctly attribute your brand.
OSWorld-Verified is saturating near 85-86%, but its harder successor, OSWorld 2.0, still tops out at 20.6% binary task completion. Here’s why.