Llama, DeepSeek, and Mistral: Documented Limitations of Open-Weight Models
Llama 4’s benchmark score came from an unreleased model, DeepSeek’s guardrails failed jailbreak tests, and Mistral hallucinates despite 256K context.
Llama 4’s benchmark score came from an unreleased model, DeepSeek’s guardrails failed jailbreak tests, and Mistral hallucinates despite 256K context.
Epoch AI’s FrontierMath math benchmark got a 42%-of-problems rewrite in June 2026. Here’s what changed, why scores jumped, and where trackers disagree.
Ahrefs, Pew, Seer Interactive, and SparkToro data on how AI Overviews cut organic click-through rates in 2026 — and where the numbers disagree.
Grok’s 2026 record: rising hallucination rates, a UK ICO/Ofcom probe into image generation, and documented Grokipedia accuracy problems.
GDPval grades AI against real professional deliverables. The top score jumped from 38.8% to 74.1% in a year, but trackers disagree on who leads now.
Q3 2026 benchmark data: knowledge tests saturating, coding scores compressing, and agentic benchmarks fragmenting into versions that disagree by 20-60 points.
No single ‘AI visibility score’ is empirically defensible yet. Here’s the metric hierarchy we’re using to measure brand AI readiness — and where the research disagrees.
Profound, Ahrefs Brand Radar, Semrush, Peec AI, and Otterly.AI each define “AI visibility” differently. Here’s what each platform’s own docs say it measures.
Public ARC-AGI-2 leaderboards show models above 90%. The official, cost-capped ARC Prize contest tied to $2M shows 24%. Here’s why.
A sourced, seven-layer framework for auditing whether AI engines can crawl, parse, cite, and correctly attribute your brand.