How ChatGPT, Perplexity, and Gemini Choose Citations Differently
ChatGPT, Perplexity, and Google’s AI Overviews all retrieve far more pages than they cite. Here’s what the data says about how each one decides.
ChatGPT, Perplexity, and Google’s AI Overviews all retrieve far more pages than they cite. Here’s what the data says about how each one decides.
HLE went from sub-10% scores in Jan 2025 to a 55-65% frontier cluster by Aug 2026 — but trackers disagree by nearly 20 points on where that frontier sits.
GA4’s AI Assistant channel, UTM tags, and server logs each capture a different slice of AI referral traffic — and each one misses most of it.
Aider’s polyglot benchmark scores coding models on 225 Exercism problems in 6 languages. GPT-5 leads at 88%, but PR-gated updates cut both ways.
A spot check of 20 B2B SaaS companies found just 6 resolve correctly in Wikidata’s entity search. Here’s the notability bar, and why it matters for AI.
Terminal-Bench tests AI agents on real Docker terminal tasks. Here’s what it measures, why v2.0 and v2.1 scores diverge, and how top trackers disagree.
Brand hallucinations stem from stale training data, entity mix-ups, or bad retrieval — and courts now hold companies liable for what chatbots say.
Vectara HHEM, Google FACTS Grounding, SimpleQA, and AA-Omniscience all claim to measure hallucination — they disagree on which model is safest, and that’s the real finding.
τ-bench grades AI agents on real tool calls, not answers. Here’s how pass^k works, how τ² and τ³-bench evolved, and why trackers disagree by 30+ points.
AI engines don’t read your brand name — they resolve it to a graph node. Here’s how entity resolution works, and where it breaks.