The AI Buyer Journey: What Prompts Buyers Actually Ask at Each Stage
64% of buyers open their first AI prompt with a category or competitor query. Here’s what G2’s own survey data shows buyers actually ask at each funnel stage.
64% of buyers open their first AI prompt with a category or competitor query. Here’s what G2’s own survey data shows buyers actually ask at each funnel stage.
MT-Bench proved LLM judges can match human agreement — but 2026 research shows position, verbosity, and self-enhancement bias run deeper than the original paper implied.
GPTBot, ClaudeBot, and PerplexityBot don’t wait for redirect chains or slow servers. A wrong status code can quietly erase a page from AI answers.
LiveCodeBench fights contamination by dating problems, not hiding them. Three trackers disagree on today’s leader — here’s what the scores mean.
ChatGPT, Perplexity, and Google’s AI Overviews all retrieve far more pages than they cite. Here’s what the data says about how each one decides.
HLE went from sub-10% scores in Jan 2025 to a 55-65% frontier cluster by Aug 2026 — but trackers disagree by nearly 20 points on where that frontier sits.
GA4’s AI Assistant channel, UTM tags, and server logs each capture a different slice of AI referral traffic — and each one misses most of it.
Aider’s polyglot benchmark scores coding models on 225 Exercism problems in 6 languages. GPT-5 leads at 88%, but PR-gated updates cut both ways.
A spot check of 20 B2B SaaS companies found just 6 resolve correctly in Wikidata’s entity search. Here’s the notability bar, and why it matters for AI.
Terminal-Bench tests AI agents on real Docker terminal tasks. Here’s what it measures, why v2.0 and v2.1 scores diverge, and how top trackers disagree.