AI Agent Benchmarks: A Map of tau-bench, OSWorld, Terminal-bench and More
tau-bench, BFCL v4, OSWorld, Terminal-Bench, WebArena, and GAIA each test a different slice of agent behavior — and their leaderboards often disagree.
tau-bench, BFCL v4, OSWorld, Terminal-Bench, WebArena, and GAIA each test a different slice of agent behavior — and their leaderboards often disagree.
GPTBot, ClaudeBot, and PerplexityBot don’t execute JavaScript. Measured crawler data shows what that means for AI visibility, and how to fix it.
Perplexity has the lowest citation error rate of 8 AI search tools in the one rigorous study available — 37%. Here’s what that number does and doesn’t mean.
HumanEval scores now cluster at 88-99% across frontier models, but three trackers disagree by 20 points. Here’s the contamination evidence and what replaced it.
Bing confirms its LLMs read schema. Google calls most AI-citation claims about it “wishful thinking.” Here’s what’s actually documented vs. marketing hype.
Gemini 3 Pro’s 10M-token window sounds huge, but MRCR v2 shows retrieval accuracy falling to ~25% at 1M tokens — trailing Claude and GPT-5.4.
MMLU is saturated and frontier labs have stopped reporting it. MMLU-Pro replaced it — but trackers disagree on who leads, and label errors linger.
GPTBot, ClaudeBot, and PerplexityBot are three bots each, not one. A sourced robots.txt reference from OpenAI, Anthropic, and Perplexity’s own docs.
Claude’s own system cards admit real limits in 2026: hallucination rates that vary 8x by methodology, 1M-token context that fails 1-in-4 retrievals, and Claude Code caps changed four times since March.
SWE-bench is four benchmarks, not one. OpenAI dropped Verified over contamination in Feb 2026 — here’s what each variant measures and where scores stand.