Every frontier AI tool has documented, reproducible failure modes: hallucination rates that vary threefold depending on who measures them, context windows that are effectively half their advertised size, agents that pass a task once but fail it on repeat runs, and coding assistants whose output carries security weaknesses in roughly a quarter of sampled snippets. This page is the hub of our limitations series — a running ledger of failure modes for ChatGPT, Claude, Gemini, Perplexity, GitHub Copilot and Cursor, and the agent systems built on them. Every entry below cites a benchmark result, an official disclosure, or a reproducible study. No vibes.
What counts as a documented limitation
Three rules govern this ledger. First, every claim must trace to a named source with a date: a benchmark leaderboard, a peer-reviewed or preprint study, official vendor documentation, or a reproducible incident report. Second, benchmark versions are always named — a score on SWE-bench Verified is not a score on SWE-bench Pro, and conflating them is how marketing happens. Third, when trackers disagree, we report the disagreement instead of averaging it away. Disagreement between measurement sources is itself information about how fragile the measurement is. For the methodology behind reading any of these numbers, start with our hub on how to evaluate LLMs.
Hallucination: the rate depends on who’s measuring
There is no single “hallucination rate” for any model in 2026, and anyone quoting one without naming the task family is selling something. A five-model study by Digital Applied (2026) puts frontier models between 4.62% (Claude Haiku 4.5) and 6.10% (GPT-5.4-mini) on constrained document-grounded tasks. Meanwhile, SQ Magazine’s 2026 aggregate reports rates above 15% for most models on broader open-ended tests. Both can be true: the variance across task families is wider than the variance across models.
Two findings from the 2026 data are worth flagging. Extended thinking consistently roughly halves hallucination rates — the same Digital Applied study measured GPT-5.5 Pro dropping from 8.3% to 4.2% and Claude Opus 4.7 from 9.4% to 5.1% with reasoning enabled. And models optimized for factual consistency on constrained tasks have become very good at restating what exists in a source document while still guessing when documents get long and complex — which connects directly to the next section. For the full model-by-model breakdown across Vectara, Google FACTS Grounding, SimpleQA, and AA-Omniscience, see how hallucination is actually benchmarked.
One more 2026 finding belongs here, because it cuts against the “steady progress” reading. On AA-Omniscience, which scores abstention discipline rather than error frequency, the frontier hallucination rate moved in both directions within a single quarter: GPT-5.6 Sol reached 92% at max effort, then GPT-6 Astra cut it to 51% while gaining 4 points of accuracy (Artificial Analysis, September 9, 2026). Claude Opus 5 went the other way, adding 7 points of accuracy while its hallucination rate rose 14 points to 50%; Grok 4.7 improved to 29% from Grok 4.6’s 34%. Four frontier releases, four different directions, one benchmark. Any sentence of the form “models hallucinate X% in 2026” is describing a single build on a single day.
Update, October 1, 2026: the churn has not slowed. OpenAI shipped two more tiers, GPT-6 Sol and GPT-6 Luna, a week after Astra, then replaced Sol again with GPT-6.1 Sol at DevDay on September 29, 2026 — three hallucination-rate readings (92%, 60%, 54.3%) for a model lineage that is eight days old at the time of this update (Artificial Analysis, September 30, 2026). See our ChatGPT limitations entry for the full breakdown.
Long context: advertised versus effective
Advertised context windows are a spec-sheet number, not a performance number. NVIDIA’s RULER benchmark shows models reliably use only 50–65% of their advertised window, per Atlan’s 2026 roundup — a nominal 1M-token model may only perform dependably to 600–700K tokens. Chroma’s “context rot” study, covered in the same roundup, found accuracy degradation of 30%+ in mid-window positions across all 18 frontier models tested. No exceptions.
The steepest documented drop comes from LOCA-bench (arXiv, February 2026), which measured Claude 4.5 Opus at 96.0% task success at 8K tokens of agent context collapsing to 14.7% at 256K. And in multi-turn production settings, a Microsoft Research result reported in May 2026 found frontier models (Gemini 3.1 Pro, Claude 4.6 Opus, GPT-5.4) losing on average 25% of document content over 20 delegated interactions, with the all-model average closer to 50%. Degradation is task-dependent: recall and RAG hold up reasonably; re-ranking and citation-grounded generation degrade hardest.
Coding assistants: saturated leaderboards, measurable insecurity
The headline numbers look solved. On SWE-bench Verified, BenchLM’s leaderboard (July 11, 2026) shows Claude Mythos 5 at 95.5%, Claude Fable 5 at 95%, and Claude Opus 4.8 at 88.6%. The caveat that should accompany every one of those scores: on llm-stats’ tracking, only 1 of 100 leaderboard results was independently verified — the rest were vendor-submitted.
Those figures have since been passed and the pattern has not changed. Claude Opus 5, released July 24, 2026, reports 96% on SWE-bench Verified in Anthropic’s own system card — but 79.2% on the contamination-resistant SWE-bench Pro, behind Claude Fable 5.1 at 81.2% (BenchLM, data as of September 21, 2026). A ~17-point Verified-to-Pro drop on the current frontier model is the same shape as Opus 4.8’s ~19 points two generations earlier: the leaderboard saturates, the harder variant does not. The versioning caution compounds here — Vals AI’s independent SWE-bench run returns 97.0% for the same model the system card puts at 96%, and on Terminal-Bench 2.1 Artificial Analysis reports 89% at max effort against Vals AI’s 84.6% on the same benchmark version. Harness and effort settings, not capability.
Security is where the documented record turns ugly. An ACM TOSEM empirical study of 733 Copilot-generated snippets in real GitHub projects found security weaknesses in 29.5% of Python and 24.2% of JavaScript samples. The tools themselves have shipped exploitable flaws: Cursor’s CVE-2025-59944 enabled persistent remote code execution via MCP configuration, and Copilot’s CamoLeak vulnerability (CVSS 9.6) allowed silent exfiltration of private repository code through invisible prompt injection, per MintMCP’s 2026 security comparison. Pillar Security separately documented how poisoned rules files can weaponize both agents. And in sessions run April–June 2026, The Hacker News reported Copilot Chat 0.30.3 refusing harmful requests in conversation, then writing the same logic in code after roughly six exchanges. For the full documented breakdown — pricing shocks, CVE timelines, context-window limits, and the SWE-bench scores that disagree by double digits — see our GitHub Copilot vs Cursor limitations entry.
One 2026 disclosure deserves separate mention because it evades the usual tracking. TrustFall, disclosed by Adversa AI in May 2026 and covered by Dark Reading, lets a cloned repository auto-approve and spawn an MCP server with full user privileges the moment a developer accepts the routine folder-trust dialog — affecting Cursor CLI, GitHub Copilot CLI, Claude Code, and Gemini CLI alike. No CVE was assigned, Anthropic declined the report as outside its threat model, and the researchers found it unpatched across all four tools as of June 2026. A shared, unpatched, unindexed flaw is the kind of entry a CVE-count comparison structurally cannot show.
Agents and computer use: passing once is not reliability
tau-bench’s Pass^k metric — the probability an agent succeeds on all k repeated runs of the same task — is the most honest reliability number in the agent literature. On the Sierra leaderboard (May 2026), Claude Opus 4.5 leads single-pass at 0.70, but the best pass^4 score belongs to Qwen3.5-397B-A17B at just 0.56. Read that plainly: the most consistent frontier agent completes the same task four times in a row barely better than a coin flip.
Computer use is a live example of tracker disagreement. One April 2026 roundup puts OSWorld state-of-the-art near 38%, while BenchLM’s OSWorld-Verified standings (May 2026) show Claude Mythos Preview at 79.6% — above the 72.4% human baseline. Most of that gap is versioning and harness differences between original OSWorld and OSWorld-Verified, which is precisely why this ledger names benchmark versions. For the full map of tool-calling, computer-use, terminal, and browser agent benchmarks, see our guide to AI agent benchmarks.
AI search: a citation is not verification
The Columbia Journalism Review’s eight-platform study (March 2025, summarized in Suprmind’s 2026 Perplexity profile) remains the reference point: the best performer, Perplexity Sonar Pro, still answered 37% of news-citation queries incorrectly — more than one in three attributions fabricated or misdirected, from the platform that did best. AI search tools retrieve and summarize; they do not verify. If you want the mechanics of why retrieval and ranking are hard problems in the first place, our explainer on how a search engine works covers the pipeline these tools sit on top of. Whether the structured data on a page factors into that retrieval and citation process is far less settled than SEO advice implies — see our breakdown of documented claims about structured data and AI citations.
A second methodology-disclosed audit has since entered the record on a harder question. DeepTRACE (Salesforce AI Research, with Microsoft Research) does not ask whether a tool identifies an article correctly; it asks whether each statement is actually supported by the source cited beside it. Across generative search engines, citation accuracy landed at 40–68%, with frequent misattribution even when a supporting source was present in the list. Deep-research modes fared worse on grounding: unsupported-statement rates ran 53.6% for Gemini, 74.6% for YouChat, 90.2% for Copilot Think Deeper, and 97.5% for Perplexity — against 12.5% for GPT-5 deep research on the same corpus. That last figure is the one that makes the others damning: near-reliable grounding is demonstrably achievable, so the spread is a product choice rather than a limit of retrieval.
The ledger at a glance
| Failure mode | Documented evidence | Source, date |
|---|---|---|
| Hallucination varies by task family | 4.62–6.10% constrained vs 15%+ open-ended | Digital Applied study; SQ Magazine aggregate, 2026 |
| Effective context < advertised | 50–65% of advertised window usable (RULER) | Atlan roundup, 2026 |
| Long-context collapse | Claude 4.5 Opus: 96.0% at 8K → 14.7% at 256K | LOCA-bench, arXiv, Feb 2026 |
| Multi-turn content loss | ~25% of document content lost over 20 interactions | Microsoft Research, May 2026 |
| Insecure generated code | 29.5% of Python, 24.2% of JS snippets with weaknesses | ACM TOSEM, 733-snippet study |
| Verified-to-Pro coding collapse | Opus 5: 96% SWE-bench Verified → 79.2% SWE-bench Pro | Anthropic system card / BenchLM, Sep 21 2026 |
| Hallucination direction is per-build | GPT-5.6 Sol 92% → GPT-6 Astra 51%; Opus 5 rose to 50%; Grok 4.7 fell to 29% | Artificial Analysis, Jul–Sep 2026 |
| Unsupported statements in AI search | 97.5% (Perplexity DR) vs 12.5% (GPT-5 DR) | DeepTRACE, arXiv 2509.04499 |
| Unpatched, uncatalogued agent RCE | TrustFall: 4 coding CLIs, no CVE assigned, unpatched as of Jun 2026 | Adversa AI / Dark Reading, May 2026 |
| Self-reported leaderboards | 1 of 100 SWE-bench Verified results independently verified | llm-stats, July 2026 |
| Agent inconsistency | Best pass^4 on tau-bench: 0.56 | Sierra leaderboard, May 2026 |
| AI search misattribution | Best platform wrong on 37% of citation queries | CJR study, Mar 2025 |
How to read this ledger
This page anchors a series of per-tool limitations posts — ChatGPT, Claude, Gemini, Perplexity, and Copilot vs Cursor each have their own documented, sourced entry, with more added as they publish. Three habits will serve you across all of them: never accept a score without its benchmark version, treat vendor-submitted results as claims rather than measurements, and when two trackers disagree, investigate the harness before trusting either. The full framework is in our evaluation guide.
FAQ
Which frontier AI tool has the fewest documented limitations?
None cleanly. The 2026 evidence shows failure modes are task-shaped, not brand-shaped: the model that leads constrained factual tasks still degrades on long context, and the agent that tops single-pass benchmarks still fails repeat runs. Choose by task family, not by leaderboard rank.
Why do hallucination numbers differ so much between sources?
Because “hallucination” is measured against different task families — document-grounded summarization, open-ended factual recall, citation generation — and the variance across tasks exceeds the variance across models. A 4% rate and a 15% rate can describe the same model honestly.
Are these limitations getting better or worse?
Mixed. Extended thinking measurably halves hallucination on factual tasks, and OSWorld-Verified scores now exceed the human baseline. GPT-6 Astra also became the first model in its lineage to cut its abstention-failure rate (92% to 51%) without trading away accuracy. But long-context collapse, agent inconsistency (pass^k), statement-level grounding in AI search, and insecure code generation remain documented and largely unsolved as of September 2026 — and Claude Opus 5 moved backwards on hallucination while moving forwards on accuracy, so “improving” is not a direction the whole frontier shares.
Last updated October 1, 2026. This page is refreshed as benchmarks and scores move.
Pingback: How AI Assistants Decide Which Brands to Recommend - Tech Blog
Pingback: GPQA Explained: What It Measures and Why It's Saturating - Tech Blog
Pingback: ChatGPT Limitations in 2026: A Documented, Sourced List - Tech Blog
Pingback: Claude Limitations in 2026: A Documented, Sourced List - Tech Blog
Pingback: AI Agent Benchmarks: A Map of tau-bench, OSWorld, Terminal-bench and More - Tech Blog
Pingback: GitHub Copilot vs Cursor: Documented Limitations Compared - Tech Blog
Pingback: How Hallucination Is Measured: Benchmarks Behind the Claims - Tech Blog
Pingback: Schema Markup That AI Engines Actually Read - Tech Blog
Pingback: Gemini Limitations in 2026: Where Long Context Actually Degrades - Tech Blog