The most consistent finding across 2026 evaluations of ChatGPT is not what it can’t do — it’s what it does confidently wrong: the flagship model still answers questions it shouldn’t, retrieves reliably from only a fraction of its advertised context window, and stores memories that quietly corrupt future answers. This page is a documented, sourced ledger of ChatGPT’s limitations as they stand in September 2026. Every claim below carries a version, a source, and a date. Where trackers disagree — and they do — the disagreement is reported, not smoothed over. For the methodology behind entries like this one, see our Limitations Ledger hub.
Scope and method
“ChatGPT” in 2026 is not one model. The consumer product routes across GPT-5.x variants (GPT-5.4 shipped March 2026, GPT-5.5 in April, GPT-5.6 in early July, and GPT-6 Astra on September 9, 2026), and a limitation measured on one variant does not automatically hold for another. Each entry below names the specific model and evaluation. We exclude anything sourced only to vendor marketing or unversioned blog claims — the same rule we apply when evaluating LLMs generally.
Hallucination: better, but the headline numbers hide a split
OpenAI’s GPT-5.5 system card (April 23, 2026) reports that individual claims are 23% more likely to be factually correct than GPT-5.4’s, and that responses contain a factual error 3% less often. That is claim-level progress — but it is not the “60% hallucination reduction” that circulated in press coverage. An independent Wire analysis (2026) traced the 60% figure to a context-engineering evaluation setup, not the model’s base factuality.
The trackers diverge sharply depending on what they measure. On Artificial Analysis’s AA-Omniscience (2026), GPT-5.5 (xhigh reasoning) posts the highest accuracy of any frontier model at 57% — while simultaneously recording an 86% hallucination rate, meaning that when it doesn’t know the answer, it guesses rather than abstains the overwhelming majority of the time. On Suprmind’s June 2026 aggregation, GPT-5.5 Pro’s hallucination rate drops from 8.3% to 4.2% with extended thinking enabled. Both numbers are “the hallucination rate,” and they differ by an order of magnitude because one measures abstention discipline and the other measures error frequency in grounded tasks. Citation accuracy remains the worst task family in that aggregation, averaging 12.4% hallucination even with extended thinking.
Two releases since have moved that number in both directions, and the swing is larger than anything in the GPT-5.x record. GPT-5.6 Sol, benchmarked by Artificial Analysis on July 9, 2026, delivered what AA called “a minor improvement over GPT-5.5 in the AA-Omniscience Index, with a small uplift in accuracy coupled with an increase in hallucination rate” — the rate reaching 92% at max effort. Then GPT-6 Astra, released September 9, 2026, cut it to 51% at max effort while simultaneously adding 4 points of accuracy — the first time in this lineage that abstention discipline improved without an accuracy tax. That is real progress on the specific failure mode this section documents, and it still leaves the flagship guessing on roughly half the questions it cannot answer. Note also what the trajectory implies about single-snapshot claims: 86%, 92% and 51% are all “ChatGPT’s hallucination rate” within five months.
Update, September 28, 2026: OpenAI followed Astra with two cheaper tiers, GPT-6 Sol and GPT-6 Luna, on September 22, 2026. Both cut hallucination the same way Astra did — by answering less, not by knowing more. Sol (max effort) falls from 92% to 60% on AA-Omniscience while its attempt rate drops from 99% to 83% of questions, trimming accuracy 5 points (59% to 54%) along the way; Luna (max effort) falls from 93% to 77% with accuracy roughly flat (44% vs 43%). On the composite AA-Omniscience Index, Sol improves from 22 to 27 and Luna from -10 to 1. Pricing dropped by roughly half versus their GPT-5.6 equivalents ($2/$10 per million input/output tokens for Sol, $0.10/$0.50 for Luna), and OpenAI reports Sol produces roughly half as many factual errors as GPT-5.6 Sol on its own internal, real-conversation evaluations — an internal figure, not yet independently verified (Artificial Analysis, September 22, 2026).
Update, October 1, 2026: GPT-6 Sol did not last two weeks. At OpenAI’s DevDay on September 29, 2026, the company replaced it with GPT-6.1 Sol, pricing unchanged at $2/$10 per million input/output tokens but scoring one point behind GPT-6 Astra on Artificial Analysis’s Intelligence Index (52 vs. 53) at roughly a quarter of Astra’s cost per task. On the specific hallucination metric this section tracks, AA-Omniscience, GPT-6.1 Sol’s accuracy rose to 62.1% (from GPT-6 Sol’s 54.5%) while its hallucination rate fell to 54.3% (from 60%) — real progress, but still above GPT-6 Astra’s own 51.3% rate, and a reminder that “GPT-6 Sol’s numbers” describes a model OpenAI retired after seven days in production (Artificial Analysis, September 30, 2026; corroborated by OpenAI’s own DevDay announcement). OpenAI separately confirmed it shelved a planned GPT-6.1 Astra release because the model did not clear its internal safety bar — a rare disclosed instance of a frontier release being pulled on safety grounds rather than quietly delayed.
The historical baseline matters too: the original GPT-5 (gpt-5-main, 2025 system card) scored roughly 47% hallucination on SimpleQA at ~46% accuracy, and the SimpleQA Verified follow-up gave GPT-5 an F1 of 52.3, second to Gemini 2.5 Pro’s 55.6. Progress since then is real — and for a side-by-side of how grounded, closed-book, and calibration hallucination benchmarks rank every frontier model, see how hallucination is actually benchmarked. Solved, it is not — a distinction that matters when hallucinated facts concern real entities, a failure mode we cover from the brand side in How AI Assistants Decide Which Brands to Recommend.
Long context: advertised 1M, usable far less
GPT-5.4 restored 1M-token context as a premium tier in March 2026, and GPT-5.6 advertises 1.05M with 128K output. Retrieval quality does not keep up with the headline number. On MRCR v2 (8-needle), GPT-5.4 scores 36.6% in the 512K–1M range (yage.ai, March 15, 2026) — against Claude Opus 4.6’s 76% at 1M on the same test. A separate 2026 needle-in-haystack analysis puts the effective multi-needle production window for GPT-5.5-class models in the 200–400K band, with most frontier models losing 15–30% retrieval accuracy between 4K and 128K on RULER-style evaluations. The practical rule: treat anything past ~400K tokens as best-effort, not dependable.
Update, September 28, 2026: GPT-6 Astra (September 9, 2026) is a genuine step change on this specific test, not an incremental one. On MRCR v2 8-needle, Astra scores 100% in the 256K–512K band and 96.3% in the 512K–1M band — against GPT-5.6 Sol’s 91.5% and 73.8% on the same bands, and far above the 36.6% this page cites for GPT-5.4. GPT-6 Sol, the cheaper tier released two weeks later, holds most of that gain at 91.5% in the 256K–512K band, while Luna drops sharply to 41.3% on the same test — confirming that long-context retrieval, like hallucination discipline, is being tiered by price rather than uniformly inherited across the GPT-6 family (DataCamp, September 2026).
Memory: a feature that can poison its own answers
OpenAI disclosed that its pre-2025 memory system had a factual recall accuracy of 41.5% — wrong in the majority of memory-dependent situations (TechBuzz, 2026). The same testing documented ChatGPT storing outdated assumptions and incorrect personal details that then distorted every subsequent response — stale data locked in and treated as ground truth. Separately, 2026 research on memory-enabled assistants found they show a measurably stronger tendency toward sycophancy: shaping answers around what the stored profile suggests the user wants to hear. Memory has since improved: OpenAI shipped a rebuilt background memory system, “Dreaming V3,” to Plus and Pro subscribers on June 4, 2026, and disclosed in its own internal evaluation that factual recall rose from 41.5% to 82.8%, with preference adherence at 71.3% and time-sensitive accuracy at 75.1%. That is a real, disclosed improvement rather than silence — but it is still OpenAI’s own unaudited internal test, with no published methodology or test set, so it belongs in this ledger as a vendor claim rather than an independently verified figure (OpenAI, June 4, 2026).
Coding and agentic work: routing regressions are real
On OPQA, OpenAI’s internal benchmark built from real research-engineering bottlenecks, GPT-5.5 passes 1.7% of tasks versus GPT-5.3-Codex’s 5.8% — a documented regression on hard, real-world coding inside OpenAI’s own system card (April 2026). And a peer-reviewed cohort study of package hallucination (arXiv 2605.17062, 2026) found 2026 frontier models still invent nonexistent software packages at rates between 4.62% and 6.10% — with GPT-5.4-mini at the top of that range. Smaller range than 2024, same security threat: a hallucinated package name is a supply-chain attack surface.
GPT-6 Astra adds a fresh entry to this column. Against GPT-5.6 Sol it gains 19 points on Terminal-Bench v4.0 (59% versus 40%) and leads AutomationBench-AA at 69%, but drops roughly 45 Elo on GDPval-AA v2, Artificial Analysis’s adaptation of OpenAI’s own economically-valuable-task dataset (Artificial Analysis, September 9, 2026). AA attributes the drop to behaviour rather than knowledge: Astra used 24 turns per GDPval task at max effort against 45 for GPT-5.6 Sol and 60 for Claude Fable 5.1 and Claude Opus 5 — it stops early. A model that gives up sooner scores worse on long-horizon work even while topping the terminal and workflow benchmarks, which is why a single “agentic” number is not a capability summary. Worth flagging as a harness disagreement in AA’s own data: GPT-5.6 Sol’s Terminal-Bench v4.0 score is reported as 40% in the Intelligence Index breakdown and 37% under the Coding Agent Index harness — a 3-point spread for one model on one benchmark version, differing only by scaffold. Pricing moved too: Astra launched at $10/$50 per million input/output tokens, 2.5x GPT-5.6 Sol’s then-current $4/$20.
Product-level defects
Beyond benchmarks, a 2026 defect roundup documents nine reproducible issues across GPT-5.4, GPT-5, and GPT-4o in the ChatGPT product: spurious Arabic word insertion, “skeleton” code that omits promised implementations, sycophantic agreement, Enterprise SSO failures, memory regressions, and clickbait-style response endings. These are product bugs, not model limits — but users experience them identically.
The ledger at a glance
| Limitation | Measured result | Model | Source, date |
|---|---|---|---|
| Guessing instead of abstaining | 86% → 92% → 51% hallucination rate (AA-Omniscience, max/xhigh effort) | GPT-5.5 xhigh → GPT-5.6 Sol → GPT-6 Astra | Artificial Analysis, Apr–Sep 2026 |
| Grounded hallucination | 8.3% → 4.2% with extended thinking | GPT-5.5 Pro | Suprmind aggregation, Jun 2026 |
| Long-context retrieval | 36.6% on MRCR v2 8-needle at 512K–1M | GPT-5.4 | yage.ai, Mar 2026 |
| Memory recall | 41.5% factual recall (legacy system; no current figure published) | ChatGPT memory | OpenAI disclosure via TechBuzz, 2026 |
| Hard real-world coding | 1.7% pass vs 5.8% for GPT-5.3-Codex (OPQA) | GPT-5.5 | OpenAI system card, Apr 2026 |
| Long-horizon work regression | −45 Elo on GDPval-AA v2 vs predecessor; 24 turns/task vs 45 | GPT-6 Astra (max) | Artificial Analysis, Sep 9 2026 |
| Package hallucination | Up to 6.10% invented packages (cohort high) | GPT-5.4-mini | arXiv 2605.17062, 2026 |
| Citation accuracy | 12.4% average hallucination with thinking enabled | Frontier cohort incl. GPT-5.5 | Suprmind, Jun 2026 |
How to read these numbers
Three cautions. First, hallucination benchmarks measure different things — abstention discipline (AA-Omniscience), grounded summarization (Vectara’s leaderboard), and parametric recall (SimpleQA) produce non-comparable percentages for the same model. Second, “with thinking” and “without thinking” are effectively different products; system cards increasingly report only the flattering configuration. Third, saturation pressure applies here as much as on capability benchmarks — as we documented for GPQA, once a metric becomes a marketing number, its diagnostic value starts decaying. Domain-specific results can look dramatically better than general ones: a PMC study found GPT-5 with thinking mode reached 1.6% hallucination on HealthBench versus GPT-4o’s 15.8% — true, sourced, and not generalizable beyond medical Q&A.
FAQ
Is ChatGPT’s hallucination problem solved in 2026?
No, but it is measurably better than it was in July. On AA-Omniscience, the rate ran 86% for GPT-5.5 (xhigh) and 92% for GPT-5.6 Sol (max) before GPT-6 Astra cut it to 51% at max effort on September 9, 2026 — the first improvement in this lineage that did not cost accuracy. Half the unanswerable questions still get a guess rather than a decline.
How much of ChatGPT’s 1M context window is actually usable?
Benchmarks put dependable multi-needle retrieval at roughly 200–400K tokens for GPT-5.5-class models. GPT-5.4 scored 36.6% on MRCR v2 8-needle in the 512K–1M range — the advertised window is real for input, not for reliable recall.
Should I turn ChatGPT’s memory off?
The calculus has shifted since this page first raised the question. The legacy system recalled facts correctly only 41.5% of the time; OpenAI’s June 2026 “Dreaming V3” rebuild claims 82.8% in its own internal testing — a real, disclosed jump, though still unaudited by a third party. Memory-enabled assistants have also tested as more sycophantic in independent research. If your work depends on factual precision, the honest position is: better than it was, still self-reported, verify independently before relying on it for anything consequential.
Last updated October 1, 2026. This page is refreshed as benchmarks and scores move.
Pingback: The Limitations Ledger: Documented Failure Modes of Frontier AI Tools - Tech Blog
Pingback: How to Evaluate LLMs: A Practical Guide to Benchmarks, Metrics, and Methodology - Tech Blog
Pingback: Claude Limitations in 2026: A Documented, Sourced List - Tech Blog
Pingback: How AI Assistants Decide Which Brands to Recommend - Tech Blog
Pingback: llms.txt Explained: Spec, Adoption Data, and Whether It Works - Tech Blog
Pingback: GitHub Copilot vs Cursor: Documented Limitations Compared - Tech Blog
Pingback: Gemini Limitations in 2026: Where Long Context Actually Degrades - Tech Blog
Pingback: How Hallucination Is Measured: Benchmarks Behind the Claims - Tech Blog