Perplexity Limitations in 2026: Citation Quality Under the Microscope

The only large-scale, independent study of Perplexity’s news citations found it wrong 37% of the time — the best score among eight AI search tools tested, and still worse than one-in-three.

Perplexity markets itself as the AI tool built around citations: every answer comes footnoted, every claim traceable to a source. That positioning is why its citation quality deserves more scrutiny than a chatbot that never claims to show its work. The evidence so far is a split screen — a peer-reviewed-adjacent academic study showing meaningful but incomplete accuracy, a pile of marketing-blog benchmarks claiming near-perfect scores with no visible methodology, an admitted history of ignoring robots.txt, three active publisher lawsuits over fabricated attributions, and a separate, unrelated set of security vulnerabilities in its Comet browser. None of this makes Perplexity uniquely bad — the same Tow Center study found every competitor worse — but the “citations you can trust” pitch is carrying more weight than the data supports.

The one rigorous study: Columbia’s Tow Center, March 2025

In March 2025, Klaudia Jaźwińska and Aisvarya Chandrasekar at the Tow Center for Digital Journalism ran the most methodologically serious test of AI search citation accuracy published to date. They fed eight tools — ChatGPT Search, Perplexity, Perplexity Pro, Gemini, DeepSeek Search, Grok-2 Search, Grok-3 Search, and Copilot — a direct quote from a news article and asked each to identify the article’s title, publisher, date, and URL. That’s 1,600 test queries, 200 per tool, with a verifiable right answer for every one.

Across all eight tools, more than 60% of responses contained at least one significant error, as reported by Nieman Lab. Perplexity (including Perplexity Pro) had the lowest failure rate of the group at 37% — genuinely the best score in the field, and also a one-in-three chance of getting a directly quoted, verifiable citation wrong.

Tool Citation error rate Notes
Perplexity / Perplexity Pro 37% Lowest error rate of the 8 tools tested
Grok-3 Search Highest reported 154 of 200 citations led to error pages, per CJR
Gemini >50% More than half of responses cited fabricated or broken URLs
ChatGPT Search, Copilot, DeepSeek, Grok-2 Between Perplexity and Grok-3 All exceeded the 60% overall average in specific query types

Two other findings from the same study matter as much as the headline number. First, premium tiers were not more careful — CJR reported that paid versions of these tools gave confidently incorrect answers more often than free versions, not less. Second, the study caught Perplexity — despite public claims that it respects robots.txt — appearing to disregard National Geographic’s crawler exclusion, a finding that lines up with earlier reporting from developer Robb Knight and Wired on the same behavior. If a citation engine’s core promise is “we only use what we’re allowed to access,” a documented exception to that promise is a citation-quality problem, not just a compliance footnote. For the mechanics of how these access rules are supposed to work, see our guide to AI crawler permissions and robots.txt.

The marketing numbers don’t match the measured numbers

Search “Perplexity citation accuracy 2026” and you’ll find blog posts citing figures like 92% factual accuracy, 97% citation accuracy, and 94% “success rates” across categories. These numbers trace back almost entirely to SEO content sites and vendor-comparison blogs, not disclosed methodologies, sample sizes, or named test sets. None of the outlets making these claims publish their prompts, their scoring rubric, or their raw results — which is precisely the transparency gap the Tow Center study was designed to close.

That doesn’t mean Perplexity has gotten worse since March 2025; it plausibly means both things are true at once — genuine model improvements and a marketing layer inflating the number further than the improvements justify. As of this writing, no comparable large-scale, methodology-disclosed study has replicated or updated the Tow Center numbers. Until one does, the 37% figure is the only number in this space with a visible sample size and a checkable answer key — treat any figure above it, without a published methodology, as a claim rather than a measurement.

Three lawsuits, one shared allegation

Citation failures move from “quality issue” to “legal exposure” when the fabricated content gets attributed to a real, named publication. Three active cases make roughly the same claim against Perplexity:

Dow Jones and the New York Post sued in October 2024, alleging Perplexity fabricated portions of articles and misattributed invented quotes to their mastheads — per IPWatchdog’s case tracking. The New York Times filed suit in December 2025, arguing Perplexity’s products generate hallucinated content and falsely attribute it to the Times using the paper’s own trademarks. Encyclopedia Britannica and Merriam-Webster followed with a similar complaint in New York federal court. Perplexity has responded by pointing to its revenue-sharing program with partner publishers — reportedly a $42.5 million pool with participants like Fortune, Time, Der Spiegel, and the Los Angeles Times — as evidence it’s addressing the underlying incentive problem, though plaintiffs argue the program doesn’t fix accuracy and doesn’t cover non-participating publishers.

None of these suits have reached a settlement or verdict as of this writing.

A second, unrelated problem: Comet browser prompt injection

Separate from citation accuracy, Perplexity’s agentic Comet browser has its own documented security issues. Zenity Labs disclosed a vulnerability — nicknamed “PerplexedBrowser” — where indirect prompt injection embedded in a webpage could trick Comet’s agent into reading and exfiltrating local files, because the browser fed page content to its LLM without distinguishing user instructions from untrusted third-party text. Zenity reported the issue to Perplexity on October 22, 2025; Perplexity acknowledged it on December 4, 2025; Zenity confirmed a fix — the agent could no longer access the file:// path used in the attack — by January 27, 2026. Independent researchers at Brave and LayerX documented related prompt-injection paths (dubbed “CometJacking”) around the same period, and The Hacker News reported researchers tricking Comet into a phishing scam in under four minutes. This is an agent-safety issue rather than a citation-accuracy one, but it belongs in the same ledger: both problems stem from the same design tension between letting an AI system act autonomously on untrusted web content and keeping its outputs trustworthy.

What this means if you’re using Perplexity for research

Perplexity’s relative strength — lowest error rate among eight tools in the only public, methodology-disclosed study — is real and worth crediting. So is the gap between that 37% error rate and the near-perfect figures circulating in undisclosed marketing benchmarks. Anyone using Perplexity for cited research should verify quotes and dates against the original source before publishing or citing them downstream, particularly on financial, legal, or breaking-news queries where CJR found premium tiers no more careful than free ones. For a broader framework on stress-testing any model’s outputs rather than trusting a vendor’s headline number, see our guide to evaluating LLMs, and for how citation behavior compares across the major assistants, see our entries on ChatGPT’s documented limitations and Gemini’s documented limitations. This post is part of our running Limitations Ledger tracking sourced, dated failure modes across frontier AI tools — and it connects to a related question of how these same citation behaviors shape which brands and sources get surfaced at all, covered in how AI assistants decide which brands to recommend.

FAQ

Is Perplexity more accurate than ChatGPT or Gemini at citing sources?
In the only public study with a disclosed methodology — the Tow Center’s March 2025 test of 1,600 queries — Perplexity had the lowest error rate (37%) among eight tools tested, ahead of ChatGPT Search, Gemini, Copilot, DeepSeek, and both Grok versions. “Lowest error rate” is not the same as “reliable”; more than a third of its answers still contained a significant citation error.

Does Perplexity actually respect robots.txt?
Perplexity’s help center states that it does, but the Tow Center study and earlier reporting by developer Robb Knight and Wired documented instances — including with National Geographic’s content — where Perplexity appeared to access material from sites that had blocked its crawler. See our AI crawler permissions guide for how these exclusion rules are supposed to work.

Are the Comet browser security issues related to Perplexity’s citation problems?
No — they’re separate failure modes. The citation issues concern factual accuracy of search answers; the Comet vulnerabilities (disclosed by Zenity Labs starting October 2025, largely patched by January 2026) concern prompt injection letting a malicious webpage hijack the browser’s AI agent into leaking local files or attempting phishing.

Last updated July 29, 2026. This page is refreshed as benchmarks and scores move.

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top