Llama, DeepSeek, and Mistral: Documented Limitations of Open-Weight Models

The three most-downloaded open-weight model families — Meta’s Llama, DeepSeek, and Mistral — each carry a documented, source-checked limitation that no leaderboard score erases: Llama 4’s headline benchmark ranking came from a customized model Meta never shipped, DeepSeek’s safety guardrails recorded a 100% jailbreak success rate in independent testing, and Mistral Large 3 pairs a 256K-token context window with below-median intelligence scores and a sub-24% SimpleQA rate. None of this means open weights are unusable — DeepSeek V4 Pro is genuinely cost-competitive and Mistral’s context handling holds up structurally. But “open” and “unverified” have overlapped often enough in 2025–2026 that each claim below needs its own citation. This entry extends our Limitations Ledger — the running record of documented failure modes across frontier AI tools — to the open-weight side of the market.

Llama 4: the LMArena score belonged to a model nobody could download

When Meta launched Llama 4 Maverick in April 2025, its blog post touted “an experimental chat version scoring ELO of 1417 on LMArena” — a result that put it near the top of the leaderboard. The catch: that variant, labeled Llama-4-Maverick-03-26-Experimental, was tuned specifically for human-preference voting (longer, emoji-heavy answers), and it was never released. The publicly downloadable Maverick weights ranked 32nd on the same leaderboard — roughly a 30-place gap between the marketed score and the shippable model (The Register, April 8, 2025; TechCrunch, April 11, 2025).

LMArena’s own maintainers said publicly that “Meta’s interpretation of our policy did not match what we expect from model providers,” and the organization subsequently banned preference-tuned submission variants (The Register, April 2025). The story resurfaced in January 2026 when former Meta chief AI scientist Yann LeCun told the Financial Times, following his November 2025 resignation, that Llama 4’s results “were fudged a little bit” through differing optimization approaches — a rare on-record admission from inside the company that produced the model. On agentic tasks, separate qualitative analysis has also found Llama 4 Maverick prone to starting with correct reasoning and degrading mid-execution as task length grows (arXiv 2512.07497). For the mechanics of how leaderboard gaming happens more broadly, see our MT-Bench and LLM-as-judge breakdown.

DeepSeek: censorship is measurable, and the jailbreak rate is the highest recorded among frontier-class models

DeepSeek’s limitations split into two separate, well-documented categories. The first is politically scoped refusal behavior: using “thought token forcing” to inspect DeepSeek R1’s reasoning traces, researchers found the model maintains an internal list of topics it will not discuss — Tiananmen Square 1989, Falun Gong, Tibet, Uyghur camps, Taiwan’s status, and Hong Kong — and refuses roughly 85% of prompts touching those subjects. South Korea’s National Intelligence Service formally flagged this asymmetric censorship pattern in its own assessment (Theori, 2026).

The second category is safety guardrail strength, and it’s separate from politics. Cisco’s red-teaming reported a 100% jailbreak success rate against DeepSeek — the company described it as unprecedented among models it classifies as frontier-capable. That result predates DeepSeek V4, so it should be read as a data point on the R1-generation model’s guardrail design, not a static verdict, but no comparably rigorous public re-test of V4’s jailbreak resistance has surfaced yet.

On raw capability, the U.S. government’s own evaluation is the most current and most heavily documented reference point. NIST’s Center for AI Standards and Innovation (CAISI) published a formal evaluation of DeepSeek V4 Pro on May 1, 2026: DeepSeek V4 lags the U.S. frontier by about eight months on CAISI’s aggregated capability measure, and — notably — scores meaningfully lower on CAISI’s held-out benchmarks (ARC-AGI-2 semi-private: 46% vs. Opus 4.6’s 63%; CAISI’s cyber benchmark CTF-Archive-Diamond: 32% vs. GPT-5.5’s 71%) than on the benchmarks DeepSeek selected for its own technical report, where it looks roughly on par with frontier U.S. models (NIST/CAISI, May 1, 2026). CAISI’s report frames this explicitly as a benchmark-selection gap, not an inference error — it reproduced DeepSeek’s own GPQA-Diamond result before running its held-out suite. On the cost side, the same report found DeepSeek V4 Pro cheaper than the comparably capable GPT-5.4 mini on 5 of 7 benchmarks, ranging from 53% less expensive to 41% more expensive depending on the task.

Downloadable weights compound the safety problem structurally: CAISI separately noted that when DeepSeek V4 Pro did refuse a narrow reverse-engineering task, a small number of repeated attempts bypassed the refusal — and because the weights are downloadable, that refusal isn’t durable once a user can copy and modify the model outside DeepSeek’s hosted service. Our explainer on how hallucination is actually measured covers the adjacent, non-political failure mode — confidently wrong outputs — that shows up across every model family, open or closed.

Mistral Large 3: a big context window doesn’t fix a below-average intelligence score

Mistral Large 3 ships a 256K-token context window and, per independent long-context testing, can operate across that full window without catastrophic degradation — a real structural strength. But on Artificial Analysis’s Intelligence Index, Mistral Large 3 scores 16 against a comparable-model median of 18, placing it below average among open-weight non-reasoning models of similar size, and reviewers additionally flag it as notably slow and verbose relative to its capability tier (Artificial Analysis, 2026).

The more concrete reliability number is on factual recall: independent evaluation puts Mistral Large 3’s SimpleQA score at roughly 23.8%, meaning it answers confidently but incorrectly on a large share of narrow factual questions outside its strongest domains (Vals AI, 2026). That’s a hallucination-calibration problem, not a context-length problem — the two are frequently conflated in vendor marketing, and Mistral’s own numbers are a clean illustration of why they shouldn’t be.

The license fine print that limits who can actually use these weights

“Open weight” doesn’t mean unrestricted. Meta’s Llama Community License requires any company that exceeded 700 million monthly active users in the month before a given Llama version’s release to separately request a license from Meta, which Meta may deny at its sole discretion — a clause aimed squarely at Llama’s largest potential competitors (WCR.legal). Separately, Meta’s license for its multimodal Llama models does not extend the grant to individuals domiciled in, or companies headquartered in, the European Union — end users of a downstream product are exempted, but EU-based developers building directly on the multimodal weights are not (Dion Wiggins, 2025). DeepSeek and Mistral’s core models carry more permissive MIT/Apache-class terms, but DeepSeek’s hosted app has separately been blocked on government devices in the U.S., Australia, Taiwan, Italy (by its data protection authority, in early 2025) and South Korea, and a bipartisan “No DeepSeek on Government Devices Act” was introduced in the U.S. Congress — a deployment restriction distinct from, but compounding, the guardrail findings above.

Where trackers disagree on the open-vs-closed gap

How large the overall open-weight deficit is depends entirely on which tracker you read, and the disagreement itself is informative. Epoch AI puts the average lag at about four months on public benchmarks (Epoch AI). CAISI’s own broader assessment puts leading open-weight PRC models roughly 6–10 months behind on non-public evaluations, narrowing to 4–7 months on cyber-specific tasks as of mid-2026 (UK AISI). Artificial Analysis’s September 2026 Intelligence Index, meanwhile, shows the current best open-weight models (GLM-5.3 and Kimi K3, not covered in this piece) sitting only 9 points behind the closed leaders — a gap description that reads far more optimistic than the CAISI figures because it’s measuring aggregate public-benchmark intelligence rather than held-out capability on cyber and agentic tasks specifically. None of these trackers are wrong; they’re measuring different things, which is exactly the kind of methodology gap covered in our guide to evaluating LLMs.

Model License type Documented limitation Source / date
Llama 4 Maverick Llama Community License (700M-MAU clause; EU multimodal restriction) LMArena’s #2 ranking came from an unreleased, preference-tuned variant; shipped weights ranked 32nd The Register / TechCrunch, Apr. 2025
DeepSeek V4 / R1 MIT (open weights); hosted app banned on government devices in multiple countries ~85% refusal rate on politically sensitive topics; 100% jailbreak success rate reported by Cisco (R1-era); V4 Pro trails U.S. frontier by ~8 months on CAISI’s held-out benchmarks NIST/CAISI, May 1, 2026; Theori, 2026
Mistral Large 3 Mistral Research/Commercial License (Apache-class for smaller models) Intelligence Index score of 16 vs. median 18; SimpleQA accuracy ~23.8% Artificial Analysis; Vals AI, 2026

FAQ

Is Llama 4 actually worse than its launch benchmarks suggested?
On LMArena specifically, yes for the shipped weights: the publicly downloadable Llama 4 Maverick ranked 32nd, not the top-tier position implied by Meta’s marketed 1417 Elo, which came from an unreleased, preference-tuned variant (The Register, April 2025). On other benchmarks — coding and general reasoning — Llama 4 trails both DeepSeek and Mistral rather than leading them.

Is DeepSeek safe to deploy for enterprise use?
That depends on the threat model. DeepSeek V4 Pro is cost-competitive and only about 8 months behind the U.S. frontier on CAISI’s own capability measure, but its refusals have been shown to be bypassable with repeated attempts, and downloaded weights make any refusal permanently less durable than a hosted API’s. Multiple governments now restrict it on official devices independent of the capability question.

Does Mistral Large 3’s 256K context window mean it hallucinates less?
No — context length and factual calibration are different properties. Mistral Large 3 handles its full 256K window without structural collapse, but its SimpleQA score of roughly 23.8% indicates a real tendency toward confident, incorrect answers on narrow factual questions outside its strongest domains.

Last updated September 17, 2026. This page is refreshed as benchmarks and scores move.

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top