GDPval Explained: Measuring AI on Real Economic Work, Not Trivia

GDPval is OpenAI’s attempt to measure AI against paid human work instead of exam questions, and on the metric that matters most — blind expert grading against real deliverables — the top model has gone from beating human professionals 38.8% of the time in September 2025 to 74.1% of the time fourteen months later, though independent trackers disagree sharply on which model actually leads today.

What GDPval actually tests

OpenAI introduced GDPval on September 25, 2025, as an evaluation built from real work products rather than synthetic questions. The full set contains 1,320 tasks (a 220-task “gold” subset is public) spanning 44 occupations drawn from the nine U.S. industries that contribute the most to GDP: professional services, finance and insurance, healthcare, manufacturing, retail, wholesale trade, government, real estate, and information. Occupations range from lawyers and financial analysts to registered nurses, audio/video technicians, and private investigators.

Each task was built by a professional averaging 14 years of experience, went through roughly five rounds of review, and asks for the kind of deliverable that occupation actually produces — an engineering blueprint, a legal brief, a customer-service email, a PDF investigation report, an Excel quote. Grading is blind: expert graders from the same occupation compare the AI output to the human original and rate it “better than,” “as good as,” or “worse than” the human’s work, per the original paper (Patwardhan et al., OpenAI, October 2025).

This is a meaningfully different design from the multiple-choice or unit-test benchmarks that dominate model cards. For a broader map of where GDPval sits next to agentic benchmarks like tau-bench and OSWorld, see our AI agent benchmarks guide.

The official trajectory: from 38.8% to 74.1%

At launch, OpenAI compared GPT-4o, o4-mini, o3, GPT-5, Claude Opus 4.1, Gemini 2.5 Pro, and Grok 4 against the 220-task gold set. Claude Opus 4.1 was the strongest performer at that point, winning or tying on just under half of tasks, with particular strength in formatting and layout; GPT-5 led on factual accuracy. OpenAI reported that scores had roughly tripled from GPT-4o (spring 2024) to GPT-5 (summer 2025).

The next jump was steep. According to OpenAI’s own release material and contemporaneous coverage (AINews, December 11, 2025), GPT-5.2 Thinking scored a 70.9% win+tie rate against human experts — up from GPT-5’s 38.8% — and GPT-5.2 Pro reached 74.1% (roughly 60% wins plus 14% ties). That is the most recent number OpenAI has published directly against its own gold set as of this writing; the company has not issued a new official win+tie score alongside GPT-5.4, GPT-5.6, or GPT-6 Astra.

Two independent trackers, same ranking, different scale — and a regression story

Because OpenAI doesn’t re-run its official grading for every competitor release, third parties have built their own GDPval implementations using the public gold set. These use a different methodology entirely: instead of grading against the human-written reference, they run models head-to-head against each other, with an LLM judge picking a winner and results aggregated into an Elo rating. That makes the two systems non-comparable in absolute terms, but useful to cross-check in relative terms.

Tracker Methodology Top model (Sept 2026) Score Runner-up
OpenAI official (evals.openai.com) Blind expert grading vs. human deliverable; absolute win+tie % GPT-5.2 Pro (last official figure, Dec 2025) 74.1% win+tie GPT-5.2 Thinking, 70.9%
Artificial Analysis GDPval-AA v2 Blind pairwise LLM judging; relative Elo, anchored to human baseline of 1,000 Claude Fable 5.1 (Adaptive Reasoning, Max Effort) Elo 1,764 Claude Opus 5 (Max Effort), Elo 1,735
BenchLM.ai GDPval-AA (normalized) Same pairwise design, rescaled to 0–100 Claude Fable 5.1 67.7% Claude Opus 5, 66.2%; GLM-5.3, 62.9%

The two third-party trackers agree on rank order — Claude Fable 5.1 first, Claude Opus 5 second, Z AI’s GLM-5.3 third — which is meaningful corroboration given they use independently built pipelines. Where it gets messier is OpenAI’s own models: Artificial Analysis’ Sept 2026 leaderboard places GPT-6 Astra’s strongest configuration at Elo 1,580 (21st) and GPT-5.6 Sol (max) at 1,624 (18th), both well behind the Anthropic and Z AI models at the top. Multiple outlets covering the GPT-6 Astra launch reported an roughly 80-Elo-point drop on GDPval-AA v2 versus its immediate predecessor — a rare case of a flagship model regressing on this specific benchmark even while improving elsewhere. Anthropic, for its part, has publicized GDPval-AA wins directly: it announced Opus 4.7 topping the leaderboard at 1,753 Elo in April 2026 and Opus 4.8 reaching 1,890 at launch in May 2026, though both figures have since been superseded as newer models entered the pool — a reminder that Elo is relative and shifts every time a new competitor is added, unlike OpenAI’s fixed-reference percentage.

Why GDPval is the closest of its peer group to saturated

Epoch AI’s February 2026 review of three “economic value” benchmarks — GDPval, the Remote Labor Index, and APEX-Agents — found GDPval the furthest along toward saturation, with top scores already at 74% versus 30% for APEX-Agents and under 5% for RLI. Epoch’s analysis also flags specific realism gaps: GDPval tasks are one-shot (no back-and-forth the way real client work involves), evaluation allows web access which introduces link rot (two of the public sample tasks already point to broken URLs), and roughly 5% of tasks assume a point in time that is now in the past, which may make them easier for models that have since learned the actual outcome. Epoch also notes that OpenAI graded its own models with web search and code interpreter tools, but evaluated Claude via the claude.ai web UI rather than Claude Code — a scaffolding choice that likely under-elicits Claude’s real capability, a concern that echoes what we’ve found writing about OSWorld’s low computer-use scores, where scaffold and tool access swing results by wide margins.

Epoch’s broader conclusion is worth quoting directly: these benchmarks “capture a narrow slice of the relevant jobs, but not the full thing,” comparable in that sense to SWE-bench — useful signal on isolated task execution, not evidence of end-to-end job automation.

What GDPval doesn’t measure

GDPval is one-shot: a single prompt, optional reference files, one deliverable, no revision loop. Real knowledge work is iterative — a lawyer revises a brief after client feedback, an analyst iterates after spotting an anomaly. It also strips out the ambiguity of real jobs: a task is handed to the model fully specified, when in practice a professional often has to figure out what the client actually needs first. OpenAI acknowledges both limitations directly in its release notes and says future versions will add interactivity and ambiguity. For the current state of frontier scores across this and other current benchmarks, see our Q3 2026 benchmark roundup.

FAQ

Is GDPval the same as the Remote Labor Index or APEX-Agents?
No. All three measure “economically valuable” digital work, but GDPval tasks are devised by experts and graded against a human reference (7-hour average completion time), RLI tasks are real freelance Upwork projects graded by expert preference (29-hour average), and APEX-Agents tasks are graded by an LLM against a rubric (2-hour average). Epoch AI’s comparison table shows top scores of 74% (GDPval), 30% (APEX-Agents), and under 5% (RLI) as of February 2026 — the gap reflects both task difficulty and realism, not just model capability.

Why do Artificial Analysis and OpenAI report such different-looking numbers for the same benchmark?
They’re measuring different things. OpenAI’s official score is an absolute win+tie rate against a fixed human reference. Artificial Analysis and BenchLM run models against each other in blind pairwise matches and convert results to a relative Elo (or normalized score), which has no fixed ceiling and shifts every time a new model joins the pool. Use OpenAI’s number to track absolute progress against human quality; use the third-party Elo boards to compare current models against each other.

Does a high GDPval score mean a model can replace a given job?
No, and OpenAI says so explicitly. GDPval tests well-specified, self-contained tasks graded in isolation from company context, prior conversation, or ambiguity resolution. Epoch AI’s assessment is that progress on GDPval should correlate with real utility as an assistant for isolated tasks, not with full occupational automation.

Last updated September 10, 2026. This page is refreshed as benchmarks and scores move.

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top