GDPval Explained: Measuring AI on Real Economic Work, Not Trivia

GDPval is OpenAI’s attempt to measure AI against paid human work instead of exam questions, and on the metric that matters most — blind expert grading against real deliverables — the top model has gone from beating human professionals 38.8% of the time in September 2025 to 74.1% of the time fourteen months later, though independent trackers disagree sharply on which model actually leads today.

What GDPval actually tests

OpenAI introduced GDPval on September 25, 2025, as an evaluation built from real work products rather than synthetic questions. The full set contains 1,320 tasks (a 220-task “gold” subset is public) spanning 44 occupations drawn from the nine U.S. industries that contribute the most to GDP: professional services, finance and insurance, healthcare, manufacturing, retail, wholesale trade, government, real estate, and information. Occupations range from lawyers and financial analysts to registered nurses, audio/video technicians, and private investigators.

Each task was built by a professional averaging 14 years of experience, went through roughly five rounds of review, and asks for the kind of deliverable that occupation actually produces — an engineering blueprint, a legal brief, a customer-service email, a PDF investigation report, an Excel quote. Grading is blind: expert graders from the same occupation compare the AI output to the human original and rate it “better than,” “as good as,” or “worse than” the human’s work, per the original paper (Patwardhan et al., OpenAI, October 2025).

This is a meaningfully different design from the multiple-choice or unit-test benchmarks that dominate model cards. For a broader map of where GDPval sits next to agentic benchmarks like tau-bench and OSWorld, see our AI agent benchmarks guide.

The official trajectory: from 38.8% to 74.1%

At launch, OpenAI compared GPT-4o, o4-mini, o3, GPT-5, Claude Opus 4.1, Gemini 2.5 Pro, and Grok 4 against the 220-task gold set. Claude Opus 4.1 was the strongest performer at that point, winning or tying on just under half of tasks, with particular strength in formatting and layout; GPT-5 led on factual accuracy. OpenAI reported that scores had roughly tripled from GPT-4o (spring 2024) to GPT-5 (summer 2025).

The next jump was steep. According to OpenAI’s own release material and contemporaneous coverage (AINews, December 11, 2025), GPT-5.2 Thinking scored a 70.9% win+tie rate against human experts — up from GPT-5’s 38.8% — and GPT-5.2 Pro reached 74.1% (roughly 60% wins plus 14% ties). That is the most recent number OpenAI has published directly against its own gold set as of this writing; the company has not issued a new official win+tie score alongside GPT-5.4, GPT-5.6, or GPT-6 Astra.

Two independent trackers, same ranking, different scale — and a regression story

Because OpenAI doesn’t re-run its official grading for every competitor release, third parties have built their own GDPval implementations using the public gold set. These use a different methodology entirely: instead of grading against the human-written reference, they run models head-to-head against each other, with an LLM judge picking a winner and results aggregated into an Elo rating. That makes the two systems non-comparable in absolute terms, but useful to cross-check in relative terms.

Tracker Methodology Top model (Sept 2026) Score Runner-up
OpenAI official (evals.openai.com) Blind expert grading vs. human deliverable; absolute win+tie % GPT-5.2 Pro (last official figure, Dec 2025) 74.1% win+tie GPT-5.2 Thinking, 70.9%
Artificial Analysis GDPval-AA v2.1 Blind pairwise LLM judging; relative Elo, anchored to human baseline of 1,000 Claude Opus 5.5 (Adaptive Reasoning, Max Effort) Elo 1,846 (new model since our last check; superseded Claude Fable 5.1’s Elo 1,764 lead)
BenchLM.ai GDPval-AA (Sep 25, 2026) Same pairwise design Claude Opus 5 Elo 1,862 Claude Opus 5.5, Elo 1,846; GLM-5.3-Flash, Elo 1,773
LLM-Stats GDPval-AA (~Sep 20, 2026) Same pairwise design Claude Fable 5.1 Elo 1,853 Claude Opus 5 (max), Elo 1,824

Update, late September 2026: the three-way corroboration we noted between trackers has broken down since a new model, Claude Opus 5.5, entered the pool. Artificial Analysis and BenchLM now both put Opus 5.5 in the top two, but they disagree on whether it or the non-“.5” Claude Opus 5 actually leads — Artificial Analysis has Opus 5.5 on top at Elo 1,846, while BenchLM’s own re-run has the plain Claude Opus 5 slightly ahead at Elo 1,862, with Opus 5.5 second at the same 1,846. A separate LLM-Stats tracker instead has Claude Fable 5.1 (the prior leader) still on top at Elo 1,853, with Claude Opus 5 second at 1,824. That is three trackers, three different leaders, all within about 40 Elo points of each other — a tighter spread than the earlier GPT-6 Astra story below, but a reminder that “the leading model on GDPval-AA” is a moving, tracker-dependent target rather than a settled fact. Where the trackers do agree: OpenAI’s own models remain well behind the Anthropic cluster. Artificial Analysis’ leaderboard placed GPT-6 Astra’s strongest configuration at Elo 1,580 (21st) and GPT-5.6 Sol (max) at 1,624 (18th) as of its September update. Multiple outlets covering the GPT-6 Astra launch reported an roughly 80-Elo-point drop on GDPval-AA v2 versus its immediate predecessor — a rare case of a flagship model regressing on this specific benchmark even while improving elsewhere. Anthropic, for its part, has publicized GDPval-AA wins directly: it announced Opus 4.7 topping the leaderboard at 1,753 Elo in April 2026 and Opus 4.8 reaching 1,890 at launch in May 2026, though both figures have since been superseded as newer models entered the pool — a reminder that Elo is relative and shifts every time a new competitor is added, unlike OpenAI’s fixed-reference percentage.

Why GDPval is the closest of its peer group to saturated

Epoch AI’s February 2026 review of three “economic value” benchmarks — GDPval, the Remote Labor Index, and APEX-Agents — found GDPval the furthest along toward saturation, with top scores already at 74% versus 30% for APEX-Agents and under 5% for RLI. Epoch’s analysis also flags specific realism gaps: GDPval tasks are one-shot (no back-and-forth the way real client work involves), evaluation allows web access which introduces link rot (two of the public sample tasks already point to broken URLs), and roughly 5% of tasks assume a point in time that is now in the past, which may make them easier for models that have since learned the actual outcome. Epoch also notes that OpenAI graded its own models with web search and code interpreter tools, but evaluated Claude via the claude.ai web UI rather than Claude Code — a scaffolding choice that likely under-elicits Claude’s real capability, a concern that echoes what we’ve found writing about OSWorld’s low computer-use scores, where scaffold and tool access swing results by wide margins. Epoch AI applies this same scrutiny to its own benchmarks, not just OpenAI’s: its FrontierMath math set went through a June 2026 correction that touched 42% of the problem set, after the original release understated the error rate.

Epoch’s broader conclusion is worth quoting directly: these benchmarks “capture a narrow slice of the relevant jobs, but not the full thing,” comparable in that sense to SWE-bench — useful signal on isolated task execution, not evidence of end-to-end job automation.

What GDPval doesn’t measure

GDPval is one-shot: a single prompt, optional reference files, one deliverable, no revision loop. Real knowledge work is iterative — a lawyer revises a brief after client feedback, an analyst iterates after spotting an anomaly. It also strips out the ambiguity of real jobs: a task is handed to the model fully specified, when in practice a professional often has to figure out what the client actually needs first. OpenAI acknowledges both limitations directly in its release notes and says future versions will add interactivity and ambiguity. For the current state of frontier scores across this and other current benchmarks, see our Q3 2026 benchmark roundup.

FAQ

Is GDPval the same as the Remote Labor Index or APEX-Agents?
No. All three measure “economically valuable” digital work, but GDPval tasks are devised by experts and graded against a human reference (7-hour average completion time), RLI tasks are real freelance Upwork projects graded by expert preference (29-hour average), and APEX-Agents tasks are graded by an LLM against a rubric (2-hour average). Epoch AI’s comparison table shows top scores of 74% (GDPval), 30% (APEX-Agents), and under 5% (RLI) as of February 2026 — the gap reflects both task difficulty and realism, not just model capability.

Why do Artificial Analysis and OpenAI report such different-looking numbers for the same benchmark?
They’re measuring different things. OpenAI’s official score is an absolute win+tie rate against a fixed human reference. Artificial Analysis and BenchLM run models against each other in blind pairwise matches and convert results to a relative Elo (or normalized score), which has no fixed ceiling and shifts every time a new model joins the pool. Use OpenAI’s number to track absolute progress against human quality; use the third-party Elo boards to compare current models against each other.

Does a high GDPval score mean a model can replace a given job?
No, and OpenAI says so explicitly. GDPval tests well-specified, self-contained tasks graded in isolation from company context, prior conversation, or ambiguity resolution. Epoch AI’s assessment is that progress on GDPval should correlate with real utility as an assistant for isolated tasks, not with full occupational automation.

Last updated September 27, 2026. This page is refreshed as benchmarks and scores move.

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top