General and professional work benchmark

GDPval

OpenAI notes tasks are one shot, so they do not test iteration or handling ambiguity, and the automated grader is less reliable than human experts.

OpenAI

At a glance

What the agent must do
The model must produce real work deliverables, such as a legal brief or care plan, for 44 occupations across the top 9 US GDP sectors.
Size
1,320 tasks; 220 task open gold subset
Scoring
Win or tie rate versus human expert deliverables in blind expert comparisons
How to run it
Gold subset on Hugging Face (openai/gdpval) with an experimental automated grader at evals.openai.com

Official sources

Other general and professional work benchmarks

Run evals on your own agent

Benchmarks compare models and agents in general. To test your own agent, see the eval frameworks comparison.

Get the weekly agent stack update

New official MCP servers, spec changes and harness releases, checked against the source. One email a week, no fluff.

Reviewed Oct 9, 2026.