General and professional work benchmark
GDPval
OpenAI notes tasks are one shot, so they do not test iteration or handling ambiguity, and the automated grader is less reliable than human experts.
OpenAI
At a glance
- What the agent must do
- The model must produce real work deliverables, such as a legal brief or care plan, for 44 occupations across the top 9 US GDP sectors.
- Size
- 1,320 tasks; 220 task open gold subset
- Scoring
- Win or tie rate versus human expert deliverables in blind expert comparisons
- How to run it
- Gold subset on Hugging Face (openai/gdpval) with an experimental automated grader at evals.openai.com
Official sources
Other general and professional work benchmarks
Run evals on your own agent
Benchmarks compare models and agents in general. To test your own agent, see the eval frameworks comparison.
Get the weekly agent stack update
New official MCP servers, spec changes and harness releases, checked against the source. One email a week, no fluff.
Reviewed Oct 9, 2026.