General and professional work benchmark
APEX-Agents
Prompts, rubrics, gold outputs and files are open source on Hugging Face under evaluation only terms that forbid training on the data.
Mercor
At a glance
- What the agent must do
- The agent must complete long, cross application tasks written by investment banking analysts, management consultants and corporate lawyers inside realistic file and tool environments.
- Size
- 480 tasks in the open dataset; leaderboard page cites 240 tasks across 31 worlds
- Scoring
- Pass@1 against expert written rubrics, plus mean score
- How to run it
- Archipelago, Docker based MCP environment, agent runner and grader
Official sources
Other general and professional work benchmarks
Run evals on your own agent
Benchmarks compare models and agents in general. To test your own agent, see the eval frameworks comparison.
Get the weekly agent stack update
New official MCP servers, spec changes and harness releases, checked against the source. One email a week, no fluff.
Reviewed Oct 9, 2026.