General and professional work benchmark

APEX-Agents

Prompts, rubrics, gold outputs and files are open source on Hugging Face under evaluation only terms that forbid training on the data.

Mercor

At a glance

What the agent must do
The agent must complete long, cross application tasks written by investment banking analysts, management consultants and corporate lawyers inside realistic file and tool environments.
Size
480 tasks in the open dataset; leaderboard page cites 240 tasks across 31 worlds
Scoring
Pass@1 against expert written rubrics, plus mean score
How to run it
Archipelago, Docker based MCP environment, agent runner and grader

Official sources

Other general and professional work benchmarks

Run evals on your own agent

Benchmarks compare models and agents in general. To test your own agent, see the eval frameworks comparison.

Get the weekly agent stack update

New official MCP servers, spec changes and harness releases, checked against the source. One email a week, no fluff.

Reviewed Oct 9, 2026.