General and professional work benchmark
Agents' Last Exam
A living benchmark built with 300+ industry experts that keeps adding tasks toward a 5,000 task target. Snorkel AI also hosts a leaderboard for it.
UC Berkeley RDI and RDI Foundation
At a glance
- What the agent must do
- The agent must complete long, economically valuable professional workflows on a computer, often in real industry software, with verifiable outcomes.
- Size
- 1,500+ task corpus across 55 sub industries; about 150 public tasks
- Scoring
- Per task score from 0 to 1 against a hidden reference, reported as pass rate and score
- How to run it
- ale_run toolkit with an experiment YAML; sandboxes on Google Cloud VMs (recommended), AWS, QEMU/KVM or local Docker
Official sources
Other general and professional work benchmarks
Run evals on your own agent
Benchmarks compare models and agents in general. To test your own agent, see the eval frameworks comparison.
Get the weekly agent stack update
New official MCP servers, spec changes and harness releases, checked against the source. One email a week, no fluff.
Reviewed Oct 9, 2026.