General and professional work benchmark
GAIA
Questions are split into 3 difficulty levels with a public validation set and a private test set. The dataset is gated and users are asked not to reshare it in crawlable form to avoid contamination.
Meta FAIR, Hugging Face and AutoGPT
At a glance
- What the agent must do
- The assistant must answer real world questions that need reasoning, web browsing, file handling and tool use, each with one short correct answer.
- Size
- 466 questions (300 test answers kept private)
- Scoring
- Quasi exact match accuracy against the ground truth answer
- How to run it
- Gated Hugging Face dataset; submit a JSONL of answers to the leaderboard space for test scoring
Official sources
Other general and professional work benchmarks
Run evals on your own agent
Benchmarks compare models and agents in general. To test your own agent, see the eval frameworks comparison.
Get the weekly agent stack update
New official MCP servers, spec changes and harness releases, checked against the source. One email a week, no fluff.
Reviewed Oct 9, 2026.