General and professional work benchmark

GAIA

Questions are split into 3 difficulty levels with a public validation set and a private test set. The dataset is gated and users are asked not to reshare it in crawlable form to avoid contamination.

Meta FAIR, Hugging Face and AutoGPT

At a glance

What the agent must do
The assistant must answer real world questions that need reasoning, web browsing, file handling and tool use, each with one short correct answer.
Size
466 questions (300 test answers kept private)
Scoring
Quasi exact match accuracy against the ground truth answer
How to run it
Gated Hugging Face dataset; submit a JSONL of answers to the leaderboard space for test scoring

Official sources

Other general and professional work benchmarks

Run evals on your own agent

Benchmarks compare models and agents in general. To test your own agent, see the eval frameworks comparison.

Get the weekly agent stack update

New official MCP servers, spec changes and harness releases, checked against the source. One email a week, no fluff.

Reviewed Oct 9, 2026.