Research and ML benchmark
Humanity's Last Exam
Published in Nature in January 2026. HLE-Rolling is a dynamic fork, and HLE-Diamond (September 2026) is a cleaned 1,000 question subset with 500 reasoning and 500 knowledge questions and recommended settings for testing with tools.
Center for AI Safety and Scale AI
At a glance
- What the agent must do
- The model or agent must answer expert level, closed ended academic questions across dozens of subjects, some with images.
- Size
- 2,500 questions; HLE-Diamond subset of 1,000
- Scoring
- Accuracy, plus calibration error
- How to run it
- pip install -r requirements.txt, then run_model_predictions.py and run_judge_results.py (model judge grading)
Official sources
Other research and ml benchmarks
Run evals on your own agent
Benchmarks compare models and agents in general. To test your own agent, see the eval frameworks comparison.
Get the weekly agent stack update
New official MCP servers, spec changes and harness releases, checked against the source. One email a week, no fluff.
Reviewed Oct 9, 2026.