Research and ML benchmark

Humanity's Last Exam

Published in Nature in January 2026. HLE-Rolling is a dynamic fork, and HLE-Diamond (September 2026) is a cleaned 1,000 question subset with 500 reasoning and 500 knowledge questions and recommended settings for testing with tools.

Center for AI Safety and Scale AI

At a glance

What the agent must do
The model or agent must answer expert level, closed ended academic questions across dozens of subjects, some with images.
Size
2,500 questions; HLE-Diamond subset of 1,000
Scoring
Accuracy, plus calibration error
How to run it
pip install -r requirements.txt, then run_model_predictions.py and run_judge_results.py (model judge grading)

Official sources

Other research and ml benchmarks

Run evals on your own agent

Benchmarks compare models and agents in general. To test your own agent, see the eval frameworks comparison.

Get the weekly agent stack update

New official MCP servers, spec changes and harness releases, checked against the source. One email a week, no fluff.

Reviewed Oct 9, 2026.