Research and ML benchmark

MLE-bench

The repo lists known issues including leaked answers and preparation bugs in specific competitions, with fixes deferred to a v2 in openai/frontier-evals. As of April 2026 the maintainers paused new leaderboard submissions while revising fairness processes.

OpenAI

At a glance

What the agent must do
The agent must work through Kaggle machine learning competitions end to end, preparing data, training models and submitting predictions.
Size
75 Kaggle competitions (22 in the Lite split)
Scoring
% of competitions earning a Kaggle style medal (Any Medal %)
How to run it
pip install -e . with Git LFS data, Kaggle API credentials and a Docker agent image; full data is about 3.3 TB, Lite about 158 GB

Official sources

Other research and ml benchmarks

Run evals on your own agent

Benchmarks compare models and agents in general. To test your own agent, see the eval frameworks comparison.

Get the weekly agent stack update

New official MCP servers, spec changes and harness releases, checked against the source. One email a week, no fluff.

Reviewed Oct 9, 2026.