Research and ML benchmark
MLE-bench
The repo lists known issues including leaked answers and preparation bugs in specific competitions, with fixes deferred to a v2 in openai/frontier-evals. As of April 2026 the maintainers paused new leaderboard submissions while revising fairness processes.
OpenAI
At a glance
- What the agent must do
- The agent must work through Kaggle machine learning competitions end to end, preparing data, training models and submitting predictions.
- Size
- 75 Kaggle competitions (22 in the Lite split)
- Scoring
- % of competitions earning a Kaggle style medal (Any Medal %)
- How to run it
- pip install -e . with Git LFS data, Kaggle API credentials and a Docker agent image; full data is about 3.3 TB, Lite about 158 GB
Official sources
Other research and ml benchmarks
Run evals on your own agent
Benchmarks compare models and agents in general. To test your own agent, see the eval frameworks comparison.
Get the weekly agent stack update
New official MCP servers, spec changes and harness releases, checked against the source. One email a week, no fluff.
Reviewed Oct 9, 2026.