Evals
AI agent evals: frameworks and benchmarks
Two different questions: how do I test my own agent before I ship it, and how good are agents in general? Eval frameworks answer the first. Public benchmarks answer the second. Every tick below links to the vendor page that shows it.
Reviewed Oct 9, 2026. Versions refresh daily. Benchmark scores are not listed because they change weekly; each entry links to its official leaderboard.
Eval frameworks for your own agents
Agent evals means the docs show scoring multi-step runs or tool calls, not only single responses.
| Framework | License | Agent evals | LLM judge | Tracing | CI | Hosted | Latest |
|---|---|---|---|---|---|---|---|
| Inspect AIUK AI Security Institute | MIT | Agent evals: supported | LLM judge: supported | Tracing: supported | ?CI: not yet verified | Self-run | 0.3.277 |
| promptfooPromptfoo (acquired by OpenAI) | MIT | Agent evals: supported | LLM judge: supported | Tracing: supported | CI: supported | Promptfoo Enterprise | 0.124.1 |
| DeepEvalConfident AI | Apache-2.0 | Agent evals: supported | LLM judge: supported | Tracing: supported | CI: supported | Confident AI | 4.2.8 |
| Pydantic EvalsPydantic | MIT | Agent evals: supported | LLM judge: supported | Tracing: supported | CI: supported | Pydantic Logfire | 2.54.0 |
| HarborHarbor Framework (creators of Terminal-Bench) | Apache-2.0 | Agent evals: supported | LLM judge: supported | Tracing: supported | ?CI: not yet verified | Harbor Hub | 0.24.0 |
| LangfuseLangfuse (ClickHouse) | MIT | Agent evals: supported | LLM judge: supported | Tracing: supported | CI: supported | Langfuse Cloud | 4.17.0 |
| OpikComet | Apache-2.0 | Agent evals: supported | LLM judge: supported | Tracing: supported | CI: supported | Opik Cloud | 2.2.96 |
| MLflow GenAI EvaluationMLflow (Databricks) | Apache-2.0 | Agent evals: supported | LLM judge: supported | Tracing: supported | CI: supported | Self-run | 3.17.0 |
| Arize PhoenixArize AI | Elastic-2.0 (source available) | Agent evals: supported | LLM judge: supported | Tracing: supported | CI: supported | Arize AX | 20.20.0 |
| W&B WeaveWeights & Biases (CoreWeave) | Apache-2.0 | Agent evals: supported | LLM judge: supported | Tracing: supported | ?CI: not yet verified | Weights & Biases | 0.53.11 |
| RagasVibrant Labs | Apache-2.0 | Agent evals: supported | LLM judge: supported | ?Tracing: not yet verified | ?CI: not yet verified | Self-run | 0.4.3 |
| Google ADK EvaluationGoogle | Apache-2.0 | Agent evals: supported | LLM judge: supported | Tracing: supported | CI: supported | Self-run | 2.11.0 |
| BraintrustBraintrust | Proprietary (open source SDKs) | Agent evals: supported | LLM judge: supported | Tracing: supported | CI: supported | Braintrust | 0.45.0 |
| LangSmith EvaluationLangChain | Proprietary (open source SDK) | Agent evals: supported | LLM judge: supported | Tracing: supported | CI: supported | LangSmith | 0.14.5 |
| Microsoft Foundry EvaluationsMicrosoft | Proprietary | Agent evals: supported | LLM judge: supported | Tracing: supported | CI: supported | Microsoft Foundry | 2.8.0 |
| Gen AI Evaluation ServiceGoogle Cloud | Proprietary | Agent evals: supported | LLM judge: supported | ?Tracing: not yet verified | ?CI: not yet verified | Gemini Enterprise Agent Platform (formerly Vertex AI) | ? |
Agent benchmarks
What each public benchmark asks an agent to do, how it is scored, and where the official leaderboard lives.
Coding
SWE-bench
2,294 tasks from 12 Python repositoriesPrinceton University (SWE-bench team)
Given a real Python repository and a GitHub issue, the agent must edit the code so the issue is fixed and the hidden tests pass.
SWE-bench Verified
500 tasksOpenAI with the SWE-bench authors (Princeton)
The agent must resolve real GitHub issues in Python repositories, using a subset of SWE-bench that human engineers screened for fair tests and clear issue text.
SWE-bench Multimodal
617 tasks in the paperSWE-bench team (Princeton University and collaborators)
The agent must fix bugs in user facing JavaScript libraries where the issue includes visual elements such as screenshots.
SWE-Bench Pro
1,865 tasksScale AI
The agent must complete long horizon bug fixes and feature work in real business, B2B and developer tool codebases.
SWE-Lancer
Over 1,400 tasks worth $1 million in payoutsOpenAI
The agent must complete real paid Upwork software tasks, either implementing fixes and features or choosing the best proposal as a manager.
Computer use
OSWorld-Verified
369 tasksXLANG Lab, University of Hong Kong (with Salesforce Research, CMU, University of Waterloo)
The agent must complete open ended tasks on a real desktop computer across web and desktop apps, operating the screen with mouse and keyboard.
OSWorld 2.0
108 tasksXLANG Lab, University of Hong Kong and collaborators
The agent must carry out long, realistic computer workflows across apps and self hosted websites, often taking a human more than an hour.
Web browsing
WebArena
812 tasksCarnegie Mellon University
The agent must complete long, realistic tasks on self hosted websites such as a shop, a forum, GitLab and a content manager.
BrowseComp
1,266 questionsOpenAI
The agent must persistently browse the internet to find hard to locate, entangled facts with short, checkable answers.
Tool use and MCP
Toolathlon (Tool Decathlon)
108 tasks across 32 apps and 604 toolsHKUST NLP
The agent must finish multi step workflows across many real apps, such as email, calendars, Kubernetes and BigQuery, mostly through MCP servers.
MCP-Universe
231 tasks across 6 domains and 11 MCP serversSalesforce AI Research
The agent must solve hard tasks by calling real world MCP servers in areas like maps, GitHub, finance, 3D design, browser automation and web search.
MCP-Bench
104 tasksAccenture
The agent must plan and run multi step, cross tool workflows on live MCP servers from fuzzy instructions that do not name the tools.
AppWorld
750 tasksStony Brook University NLP
The agent must write interactive code against 457 APIs across 9 simulated everyday apps to complete tasks for simulated users.
Berkeley Function Calling Leaderboard (BFCL)
Several thousand test entriesUC Berkeley (Gorilla LLM team)
The model must call functions and tools correctly, including multi turn use, memory, web search and avoiding hallucinated calls.
Research and ML
MLE-bench
75 Kaggle competitionsOpenAI
The agent must work through Kaggle machine learning competitions end to end, preparing data, training models and submitting predictions.
Humanity's Last Exam
2,500 questionsCenter for AI Safety and Scale AI
The model or agent must answer expert level, closed ended academic questions across dozens of subjects, some with images.
General and professional work
Harbor-Index
82 tasksHarbor and Terminal-Bench team (Stanford University, Laude Institute)
The agent must solve a compact, hard mix of tasks drawn from many existing agent benchmarks, run through one standard harness.
GAIA
466 questionsMeta FAIR, Hugging Face and AutoGPT
The assistant must answer real world questions that need reasoning, web browsing, file handling and tool use, each with one short correct answer.
METR Time Horizons
METR
Agents attempt software, ML and cybersecurity tasks of known human duration, and the result is the task length at which they succeed half the time.
GDPval
1,320 tasksOpenAI
The model must produce real work deliverables, such as a legal brief or care plan, for 44 occupations across the top 9 US GDP sectors.
Agents' Last Exam
1,500+ task corpus across 55 sub industriesUC Berkeley RDI and RDI Foundation
The agent must complete long, economically valuable professional workflows on a computer, often in real industry software, with verifiable outcomes.
APEX-Agents
480 tasks in the open datasetMercor
The agent must complete long, cross application tasks written by investment banking analysts, management consultants and corporate lawyers inside realistic file and tool environments.
How to pick
- Testing a prompt or agent in CI with assertions: promptfoo, DeepEval or Pydantic Evals.
- Research-grade agent evals with sandboxes and many benchmarks ready to run: Inspect AI or Harbor.
- Tracing production traffic and scoring it later: Langfuse, Opik, Arize Phoenix, LangSmith or Braintrust.
- Built on a framework already: Google ADK, MLflow and Microsoft Foundry include their own evaluators.
- Reading benchmark claims: check which version and subset was used. SWE-bench Verified, OSWorld and Terminal-Bench all changed in ways that make old and new scores hard to compare.
Get the weekly agent stack update
New official MCP servers, spec changes and harness releases, checked against the source. One email a week, no fluff.