Evals

AI agent evals: frameworks and benchmarks

Two different questions: how do I test my own agent before I ship it, and how good are agents in general? Eval frameworks answer the first. Public benchmarks answer the second. Every tick below links to the vendor page that shows it.

Reviewed Oct 9, 2026. Versions refresh daily. Benchmark scores are not listed because they change weekly; each entry links to its official leaderboard.

Eval frameworks for your own agents

Agent evals means the docs show scoring multi-step runs or tool calls, not only single responses.

FrameworkLicenseAgent evalsLLM judgeTracingCIHostedLatest
Inspect AIUK AI Security InstituteMITAgent evals: supportedLLM judge: supportedTracing: supported?CI: not yet verifiedSelf-run0.3.277
promptfooPromptfoo (acquired by OpenAI)MITAgent evals: supportedLLM judge: supportedTracing: supportedCI: supportedPromptfoo Enterprise0.124.1
DeepEvalConfident AIApache-2.0Agent evals: supportedLLM judge: supportedTracing: supportedCI: supportedConfident AI4.2.8
Pydantic EvalsPydanticMITAgent evals: supportedLLM judge: supportedTracing: supportedCI: supportedPydantic Logfire2.54.0
HarborHarbor Framework (creators of Terminal-Bench)Apache-2.0Agent evals: supportedLLM judge: supportedTracing: supported?CI: not yet verifiedHarbor Hub0.24.0
LangfuseLangfuse (ClickHouse)MITAgent evals: supportedLLM judge: supportedTracing: supportedCI: supportedLangfuse Cloud4.17.0
OpikCometApache-2.0Agent evals: supportedLLM judge: supportedTracing: supportedCI: supportedOpik Cloud2.2.96
MLflow GenAI EvaluationMLflow (Databricks)Apache-2.0Agent evals: supportedLLM judge: supportedTracing: supportedCI: supportedSelf-run3.17.0
Arize PhoenixArize AIElastic-2.0 (source available)Agent evals: supportedLLM judge: supportedTracing: supportedCI: supportedArize AX20.20.0
W&B WeaveWeights & Biases (CoreWeave)Apache-2.0Agent evals: supportedLLM judge: supportedTracing: supported?CI: not yet verifiedWeights & Biases0.53.11
RagasVibrant LabsApache-2.0Agent evals: supportedLLM judge: supported?Tracing: not yet verified?CI: not yet verifiedSelf-run0.4.3
Google ADK EvaluationGoogleApache-2.0Agent evals: supportedLLM judge: supportedTracing: supportedCI: supportedSelf-run2.11.0
BraintrustBraintrustProprietary (open source SDKs)Agent evals: supportedLLM judge: supportedTracing: supportedCI: supportedBraintrust0.45.0
LangSmith EvaluationLangChainProprietary (open source SDK)Agent evals: supportedLLM judge: supportedTracing: supportedCI: supportedLangSmith0.14.5
Microsoft Foundry EvaluationsMicrosoftProprietaryAgent evals: supportedLLM judge: supportedTracing: supportedCI: supportedMicrosoft Foundry2.8.0
Gen AI Evaluation ServiceGoogle CloudProprietaryAgent evals: supportedLLM judge: supported?Tracing: not yet verified?CI: not yet verifiedGemini Enterprise Agent Platform (formerly Vertex AI)?

Agent benchmarks

What each public benchmark asks an agent to do, how it is scored, and where the official leaderboard lives.

Coding

Terminal

Computer use

Web browsing

Tool use and MCP

Customer service

Research and ML

General and professional work

How to pick

  • Testing a prompt or agent in CI with assertions: promptfoo, DeepEval or Pydantic Evals.
  • Research-grade agent evals with sandboxes and many benchmarks ready to run: Inspect AI or Harbor.
  • Tracing production traffic and scoring it later: Langfuse, Opik, Arize Phoenix, LangSmith or Braintrust.
  • Built on a framework already: Google ADK, MLflow and Microsoft Foundry include their own evaluators.
  • Reading benchmark claims: check which version and subset was used. SWE-bench Verified, OSWorld and Terminal-Bench all changed in ways that make old and new scores hard to compare.

Get the weekly agent stack update

New official MCP servers, spec changes and harness releases, checked against the source. One email a week, no fluff.