Confident AI

DeepEval

Pytest-style Python framework for unit testing LLM apps and agents with research-backed metrics.

Apache-2.0Pythonv4.2.8

What it supports

  • Evaluates multi-step agents and tool calls: supportedEvaluates multi-step agents and tool callsDocs
  • LLM as a judge scoring: supportedLLM as a judge scoringDocs
  • Tracing: supportedTracingDocs
  • Runs in CI: supportedRuns in CIDocs

Install

pip install deepeval

Links

Other eval frameworks

Get the weekly agent stack update

New official MCP servers, spec changes and harness releases, checked against the source. One email a week, no fluff.

Reviewed Oct 9, 2026.