MLflow (Databricks)

MLflow GenAI Evaluation

Evaluation, LLM judges and tracing for GenAI apps and agents within the open source MLflow platform.

Apache-2.0PythonTypeScriptv3.17.0

What it supports

  • Evaluates multi-step agents and tool calls: supportedEvaluates multi-step agents and tool callsDocs
  • LLM as a judge scoring: supportedLLM as a judge scoringDocs
  • Tracing: supportedTracingDocs
  • Runs in CI: supportedRuns in CIDocs

Install

pip install mlflow

Links

Other eval frameworks

Get the weekly agent stack update

New official MCP servers, spec changes and harness releases, checked against the source. One email a week, no fluff.

Reviewed Oct 9, 2026.