Braintrust

Braintrust

Hosted platform for tracing AI apps and agents, scoring them with experiments, and monitoring quality in production.

Proprietary (open source SDKs)PythonTypeScriptGoJavaRubyC#v0.45.0

What it supports

  • Evaluates multi-step agents and tool calls: supportedEvaluates multi-step agents and tool callsDocs
  • LLM as a judge scoring: supportedLLM as a judge scoringDocs
  • Tracing: supportedTracingDocs
  • Runs in CI: supportedRuns in CIDocs

Install

pip install braintrust

Links

Other eval frameworks

Get the weekly agent stack update

New official MCP servers, spec changes and harness releases, checked against the source. One email a week, no fluff.

Reviewed Oct 9, 2026.