Microsoft

Microsoft Foundry Evaluations

Built-in evaluators and cloud evaluation runs in Microsoft Foundry for scoring agents on task adherence and tool use.

ProprietaryPythonv2.8.0

What it supports

  • Evaluates multi-step agents and tool calls: supportedEvaluates multi-step agents and tool callsDocs
  • LLM as a judge scoring: supportedLLM as a judge scoringDocs
  • Tracing: supportedTracingDocs
  • Runs in CI: supportedRuns in CIDocs

Install

pip install azure-ai-projects

Links

Official documentationHosted option: Microsoft FoundryPricing

Other eval frameworks

Get the weekly agent stack update

New official MCP servers, spec changes and harness releases, checked against the source. One email a week, no fluff.

Reviewed Oct 9, 2026.