UK AI Security Institute

Inspect AI

Python framework for writing and running LLM and agent evaluations, with built-in agents, tools, sandboxing, scorers and a log viewer.

MITPythonv0.3.277

What it supports

  • Evaluates multi-step agents and tool calls: supportedEvaluates multi-step agents and tool callsDocs
  • LLM as a judge scoring: supportedLLM as a judge scoringDocs
  • Tracing: supportedTracingDocs
  • ?Runs in CI: not yet verifiedRuns in CI

Install

pip install inspect-ai

Links

Other eval frameworks

Get the weekly agent stack update

New official MCP servers, spec changes and harness releases, checked against the source. One email a week, no fluff.

Reviewed Oct 9, 2026.