LangChain

LangSmith Evaluation

Evaluation and observability features of the LangSmith platform for testing agents on datasets and scoring production traces.

Proprietary (open source SDK)PythonTypeScriptv0.14.5

What it supports

  • Evaluates multi-step agents and tool calls: supportedEvaluates multi-step agents and tool callsDocs
  • LLM as a judge scoring: supportedLLM as a judge scoringDocs
  • Tracing: supportedTracingDocs
  • Runs in CI: supportedRuns in CIDocs

Install

pip install langsmith

Links

Other eval frameworks

Get the weekly agent stack update

New official MCP servers, spec changes and harness releases, checked against the source. One email a week, no fluff.

Reviewed Oct 9, 2026.