All updatesJuly 27, 2026

Review traces with AI evaluators

Every trace can now be reviewed automatically, not just the ones your team has time to open.

How it works

Build custom AI evaluators with your own rubric and verdict schema. Use an LLM judge for a fast one-pass review, or an agent judge when the evaluation needs to inspect span data, call tools, or check external context first.

Agent judges can use connected integrations, custom MCP servers, or allowlisted HTTP fetches during review.

Test an evaluator against a real trace in the playground, then assign it per question alongside your human reviewers.

Evals now live at /evals, with old /feedback links redirecting automatically. Trace and span eval results are also exposed to agents querying over MCP, so evaluation outcomes can feed back into debugging and investigation workflows.

Read the AI Evaluators docs