book an evals working session with us
All updates

July 27, 2026

Review traces with AI evaluators

Every trace can now be reviewed automatically, not just the ones your team has time to open.

Updates
  1. Write detections in codeSeptember 7, 2026
  2. Review traces with AI evaluatorsJuly 27, 2026
  3. Replay an agent run step by stepJuly 1, 2026
  4. See cost, latency, and errors across every runJune 16, 2026
  5. Triage in Your EditorMay 29, 2026
  6. Import Your Traces from LangSmith and BraintrustMay 28, 2026
  7. Ask your trace data anythingMay 1, 2026
  8. Keep sensitive data out of your tracesApril 12, 2026
  9. Version your prompts like codeMarch 5, 2026
  10. Instrument multi-turn conversationsJanuary 20, 2026
  11. Collaborate on traces with your teamDecember 10, 2025
  12. See exactly what your AI agent didNovember 18, 2025
  13. Catch issues before your users doNovember 5, 2025
  14. Zero-config observability for every frameworkOctober 22, 2025

How it works

Build custom AI evaluators with your own rubric and verdict schema. Use an LLM judge for a fast one-pass review, or an agent judge when the evaluation needs to inspect span data, call tools, or check external context first.

Agent judges can use connected integrations, custom MCP servers, or allowlisted HTTP fetches during review.

Test an evaluator against a real trace in the playground, then assign it per question alongside your human reviewers.

Evals now live at /evals, with old /feedback links redirecting automatically. Trace and span eval results are also exposed to agents querying over MCP, so evaluation outcomes can feed back into debugging and investigation workflows.

Read the AI Evaluators docs