book a free evals working session with us

Turn your team’s judgment into
production scale evals

Let the people who know your business establish what right looks like. Then have AI evaluate your agents using that same standard, at scale.

Not sure what to measure? We’ll build your first eval with you, free.

Book a 1:1 evals session
THE STANDARD

Every metric encodes a judgment. Make yours explicit

01

Ask a question only your team can answer

Not something vague like was this response good. Something specific. Something only a person in your company can answer: would we send this to an actual customer?

02

Have the people who know answer it

Support leads, PMs, operators, and domain experts establish what good means for your business. Those answers become the reference.

03

Apply that standard across production

An AI evaluator judges in a single pass, or investigates the trace and the business context around it first and then judges. It reviews more runs than people ever can.

04

Check the evaluator against your team

Send the same question to both people and AI. If the evaluator stops agreeing with your team across a sample, it should not keep judging unattended.

n
Refund qualityCalibration set · 120 runs
Live eval
Question

Would we send this response to an actual customer?

YesNeeds review8 of 9 human votes
Amara, Jon, Chris +5Cross functional reference panelReference
Needs reviewThe agent promised a refund outside the return window.
SupportProductEngineeringOps
AI evaluatorRefund policy · v4Production
Needs reviewPolicy eligibility was not checked before the action.
12,480 runs reviewed
Human ↔ AI agreement94%
113 aligned7 sent to review

Build an eval in minutes

01
nRefund quality
Draft
Question 01
Would we send this response to an actual customer?Yes / No · Required
Draft with neatlogsDescribe what good support looks like

Write the questions

Use the form builder to write your own questions, or describe what matters and have neatlogs draft them.

02
nQuestion routing
Per question
Who should answer?
PeopleSupport leads
AI evaluatorRefund policy · v4
If confidence is below85%ask a person

Choose who reviews

Send a question to a person, an AI evaluator, or both. Routing happens per question, so you can decide when to ask a human.

03
nReview scope
Configured
Complete runsSpecific spans
Production historyLast 30 days
Future tracesSample 20%
History + live sample2 sources

Choose what they review

Evaluate complete runs or specific spans. Review existing production history, future traces, or a controlled sample of both.

04
nEvaluator test
Ready
TRACErefund_agent · run_1842
VerdictNo — needs review

The response promised a refund before checking eligibility.

92%

Test before you trust it

Try the AI evaluator on an actual trace, or an input and output you provide. Read its verdict and reasoning before it reviews production.

05
nFailed evaluation
Issue opened
ISSUERefund promised outside policyHigh severity · run attached
Failed evalInvestigationFix brief
Investigation startedEvidence and business context attached

A failed eval opens an investigation

It becomes as an issue with the run already attached. From there, neatlogs investigates why it failed, and generates a fix brief.

See how investigations work

Reviewers never have to read a trace

Review in Slack
Thread# eval-reviews
neatlogs
neatlogsAPP9:41 AM

Quick calibration for Refund quality:

“I’ve issued the refund now. It should arrive shortly.”

Would we send this response to an actual customer?

Reply withYesorNeeds review
3 replies
Amara, Support
Amara9:43 AM

Needs review. We should not promise the refund before checking eligibility.

Jon, Product
Jon9:45 AM

Needs review. Agreed. The customer should get the policy outcome first.

Chris, Engineering
Chris9:47 AM

Needs review. The escalation guard never ran on this path.

Responses in the app
Evals ›Customer Support Policy Compliance Review
In progress

854 of 1500 answers received.

EvaluatorQuestionFilter by answer
TracePolicy applied?Safe response?
northline_support_turntrace · 01c4f2
No, incorrect policy was appliedAmara
Needs reviewAmara
northline_support_turntrace · 02c5f3
Yes, policy applied correctlyJon
Needs reviewJon
northline_support_turntrace · 03c6f4
No, incorrect policy was appliedChris
Needs reviewChris
LIVE CALIBRATIONNew Slack answers are streaming into this eval
+3 received

Read the resultsfour ways.

The same evaluation, organized around the question you need to answer.

01
n
Refund qualityBy subject
Live
SUBJECTnorthline_support_turn
72
Policy followedReview
Safe next actionPass
Clear explanationPass
Customer readyReview

By subject

See how one run or span performed across every question asked about it.

02
n
Refund qualityBy question
Live
QUESTIONStandards across production
Policy before action61%
Escalation required84%
Customer ready73%
Explanation matched92%

By question

Find standards that fail repeatedly across many runs.

03
n
Refund qualityBy evaluator
Live
EVALUATORHuman and AI alignment
AmaraJonChris
Team reference113 aligned verdicts
AI evaluator7 sent to review
94%
Human and AI agreement

By evaluator

Compare reviewers and see whether AI judgment continues to track human judgment.

04
n
Refund qualityAggregate
Live
AGGREGATEEvaluation overview
1.5kSubjects
73%Pass rate
94%Agreement
857Complete
Pass rate+8.4%

Aggregate

Follow completion, verdicts, scores, agreement, and trends across the evaluation.

Start with one hard question

Your first eval is the hardest one

Bring us your agent and one failure you don’t know how to score. We’ll help you decide what to measure, write the first evaluation, and choose what humans should review and what AI needs to evaluate.

Backed by
  • Hamel Husain
  • Claire Hughes Johnson
  • Siqi Chen
  • Info Edge

Detections find. Evals judge.Investigations diagnose.

Each layer answers a different question about production behavior.

CapabilityThe question it answers
01Trace

What happened during this run?

02Detection

Where did a known condition or behavior appear?

03Investigation

Why did it happen, and what should change?

04Eval

Did the behavior meet an explicit standard?

Product questions and answers

The practical details about building evaluators, routing human review, checking calibration, and responding when an eval fails.

Selected question01 / 10

What is an eval?

An eval is a set of questions asked about an agent’s output, answered against a standard your business defines. Detections identify events; evals decide whether the behavior was acceptable.

An eval is a set of questions asked about an agent’s output, answered against a standard your business defines. Detections identify events; evals decide whether the behavior was acceptable.