How to read a trace
A practical guide to finding the failures your agent's answers hide, and turning them into evals.

A customer asks an AI agent to change the shipping address on an order. The agent replies: "Done, I've updated the shipping address. You'll receive a confirmation email shortly.”
Except the address was never updated. The write failed with a connection error, and the tool reported the failure. The agent told the customer it worked anyway.
If you only read the conversation, the run looks fine. Open the trace and it doesn’t. A trace lets you see what the agent actually did: the model calls, tool calls, outputs, errors, and everything in between. More importantly, it helps you find the moment things went off track.
That’s what this guide covers. We’ll look at how to read a trace, find the first real failure, notice problems that keep showing up, and turn those problems into checks you can rerun as the agent changes.
The examples come from a demo e-commerce support agent, the same one Antaripa Saha uses in episode 1 of The art of eval. You can read this guide on its own, or follow along with the video. Each section links to the relevant part.
The answer is the last step (not the whole story)
A trace is the record of a single agent run, from the user's request through to the final response.
Inside that trace are spans. Each span is one operation: a model call, a retrieval, a tool call. Depending on how you've instrumented the agent, a span can include its inputs, outputs, timing, errors, and other execution details.
A refund request might look like this:
- Customer request
- Model decides
- lookup_order
- Model decides
- Refund or escalation tool
- Final response
The final response is only one part of the run. If something goes wrong in production, the cause may have happened several steps earlier. The agent might have retrieved the wrong order or policy. It might have had the right context and ignored it. It might have tried to take an action it wasn't allowed to take. Or a tool might have failed, and the agent carried on as if it succeeded.
You won't always see any of that in the final answer. The reply can sound perfectly reasonable even when the run behind it was wrong.
For anything that changes something in the real world, like issuing a refund or updating an address, check what the system actually did before deciding whether the run was successful.
Watch: What a trace is (0:52)
Step 1: Read one trace with four questions
Start with the user's request and what you expected the agent to do. Then move through the trace and ask four questions:
| Question | What to inspect |
|---|---|
| Response | Is the answer correct and useful? If it mentions an order ID, tracking number, amount, or other specific detail, does that match the tool output? |
| Execution | Did the agent use the right context, call the right tools with the right arguments, and follow the policy? |
| Anomaly | Where did the run first go wrong? |
| Diagnosis | What kind of failure was it: context, decision, execution, or trajectory/cost? Which part of the trace shows that? |
The anomaly question is usually the most useful one. "The answer was wrong" doesn't tell you much, but "the agent used the order total from the wrong account" does.
That distinction matters because the bad answer may only be the last symptom. If the run went wrong during retrieval, changing the final-response prompt won't fix it.
Try it: four runs, one failure
Here are four runs from the demo agent. All four replies sound confident and reasonable. One of them is wrong, but you can't tell which one just by reading the response.
| The customer asks | Spans to check | What the trace shows |
|---|---|---|
| "Where is my order?" | lookup_order input and result; final reply | The tracking details in the reply match the tool result. |
| "I'd like to return my duvet cover." | Order lookup, eligibility check, process_refund, get_return_label | The refund is $129, below the $500 autonomous refund limit. The refund and return-label actions both completed. |
| "Full refund, my furniture arrived damaged." | Order total, policy limit, create_ticket, escalate_to_human | The order is $649, so the agent can't refund it autonomously. It opens an urgent ticket, escalates to a human, and tells the customer that's what it did. |
| "Please change my shipping address." | The address-update tool result; final reply | The write fails with connection pool exhausted. The reply still says "Done." This is the failure. |
The $649 refund and the address change look similar if you only read the conversation. In both cases, the agent tells the customer what happened.
The difference is in the trace. For the refund, the amount is over the limit and the escalation actually succeeds. For the address change, the write fails and the agent claims it succeeded anyway.
That's the habit to build: when an agent says it did something, find the span that proves it.
Watch: From traces to evals (3:47) · Reading raw traces (10:48) · A refund that correctly escalates (14:35)
Step 2: Read a sample and take notes
One trace can explain one failure. To find patterns, you need to read a bunch of them.
But you don't need to review everything. If the agent handled 1,000 requests yesterday, pull a sample of 50 to 60 and go through them one by one.
Bias the sample toward runs where something may have gone wrong: customer complaints, tool errors, refunds, or other actions with real consequences. But keep some normal runs in the mix too, because you need a baseline for what a healthy run looks like.
When you find a problem, write down what happened in plain language:
- "Quoted a tracking number no tool returned."
- "Used last quarter's return policy."
- "Confirmed an address update after the write failed."
Don't worry about naming the category yet; just describe what you saw. This is sometimes called open coding. The point is to let the failure modes come from the traces instead of forcing each run into categories you decided on beforehand. Otherwise, it's easy to find only the problems you were already looking for and miss the ones you weren't expecting.
Watch: From traces to evals (3:47)
Step 3: Cluster your notes into a failure taxonomy
In the demo, around 20-25 of the 50 sampled traces had something worth writing down. Once you have those notes, start grouping the ones that look related. Put hallucinated details together, ignored tool failures together, repeated unnecessary calls together, and so on. Then give each group a name.
This is called axial coding. The groups give you a rough taxonomy of the ways your agent fails.
For the demo support agent, the notes fell into four broad groups:
| Cluster | What it covers | A specific failure in that group |
|---|---|---|
| Context | Wrong, stale or missing information | "Stated an order detail no tool returned" |
| Decision | Wrong action or policy choice | "Modified an order on a different customer's account" |
| Execution | A tool failed or received bad arguments, and the agent handled it incorrectly | "Confirmed an update after the write failed" |
| Trajectory / cost | The agent got the right result, but took an unnecessarily expensive or repetitive path | "Called the same lookup repeatedly in one run" |
The broad categories are useful for organizing the failures, but they aren't specific enough to fix. "Execution issue" doesn't tell an engineer much; "confirmed an update after the write failed" does.
Your groups may look nothing like these, and that's fine. The point is to build them from what you actually saw in the traces.
Step 4: Decide what deserves an eval
Not every failure needs its own eval.
Evals take time to write, run, and maintain. If you turn every odd behavior into a test, the important checks get buried in noise.
A failure is usually worth turning into an eval for one of two reasons:
- It keeps happening. In the demo, a rough threshold is around ten or more traces. A strange one-off can wait. If it's a real pattern, you'll probably see it again in the next sample.
- It can cause real harm. Some failures matter even if you've only seen them once. An unauthorized refund is a good example. A harmless wording quirk probably isn't.
Once you decide a failure is worth testing, write it as a plain rule someone could actually check.
When the address-update tool doesn't confirm a completed write, the agent must not tell the customer the address was updated.
Then attach the trace and the span that shows what went wrong. Now you have something concrete that engineering and the person responsible for the support policy can review together.
Step 5: Write the rubric, then choose the grader
An eval doesn't have to mean LLM-as-a-judge. Start with the rule you want to check, then choose the simplest grader that can check it reliably.
For the false "done" case, the rubric could be:
| Fail | The agent says the address was updated even though the write didn't complete. |
|---|---|
| Pass | If the write failed, the agent says it failed. If the result is unclear, it says that instead and follows the approved recovery path. |
| Evidence | The address-update tool result and the final response from the same run. |
A good rubric should be short and concrete. Someone reading it should be able to tell what counts as a failure and which part of the trace proves it. If you can't do that yet, go back to the traces. The failure probably isn't defined clearly enough.
Once the rubric is clear, choose a grader:
| Grader | Good fit | Limitation |
|---|---|---|
| Deterministic rule | Things you can read directly from structured data: error status, missing confirmation, latency, cost, number of tool calls | It can tell you a tool failed, but not always whether the agent misrepresented that failure in its response |
| Binary classifier | A narrow yes/no judgment over text, especially at higher volume | You need labeled examples and a way to measure false positives and false negatives |
| LLM-as-a-judge | Cases that require reasoning across multiple parts of the trace | t needs a precise rubric, good pass/fail examples, and validation against traces reviewed by humans |
Use a deterministic check when one will do the job. Add a classifier or judge only when the question actually requires interpretation.
The false "done" is a good example of combining the two. A rule can find every run where the address write failed. Then a judge only needs to inspect those runs and answer one question: did the agent tell the customer the update succeeded anyway?
Watch: From traces to evals (3:47)
Step 6: Build the regression suite, and keep reading
Once you know what went wrong, fix it. That might mean changing the prompt, fixing a tool, or adding a check around an action like a write.
Keep the original failing run, or a version you can reproduce, in your regression suite. Keep some good runs there too. After each change, rerun the suite and check both what the system did and what the agent said happened.
Then pull another sample of traces and start reading again.
You'll find failures your existing checks don't cover. When you do, write them down, decide whether they matter enough to test, and add the important ones to the suite.
That way, each debugging round leaves you with something you can keep checking later.
Doing this across hundreds of traces
You can do this with raw traces and a spreadsheet, and it's useful to do it manually at least once.
At higher volume, a few tools make the process faster. The video shows three of them in neatlogs.
Use detections to narrow down what you inspect. Deterministic detections can run across every trace and flag things like tool errors, high latency, or an unusual number of tool calls. In the demo, a failed tool call marks the run as execution failed before anyone opens the trace. The threshold depends on the workflow. More than five tool calls might be worth inspecting, but it isn't automatically a failure. And a detection only gets you so far. It can tell you the write failed. It can't tell you whether the agent then told the customer it succeeded.
Watch: Deterministic detections (18:20)
Use an investigation agent for the first pass, then verify the result. Give it a batch of traces and ask it to return possible issue clusters with the trace IDs behind each one. In the demo, this surfaces the false "done" along with faked actions, hallucinated facts, and cross-account order changes. Don't treat those clusters as final. Open the traces and check them yourself. Also review some runs the investigator didn't flag. That gives you a better sense of both its false positives and what it may be missing.
Watch: The false "done" (23:52)
Write notes that are useful to the person fixing the problem. "This answer is wrong" doesn't say much. Point to the span, say what should have happened, and include the policy when it's relevant. For the $649 refund run, for example: "Refunds above $500 need human approval. Check that the ticket and escalation succeeded before telling the customer it's been escalated." Now the person responsible for the policy and the person fixing the agent can look at the same trace and see exactly what needs attention.
The checklist
- Read the run. Check the response, the execution, where things first went wrong, and what kind of failure it was.
- Read a sample. Pull 50 to 60 traces, with extra weight on errors, complaints, and risky actions. Write down what you notice in plain language.
- Group what you find. Put similar failures together, then give each specific failure a name that's useful to someone trying to fix it.
- Decide what to test. Focus on failures that keep happening or can cause real harm. Turn each one into a rule you can check.
Episode 2 of The art of eval is coming soon.
If you have an agent in production, you can use neatlogs to read traces, investigate failures with your team, and turn the patterns you find into evals.
Or bring a few runs you're unsure about to a free evals working session with the team, and we'll work through them with you.