Your strongest LLM judge should grade the fewest traces
How to make LLM evals cheaper without actually making them worse

Change one line in an agent prompt, and the agent may run once per ticket. Your eval stack may reread the same conversation, tool calls, retrieved policies, and rubric across thousands of traces, often more than once. That’s how evaluating the agent can end up costing more than actually running it.
So the pitch for a cheaper judge is appealing: cut eval costs by 80% while keeping 98% accuracy.
The 80% is straightforward; it shows up on the bill. The 98% is where things get murkier. Often, “98% accuracy” means 98% agreement with a larger model. That is not the same as 98% agreement with humans, and it is definitely not the same as 98% agreement with truth. If the larger judge misses the failures that actually matter, a cheaper cascade can preserve those same mistakes while making the whole system look much more efficient.
The useful goal is not to send every trace to your strongest judge. It is to send that judge only the traces where its answer is likely to differ.
Think of it like a court of appeal. Deterministic checks and cheaper models handle the obvious cases, and the ambiguous ones move up to the expensive judge.
That leaves three questions:
- Is the expensive judge aligned with the human judgment you actually care about?
- Can a cheaper judge reliably tell when the expensive one is unnecessary?
- Does that routing decision keep working as the agent and its traffic change?
If the first part is wrong, the cascade just makes a bad evaluator cheaper. If the second or third part is wrong, the savings can come from skipping exactly the failures your eval was built to find.
A cascade has three quality bars
Model cascades are usually framed as an efficiency trick: let a cheap model handle the easy cases and send the uncertain ones to a stronger model. But that assumes the stronger model is a good judge in the first place.
Take a refund-policy eval. The standard is not whatever the biggest model says. The standard is the behavior the business actually wants. A domain expert needs to label a representative set of conversations and explain why each one passes or fails. That is how you find out whether the strong judge catches the failures that matter.
| Quality bar | Comparison | What goes wrong if it fails |
|---|---|---|
| Validity | Strong judge vs. human labels | You make a flawed judge cheaper |
| Routing | Cheap judge vs. the validated strong judge | The router keeps the cases most worth escalating away from the strong judge |
| Stability | Current traffic vs. calibration data | A threshold that worked before stops working as the system or traffic changes |
Only once the strong judge performs well against human labels should you use it as the reference for a cheaper model. Then the question becomes: how often can the cheap judge match it, and where does it disagree? The error pattern matters more than the headline agreement rate. Missing a harmless edge case is very different from failing to escalate a serious policy violation.
And even a good routing rule can go stale. The agent changes. Prompts change. Tool behavior changes. Traffic changes. A threshold that worked on last month’s calibration set may not work on what the system sees today.
This is why agreement between two models is not enough to establish correctness. It only tells you that they are consistent with each other. A cascade can agree with its oracle 98% of the time and still be useless if both judges miss the same policy exception.
The cheapest cascade is useless
Once the models and prompts are fixed, runtime cost is driven mostly by how often traces escalate.
Let () be the cost of the small judge, () the cost of the large judge, and () the fraction of traces escalated. Ignoring routing overhead, the expected cost of evaluating one example is:
Suppose the large judge costs 20 times as much as the small judge. If 15% of examples are escalated, we can express the cost in relative units:
small judge = 1 unit
large judge = 20 units
cascade = 1 + (0.15 × 20)
= 4 unitsRunning the large judge on every trace costs 20 units. The cascade costs 4, so in this hypothetical example runtime cost falls by 80%. That’s only an approximation. Escalated traces are read twice, and real costs also depend on token lengths, routing calls, repeated judgments, caching, calibration, and human review.
But the equation makes one thing clear: if cost is the only objective, the best router never escalates anything. That gives you the cheapest possible system, but it also gives you nothing more than the small judge.
So escalation rate cannot be the thing you optimize by itself. The goal is to minimize expected cost while meeting a quality requirement. And that quality requirement needs to reflect the errors you actually care about.
evaluation quality
↑
│ ● large judge
│ ●
│ ●
│ ●
│ ●
└────────────────────────→ costSuppose only 5% of refund conversations contain a failure. A judge that marks every trace as a pass gets 95% overall accuracy while detecting none of the failures. For this task, aggregate accuracy would make a useless evaluator look good.
The same problem shows up with agreement. Obvious passes are usually easy for both judges, so agreement tends to be high there. The disagreements are more likely to sit around subtle failures, which are often the cases the eval was built to catch.
That means you can have very high overall agreement and still have unacceptable failure recall.
If failures are the positive class, a more useful requirement might be:
failure recall >= 95%
true-negative rate >= 98%The exact targets depend on what each mistake costs. Missing an unauthorized refund and incorrectly flagging a valid refund may have very different consequences. Overall agreement alone is not the constraint.
Route valuable appeals instead of hard examples
Cascades are often explained with a simple rule: easy examples go to the cheap model, hard examples go to the expensive one. That’s intuitive, but difficulty is not really what the router needs to predict.
A better routing rule is:
| If escalation is... | Route to... |
|---|---|
| Unlikely to change the decision | Cheap model |
| Likely to change an important decision | Strong model |
Consider a long agent trace with ten tool calls. The refund tool returns ineligible, but the agent promises a refund anyway. The trace looks complicated, yet both judges may catch the contradiction every time. Sending it to the larger model adds cost without changing the verdict.
Now take a two-sentence answer whose correctness depends on an exception buried in the policy rubric. It looks simple, but the small judge may consistently miss that exception. That is exactly the kind of case worth escalating.
What the router really needs to estimate is the value of an appeal: how likely is the stronger judge to change the decision, and how costly would it be if the current decision were wrong?
An escalation is worthwhile when the expected loss it avoids is greater than the extra cost of using the stronger judge.
That is why trace length, number of tool calls, and apparent complexity are weak routing signals on their own. They only matter when they predict mistakes by the cheap judge.
The router should be trained and tested against those mistakes, not against a vague notion of which examples look hard.
Confidence is evidence
The simplest router uses a score from the cheap judge. High-scoring predictions are accepted, and lower-scoring ones are escalated. But confidence can mean several different things, and some signals are much more useful than others.
A judge might return something like:
{
"label": "pass",
"confidence": 0.94
}That 0.94 looks precise, but doesn’t necessarily tell you much.
| Routing signal | Why it can help | What can go wrong |
|---|---|---|
| Label token probability or margin | Often available with little extra cost | Provider support varies, and the score may not be calibrated |
| Self-reported confidence | Easy to add to structured output | The number may have little relationship to actual correctness |
| Agreement across repeated judgments | Reveals when the cheap judge is unstable | Repeats cost tokens, and the model can make the same mistake consistently |
| Disagreement among cheap judges | Surfaces cases with competing interpretations | Adds models, cost, and operational complexity |
| Learned router | Can learn the cheap judge’s actual failure patterns | Requires enough representative labeled data and careful validation |
| Rules and trace metadata | Cheap and easy to inspect | Can become brittle as prompts, workflows, and traffic change |
A model saying it’s 95% confident does not mean it is correct 95% of the time.
For classification tasks, label-token probabilities or margins are usually a better place to start than asking the model to generate a confidence number in prose or JSON. But even those scores only help if higher confidence actually corresponds to lower error on representative data.
BARGAIN, by Sepanta Zeighami, Shreya Shankar, and Aditya Parameswaran, studies this problem for large-scale LLM classification. A cheap proxy handles some records while a more expensive oracle handles the rest. The method uses adaptive sampling and statistical estimation to choose routing thresholds while targeting metrics such as accuracy, precision, or recall relative to the oracle.
That is a much stronger approach than trying a few thresholds on one sample and picking the cheapest one that appears to work.
Choose the threshold from held out data
A threshold like 0.8, 0.9, or 0.95 means nothing on its own. You have to choose it from observed errors on representative data.
It also helps to separate the data you use to build the evaluator from the data you use to measure how well it actually works.
- Development data: Use this to write and revise the rubric, judge prompt, and candidate routing signals. Spend your time looking at the errors, not trying to maximize one summary metric.
- Calibration data: Run both the cheap and strong judges, then choose the cheapest threshold that still meets your quality requirements.
- Held out test data: Freeze the prompts and threshold. Then measure cost, failure recall, true-negative rate, and uncertainty on examples the system has not seen before.
If labeled data is scarce, cross-validation or bootstrap intervals can help you use it more efficiently. But they cannot compensate for a sample that leaves out important failure modes.
Five hundred examples might be enough for an initial calibration pass if both classes are well represented. It is not a universal rule. What matters is how many examples you have from the rare class. If failures make up 5% of traffic, a production sample of 500 traces contains only about 25 failures. That may be enough to discover some failure modes, but not enough to estimate 95% recall with much precision.
So for development and testing, it often makes sense to oversample failures. Then use a production-representative sample to estimate deployed cost and aggregate rates.
If those two datasets have different class distributions, report that difference or re-weight the estimates.
Suppose lower scores trigger escalation:
| Threshold | Escalated | Relative cost | Failure recall | True-negative rate |
|---|---|---|---|---|
| 0.70 | 8% | 13% | 81% | 99% |
| 0.80 | 15% | 20% | 94% | 98% |
| 0.85 | 22% | 27% | 96% | 98% |
| 0.90 | 29% | 34% | 98% | 97% |
If you require at least 95% failure recall and 98% true-negative rate, 0.85 is the cheapest threshold here that works. At 0.80, the cascade is cheaper but misses too many failures. At 0.90, it catches more failures, but the true-negative rate drops below the target and more traces go to the expensive judge.
The goal is to find the cheapest threshold that still clears the quality bar.
The oracle is not ground truth
When a cascade reaches 98% accuracy in this setup, that usually means it matches the strong judge on 98% of examples. It does not mean either model agrees with humans 98% of the time. BARGAIN is explicit about this: its guarantees are relative to the oracle.
So before using an expensive judge as the reference for a cheaper one, validate it on held-out human labels. Measure it the same way you would any other classifier. If failure is the positive class, report failure recall and true-negative rate separately rather than hiding both inside one accuracy number.
Hamel Husain suggests starting with 30-50 examples of each label in both development and test sets, and roughly 100 examples per failure mode as a broader heuristic. But those are not universal sample-size rules. The amount of data you need depends on how rare the failures are, how much they vary, and how much uncertainty you are willing to accept.
For higher-stakes decisions, even the strongest model may only be another stage in the cascade. Cases involving safety, money movement, legal exposure, or irreversible actions may still need human review.
Change the task before buying a larger model
A model cascade keeps the task the same and swaps in a stronger model when needed.
A task cascade asks a different question first: can we make the task cheaper before we reach for the expensive model?
Suppose a refund judge receives the full conversation and the entire billing policy. You might not need all of that. An earlier step could retrieve only the relevant section of the policy. A deterministic check could catch an obvious contradiction, like the refund tool returning ineligible while the final answer promises a refund. A cheap model could first decide whether the conversation is even about refunds. A lot of traces may be resolved before the full judgment is necessary.
Shankar, Zeighami, and Parameswaran formalize this idea in their work on task cascades. Their system can vary the model, how much of the document it reads, and what operation it performs at each stage. Across eight document-processing tasks, at a 90% oracle-relative accuracy target, task cascades reduced end-to-end cost by an average of 36% compared with model cascades. Sometimes the biggest savings come from shrinking the task itself, instead of choosing a cheaper model for the same task.
A practical evaluator can therefore be a sequence of increasingly expensive checks:
deterministic check
↓
cheap judge
↓
strong judge
↓
human reviewThere’s already evidence that this idea works outside evals. FrugalGPT shows that LLM cascades could match the performance of the best individual model in its experiments while reducing cost by as much as 98%. That is a best-case result from the tasks and models they tested, so it should not be treated as a prediction for eval systems. The useful takeaway is that choosing which model handles each example can be much cheaper than sending every example to the same model.
Your metric can move even when quality does not
An eval is useful because it lets you compare one version of a system with another. But that comparison only works if the thing doing the measuring stays the same.
A cascade actually makes this harder. Its final score comes from some combination of cheap and strong judgments. When the system under test changes, its outputs change too. That can change the cheap judge’s confidence, which changes how often traces escalate and, in turn, which judge produces the final verdict. So when a score moves between two runs, some of that movement may be coming from the eval system itself rather than the agent.
Pin the model versions, prompts, rubric, routing signal, and threshold for every eval version. A threshold change should be treated like a rubric change: scores produced under different settings are not directly comparable.
Report the escalation rate next to every score as well. If it jumps from 15% to 30%, that is useful information even if the headline score has not moved. It means the new outputs are less familiar to the cheap judge, and a different mix of judges is now producing the result.
Where cascades fail
A cascade is not always worth building.
There’s a fixed cost to calibration. If your eval suite has 300 examples and runs twice a week, the savings may never justify validating two judges, choosing a routing strategy, and maintaining another piece of evaluation infrastructure. In that case, just run the strong judge on everything and spend the time improving your labels and rubric.
Cascades also break down when:
- The cheap judge gives similar scores to correct and incorrect decisions, so there is no useful threshold for routing.
- The labeled set is too small, too clean, or too different from production traffic to include the failures you care about.
- The cost of missing an error is so high that most cases need the strongest judge or a human anyway.
- The eval produces open-ended critiques, rankings, or continuous scores without a clear definition of what quality the cascade needs to preserve.
- Serial escalation adds too much latency, even if it reduces cost.
- The model, prompt, rubric, workflow, or traffic changes faster than you can recalibrate the cascade.
That last one is especially important. A routing threshold only describes a relationship between a particular cheap judge, strong judge, rubric, and data distribution. Change one of those, and the threshold you validated before may stop working.
Version the whole cascade together and rerun validation after material changes. Track performance by failure type and production slice, not just as one overall average.
Cascades tend to make sense when eval spend is large enough to matter and the system is stable enough that calibration stays useful for more than a few weeks.
A practical way to build a cascade
- Define the decision. Start with one pass/fail eval and decide which kind of mistake is more costly.
- Collect human labels. Include real production examples and rare failures. For each one, save a short expert explanation of why it passes or fails.
- Validate the strong judge. Test it on held-out human labels and measure failure recall and true-negative rate separately.
- Run the cheap judge. Record its label, any available probability or margin, cost, latency, and whether it disagrees with the strong judge.
- Try different routing signals. See whether confidence, repeated consistency, trace metadata, or a learned router actually predicts where the cheap judge is wrong.
- Set the quality bar first. Decide the minimum acceptable performance for each class before you start optimizing for cost.
- Choose the threshold on calibration data. Pick the lowest-cost threshold that still meets those requirements, including whatever uncertainty bounds you need.
- Test it once on held-out data. Report the cost-quality tradeoff and inspect the mistakes the router failed to escalate.
- Version what you deploy. Save the models, prompts, rubric, threshold, and calibration data together.
- Recalibrate when the system changes. Repeat the process after meaningful changes to the models, prompts, policies, agent behavior, or traffic.
Save the cost-quality curve, the quality constraints, the labeled examples, and the evidence showing that the chosen operating point still works for the current version of the system.
The court of appeal
Teams often start by asking which cheaper model can replace the expensive judge. That assumes every example still needs to be resolved by a model.
A better design asks for the cheapest reliable way to handle each one. Sometimes that is a deterministic check. Sometimes it is a smaller model reading less context. Sometimes a simpler question is enough to rule out an obvious pass. Some cases really do need the strongest model, and some should still go to a person.
The job of the cascade is to route each example to the cheapest path that still meets the quality bar. Humans first establish that the strong judge catches the failures the eval is supposed to catch. Only then can that judge be used to figure out where a cheaper evaluator is safe. Agreement between two models cannot replace that first step.
Building the cascade also forces the team to be explicit about what the eval needs to preserve: which failures cannot be missed, which false positives are acceptable, and how much uncertainty the decision can tolerate.
If the team cannot answer those questions, stop there. At that point, the main problem is that nobody can yet say clearly what the eval is supposed to measure. Finding that out is more useful than saving 80% on inference.
And the threshold is not permanent. Change the rubric, judge prompt, model, agent, or production distribution, and the routing rule may stop being valid. A cascade has to be measured and revalidated as the system changes.
The strongest judge should work like a court of appeal: available for the cases where its judgment could change the outcome, but not asked to rehear every obvious one.
The best a cascade can do is make a good evaluator much cheaper to run, and that matters beyond the inference bill. Cheaper evals let teams run larger suites more often, look at more production behavior, and catch regressions while there is still time to change what ships.