Self-healing is self-defeating
How an agent that fixes every mistake can slowly erase your strategy.

Imagine two expense-management companies, Atlas and Relay. They sell nearly the same product to similar customers. They work with the same banks, and even use the same LLM to review expenses and answer policy questions.
But they disagree about one important thing.
- Atlas believes context matters more than the rule. If an employee made a legitimate purchase but filed it incorrectly, the agent should help them work through it. Atlas charges more partly because it is willing to absorb that ambiguity. Judgment is part of the service.
- Relay believes the rule matters more than the context. Policies are explicit, answers are fast, and exceptions are unusual. Customers choose Relay because they know the same rule will be applied every time. Consistency is part of the product.
Neither approach is obviously better. They’re just two different answers to the same question: what should a good expense-management company do?
Now give both the same self-healing agent system
When the agent gets something wrong, the self-healing system finds the cause, rewrites the instructions, adds a test, ships the change, and watches to see whether the failure happens again.
Both companies give it sensible goals: fewer mistakes, faster resolutions, happier customers, more consistency. The usual stuff.
None of those goals says what Atlas should preserve that Relay should not. None of them says which tradeoffs are intentional, or which kinds of friction are actually part of the product. And that is where the problem begins.
At first, the autonomy is almost entirely useful
The system catches broken tool calls, bad retrievals, contradictory instructions, answers that ramble, and answers that invent policy. Which is great, because nobody at Atlas or Relay wants an agent that forgets context or calculates a refund incorrectly.
These are easy failures, and everyone agrees they are failures. Then the fixes become a little less obvious.
At Atlas, the agent spends a long time investigating unusual claims. Sometimes it makes exceptions that are difficult to explain from the written policy alone. Then the evaluator sees inconsistency, and the system tightens the rules. As a result, resolution gets faster, policy adherence improves, and fewer claims receive special treatment.
At Relay, the opposite problem appears. Strict refusals lead to repeat contacts and frustrated customers. The system softens the responses and sends more borderline cases to a human. As a result, satisfaction improves, fewer customers come back with the same problem, and more claims start receiving manual review.
Every incremental change makes sense: Atlas gets faster and Relay gets more flexible. The metrics improve at both companies, so there’s less reason for anyone to look closely at what’s actually happening.
Six months later, errors are way down, and resolution times are better. Nobody is looking closely or approving individual changes anymore. Atlas and Relay are starting to behave alike.
The line between an error and a choice
We talk about reliability as if correct were always measurable.
Sometimes it is. The agent pulled the right customer record. The arithmetic is right. The policy it cited actually exists. Private data stayed private. Atlas and Relay would agree on all of those.
But once an agent does meaningful work, correctness contains choices. Should it make the exception or enforce the rule? Spend another minute investigating, or give the customer an answer now? These are trade-offs, and companies answer them in pricing and policy long before anyone writes an eval rubric.
Which is why Atlas and Relay can look at the same behavior and want completely opposite fixes. Atlas sees a long investigation and thinks: this is the care people pay us for. Relay sees it and thinks: our product should never require this much interpretation.
The trace can tell both companies exactly what happened. It cannot tell them which company they want to be.
The judgment has to come from somewhere
A self-healing system still needs some definition of better.
There are roughly three places it can get one:
- You define it. Atlas can decide that customer trust matters more than handling time; Relay can put a ceiling on exceptions. The loop runs freely inside boundaries someone chose. And someone still decides what happens when the metrics conflict, and redraws them when the business changes.
- You learn it from history. The system reads old tickets, approvals, and overrides, and learns how the company behaved before. But history can hold expired policies, inconsistent managers, one-off concessions, and decisions nobody would make today.
- You let the model decide. It weighs trust against cost, risk against speed, and chooses what the company ought to do. It might choose well, but it’s now setting the objective (instead of improving against it). That means judgment still happens, but it’s now moved into the machine.
The obvious reply is that the first option should be enough. Nobody says a compiler lacks autonomy because a human wrote the program. Atlas defines its objective, Relay defines its own, and the system runs free inside them. That’s how good software usually works and for a while, it works here too.
The problem is that a specification can only settle trade-offs you thought to specify. The decisions that reveal what a company really values tend to appear in cases nobody planned for: the expense that is technically ineligible but obviously reasonable, or the customer who costs you money today but may matter enormously tomorrow. Those are often the cases where one company’s answer differs from another’s. You can write more rules to include them, but a specification detailed enough to resolve every future trade-off would need to contain the answer to situations the company has never encountered.
So the first option works as long as reality stays inside the boundaries you already drew. Eventually something new happens, and then the system can either ask someone what to do next, or it can decide.
And if it decides, the third option has arrived. Specification degrades at the rate the business encounters new things, which, for interesting businesses, happens all the time.
The median is good enough
The easy way around this is to give the model a broad goal (like improving customer experience, or maximizing resolution), and let its general intelligence handle everything you did not specify.
The result will usually be good, and that’s not always a good thing. Modern models already know what a reasonable apology sounds like. They know when a refund feels fair, and when a refusal seems safe. Wherever company-specific intent is missing, the agent is pulled toward the same priors: the safest tone and the median exception. Atlas and Relay won’t start producing identical words, but may start making increasingly similar choices. That is how a distinctive product can become best in class and less distinctive at the same time.
Better models can make this harder to notice. A weak model exposes missing judgment by getting stuck. It flails, produces something obviously wrong, escalates, and forces someone to notice that nobody ever told it what this company considers correct. The stronger the model, the harder this is to see. Its answer scores well against generic evaluators, so similar answers accumulate.
Hamel Husain says: if a system can identify and fix every problem for you, the same capability will eventually be available to everyone else. If you compete on price, uptime, or strict adherence to a standard, the median is where you want to be. But if customers choose Atlas because Atlas exercises judgment differently from Relay, that difference is why anyone pays Atlas. Remove that difference, and Atlas has no product left.
There’s another way to avoid specifying behavior somewhat. Instead of telling the system what the agent should do, tell it what the business wants: higher retention, more revenue, faster resolution, better satisfaction, lower risk. Then let the loop find what works.
Atlas starts granting more exceptions, and satisfaction improves immediately. Then support costs rise a month later, and fraud increases two quarters after that. A large customer renews because employees love the flexibility. Finance says the policy is becoming impossible to administer.
Did the change work? It depends on which outcome matters most, how much each one matters, and how far into the future you’re willing to look. Those trade-offs still have to live somewhere.
There’s a world in which this argument fails
If companies run highly autonomous loops against broad objectives for 2-3 years and still remain meaningfully different from their competitors on the cases they would once have decided differently, then the pull toward the median is weaker than I think.
I would want to know that. We don’t have that evidence yet, because nobody has been running these systems for long enough.
In the meantime, there is a simpler test you can run on your own agent. Find the situations where your company and your closest competitor would make meaningfully different decisions. Then compare how your agent handled those cases a year ago with how it handles them now.
Has it become more like your company, or more like a very good version of everyone else?
If that gap has narrowed, your dashboards probably will not tell you. The metrics can improve either way.
“Ask just once”
None of this means a person should approve every run.
Imagine a reimbursement request arrives at Atlas two days after the deadline. The policy says: reject it. The employee was traveling for a family emergency and the receipt is valid. A finance lead reads the case, thinks about it, and approves the exception.
That decision contains a small piece of how Atlas thinks: deadlines matter but they aren’t the point of the policy. A documented emergency can justify an exception, and the exception only applies to otherwise valid expenses.
Most organizations loses these pieces. They’re scattered across Slack threads, emails, documents, spreadsheets, and people’s heads. The reasoning isn’t documented, or made available to the model. A month later, something similar happens, and another person reconstructs the same judgment from scratch. Then that reasoning disappears again.
A better system treats the decision as an asset. The first time an unfamiliar case appears, it gathers the evidence and asks the right person. The second time, it retrieves the precedent. Eventually, it handles the case without interrupting anyone (unless something material has changed).
Everything around the judgment should be automated: find the failure, connect the feedback to the run, gather the policy and precedent, propose a diagnosis, draft the fix, generate tests, ship the approved change, watch whether it held. Only when the evidence runs out and a genuinely new choice appears should the system find the person with the authority to make it.
That’s the realistic, more honest version of a self-healing loop:
- Automate everything that isn’t a choice.
- Ask someone when it is.
- Remember the answer.
That’s what we are building at neatlogs: a system that automates the work around judgment, then turns each judgment into context the agent can use again. The goal is an agent that only needs you once.