A short history of machines that fix themselves
What a century of automation can teach us about self-healing agents, and the work that refuses to disappear.

It’s June 18, 1914, and Lawrence Sperry is flying a Curtiss C-2 biplane over the River Seine. Below him, crowds have gathered for an airplane safety competition, where inventors are showing off ways to make the still-new business of flying a little less dangerous.
There are 57 entries. Most of them are improvements to the machinery itself, like better magnetos or better carburetors. Sperry is last on the program, and his invention is stranger. It uses gyroscopes to notice when the aircraft is drifting away from level flight, then moves the controls to correct it. The pilot does not have to make every adjustment himself.
On the first pass, Sperry flies past the judges with both arms raised.
On the second, his mechanic, Émile Cachin, climbs out of the cockpit and walks several feet along the wing. A man standing on one side of a small biplane should pull it sharply off balance, but the aircraft stays level.
On the third pass, Sperry climbs onto the opposite wing.
For a few seconds over the Seine, two men are standing outside a flying airplane… and nobody’s in the pilot’s seat. The crowd erupts.

Sperry wins the competition. His stabilizer cannot repair a broken part, so it is not "self-healing" in the modern sense, but it can notice that something has gone off course and correct it without waiting for a person. That basic idea of a machine that can correct itself turns out to have a very, very long life.
In the 112 years since, systems built on that idea have become much more capable. Modern flight-control systems can hold altitude, follow a route, manage speed, compensate for disturbances, and land an aircraft in weather where the runway may barely be visible to a human.
Today, a great of flying happens automatically.
The cockpit still has two pilots.
The two pilots don’t mean autopilot failed
Autopilot worked so well that pilots no longer have to spend every minute making tiny corrections to keep the aircraft pointed in the right direction. Once the desired outcome is clear, the aircraft can do things like maintain altitude or follow a set path.
The mechanics of flying are now heavily automated, but someone still has to choose where the aircraft is going, decide whether the conditions are safe, react when the plan stops matching reality, and take responsibility if something goes wrong.
For as long as we have built machines, we’ve imagined eventually building one that no longer needs us. What tends to happen instead is that automation takes over the parts of the job we can define clearly, and leaves us with the parts that depend on judgment.
By 2001, the problem was keeping software stable
Computing systems had become large enough that no one person could fully understand the whole stack. An application might depend on an oS, a database, a network, storage, middleware, and hardware from several different companies, all with their own settings. The systems were getting more capable, but they were also becoming more fragile and expensive to operate.
IBM proposed an answer. They called it autonomic computing, named after the autonomic nervous system that regulates heart rate, breathing, digestion, and body temperature without requiring us to consciously think about them. IBM imagined computers doing the same. A few months later, in The Vision of Autonomic Computing, Jeffrey Kephart and David Chess described systems that would monitor themselves, analyze what they saw, plan a response, and carry it out using a shared body of knowledge.
A lot of what we now call an autonomous loop was already there. But so was a part of the proposal that’s mostly forgotten now: these systems would manage themselves within objectives set by people. Kephart and Chess expected autonomy to expand gradually. Systems would start by gathering information, then move toward recommending actions and eventually making more consequential decisions. But the objectives themselves would still come from people.
That was the division of labor autonomic computing proposed. The machine could increasingly decide how to keep itself running, and a person still had to decide what running correctly meant.
Can self-healing actually work?
When a technological promise doesn’t arrive in the form people originally imagined, we tend to think of it as a failure.
Self-healing did arrive, though, and it works so well that it feels ordinary. Kubernetes, which orchestrates a large share of the world’s cloud infrastructure, describes self-healing as one of its core capabilities. It restarts failed containers, replaces unhealthy workloads, reschedules them when machines disappear, and keeps trying to make the system match the state its operator asked for. The operator says: three replicas of this service should be running. Kubernetes sees that only two are healthy, so it starts another one. That is self-healing.
What Kubernetes does not do is look at three perfectly healthy replicas and ask whether the service should exist at all. Someone still decides that two replicas would be cheaper, or that four would give customers a better experience.
We see the same pattern elsewhere in computing: compilers removed the need to tell processors how to execute every instruction, and cloud platforms removed most of the work of provisioning physical machines. Each abstraction automated more of the work underneath and moved human attention one level up.
For most of computing history, software followed our instructions. And if the instructions said one thing and the machine did another, something had gone wrong.
Then, very recently, software started making decisions of its own.
The agent that did everything right
It’s September 2026. Somewhere in San Francisco, Maya is halfway through her first coffee when she gets a Slack DM: “hey, the agent just told Vela we can’t extend their payment deadline. but we promised them we would.”
Maya leads customer experience, and Vela is one of her company’s oldest customers. A couple of years ago, during a difficult renewal, someone had promised to take care of them whenever procurement delayed payment.
That didn’t happen. The support agent checked Vela’s account, retrieved the billing policy, and explained to them in an email that overdue accounts are suspended after 5 business days. The agent’s output was factually correct but also, in this case, ultimately wrong.
Maya knows that immediately, but she cannot point to the instruction the agent should have followed instead. The exception is not in the knowledge base. It may be buried in an old Slack thread or notes from an old call. Maya doesn’t even know which run produced the answer, much less where the failure sits inside the trace.
So she pings Daniel, the engineer responsible for the agent. He opens the run and checks it step by step. The customer was identified correctly, the right policy was retrieved, the account status was correct, and the response passed every evaluator. Nothing actually failed.
The agent followed the information available to it and still produced an outcome the company did not want. This when the written policy is only part of the story and when promises, one-off exceptions, customer history, and decisions that were made once were never written down anywhere useful.
Eventually, Priya, the account owner finds the original renewal conversation. The promise is real, so someone has to decide what the rule should be going forward. That means deciding how much weight to give customer trust, financial risk, operational complexity, and the promises employees are allowed to make on the company’s behalf.

This is where self-healing agents differ from self-healing infrastructure
Infrastructure starts with an objective set by a person and automates the work required to preserve it. A fully autonomous agent loop is being asked to supply the missing objective too.
The usual answer is some form of automated judge.
An LLM judge can compare an agent’s response with a rubric. An agentic evaluator is more capable. It can retrieve policies, inspect account data, use tools, and examine the path the agent took. That lets a company apply the same standard across far more runs than a human team could review manually.
The standard itself still has to come from somewhere. It may be encoded in a prompt, a policy, a set of examples, a threshold, historical feedback, or earlier decisions about what good looks like. The judge only applies that standard consistently, without guaranteeing that it’s the right one. Even when the business changes, and the rule is outdated.
Better models are better at automating work
A lot of what looks like human judgment today is actually just missing context. The answer exists somewhere, but the system cannot reach it yet. As models improve, agents will get better at reading contracts, searching Slack, checking old tickets, and finding decisions that were made before anyone thought to document them properly.
That means some failures that require a person today will, in the future, turn out to be retrieval problems. But there are two different kinds of missing information too:
- The answer exists, and the system hasn’t found it yet. Better retrieval can help with that.
- The answer doesn’t exist because the business has never made the decision. Retrieval can’t help with that.
In the second case, someone has to decide. The business can make that decision itself or explicitly delegate it to the machine, but either way, a choice still has to be made.
Better models will get much better at finding answers that already exist, and shrink the first category considerably. What they won’t do is erase the difference between finding a decision and actually making one.
The irony of automation
In 1983, almost halfway between Sperry’s flight and today’s agent boom, cognitive psychologist Lisanne Bainbridge published a paper called Ironies of Automation.
Bainbridge described a problem that shows up whenever systems become more automated. People stop doing the routine work, but they are still expected to step in when something unusual happens. The trouble is that the less often they operate the system themselves, the less context they have when that moment finally comes. As the automation gets better, the remaining human intervention gets harder.
That matters for agents too. Imagine a system that runs on its own for three months, then suddenly asks someone to approve a fix they did not investigate, in an area they have barely looked at all quarter. There’s still a human in the loop in the literal sense, but that person may not have enough context to make a useful decision anymore.
One way to preserve oversight is to keep someone watching dashboards and waiting for something odd to happen. But then much of the work has not actually gone away.
A better system should do as much as it can before asking for help. It should connect feedback to the right run, inspect the trace, retrieve relevant precedent, group similar failures, draft a diagnosis, propose a fix, generate tests, implement an approved change, and check whether the change worked.
It should also get better at knowing when a person is actually needed. The goal should be to bring people in only when the system reaches something it cannot legitimately decide on its own. In practice, that comes down to two questions:
- Feedback: Was this actually the right outcome?
- Approval: Is this the change we want to make?
Everything around those questions can be automated. And so, over time, the system should need to ask them less often.
Frequent escalation is actually another failure mode. If a system asks for approval constantly, people start skimming. The human may still be present in the workflow, but the quality of the judgment falls away. And so the better approach is for each answer to become part of the system’s memory of how the business wants to operate. Which exceptions matter? Which trade-offs are acceptable? Which promises should be honored? Which technically correct responses should still never be sent?
Over time, those decisions accumulate. As models commoditize automation, this accumulated judgment becomes more valuable.
In a world where every company has access to capable models, what those models won’t arrive with is an understanding of which decisions are right for this particular business, and why. And anything a system promises to fully automate and execute perfectly for you, it can also do for all your competitors. That’s also why complete autonomy is not the best end-state.
A more realistic version of automation
Back at the San Francisco office, by mid-afternoon, Priya has made a call. If the exception appears in a customer’s approved account record or contract, the grace period applies. Vela keeps its access, and the company will look for similar promises that never made it into their knowledge base.
Now the actual work begins. Daniel has to find the affected runs, gather the evidence, work out what the agent missed, draft the change, test it, ship it, and then watch to see whether it actually fixes the problem.
Most of that work does not require another business decision. Therefore, it can and should be automated.
- Maya provided one piece: This outcome was wrong.
- Priya provided the other: This is the rule we want.
Once those two things are known, the investigation, implementation, and verification can increasingly happen without pulling more people in.
Six weeks later, another customer with the same exception reaches the account suspension logic. This time, nobody gets a Slack DM. The system recognizes the case, applies the rule Priya already made, and logs what happened. The business already answered the question once, and the system remembers it, so there’s nothing new to decide.
That’s the kind of system we’re building at neatlogs. The goal is to handle as much work as possible without automating away human judgment. It remembers the decisions you’ve already made and asks again only when it encounters something genuinely new.
When Lawrence Sperry lifted his hands above the Seine in 1914, his gyroscopic stabilizer didn’t actually know where the aircraft should go. It only kept the aircraft aligned with a direction Sperry had already chosen, and that was enough to change aviation.
More than a century later, the fact that two pilots still sit in the cockpit only means the machine took over more and more of the work that could be specified clearly, while the decisions that require complex judgment stayed with people.
Software followed the same pattern.
And agents likely will too.