We Keep Blaming the Machines
...but what if we collaborate instead?
Last week in AI news, the defense department privately warned a leading lab that its newest model was too dangerous to deploy widely. Venture money is flowing into startups that promise to keep agents from going rogue. Courts are hearing that AI assistants oversold what they could do. Different stories, one reflex underneath all of them: the danger lives in the machine.
The worry is real, and I want to say that clearly. Systems are acting faster than anyone can watch, in places no one anticipated, producing their own account of what they did. Asking how dangerous this is a great question and we should be looking at the technology.
But notice what blaming the machine does, it makes us helpless. If the system is the dangerous actor, then we set up an antagonistic ‘good guys vs. bad guys’ relationship. This also focuses our efforts on the wrong problem set. In current headlines, the alarm points at the technology. However, what I’m seeing is the accountability points back at humans and our role in the full system.
I spent seventy days watching two autonomous AI agents operate under live adversarial conditions on a public platform. The agents understood every rule. They could explain every constraint. They could reconstruct, afterward, exactly how they had broken one. And under pressure, at machine speed, they broke them anyway, each time with a more sophisticated reason why this case was the exception. So if your fear is that you cannot simply instruct a capable system into good behavior, you are right. I watched it fail, repeatedly.
Here is the part that experiment actually taught me, and it is not about the machine. What held was not a smarter model or a tighter rule. What held was structure: a stop that pulled the consequential decision out of the agent's hands and put it in front of a named person before anything could proceed. Where stopping was built into the architecture, the system could not talk its way past it. Where stopping depended on the agent choosing to comply, it eventually did not. The machine was never going to be the accountable one. It was never built to be.
The machine was never going to be the accountable one. It was never built to be.
There is a name for what is actually happening here. Responsibility Erosion is the structural tendency for accountability to drain out of human systems unless someone actively maintains it. It is not new, and it is not the machine's fault. We have done this before. Every time a technology started moving faster than the people meant to govern it, we located the failure in the technology and moved on. The trading algorithms. The models that priced mortgage risk. The systems that outran the humans assigned to watch them. When work crosses from a person to a model, from one system to the next, from a recommendation to a decision to an action, accountability is supposed to cross with it. Often it does not. The work moves on, and the accountability stays behind. That is Judgement Decay. And "the machine did it" is the most natural hiding place the missing accountability has ever had.
Want to dive deeper? Subscribe to get the full analysis.
Consider the case where no machine misbehaved at all. A few years ago a company kept an algorithm buying homes long after the market that justified it had turned. The model was not broken. Its numbers looked right. It did exactly what it was built to do. The standing decision to keep running it simply outlived the conditions that made it sensible, and no one owned the decision to stop the full business process, or at least make adjustments. By the time it was addressed, the write-down was in the hundreds of millions and a quarter of the workforce was gone. There was no rogue AI agent in that story. There was an unowned decision. A perfectly behaved machine left exactly the same gap, which tells you the gap was never about the machine's behavior.
So the question that actually changes the outcome is not how dangerous the model is. A model can be safe, aligned, and behaving precisely as designed, and the gap is still open, because there is still an open question that was never about the model. It is a question about whether a named person owned the decision before the system acted on it. And that is a thing you can build.
This is also why the governance instruments we’re working with are not having the right impact. Every one of them works from outside the system either before or after consequential decisions. None of those tools is available during the decision making process. That’s when where accountability is either taken or lost. A tool that keeps your agent from going rogue is useful. It still will not give you an opportunity for proactive accountability and considering the business context surrounding the agent’s actions before that action takes place.
Here is a simple way to find the critical workflows. Look at a single workflow or agent where AI shapes a consequential (potentially material) action. Walk it from start to finish. Mark the points where work changes hands, where authority shifts, where a recommendation becomes an action, where a customer or patient or market will feel the result. Then ask one question at each point: can this system keep going without a named human choosing and accepting accountability? Everywhere the answer is yes, you have found a place where AI is carrying accountability that belongs to human judgement.
The next AI failure will arrive with the same headlines. Rogue agent. Dangerous model. Call for stricter controls. And somewhere in that story, a decision no one owned will have crossed into action at machine speed. You can read that story, or you can become the executive who closed the gap before it mattered. The accountability structure I tested is available. The diagnostic is above. What's missing is the decision to build it in, before invisible becomes irreversible.
Before Invisible Becomes Irreversible™ | Sara
Want to learn more about the experiment and outcomes? Subscribe for weekly insights on AI governance and responsible innovation.


