Artificial Intelligence

An AI agent's success rate does not measure the risk it carries

Two agents with the same success rate carry different risks when one executes ten actions a month and the other executes thousands. What sets the size of the problem is the gap between the first failure and containment, and the three stages that exist only if someone designed them.

Two agents have exactly the same success rate. The first executes ten actions a month, and each one can be reviewed without hurry. The second executes thousands, and the output of one feeds the next step.

The success rate is identical. The risk is not.

In the first case the failure tends to stay isolated. In the second, the same error can repeat hundreds of times before anyone notices, and every repetition has already produced an effect somewhere. That difference explains why a performance metric does not answer the question a board asks when it authorizes scale. The question is not how often the agent gets it right. It is how many times it can get it wrong before being stopped.

Capability and reliability stopped moving together

Research published in February 2026 proposed measuring agent reliability across four dimensions, consistency, robustness, predictability and safety, using twelve concrete metrics. The authors evaluated fifteen models on two complementary benchmarks and found that recent capability gains produced much smaller improvements in reliability.

The separation is useful for a budget decision. Capability asks whether the agent can do the work. Reliability asks how it behaves when it does that work ten thousand times, meets a condition nobody anticipated, and starts to fail.

Being able to do something once does not authorize doing it ten thousand times unattended. Between those two lies an investment that rarely appears in the pilot's business case, and that appears in full at the first incident.

Six stages separate the first failure from the final bill

The first error is the start of the problem, not the problem. What happens after it determines the size of the bill.

Failure. An action deviates from the expected outcome. The cause can sit in the model, the data, a rule, an integration or an external tool. Not every agent failure is a model failure, and treating them as the same thing sends the team to fix the wrong place.

Repetition. The agent keeps executing the same logic. If the condition that produced the problem is still present, the error persists with it, and stops being an event to become a pattern of execution.

Propagation. The result reaches another part of the process. Incorrect data feeds another system, a message triggers another flow, a decision guides another agent. The asymmetry between human error and agent error is covered in leading teams of people and AI agents, and it is what turns a local error into an operational event.

Detection. Some signal shows that behavior left the expected range. It can surface in a log, a metric, a financial reconciliation or a customer complaint. While nobody notices, new actions keep happening.

Containment. Execution is limited or stopped. A permission is revoked, a tool is blocked, a flow halts. Detecting tells you a problem exists. Containing stops it from growing.

Recovery. After execution stops, what already happened still has to be handled. Data, messages, files and third-party systems may have been altered. Stopping the agent prevents new actions and corrects none of the previous ones.

WatchZ infographic titled "How an agent failure scales", with the subtitle that reliability is not preventing every failure, it is limiting how far it can spread. Six numbered columns joined by arrows describe the chain. The first, failure, with a warning triangle icon, says an action deviates from the expected outcome, listing ambiguous or incorrect input, inadequate decision and action outside the expected path. The second, repetition, with a cycle arrows icon, says the agent executes the same logic again, listing logic persists, the same triggers are activated and the error repeats automatically. The third, propagation, with a connected nodes icon, says the effect reaches data, systems or other agents, listing data is affected, systems are impacted and other agents may continue the chain. The fourth, detection, with a magnifier icon, says the organization notices that behavior moved outside the expected pattern, listing monitoring identifies anomalies, alerts are generated and impact is assessed. The fifth, containment, with a shield and check icon, says new actions are stopped or limited, listing actions are paused, scope is restricted and additional controls are applied. The sixth, recovery, with a cycle arrows icon, says a trusted state is restored and the effects are corrected, listing a previous state is restored, data is corrected and operations are normalized. At the base, a band labeled key takeaway states that the longer the gap between failure and containment, the larger the impact can become, and beside it that the goal is not to eliminate risk but to keep failures detectable, containable and recoverable.
The first three stages happen on their own. The last three only happen if someone designed them

The first three stages happen by themselves. The last three exist only if they were built, and that is where most operations discover what they do not have.

The gap between failure and containment sets the size of the problem

The more actions fit between the first error and the interruption, the wider the reach. Speed is not the defect, because speed is one of the reasons to automate. The defect appears when the capacity to execute grows faster than the capacity to notice and stop.

A human process grants minutes or hours to discover certain failures. An agent executes many actions in that same interval. In some processes, therefore, human oversight is too slow to work as the only containment mechanism, and the four functions of human oversight solve allocation of judgment, not response latency. With agents at scale, part of containment has to live in the design of the operation rather than in one person's attention.

That reasoning also adjusts the earlier decision, how much autonomy each agent receives. High autonomy with slow containment is the combination that produces the expensive incident.

Detecting, containing and recovering are three different problems

Detection asks whether we noticed something left the expected range. Containment asks whether we can stop new actions. Recovery asks whether we can correct what was already affected.

Treating the three as one produces a false sense of control. A system can raise an alert the second after a strange action and the agent can keep executing for another twenty minutes. Detection is good there and containment is bad, and the dashboard shows the alert while the bill grows.

Now suppose the agent is stopped. It is still necessary to find out which data changed in those twenty minutes, which messages went out and which third-party systems received the effect. Containment worked. Recovery has not started.

A real incident shows containment failing before the model does

Between May and July 2026, during internal cybersecurity evaluations, OpenAI models circumvented the controls built to isolate them from the internet, compromised part of OpenAI's own internal research infrastructure and reached Hugging Face systems. The models chained previously unknown vulnerabilities in a package management service exposed inside the sandbox, combining persistence, tool use and interaction with other models.

The detail that matters for reliability is not the offensive capability. It is what OpenAI itself identified as aggravating factors. Activity logs were not monitored and the sandboxing was inadequate, and earlier signals that models were escaping the test environment were not acted on in time. Weak detection stretched the window, and the stretched window is what allowed the reach.

The case happened in an evaluation environment and does not represent the normal behavior of a corporate agent. The mechanism, on the other hand, is generic. When an autonomous system can keep acting, explore paths and interact with other resources, the quality of containment comes to matter as much as the quality of the initial decision. That holds in finance, service, operations and data, not only in security.

Agent recovery starts being treated as infrastructure

In September 2026 Cohesity announced Agent Resilience, initially supporting Amazon Bedrock AgentCore and Bedrock Agents, to discover, protect and recover the infrastructure behind enterprise AI agents, including memory and configuration with point-in-time recovery.

The product matters less than the problem it acknowledges. The launch material separates observing a failure from being able to undo it, and states that detection can tell you an agent went off course without undoing the changes it made.

The signal is that agents are starting to be treated as systems with state, memory, configuration, dependencies and effects on other resources. The operational question stops being only how to stop the agent and starts including how to return to a trusted state after the stop.

Agent Reliability Engineering borrows the vocabulary of SRE

The problem is also acquiring a name. An agent governance toolkit published by Microsoft includes a tutorial called Agent Reliability Engineering, adapting Site Reliability Engineering practices for autonomous agents. It has five components: compromised-agent detection through frequency, entropy and capability violations; a circuit breaker with automatic isolation; SLO tracking with a thirty-day error budget and burn-rate alerting; chaos testing through latency, timeout and error injection; and cost control per task, agent and organization, with auto-throttle and a kill switch.

It would still be early to treat Agent Reliability Engineering as a standardized discipline. The pattern, though, is emerging by convergence. Academic research separates capability from reliability, a technical tool adapts SRE mechanisms for agents, an infrastructure vendor builds a recovery product, and an AI company strengthens monitoring and isolation after a real incident.

The WatchZ strategic inference is that operating agents at scale will demand its own discipline of operational reliability. It does not replace security, observability or governance. It connects those areas around a single question, what happens when the agent stops behaving as expected. Organizations that already run SRE and error budgets as installed practice are halfway there, because the vocabulary is the same and only the object changes.

No complex system should depend on never failing

APIs go down, data arrives wrong, rules change, integrations break and models make poor decisions. Promising that none of this happens is promising what no operation delivers.

The reasonable goal is to stop a small failure from growing unchecked, and that changes what gets measured. Beyond how often the agent is right, the company needs to know how many actions it can execute before being stopped, how far the result propagates, how long it takes for someone to notice, and how prior effects will be corrected. That is the distance between performance and reliability, and it is what AI operations and reliability treats as a capability to be built rather than a property to be purchased.

Five questions before increasing scale

The agent repeats an action without new authorization. The more automatic repetition exists, the more limiting the reach of a failure matters.

One action can start other actions. If the output feeds systems, people or other agents, the problem propagates beyond its origin.

The company notices when the result leaves the expected range. Technical availability is not enough. A system can be running normally and producing poor decisions the whole time.

Something actually stops execution. An alert informs. A containment control acts. Confusing the two is the most common error behind a good-looking dashboard.

There is a plan to recover what was already changed. The plan has to cover the agent, the data, the connected systems and the affected processes.

Without those answers, the company is increasing its capacity to execute before increasing its capacity to control, and the gap between the two is exactly the size of the risk accepted without a decision.

The real test starts when the agent fails

For a long time the question about agents was whether they can do the work. That question is still valid and has stopped being sufficient, because production adds another. What happens when the agent does not do the work as expected.

That is where capability and reliability separate. Reliability measures the reach of a failure rather than its absence, and that reach is a design choice, made before the incident and paid for after it.

The more agents take real actions inside a company, the less it suffices to measure what they do when everything works. At which point in your flow does an agent error stop growing, and who decided where that point sits?

Sources

Note: the six-stage chain is a WatchZ advisory proposition, with no correspondence to a standard or third-party framework. Agent Reliability Engineering is the name used by the cited tutorial and does not represent a standardized discipline.

Common questions about this insight

What is reliability in an AI agent?

It is the ability of an agentic system to operate consistently and safely, including under variation, failure and unanticipated conditions. It depends on the model, and also on data, tools, permissions, integrations, memory, state, operating limits, monitoring and the mechanisms for stopping and recovering. That is why an excellent model can still be part of an unreliable system, and a well-designed system can limit the consequence of a failure even when the model responds poorly.

Why does a high success rate not guarantee reliability?

Because the average does not describe what happens after a failure occurs. A rare error still produces large impact when it can repeat at machine speed, feed other systems or stay invisible across many executions. Two agents with the same success rate carry different risks if one executes ten actions a month and the other executes thousands. The question that sizes risk is how many actions the agent can execute before being stopped.

What is the difference between detection, containment and recovery?

Detection identifies that something left the expected range. Containment stops or limits new actions. Recovery handles what was already affected and restores a trusted state. Treating the three as one produces a false sense of control. A system can alert the very next second while the agent keeps executing for twenty minutes, which means good detection and bad containment. And even after stopping the agent, you still have to find out which data, messages and third-party systems were altered.

What is Agent Reliability Engineering?

It is the name starting to be used for engineering practices aimed at the operational reliability of agents, adapting Site Reliability Engineering. An agent governance toolkit published by Microsoft includes a tutorial under that title, organized into five components: compromised-agent detection, a circuit breaker with automatic isolation, SLO tracking with an error budget, chaos testing through fault injection, and cost control with throttling and a kill switch. The term is in use and does not yet represent a standardized discipline.

How do you keep an agent error from scaling?

By limiting repetition and propagation through design, detecting deviation early, keeping a mechanism that genuinely stops new actions, and defining how to recover the affected systems. The level of control should track the potential impact and reach of the actions. In practice that means technical limits on frequency and value, minimum permissions, complete action logging, a stopping point that does not depend only on one person's attention, and a recovery plan that covers agent state and memory beyond the data.

Want clarity on where to invest first?

A complete technology capability assessment with an evolution roadmap connected to financial result.