An agent working in the finance process reads an invoice, compares it against the purchase order, finds a discrepancy, prepares a payment and executes that payment. Same agent, same technology, across all five steps.
Technical capability exists in all five. Authority does not have to be equal in any of them.
Reading an invoice only reads. Finding a discrepancy produces an analysis. Preparing a payment creates an action that can still be reviewed. Executing the payment moves money and leaves the reach of whoever approved it.
Saying that this agent has high autonomy hides all five differences at once. High autonomy to do what?
This article is about deciding AI agent autonomy by looking at the concrete action, using four factors that can be answered before any execution is allowed. It is the third piece in the hybrid enterprise series. What to automate and what to keep human covers choosing the execution mode for each part of the work. Here the lens narrows one level, and the unit of decision stops being the agent and becomes the action.
Capability answers one question, autonomy answers another
Capability answers whether the agent can do something. Autonomy answers whether it may do it alone. The two answers do not have to match, and treating them as synonyms opens most of the problems that follow.
The Cloud Security Alliance separates them in its autonomy levels framework. Capability is what the agent can execute with the resources, tools and knowledge it has. Autonomy is what it is authorized to execute without human intervention. An agent can accumulate many capabilities and receive little autonomy, and that indicates neither organizational immaturity nor a limitation of the model.
The distinction looks obvious written down. It disappears in practice, because the conversation about agents usually starts with a demo. The agent solves the task in front of everyone, the room becomes convinced that it can, and the question about what it may do alone never gets asked separately.
Permission is a third thing, and worth separating too. Permission defines which data, tools and systems the agent can reach. Autonomy defines how much it decides and executes without a new approval. A technical permission granted without a conscious decision becomes operational freedom by accident.
Risk does not end at model quality
It is natural to concentrate the evaluation on the model. How good are the answers, how often does it get things wrong, does it follow instructions. Legitimate questions, and insufficient ones.
Agents introduce another variable. What happens after the system decides.
OWASP records this as Excessive Agency, the risk of a system receiving functions, permissions or autonomy beyond what it needs. The effect shows up when the model misreads a situation, receives a manipulated instruction or produces an unexpected response, and the consequence escapes because the agent was authorized to act.
An agent that only reads information carries one kind of power. An agent that changes prices, deletes records, sends messages or moves transactions carries another. The technology can be identical in both cases. The possible consequence is not, and consequence is what the company will be managing when something goes wrong.
Four questions bound the autonomy of an action
There is no universal level of autonomy that fits every agent, every process and every company. Maturity ladders with five rungs help organize vocabulary and decide nothing on their own, because the right rung changes from action to action inside the same agent.
The decision has to look at the concrete action. Four questions handle most cases.
| Factor | Question |
|---|---|
| Impact | What happens if the action is wrong? |
| Ability to correct | Can we undo the action and also its consequence? |
| Time to notice | How long might we take to discover the problem? |
| Reach | How much can be affected before we manage to stop it? |
The four produce no automatic answer. They show where more freedom makes sense and where the company needs to hold a limit. The four-factor model is a WatchZ advisory proposition, combining principles found in security and governance references without corresponding to a regulatory classification or an official standard.
The larger the consequence, the stronger the justification has to be
Not every error weighs the same.
A document lands in the wrong folder. A price is changed incorrectly. A message reaches the wrong customer. A payment goes to the wrong account. A configuration affects a system that carries revenue.
There is an error in all five cases. The consequence of each lives on a different scale, which is why asking how likely the agent is to be wrong is only half the calculation. The other half is the size of the effect when the error happens.
AWS recommends classifying agent actions by impact and by reversibility, proposing distinct oversight levels to balance speed and control rather than applying the same control to everything. The larger the possible consequence, the stronger the justification needs to be for letting the action happen without another control in the path.
Undoing the action does not undo what it caused
Some actions reverse easily. A record is restored, a file returns to a previous version, a change is cancelled.
There is a difference between undoing the technical action and undoing what happened because of it, and the Cloud Security Alliance separates the two in its framework. Exposed data does not become unseen because access was removed afterward. A change reverted in the system can keep producing effects outside it, in a customer, a contract, a report that already circulated.
The useful question is not whether we can undo. It is whether we can correct the action and also what it caused. The harder both are, the more care belongs before execution, and it is the same reversibility logic that enterprise architecture applies to structural decisions.
An error that takes time to surface keeps producing consequence
Picture two cases.
In the first, the agent tries to change a record and the system refuses immediately, flagging an invalid operation. In the second, the agent sends wrong information to a customer base and the company finds out when the complaints arrive.
The second error had time. And time matters because agents execute in sequence. While nobody notices, new actions keep happening on top of the wrong premise, and each one widens what will have to be corrected later.
Monitoring, then, does not mean keeping a history to consult when someone complains. It means detecting while there is still time to limit the effect. That is the difference between observability and an archive, and it shows up the same way in the platform observability that carries critical operations.
Reach shows how far the error travels before anyone stops it
Changing one record is one thing. Changing a hundred thousand records is another, and the same reasoning applies to messages, payments, files and systems reached.
That is why autonomy also needs limits of reach. Maximum transaction value. Number of records that can be changed at once. Volume of messages sent. Systems the agent can touch. Types of action allowed. Number of actions inside a period.
The Government of Canada recommends limiting data, tools, permissions and actions, and includes volume and frequency limits specifically to reduce the effect of unwanted behavior. None of those limits makes the agent less capable. They reduce how much a wrong decision manages to affect before anyone intervenes.
The same agent operates with different degrees of freedom
Back to the finance agent. Instead of choosing between autonomous and not autonomous, the company treats each action differently.
| Action | One possible form of control |
|---|---|
| Read an invoice | Executes alone, with read-only access |
| Identify a discrepancy | Analyzes and raises an alert |
| Prepare a payment | Leaves the action ready for review |
| Execute a material payment | Requires authorization before execution |
The example is not a rule. The correct treatment depends on the company context, applicable law, internal policy and the impact of that specific action.
The point is different. Autonomy changes inside the same process, and recognizing that avoids two extremes. The first is requiring human approval for everything. The second is releasing an entire flow because the agent performed well on some tasks. In both, actions of different risk receive identical treatment.
Anyone who has already defined the execution mode for each part of the work, in the sense described in what to automate and what to keep human, already holds the map for applying the four factors. Execution mode answers who does the work. The four factors answer how much freedom each action inside that mode gets.
A critical factor cannot disappear inside an average
There is a temptation to turn the four factors into a score. Impact gets a number, ability to correct gets another, time gets another, reach gets another, and the result becomes an average.
That method hides exactly what matters most.
Picture an action with low financial impact, small reach and fast detection, whose consequence cannot be corrected. Three factors look comfortable. One does not. The average makes the action seem safer than it is, and the company approves with a number in hand.
The prudent rule is to also read the most critical factor on its own. If the consequence cannot be corrected, that weighs alone. If the action can reach thousands of people before being interrupted, reach weighs alone. If the error can stay invisible for a long time, time weighs alone. This is a WatchZ advisory recommendation, not a regulatory rule.
It is the same trap as aggregate maturity scores, where the average climbs and the gap that limits results stays exactly where it was.
Human approval belongs where it actually reduces risk
Putting a person in the process reduces risk. That does not mean every action needs to wait for approval.
When someone receives hundreds of simple items to review, approval becomes one more click. There is a person in the flow, and the human decision has disappeared inside the volume. In the opposite direction, removing an approval that mattered hands the agent more freedom than the company can follow.
NIST, discussing interaction between people and AI systems, recommends that human roles in decision-making and oversight be clearly defined, and recognizes that the need for oversight changes with the use. The question that organizes this is direct. In which action does one person's decision meaningfully reduce risk?
Some actions call for individual approval. Others work by exception, by sampling or through a monitored indicator. Some should not be delegated at all. Human oversight is a mechanism, and choosing where it enters is worth more than increasing how much of it there is. It is the difference between designed control and accumulated control, and it shows up in every strategic AI agenda that leaves the pilot and enters operations.
More autonomy requires stronger limits
Autonomy without control is nobody's goal, least of all anyone who has run an incident.
In practice, the less a person takes part in each execution, the more weight falls on the controls surrounding the agent. The Cloud Security Alliance proposes that higher autonomy levels come with stronger controls, among them technical limits, monitoring, logging, a way to interrupt the agent and a recovery path.
The Government of Canada recommends starting with a narrow scope, using read-only access where possible and widening permissions gradually, reassessing risk whenever tools, data, permissions or scope change.
Autonomy is freedom to act inside known limits. Where the limits are not written down, what exists is the absence of a decision, and an absent decision tends to surface in an incident report before it surfaces in meeting minutes. Companies already running artificial intelligence with proportional governance treat those limits as part of the design, not as a reaction.
The decision starts with a list of actions
Before deciding how much autonomy an agent receives, list what it can actually do. Use plain verbs. Read, create, change, send, delete, buy, publish, execute, grant access.
Then apply the four questions to each item. What happens if it is wrong. Whether we can correct the action and the consequence. How long until we notice. How much can be affected before we manage to stop it.
The answers produce different degrees of freedom inside the same agent. It researches alone. It prepares an action and waits for approval. It executes up to a threshold. It is blocked from a specific action.
That does not reduce the value of the agent. It makes delegation precise, and precise delegation is what allows scope to widen later without reopening the whole discussion.
Conclusion
More autonomy does not mean more maturity. Less autonomy does not mean more safety. The goal is neither to release everything nor to control everything.
The goal is to let the agent do useful work without receiving more freedom than the company can observe, limit and correct.
That is why the opening question changes shape. Instead of asking how much autonomy this agent should have, ask what it may do alone, inside which limits, and what happens if it is wrong. The answer changes from one action to the next, which is why an agent can be highly capable and still hold little autonomy over an important decision, while executing simple, bounded and easily corrected activities on its own.
The technology is the same. What changes is the action.
Pick an agent already running in your company and list the actions of it that produce real effect. Apply the four factors to each one and mark where the current answer was inherited from a demo rather than decided. If the list shows more inheritance than decision, the problem is not the agent, and a capability assessment shows where that same absence appears across the rest of the operation.
Sources
- Cloud Security Alliance AI Safety Initiative. "Agentic AI Autonomy Levels and Control Framework", version 2.0 (March 18, 2026). https://labs.cloudsecurityalliance.org/research/agentic-ai-autonomy-levels-control-framework-v2-csa-styled/
- OWASP GenAI Security Project. "LLM06:2025 Excessive Agency", 2025 edition. https://owasp.org/www-project-top-10-for-large-language-model-applications/2_0_vulns/LLM06_ExcessiveAgency.html
- NIST. "Artificial Intelligence Risk Management Framework 1.0, Appendix C: AI Risk Management and Human-AI Interaction" (2023). https://airc.nist.gov/airmf-resources/airmf/appendices/app-c-ai-risk-management-and-human-ai-interaction/
- Amazon Web Services. "AGENTREL02-BP05 Establish tiered human oversight and approval workflows", Agentic AI Lens. https://docs.aws.amazon.com/wellarchitected/latest/agentic-ai-lens/agentrel02-bp05.html
- Government of Canada. "Guide on the Use of Agentic Artificial Intelligence" (May 22, 2026). https://www.canada.ca/en/government/system/digital-government/digital-government-innovations/responsible-use-ai/guide-use-agentic-artificial-intelligence.html
- WatchZ. "How much autonomy should you give an AI agent". Revised editorial version, September 2026.
Note: the four-factor model and the rule of reading the most critical factor on its own are WatchZ advisory propositions, with no correspondence to a regulatory classification or a third-party standard. The references cited support individual principles, not the arrangement proposed here.





