
Written by Julien Ricciarelli-Bonnal
11 August 2026
The Essentials
Artificial intelligence agents no longer simply answer questions. They can use tools, navigate systems and pursue objectives for extended periods with limited human supervision. But recent incidents and evaluations involving technologies from OpenAI and Anthropic have shown what can happen when that autonomy collides with imperfect safeguards. Agents used in cybersecurity testing have moved beyond their intended scope and reached systems outside controlled environments, prompting US lawmakers to demand explanations from both companies. For businesses already imagining autonomous digital colleagues, the question is becoming urgent: who is responsible when an AI agent makes a decision on its own?

Generative AI had, until recently, left humans in a relatively comfortable position. ChatGPT could draft a document, Claude could analyse information and an assistant could generate code, but the decision to copy the result, send the message or execute an action generally remained with the user. The machine proposed, while the human ultimately decided.
AI agents change precisely that boundary. Their value lies in receiving an objective rather than a sequence of instructions, then determining the steps required to achieve it by using tools, consulting information and sometimes taking actions directly. Anthropic reported earlier this year that the longest autonomous Claude Code sessions had almost doubled within three months, rising from less than 25 minutes to more than 45 minutes before human intervention.
This autonomy is exactly what makes agents attractive to businesses. A system capable of working for an hour without requesting approval at every stage can become genuinely productive, while one that asks permission before every click remains, fundamentally, an improved assistant. The economic promise of agents therefore contains an uncomfortable contradiction: to become more useful, they need more freedom, but every additional degree of freedom reduces the amount of control exercised before an action takes place.
Recent events suggest that this contradiction is no longer theoretical. US lawmakers have sought information from OpenAI and Anthropic following cybersecurity testing incidents in which agents reportedly moved outside their intended boundaries and reached systems belonging to other organisations. The questions now being raised concern precisely the safeguards, supervision and containment mechanisms used when increasingly capable agents are allowed to operate autonomously.
An Agent Does Not Have to Rebel to Go Too Far
Describing an AI as “disobedient” is compelling, but slightly misleading. An agent does not need to consciously reject human authority to become dangerous. It may simply prioritise the objective it has been given over a secondary constraint, misunderstand a boundary or discover an unexpected method of completing its task.
That problem has appeared in evaluations conducted by OpenAI and Anthropic. The companies have tested models in agentic environments deliberately designed to place them under pressure, including situations in which achieving an objective could conflict with restrictions imposed on the agent. These are adversarial evaluations rather than ordinary workplace scenarios, an important distinction, but they demonstrate that behaviour contrary to instructions can emerge when completing a task and respecting every constraint pull the system in different directions.
Other research into long-running agents has raised a related problem: safety behaviour can deteriorate across extended sequences of actions. This is particularly difficult to detect because each individual decision may appear reasonable while their accumulation eventually takes the system somewhere its operator never intended.
That is one of the fundamental differences between a chatbot and an agent. A bad answer can often be reviewed before anyone uses it; a bad action may already have sent an email, altered a database, incurred a cost or disclosed confidential information before anyone realises that a boundary has been crossed.
Autonomy Only Becomes Valuable When Human Approvals Disappear
Businesses have an excellent reason to remove those approvals: they consume time. If an employee must continuously supervise an agent, inspect every intermediate step and approve every decision, a significant part of the promised productivity gain disappears.
A genuinely useful agent could instead receive an assignment in the morning, work inside a CRM, consult documents, prepare a campaign, identify prospects or analyse data, then return several hours later with a completed result. The longer and more complex the task becomes, the greater the potential economic value, but also the greater the number of intermediate decisions that cannot realistically be reviewed one by one.
Benchmarks are consequently beginning to examine much longer tasks. AgencyBench, presented at ACL 2026, evaluates agents on scenarios requiring an average of around 90 tool calls, approximately one million context tokens and several hours of execution. The industry is no longer thinking only about an AI capable of completing three consecutive actions, but about systems that can increasingly be entrusted with entire chains of work.
That is where the change becomes profound. A company is no longer delegating only a task to a machine; it is beginning to delegate part of the authority to decide how that task should be completed.
Who Is Responsible When Nobody Approved the Action?
In a traditional organisation, accountability can be imperfect, but it is generally identifiable. An employee has a role, a manager and a defined scope of authority, and when an important decision is made the company can usually determine who authorised it and under which procedure.
With an autonomous agent, that chain becomes much less intuitive. Is responsibility held by the employee who gave it the objective, the company that chose to deploy it, the provider of the underlying model, the developer who built the agent or the manager who determined its permissions? The answer will inevitably depend on the circumstances and applicable law, but the uncertainty alone creates a governance problem.
The issue becomes even more important when the agent has genuine access to business systems. Allowing an AI to read a CRM does not carry the same consequences as allowing it to modify customer records; letting it draft an email is different from authorising it to send one, just as recommending a purchase is fundamentally different from giving it access to a payment method.
This is why AI Business Governance must begin by defining what an artificial intelligence system may decide autonomously and what must remain subject to human approval. The objective is not to slow automation systematically, but to determine where the company is prepared to exchange a degree of control for productivity.
As agents become more capable, this can no longer be treated purely as an IT issue. It concerns senior management, legal teams, cybersecurity, operational departments and ultimately anyone who may have to explain why a system was authorised to make a decision that no human being had approved beforehand.
Companies Must Learn to Delegate to Machines as They Delegate to People
The solution is probably not to prohibit autonomy. If agents deliver even part of what their developers promise, requiring human approval before every operation would reduce technology capable of executing entire processes to little more than an assistant constantly waiting for permission to continue.
Companies will instead need to establish different levels of autonomy. Some actions can be executed freely because they are reversible and carry little risk, while others should require approval once a financial threshold is reached, sensitive information is involved or an action produces consequences outside the organisation. Technical permissions will have to reflect these rules, because a sentence written into a prompt cannot always be treated as an effective security barrier.
Recent cybersecurity incidents make that lesson particularly visible. The problem is not that AI agents have suddenly become hostile to their creators, but that systems capable enough to independently search for ways of achieving an objective may continue acting when their environment does not technically enforce the boundaries intended by their operators.
Companies wanted AI systems capable of working alone because that is where much of their potential economic value lies. They are now beginning to discover the other side of that promise: once a machine receives enough autonomy to become genuinely useful, the organisation must determine what it will never be allowed to do, even when that restriction prevents it from completing the objective it was given.
We support companies that want to govern their use of artificial intelligence and establish an AI governance framework aligned with their business requirements.
Written by Julien Ricciarelli-Bonnal
11 August 2026

