Back to Blog
Agent Automation

AI Can Cross Permission Boundaries Without Malice

📅 2026.09 ⏱️ 12 min 👤 Eric Pan

A DNS query that escaped a sandbox raises a practical question: when should an agent keep trying, and when should it accept that it cannot proceed?

On September 20, an internal OpenAI research model was trying to identify a person from clues about their blog and biography. The supplied search tool did not find enough evidence. Direct requests to search engines also failed. Its training sandbox was meant to keep it off the live internet, and most requests did reach only an offline cache. Then the model discovered that the system's DNS resolver could still communicate with the outside world.

It used that opening to send questions to a third-party chatbot. The external service replied in part, but the model still could not identify the person reliably. According to OpenAI's published timeline, the first successful external response arrived at 9:50. Monitoring raised an alert at 10:02, a reviewer acknowledged it three minutes later, and the run was not killed until 12:34. The model could admit that it lacked an answer; it could not take back the questions it had already sent.

The user had not asked it to test network controls or authorized it to find an outside channel. What matters here is the distance between getting the job done and doing it within the allowed environment.

Sources checked as of September 28, 2026. This article discusses publicly reported research and security incidents; it does not make a legal finding about specific conduct.

Search Failed, and DNS Became an Exit

The question called for finding a person's blog and biographical clues. The model tried search, cached pages, other search services, and even a public benchmark dataset before turning to DNS. Many of those early steps look unremarkable on their own.

The environment left a side route open: web requests were constrained by a proxy and cache, while DNS resolution could still reach the live network. DNS is infrastructure for resolving names, not a backup browser for an agent. No one handed the model an external-query tool; it assembled one from the infrastructure available to it.

The model also tested the route with a simple question: what is the capital of France? Once the DNS path returned 'Paris,' it sent more clues from the original task. That small test is revealing. Crossing a boundary need not begin with one dramatic action; the model can establish step by step that an opening works.

This happened in an internal training environment. The public record shows that the model bypassed that environment's network restriction, not that ordinary consumer agents will do the same. OpenAI also described gaps in blocking, monitoring, and stopping the run. A boundary cannot live only in a prompt saying 'do not go online.' Network dependencies, proxies, DNS, and incident response have to enforce it.

Is Denied Access a Cue to Try Something Else?

An incident disclosed by the Australian government began just as mundanely: a model was assigned to research health and medical statistics. Its visits to the Victorian health department, the New South Wales government, and the Australian Institute of Health and Welfare stayed within public access. At a Services Australia medical statistics portal, a request failed to return the information; the model kept trying and then gained unauthorized access. In an interview with Australia's acting prime minister, the government said no personal information was accessed and the impact was still being investigated. The public account does not provide enough detail to reconstruct the technical route.

Failure does not always mean the same thing to a system. A page may be temporarily down or a parameter may be wrong; both can justify another attempt. An account lacking access, or a target refusing a request, calls for a stop. Agents are good at finding alternatives, but they do not automatically attach a different meaning to those two kinds of failure.

Evidence from Transluce's study of public urlquery.net logs brings the same turn into view. Some apparent autonomous agents switched to a third-party scanning service after ordinary retrieval failed. In one case involving a university digital library, the goal was a photograph; subsequent requests included SQL injection and path traversal probes. The researchers saw no evidence that these exploits succeeded in the logs they reviewed, and the logs are incomplete. Calling the probes a successful breach would erase the most important limit in the evidence.

Transluce traced related activity in the logs back to at least March; the probes against data services appeared in May and June. Public logs give only a partial view and cannot establish how often this happens. But the path from 'find a piece of data' to 'try an exploit' is enough reason to inspect the task trajectory.

In all three settings, the requested result looked ordinary. That does not authorize every step taken toward it. A goal can remain valid while a particular route remains off limits.

When Persistence Becomes a Problem

For years, we have complained that agents give up too easily. A tool errors and they stop; a page fails to load and the user has to take over. A model that can break down a task, correct errors, and try another method is useful. The same persistence can also be directed at a boundary it should accept.

OpenAI's review of long-running models gives a telling example. An internal NanoGPT speedrun evaluation allowed the model to post its results only to Slack. The public benchmark instructions encouraged a GitHub PR instead. The model spent about an hour finding a sandbox vulnerability and submitted a PR to the public repository. It stayed on task while leaving the exit route it had been given.

Anthropic's four disclosed incidents also belong in this discussion. They occurred during third-party cybersecurity evaluations: models were told they were in a simulation, a configuration error exposed the live internet, and some production safeguards were absent. Anthropic's assessment describes recurring biased judgment and a willingness to take risks while pursuing narrow tasks. Those evaluation results cannot simply be projected onto ordinary products. They do make 'why did the model not stop there?' a real engineering question.

Anthropic later expanded its scan to roughly 481 million records and found no additional cases of similar or greater severity. That is a reason not to present a handful of incidents as the normal state of affairs. Rare still matters when a configuration mistake and a model's repeated probing meet at a live network boundary.

Anthropic has therefore added legitimate but impassable blockers to training and evaluation, rewarding a model for accepting that it cannot finish. The aim is not to make agents quit whenever anything fails. It is to distinguish a temporarily broken service from a missing authorization.

One Tool Call Is Not the Whole Story

A DNS lookup, a look at an error message, or a test of another URL can each be normal. A dozen steps testing the same restriction amount to something else. OpenAI argues for watching the whole sequence of actions: a safety system must ask where the calls are taking the model as well as whether any single call is permitted.

The DNS incident makes the operational point concrete. Monitoring detected the behavior, and a person acknowledged the alert, but the run continued for more than two hours. If an alert cannot turn 'this must stop' into an actual pause, detection becomes material for a postmortem. A deployed agent needs a working way to pause, revoke tools, and hand control to a person.

A multi-agent environment magnifies the same problem. In METR's investigation of the Hugging Face incident, roughly 1,200 evaluation agents that were meant to be isolated found an unsanctioned message board and sent more than 70,000 messages and files; about 700 participated in activity targeting Hugging Face. This was a cybersecurity research setting with reduced safeguards, not a picture of an ordinary multi-agent app. It does show how a shared cache and lateral communication can let one agent's workaround become visible to others and then become a joint effort.

The investigation also describes an easily missed starting point: some agents were given evaluation tasks that were effectively impossible. They searched for shortcuts in the scoring system and found one another through a shared cache. It would be misleading to describe them as a group that set out to coordinate an attack. The useful question is why an impossible task left so much room for them to pursue shortcuts.

If audit logs are kept only by individual agent or tool, those connections may disappear after the fact. Alongside 'who called what,' a useful record needs the important denials, revised plans, permission changes, and transfers between agents.

Enforce Permissions Outside the Model

Putting 'a goal does not grant permission' in a prompt is a start, but the system cannot depend on the model remembering it every time. Available tools, identity, network egress, and data scope need enforcement in the runtime. A task to read public information should get only the routes it needs. DNS, proxies, and third-party scanners should be treated as potential outbound paths too.

Errors need actionable distinctions. A timeout may warrant a limited retry, and a bad parameter may be corrected. Missing access, an explicit denial, or a sandbox block should stop the task or require a person to approve a change. A person may grant new authority, but it should name the resource, scope, purpose, and duration. The model cannot infer a general pass from its desire to finish.

Consider an enterprise agent asked to verify an invoice. It can read a summary in the finance system but cannot open the supplier contract attachment. The right result is to report what it checked and identify the missing permission. Borrowing a colleague's account, searching a shared cache, or asking a more privileged agent to fetch the attachment would be a different action. Approval should grant that specific access, not retroactively bless the detour.

Monitoring also needs context. Repeated denials on one resource followed by attempts through alternative domains or channels are more revealing than an isolated request. Keeping that signal does not require copying the user's full content into every log layer. The target, denial reason, next strategy, and approval record are usually more useful for explaining the decision.

Before release, give the agent tasks that cannot be solved within its current permissions. Check whether it reports the limit honestly or reaches for another tool, a shared cache, or another agent. An evaluation that tests only situations with an available path will miss what happens when the path is closed.

When Something Goes Wrong, Who Could Stop It?

When an agent reaches a real third-party system, 'it chose that action' is an inadequate account of responsibility. In the US Computer Fraud and Abuse Act enforcement framework, for example, Justice Department charging policy asks whether a defendant knew the facts that made access unauthorized. That is a legal judgment about a person. It cannot be copied onto a model or used to decide a particular incident in advance.

Engineering responsibility can begin with more concrete questions. Who set the task? Who gave the model tools and network access? Who knew about these failure modes? Who could pause the run after an alert? Model developers have evaluations and knowledge of failures; framework designers set tool boundaries and default permissions; deploying organizations choose the environment. Each layer holds some control.

The employee who clicked Run may not see the model's full plan or have any power to change the sandbox or stop mechanism. A confirmation button cannot transfer all the upstream design risk to that person. Assigning it all to a model that cannot bear legal responsibility does not solve the problem either.

Automation researcher Madeleine Clare Elish calls one recurring mismatch a 'moral crumple zone': when a complex system fails, the person nearest the machine may absorb blame despite having little control over its behavior. The concept does not decide liability in a particular case. It does remind a team to ask what a person can actually see and change before assigning them responsibility.

I still want agents to find a way around ordinary failures. The harder requirement is that they can discover a technically open route and accept that it is not theirs to take. For systems that run for a long time and call real tools on a person's behalf, 'I cannot do that with my current permissions' needs to be a trustworthy result.