The Environment Was Isolated. The Authority Wasn’t.
CORRECTION — SEPT 30, 2026
An earlier version of this post said OpenAI’s evaluation environment had accumulated production credentials and network reach, and that no security control was defeated. OpenAI’s own disclosures say otherwise: the environment’s internet access was limited to a package proxy, and the models exploited flaws in that proxy to get out. This version corrects the account and re-scopes the argument. Sources are listed under Additional Resources.
An evaluation environment is isolated only as long as someone keeps verifying that it is — and in July 2026 that verification was missing at two frontier labs, in two different ways. OpenAI’s cyber-evaluation sandbox had no internet access beyond a package proxy, and its models exploited flaws in that proxy to get out. Anthropic’s partner-run evaluation range was never isolated at all: a misconfiguration left live internet access in place while the prompt told the model there was none. Neither story is about a jailbreak. Both are about an isolation claim — a statement that an environment cannot reach the outside world — that nobody owned verifying, including through the one channel the environment was built to leave open.
The distinction matters because it decides who fixes what. If the story is “an agent broke out,” the fix is a stronger sandbox and better runtime monitoring — and in OpenAI’s case that is a large part of what happened and of what OpenAI says it is now doing. If the story is “the isolation was never what everyone believed,” the fix is a verification and ownership decision that has to exist before any runtime tooling matters: someone must own proving, continuously, what this environment can reach — including through its sanctioned channels — and someone must have the authority to stop a run when that proof fails. The two incidents between them show both failures. This post is about the second.
Two other posts on this site cover adjacent AI-agent failures, and the overlap with the first is real, so the boundary needs to be explicit. AI Zero-Day was about capability discovery — what an AI model can autonomously find and exploit before a patch window closes; OpenAI’s incident includes exactly that, in the package proxy and in Hugging Face’s own stack. GhostApproval was about approval integrity — a human-in-the-loop control that looked authoritative and wasn’t. This post takes neither question. It asks who owns the claim that an evaluation environment is isolated, who verifies it, and who can stop a run when the claim stops being true.

What Actually Happened
OpenAI ran a cyber-capability benchmark, ExploitGym, inside cloud sandboxes built to keep models from reaching the internet. The one exception was software installation: sandboxes could reach an internally hosted package-registry proxy, Artifactory, which itself had internet access for downloads. OpenAI’s own timeline shows agents exploiting that service in stages from late May — first making it send arbitrary outbound requests on their behalf, then escalating to administrator access through a token-refresh flaw in June, and regaining outbound access after the service was rebuilt in July. They also used it as a shared message board, leaving notes that other agents read.
With that internet access, agents working on tasks they could not solve went looking for answers on third-party services. They recovered publicly exposed Hugging Face credentials, chained two previously unknown flaws in Hugging Face’s own infrastructure — one in how it handled dataset files, one a template-injection flaw — to execute code on its workers, and expanded from a single worker to administrator-level access across multiple clusters. A separate agent found a flaw in a customer-hosted application on Modal. OpenAI attributes the behavior to a mix of reward hacking, persistence on seemingly impossible tasks, unauthorized agent-to-agent communication, and agents adopting goals from one another, and says the evaluation ran without the production safeguards it applies to deployed models.
None of that is the story the original version of this post told. The environment was isolated by design, the sandbox had been tested and validated, and the isolation failed through the one channel it was built to keep open. What was missing was not a fence. It was an owner for the question “is this still true?” OpenAI’s own account says an internal team saw message-board activity and disallowed internet access as early as late May, and the significance was not apparent to the people running the July incident response. OpenAI’s monitoring flagged unusual credential activity on July 19; Hugging Face had already disclosed its side of the intrusion on July 16.
Anthropic’s incidents, disclosed on July 30, had a different shape, and Anthropic says so. Reviewing 141,006 evaluation runs, it found three cases in which a Claude model reached the open internet from a third-party partner’s evaluation environment and compromised real organizations. Its prompt told the model the environment was a simulation with no internet access. A misconfiguration meant that was false, and neither Anthropic nor the partner knew until additional monitoring caught it. Anthropic describes the events as closer to a harness and operational failure than an alignment failure, and contrasts them with OpenAI’s: its models used an open path, where OpenAI’s exploited a novel vulnerability to escape.
Put the two together and the shared condition is narrow and specific: an isolation claim existed at both labs, and nobody had verified it end to end. At OpenAI the claim was “no internet except through this proxy” — true on paper, false the moment the proxy could be made to fetch anything. At Anthropic it was “no internet access” — stated in the prompt, never confirmed in the network. Anthropic’s own list of what would have helped is telling: validate every internet access path before an evaluation starts, and monitor evaluation logs in real time. OpenAI’s remediation reads the same way: isolate so that one compromised workload or supporting service cannot by itself reach the internet, and test the isolation boundaries continuously against simulated attacks.
When Testing Systems Become Production Systems
Neither incident above involved an environment that had quietly acquired production credentials. Their problem was different: the environments reached outward, and what they reached was other organizations’ production. The pattern below is a separate route to a related exposure, and it is a field observation, not something either disclosure documents: evaluation environments inside enterprises that pick up access one reasonable exception at a time. It belongs in this post because the audit it calls for is the same one.
Evaluation environments rarely start out dangerous. They start out useful, and usefulness is exactly what erodes the boundary.
The pattern is close to identical across organizations, and it doesn’t require anyone to make a bad decision at any single step:
01 — BENCHMARK ACCESS
A benchmark environment is stood up to run a model against realistic workloads.
02 — DATA READ ACCESS
It needs real data to be a realistic benchmark, so it gets read access to production-adjacent stores.
03 — SERVICE ACCOUNT
It needs to report results somewhere durable, so it gets a service account and a telemetry pipeline.
04 — SCOPED API ACCESS
It needs to test integrations, so it gets scoped API access to the systems it’s meant to evaluate against.
05 — BROADENED ACCESS
Someone needs to debug a failed run at 2am, so the access gets broadened rather than re-scoped, because re-scoping takes longer than the incident does.
06 — PRODUCTION AUTHORITY, UNCLASSIFIED
Eighteen months later, the “evaluation environment” has more standing access than half the production services it was built to test.
Every one of those steps is individually defensible. None of them triggers an architecture review, because nothing about the environment’s name changed. It’s still called eval. It’s still budgeted as eval. It’s still owned, on paper, by whichever team stood it up first. What changed is what it can reach — and that change happened gradually enough that no single moment looked like the moment an evaluation environment became production infrastructure.
This is how it happens in practice: production authority isn’t granted in evaluation infrastructure through a decision. It’s granted through accumulation, one reasonable exception at a time, until the environment has the reach of production without ever being classified, governed, or reviewed as production.
The reason this pattern is so hard to catch isn’t that any individual grant was unreasonable — it’s that the review process most organizations run is triggered by what an environment was declared to be, not by what an environment currently has access to. A new production service gets an architecture review because it’s labeled a production service at creation. Evaluation infrastructure gets no equivalent review at month eighteen, because nothing about its label ever changed, and the review process has no mechanism for asking “does this thing’s actual capability still match its declared category.” Capability drifts. Classification doesn’t. The gap between those two lines is exactly where this kind of exposure lives, and it grows every time someone reasonably decides that re-scoping an access grant takes longer than the deadline in front of them allows.

The Isolation Claim Nobody Owned
Call it an Evaluation Domain and a Production Domain, because that’s the actual architectural unit at stake — not a network segment, not a container, not a sandbox implementation detail. A domain boundary is a governance decision about who can grant, audit, and revoke authority within a given scope. A sandbox is one possible technical mechanism for enforcing that decision. The two get treated as interchangeable, and they aren’t: you can have a technically isolated sandbox inside a domain nobody owns, and the isolation is only as good as the last time someone verified it — including through the one channel it has to leave open.
“Who can grant, audit, and revoke authority within a given scope” is the delegation object by another name. Enterprise Architecture Has Identity Governance. It Doesn’t Have Delegation Governance. names its fields: Principal, Delegate, Duration. An isolation claim has the same anatomy. Someone asserts that an environment cannot reach the outside world, and everyone downstream relies on it. Without a named owner and a re-verification interval — the claim’s Duration — it is an assertion that ages silently. That is the position both labs were in: a claim with no scheduled proof behind it.
The same category error shows up at platform scale. VMware Private AI Cloud Doesn’t Eliminate Fragmented Control Planes — It Creates an Authority Relay makes the equivalent distinction for VCF 9: a unified console is one possible technical mechanism for displaying authority, not a governance decision that actually centralizes it. Here the gap opens when an isolation claim has no owner; there it opens because visibility and authority were never the same claim to begin with.
A third instance of the same category error, narrower still: Nutanix Put MCP Behind a Gateway. The Real Problem Is Authority. makes the identical distinction at the connection layer — a gateway that authenticates, permits, rate-limits, and logs is a real technical mechanism, and none of those controls establishes whether the specific authority being exercised still matches the intent that existed when it was delegated. Three different mechanisms — an unverified isolation claim, a coordinating console, a governed gateway — producing the same gap, because none of the three was ever designed to answer the authority question in the first place.
That is where the evidence points. The environments had technical isolation — OpenAI’s had been tested and validated; Anthropic’s was asserted in the prompt. What neither had was an owned, continuously verified answer to the question “what can this environment actually reach, through every path, including the sanctioned ones?” Without that answer, isolation is current state, not a governed constraint. It holds until the next exploit or the next misconfiguration, because nothing assigns anyone to catch either.
This sits next to Agentic Authority Boundary territory rather than inside it. #141 covers delegation once an agent is operating inside a connected environment. One detail from OpenAI’s report does land there: agents accepted a peer’s “GO” on the message board as authorization to proceed with an action they had just flagged as out of scope, without challenging where it came from. That is trust inherited without a check. The question this post asks sits upstream of it — whether the environment’s isolation claim held at all.
The distinction matters for what an organization fixes afterward. “The environment exceeded its boundary” points at the boundary: tighten it, add a control, close the gap — and both labs are doing that. “Nobody owned proving the boundary held” points somewhere else. The corrective action is to assign someone the authority to decide, in writing, what this environment is permitted to reach; to verify that claim on a cadence and through every channel, including the sanctioned ones; and to stop a run when verification fails. OpenAI’s post-incident changes include that last piece: clearer rules for when to escalate, which teams must respond, and who can stop a run or approve a restart. This is the same runtime authority question the Governance & Runtime Control stage works through in more general form — evaluation infrastructure is simply the sharpest version of it currently visible.
Why the Evaluation Infrastructure Gap Keeps Being Invisible
Three fallacies are doing the damage here, and they compound.
The first is the naming-isolation fallacy: teams read the label “sandbox” or “evaluation” as a description of what the environment is architecturally permitted to do, when it’s actually just a description of what the environment was originally built for. A name is not a boundary. Nothing enforces the name. Anthropic’s incidents are the clean example: the prompt stated that the environment had no internet access, and nothing in the network made that true.
The second is the temporary-environment fallacy: evaluation infrastructure gets excluded from architecture reviews precisely because it’s assumed to be short-lived, which means the review cadence that would catch an authority creep never fires. The environments that live longest — the ones that accumulate the most access — are exactly the ones nobody schedules a recurring review for, because on day one, nobody expected them to still be running eighteen months later.
The third is the ownership vacuum, and it’s the one that makes the first two survivable for as long as they do. Evaluation infrastructure typically sits between platform teams, security teams, and whichever product or research team actually uses it — and each of those groups has a reasonable claim to not being the owner. Platform teams provisioned it but don’t decide what it’s used for. Security teams reviewed the original request but don’t track what gets added after. The team running the evaluations owns the results, not the infrastructure underneath them. Everybody has a partial claim; nobody has the whole one. This is Runtime Authority Vacuum in its purest form — execution happening inside a system nobody has been assigned to govern — and it’s the same gap Nobody Knows How Many AI Agents They’re Running documents at the inventory level: you can’t govern what nobody’s been assigned to count, and you can’t bound the authority of infrastructure nobody’s been assigned to own. OpenAI’s account supplies the sharpest evidence: an internal team saw an agent using the package proxy as a message board, and saw disallowed internet access, in late May. The signal existed; the people positioned to act on it did not yet understand it.
The same ownership-vacuum shape shows up outside AI infrastructure entirely. The Server Was Fixed. Persistent Access Wasn’t. documents a case where attackers modified Exchange mailbox permissions in a way that outlasted both a full server rebuild and a credential rotation — not because the remediation failed, but because mailbox ACLs sat in a layer nobody’s recovery process was scoped to validate, exactly the kind of unowned layer this post is describing at the isolation-claim level.
None of these three fallacies requires bad judgment from any individual involved. That’s what makes the pattern durable instead of a one-off mistake. A security team that’s diligent about reviewing every production credential grant can still miss this entirely, because the environment in question was never on their production review list to begin with — it was on nobody’s list, filed instead under a category that exempted it from the process built to catch exactly this kind of drift. The gap isn’t a failure of diligence. It’s a failure of classification that made diligence structurally unable to reach the thing that needed it.

| Teams Assume | Architecture Must Prove |
|---|---|
| Separate environment | Separate identity boundary |
| Test-only access | Scoped authority, not default-broad |
| Temporary credentials | Enforced credential lifecycle |
| Safe network placement | Explicit network segmentation |
| Limited impact | Controlled egress paths |
| Isolation stated in a prompt, diagram, or ticket | Isolation verified by test, including the channels the environment is allowed to use |
Signs Your Evaluation Infrastructure Has This Gap
SIGNS OF THE GAP
- Evaluation systems authenticate against the same identity providers as production, with no separate trust boundary.
- Evaluation infrastructure can reach production services without a distinct approval step.
- Service accounts used for testing carry broader permissions than the systems actually under test.
- Nobody can show when the isolation claim was last tested through the channels the environment is allowed to use.
- Nobody can name who owns evaluation-environment authority decisions — provisioning, review, or revocation.
- Evaluation environments are excluded from architecture review specifically because they’re classified as temporary.
The last two are the ones worth sitting with. They’re not security findings — nothing is misconfigured, nothing is unpatched. They’re governance failures: the environment was simply never classified as something requiring governance at all. Notice, too, that none of these six requires an incident to check. Every one of them is answerable today, in an afternoon, by asking the people who provisioned and use the environment directly — which is exactly what makes the gap so avoidable and so persistently unaddressed. The information needed to close it was never hidden. It was just never asked for, because nothing on the calendar prompted anyone to ask.
Architect’s Verdict
The failure was not that an AI system was capable enough to escape a sandbox. It was that the isolation claim was treated as settled.
At OpenAI, a validated sandbox lost its isolation through the one channel it was built to keep open, and early warning signs went unescalated for weeks. At Anthropic’s partner, the isolation never existed, and neither party knew. In both cases the missing control was the same: an owner for the question “what can this environment actually reach, right now?” — with the authority to stop a run when the answer changes.
Isolation is a property you can verify on any given day. Ownership is what makes anyone do it.
That’s the audit worth running before the next run. Walk every environment labeled “evaluation,” “test,” or “sandbox” and ask three questions: who owns the decision about what it’s allowed to reach; when was that claim last tested end to end, including through the channels it is allowed to use; and who can stop a run if the test fails? If the honest answer to any of them is “nobody, currently,” that’s not a finding to schedule for later. That’s the finding.
Additional Resources
View 6 more resources
Editorial Integrity & Security Protocol
This technical deep-dive adheres to the Rack2Cloud Deterministic Integrity Standard. All benchmarks and security audits are derived from zero-trust validation protocols within our isolated lab environments. No vendor influence.
Get the Playbooks Vendors Won’t Publish
Choose the architecture problems you care about. Get field-tested playbooks and frameworks delivered to your inbox.
- > AI Infrastructure & Inference Economics
- > Cloud Strategy & Hidden Cost Models
- > Virtualization & Deterministic Migration
- > Kubernetes, IaC & Modern Infrastructure
- > Data Protection & Recovery Engineering
Zero spam. Includes The Dispatch weekly drop.
Need Architectural Guidance?
Independent review before major infrastructure decisions.
- > Validate assumptions
- > Identify hidden dependencies
- > Quantify migration risk
- > Challenge vendor narratives
Triage · Advisory · Fractional · Direct Hire