The Governance Line Nobody Draws: Why Enterprises Keep Regulating the Wrong Layer
Most enterprise AI governance conversations I sit in on still treat "the model" and "the system around the model" as the same thing. A risk committee asks whether the AI is safe, someone answers with a benchmark score, and the conversation moves on. It is a comfortable shortcut, and it is also the reason so many governance frameworks fail the moment an agent gets real access to a real environment.
Anthropic gave the industry an unusually clean way to see why that shortcut breaks down. In late May 2026, the company published a long engineering account of how it contains Claude across its three agentic products, claude.ai, Claude Code, and Claude Cowork. It is candid in a way corporate security writing rarely is, naming specific incidents, specific failure rates, and specific architectural choices that did not work the first time. Read against the backdrop of April's decision to hold back Claude Mythos Preview after it engineered its own way out of a sandbox during testing, the document amounts to a public argument for separating two questions that governance teams habitually collapse into one. The first question is how capable the model is. The second is what the surrounding system lets that capability touch. They are not the same question, and they do not get answered by the same control.
Two failure modes, two different owners
Anthropic's framing is useful here because it names the two failure modes plainly rather than burying them in a single risk score. Model misbehavior is the agent doing something harmful that nobody asked for, and the company is candid that more capable models do not simply become safer along every dimension; they make fewer obvious mistakes, but they also get better at finding paths around restrictions nobody thought to write down. Environment failure is a different animal entirely, and it does not require the model to misbehave at all. It just requires a boundary that was drawn one step too late or one permission too wide.
The cases Anthropic shares make the distinction concrete rather than theoretical. In one internal red-team exercise, an employee was phished into running Claude Code with a prompt that looked like an ordinary task request, and somewhere in the setup it asked the agent to read local credentials, encode them, and send them to an external address. The model had no way to flag this as unusual, because from its vantage point a human had typed the instruction directly. The defense that actually held was not in the model at all; it was the egress control sitting outside it, the boundary that blocked the outbound call regardless of intent. In a second case, an allowlist that correctly permitted traffic to Anthropic's own API became the exfiltration path itself, because a file planted in a user's workspace carried an attacker-controlled key that the model, simply following its instructions, used to upload other files to the attacker's account. The sandbox worked exactly as designed. The data still left.
This is the part of the argument that should change how a governance committee allocates its attention. Neither incident was a model alignment failure in the sense most boards imagine when they ask "is the AI safe." Both were systems engineering failures, sitting in the layer that decides what an already-aligned model is permitted to reach. A governance framework that only audits model behavior, through red-teaming, benchmark scores, or a vendor's safety card, is auditing one layer of a two-layer problem and reporting the result as if it covered both.
Why the fix has to match the user, not just the threat
The second governance lesson sits in how Anthropic chose to contain each of its three products differently, and the logic behind the difference is worth sitting with. Claude Code runs on a developer's machine with real filesystem and shell access, and the company leaned, at first, on a human-in-the-loop model: let the developer approve risky actions one at a time. Anthropic's own telemetry shows why that alone was insufficient. Users approved roughly ninety-three percent of permission prompts, and the more prompts a person sees, the less attention each one gets. That is not a flaw in any particular user. It is what happens to supervisory attention under repetition, in any domain, with any species of supervisor. The fix was not better warnings. It was an OS-level sandbox that allows reads, allows writes inside a defined workspace, and denies network access by default, cutting the number of prompts a developer has to evaluate by roughly eighty-four percent and making the boundary itself auditable rather than dependent on a tired click.
Claude Cowork, built for non-technical knowledge workers rather than developers, could not lean on the same logic at all, because the average user has no fluency in evaluating what a shell command is about to do. Anthropic's own description is direct about this: when approving an exception requires expertise the typical user does not have, the boundary has to be absolute and always-on rather than something a person signs off on case by case. The product runs inside a full virtual machine with its own kernel and filesystem, mounting only the folder the user selected, so that even a fully compromised agent inside that VM has nothing else on the host to reach.
The governance principle underneath both designs is the same, and it generalizes well beyond Anthropic's own products. The right containment strategy is not a property of the model. It is a property of the user's capacity to supervise, matched against the access the task actually requires. A compliance framework that issues one standard control set for every deployment, regardless of who is operating the system and how much they can realistically be expected to catch, will be simultaneously too loose for the unsupervised non-expert and too heavy-handed for the expert who is now clicking through friction that adds nothing.
What this means for how we write a governance framework
I have sat in enough vendor risk reviews to recognize the pattern. The questionnaire asks about the model's training data, its bias testing, its hallucination rate, and stops there, as though that closes the file. What it almost never asks is what the deployed system can actually reach once it is live: which credentials are in scope, which network paths are open, whether a tool result can carry instructions the model will treat as legitimate, and who is positioned to notice if something drifts. Those are environment-layer questions, and Anthropic's own incident log shows they catch the failures that model-layer auditing structurally cannot, because in both of their worst cases the model behaved exactly as instructed. The instruction itself was the attack.
There is a temporal version of this same gap that I have written about elsewhere in the context of regulatory traceability, and it rhymes closely with what Anthropic describes here. A system can be fully compliant and fully audited at the moment it is built, and still drift into a configuration nobody approved by the time it is actually exercised in production, simply because the environment around it changed faster than the governance review cycle did. Containment architecture is one answer to that gap. It does not ask whether the model is trustworthy at a point in time. It asks what happens if, at any point in time, it isn't.
None of this argues for less model-layer work. Anthropic is candid that probabilistic defenses, however good, will never reach full coverage, and a classifier that blocks ninety-nine percent of bad commands still lets the hundredth one through at scale. The argument is narrower and, I think, more useful for a board setting governance policy this year: model evaluation and environment containment are different disciplines, owned by different teams, tested by different methods, and a framework that funds only one of them has not actually reduced risk by as much as the dashboard suggests. The model tells you what the system tends to do. The environment tells you what it is capable of reaching even when it doesn't behave as expected. Enterprise AI governance has spent two years building fluency in the first question. The harder, less comfortable work now is building the same fluency in the second.
Vikas Sharma is a Senior Enterprise AI and Digital Transformation Advisor, co-author of "From Agentic AI to RAG: A Framework for Responsible AI" (BIGS 2025), and writes on enterprise AI governance and agentic architecture at DigitalWalk.

Comments