Browser Session Isolation for Autonomous Web Agents
Prompt injection poses risks that traditional browser sandboxing wasn't designed to address.

Browser sandboxing was built to stop malicious code from reaching a human's machine. Autonomous web agents face a threat that works by talking the agent into misbehaving, using persuasion rather than any technique that breaks out of a sandbox. That distinction, not a gap in engineering effort, is why the isolation model built for human browsing cannot be carried over unchanged to agents that read pages, make decisions, and act on their own.
Why browser sandboxing does not protect autonomous agents
Remote browser isolation was built around a simple premise: a person sits at an endpoint, and the content they view online might carry something dangerous. The remote service executes the web page, then sends the person only a safe representation of it, whether through pixel streaming, DOM mirroring, or Network Vector Rendering. Active code runs somewhere else and never touches the user's device. The controls layered on top of this model, blocking downloads, disabling clipboard access, restricting printing and keyboard input, all assume a human is the one looking at the screen and the one who needs protecting from what the page might try to do to their machine.
Browser process sandboxing works on a related but distinct problem: containing what a web page's own code can do once it's running. WebKit2 was built the same way, as a split-process model where the web content process runs isolated from the UI process. These are real and effective boundaries, but they govern what a renderer can do to its host, not what an agent reading that renderer's output decides to do next.
An autonomous agent breaks both halves of that model at once. It isn't a human glancing at a sanitized screen, and the risk it carries doesn't run from the page toward a device that needs shielding. The question that matters is what the agent does with the words on the page, and neither RBI nor Chromium's site isolation was ever built to answer that question.
Prompt injection at the semantic layer, the attack browser sandboxing was never built to stop
The attack that actually threatens autonomous web agents is an instruction, embedded in ordinary page content, that the agent reads and obeys as though a legitimate operator had issued it. The agent doesn't need to be hacked in any conventional sense. It only needs to be persuaded, and persuasion travels down the exact same channel as the legitimate content the agent was sent to retrieve.
Process isolation is irrelevant to the attack, not merely insufficient against it. A renderer sandbox stops a script from reading another process's memory. It does nothing when the attack is a sentence of text that the agent's own reasoning model treats as a command, using permissions the agent already legitimately holds.
Picture an agent working through an e-commerce task, moving from listing to listing to compare prices or place an order. The agent carries out these instructions using its own valid session, so nothing about the request looks anomalous from the outside. No credential was stolen. No exploit was run. The agent simply did what the page told it to do.
Research out of ETH Zurich and Anthropic locates the structural cause of this problem with some precision. It is a property of how any agent that reads web pages has to perceive its environment in the first place, and it is why the problem resists being patched away.
Current defenses fall into two camps, and neither closes the gap alone. Defenses that do offer strict guarantees isolate the agent from untrusted content so thoroughly that the agent loses most of its ability to actually read and use the page, which defeats the purpose of deploying it. The lesson that follows is architectural: no filter will catch every injected instruction, so the system built around the agent has to assume some instructions will get through and has to limit what they can do once they have.
What Untrusted Content Masking does differently
Untrusted Content Masking, developed by researchers at ETH Zurich together with Anthropic, is the most rigorous published attempt to put a real trust boundary back into the semantic layer. It starts from an observation about the structure of a web page rather than its content: a page's Document Object Model carries enough information to tell trusted regions apart from untrusted ones without the defense ever having to read what the untrusted regions actually say. A navigation bar built by the site operator and a comment submitted by an anonymous user occupy recognizably different positions in a page's structure, and UCM uses that structural signal directly.
The mechanism redacts untrusted regions of the DOM before the page ever reaches the agent's planning model, replacing them with labeled placeholders. When a task genuinely requires reading what's inside one of those masked regions, the agent routes a query to a Quarantined Model, a separate, isolated model that reads the hidden content directly and returns only a structured response: a boolean, an integer, an enum value, a date, a float. The interface doesn't accept free-form output, so nothing resembling a free-form instruction can travel back out of it.
What makes UCM worth taking seriously is also what limits it. The Quarantined Model interaction is routed through a sandboxed interface that enforces strict privilege separation, so the security guarantee UCM offers already assumes some process or compute boundary exists around that Q-Model and its interaction layer. UCM is a semantic-layer defense that depends on a compute-layer boundary to hold; it does not replace one. And its guarantee covers a specific failure mode: it stops an untrusted instruction from reaching the planning model directly. It does not cover what happens if an injected instruction somehow survives the Q-Model's structured output constraints, and it does not address credential misuse inside an otherwise legitimate session, data crossing between tenants sharing infrastructure, or the broader surface a persistent agent exposes once it has filesystem and network access of its own. Those are compute and infrastructure problems, and they need a compute and infrastructure answer.
Per-session compute isolation as the architectural boundary for autonomous agent browser sessions
Containing what a compromised agent can do requires a boundary enforced beneath the layer where the agent itself reasons, acts, and makes decisions. So you can't rely on process or container separation alone; you need hardware-level virtualization instead.
Containers share their host's kernel. A widely used container runtime has documented escapes, so shared-kernel isolation is not an absolute boundary, and that matters directly for any workload running instructions an attacker partially controls.
The case for microVMs rests on where the enforcement actually happens. Containers, egress denylists, and permission prompts all live in the same space the agent itself reasons in: userspace, application logic, configuration. A microVM's isolation, by contrast, is enforced by the hardware underneath, a layer the agent has no way to perceive or influence regardless of what instructions it's been fed. gVisor sits between these two positions: its user-space kernel, Sentry, intercepts system calls and gives stronger isolation than a bare container, at the cost of added overhead on host I/O and isolation still weaker than full hardware virtualization. Firecracker was built to keep boot time and memory overhead low, so launching a fresh VM per session is operationally realistic, not just a theoretical nicety.
Per-session isolation means that each agent run gets its own browser process, its own filesystem, its own memory allocation, and its own lifecycle, and when that session ends, the platform destroys its writable state unless a policy explicitly says to keep it. If a session gets compromised mid-run, it has nothing left to contaminate once it closes. A second control layer sits alongside this one: default-deny network egress limits what a compromised agent can reach even while it's live, and credentials scoped to the specific task, injected at the start of the job rather than held statically across sessions, mean a hijacked agent can't take a valid login and use it to do something the original request never intended.
That kind of orchestration needs per-session isolation, separation between tenants, credential controls, network policy, and audit records that can be traced back to a specific session, all enforced together.
Snapshot-restore and scoped credentials for per-session VM isolation at scale
The obvious objection to running a dedicated microVM per session is cost: doing that for every agent run sounds like paying for a fleet of idle machines. The objection is reasonable, and it has a real answer in how modern snapshot-restore scheduling and credential injection are built.
Agents spend most of their time waiting, on a user's next instruction, on a network response, on a signal from an orchestrating service. A platform that keeps live compute running for every session through all of that idle time pays for compute it isn't using. Two different models have emerged to deal with this, and the comparison between them matters. Perpetual scale-to-zero instead pauses the sandbox and keeps its filesystem and memory state intact, so idle CPU and memory charges drop to nothing while a fast resume still works, and the isolation boundary holds without forcing every returning user down the slow, cold path.
How often users return in practice is what actually decides which model pays off. Copy-on-write storage extends the same idea in a different direction: it lets a platform fork a running agent's live state into separate, divergent trajectories without duplicating the full VM image each time, so you can explore several possible next steps or evaluate competing plans before committing to one.
Credentials need the same kind of dynamic handling. Assigning one IAM role per tenant works until the tenant count climbs into the thousands, at which point the role sprawl becomes unmanageable on its own. Benchling's deployment, described in the AWS Machine Learning Blog, shows this pattern running at real scale: AI agent-generated code executing across thousands of life-sciences tenants, more than 250 of them active in a given week, with credentials injected per job through AWS STS and no static role ever accumulating behind the scenes.
Multi-tenant agent fleet patterns and the trust boundaries each one enforces
Running a single isolated session is one problem. Running a fleet of them across many customers is another, and it forces an explicit decision about where the isolation boundary between tenants actually sits. No single pattern wins on security, cost, and operational simplicity all at once, so the right choice depends on what must never be allowed to cross from one tenant's session into another's.
Three patterns cover most of the design space. A silo architecture gives every tenant dedicated infrastructure: access is enforced through IAM resource-based policies alongside network controls such as VPC isolation, with compute kept separate between tenants. A bridge architecture makes the tenancy decision independently at each layer of the system, siloing the parts that handle tenant-specific API traffic, for instance, while running a pooled agent runtime underneath where every session still executes in its own isolated microVM.
Axonius illustrates the silo end of this spectrum in production: the company deploys a dedicated agent per customer, runs each user session on its own isolated microVM, stores container images per tenant separately in Amazon ECR, and handles provisioning and teardown per customer through AWS CloudFormation. The underlying question that should drive the choice between these patterns is simple to state and consequential to answer: what happens if one tenant's agent reads or influences another tenant's session? For data as sensitive as Benchling's life-sciences records or Axonius's cybersecurity tooling, the answer points toward silo or bridge architectures, not pool.
Across all three patterns, the same underlying primitive keeps reappearing: one persistent, isolated agent per user or tenant, with credentials and state scoped tightly to that identity. That arrangement is what makes an agent's behavior attributable to a specific session, auditable after the fact, and recoverable if something goes wrong.
Implications of the isolation boundary for agent-native product developers
Building an agent-native product today is a choice between infrastructure that enforces the right trust boundaries by default and infrastructure that leaves a developer to reconstruct those boundaries by hand, at every layer, for every tenant. Per-session microVM isolation, scoped and short-lived credentials, default-deny network policy, and semantic-layer defenses like Untrusted Content Masking address different parts of the same threat, and none of them substitutes for the others. An agent that reads the open web, authenticates into real business systems, and acts without a human checking every step needs all of these layers working together: a compute boundary enforced below where the agent reasons, credentials narrow enough that a hijacked session can't be stretched beyond its original task, and a masking layer that keeps adversarial text out of the planning model from the start. Treating any one of these as sufficient on its own leaves a system protected against the attack it was built to stop and exposed to the one that actually occurs.


