Kernel-Level Sandboxing for AI-Written Code Execution
Kernel isolation protects AI code execution from container escapes and supply-chain attacks.

AI-generated code is attacker-influenceable by construction, and that fact alone should change how every engineering team thinks about where it runs. Code written by a developer has a human author behind it, someone who can be asked why a line exists and held responsible for what it does. Code generated by an agent follows a different path entirely: it comes out of prompts, tool results, and retrieved content, any of which an attacker can shape through prompt injection.
This is the baseline condition for any agent that touches external content, calls APIs, or processes user input before generating code, not a narrow edge case that only applies to agents built carelessly. An agent summarizing a support ticket, pulling data from a third-party API, or reading a document a user uploaded is all, by definition, exposed to this channel. A well-intentioned agent executing what looks like its own logic may in fact be executing logic an attacker planted several steps upstream, and nobody in the pipeline, human or model, ever reviewed it.
The consequence for sandbox design is specific and non-negotiable: the threat model has to assume the code itself might be hostile, not just buggy. That's a different engineering problem than catching an off-by-one error or an unhandled exception, and it's the distinction that decides which isolation architecture actually fits the job. The case for sandboxing agent-generated code touches several concerns at once: security, resource isolation, predictable environments, the ability for sessions to persist through long-running tasks, and running many tenants on shared infrastructure. Security comes first among these, because an agent that generates malicious or merely buggy code can, absent containment, reach sensitive data or alter files it was never meant to touch. Every other sandboxing concern assumes this one is handled.
What shared-kernel containers protect against (and what they don't)
Container isolation was built for a world of trusted code: CI pipelines running code a developer wrote and reviewed, cached build artifacts, workloads where the content was never in question, only the resource contention. Containers isolate processes from one another at the operating system level, but every container on a host still shares that host's kernel, and every syscall a containerized process makes lands directly on that shared surface. That design is efficient and battle-tested for the workloads it was built for. It is the wrong boundary for code whose content an attacker can influence.
The specific failure mode is the container escape: a flaw that lets a process break out of its container and reach the host kernel directly. A kernel vulnerability reachable from inside a container doesn't stay contained to the tenant that triggered it. It opens a path to the host itself, and from the host, to every other workload sharing that kernel. That is the blast radius under discussion when untrusted, AI-generated code runs on shared-kernel infrastructure: not one compromised task, but the machine underneath it and everything else running there.
Container design assumes the code running inside it is trustworthy, and that is the exact assumption AI-generated code breaks. Containers remain a reasonable choice for cached build steps and code a developer wrote and owns, because in that setting the premise holds. The premise stops holding the moment the code's content is shaped by inputs an attacker can reach, whether through a prompt, a retrieved document, or a tool result standing between the user and the model. Most workloads a development team already runs in containers are not attacker-influenceable in this way, so the objection that containers are "good enough" for most things isn't wrong on its own terms. It just doesn't apply here. Once code generation depends on external, attacker-reachable input, the shared-kernel boundary is the wrong place to put the trust, regardless of how well that boundary performs for everything else a team runs.
The isolation spectrum: containers, gVisor, and microVMs compared
Three distinct mechanisms sit on a spectrum running from shared-kernel to fully virtualized, and each one answers the same underlying question differently: where should the line between "trusted" and "untrusted" actually sit. Hardened containers push that line further out using seccomp syscall filtering and namespace restrictions, narrowing what a workload can do without changing the fact that the kernel underneath remains shared. That narrows the attack surface without closing it.
gVisor moves the line further by intercepting a workload's syscalls in user space and acting as a guest kernel of its own, so the application underneath never touches the host kernel directly. That's a meaningful improvement over bare containers, but the interception layer is itself software, and software carries bugs just as the host kernel does. gVisor also adds real I/O overhead, with sequential large-file I/O running 30 to 50 percent slower, a cost that matters directly for agents that process datasets or move files as part of their work.
MicroVMs draw the line in a fundamentally different place. Technologies like Firecracker, and orchestration layers such as Kata Containers that run atop VMMs including Firecracker, Cloud Hypervisor, or QEMU, give each workload its own stripped-down Linux kernel, booted inside a KVM-backed VM. An exploit running inside that microVM has to break through KVM and the hardware virtualization boundary to reach the host, a different category of barrier than crossing a shared kernel. Firecracker's architecture reinforces this isolation at the process level too: each microVM gets its own VMM process rather than sharing a single daemon across every microVM on the host, so compromising one VMM doesn't hand an attacker the rest of the fleet. That design pays off both in security and in how predictably operators can reason about failure.
The assumption that virtualization necessarily means slow I/O doesn't hold up under the numbers. Firecracker's overhead on sequential large-file I/O comes in substantially lower than gVisor's 30 to 50 percent penalty, which matters directly for agents reading and writing datasets, logs, or generated files as part of their work. One real constraint remains on the microVM side: Firecracker has no PCIe passthrough, so workloads that need direct hardware access inside the sandbox can't run on Firecracker-based isolation today. Passthrough support has been explored on a feature branch and may land in a future release, but as of now, GPU-inside-sandbox workloads have to look elsewhere for isolation.
MicroVM cold-start latency is no longer a blocking objection
The strongest historical argument against microVMs has been speed: VMs take seconds to boot, and agents make tool calls inside an inner loop where every turn's latency compounds across the whole session. A two-second sandbox startup, repeated across dozens of tool calls in a single agent run, makes a product feel broken even when the underlying logic is sound. That objection was legitimate for years. It no longer holds, because snapshot-restore has solved the slow boot times that caused it.
The mechanism is straightforward. A snapshot captures a microVM after its environment is fully set up, Python runtime installed, dependencies resolved, servers already running, and every subsequent invocation restores from that snapshot instead of booting the machine from zero. The environment the agent needs is already sitting there in memory, ready to resume rather than ready to boot. This same mechanism extends to concurrency: multiple parallel runs can fork from one shared snapshot, each landing in its own isolated VM without stepping on another fork's state, which is how production agent fleets absorb bursts of simultaneous tool calls without paying a full boot cost for every one of them.
The latency figures across current implementations illustrate the range rather than ranking any single approach. Resume-from-standby can land under 25 milliseconds, with standby triggered automatically after 15 seconds of network inactivity at no compute cost. Firecracker microVM cold starts have been measured at 150 milliseconds. A single-executable microVM project called SmolVM, launched April 17, 2026, achieves sub-200-millisecond cold starts running on Hypervisor.framework and libkrun, aimed at macOS and Windows environments where Firecracker can't run, while still supporting Linux through KVM. Checkpoint-and-restore elsewhere has been clocked at roughly 300 milliseconds with no compute charge while idle. A gap still exists between Firecracker's 150-millisecond-to-2-second cold-start range and browser isolates starting in under 50 milliseconds. The gap still means the number that counts in production is restore latency, not cold-start latency, once snapshot-restore is actually in use.
What kernel-level isolation enforces inside the sandbox
A microVM boundary solves the problem of host escape. It does nothing to constrain what the agent does once it's running inside that boundary, and that gap means isolation at the kernel level is necessary but not sufficient on its own. A complete security posture requires five additional controls layered on top of the microVM boundary.
Network egress needs to default to deny, with outbound traffic permitted only through an explicit allowlist, closing off the paths an agent could otherwise use to exfiltrate data or call back to infrastructure an attacker controls. The filesystem surface an agent can write to should be bounded deliberately, because unrestricted write access inside the VM still lets malicious state persist even when the VM itself never escapes. Package installation needs scrutiny too: supply chain attacks arrive through pip, npm, and similar registries, and a sandbox that lets an agent install anything without inspection has simply moved the attack surface one layer inward rather than removing it. Resource limits on CPU and memory keep a compromised or runaway workload from degrading the host's scheduler for every other tenant sharing it.
A fifth control sits outside the kernel layer. Prompt injection delivered through tool results, an API response or a retrieved document crafted to redirect the agent's next code-generation step, is an application-layer problem that no amount of kernel isolation touches. This is where MCP currently has a real gap: it lacks an enforceable security system of its own for the servers agents connect to, so protection depends on the integrity of whoever built the MCP server and whatever ad hoc controls the host happens to add. The tooling layer that feeds instructions into the agent is itself an injection surface that kernel isolation was never designed to cover.
Research is beginning to address this specific gap. AgentBound, published at FSE 2026, is described as the first access control framework offering capability-constrained execution for MCP ecosystems, drawing its design from Android's permission model. Its policies can reportedly be generated automatically from source code, and the paper's authors report that it blocks a majority of security threats from malicious MCP servers with negligible enforcement overhead. Peer review of the work has flagged that the accuracy figure needs a fuller error breakdown before its blocking claims can be taken at face value, which places AgentBound as a promising early research direction rather than a finished, deployable answer to the MCP security gap.
How persistent state interacts with the isolation boundary
Production agent architectures are rarely stateless, and that fact puts direct pressure on how the sandbox boundary gets designed. Patterns like chaining tool calls, orchestrator-worker setups, reflection loops, and human-in-the-loop gates all depend on an agent carrying context from one turn into the next. Treating each invocation as an isolated, memoryless event doesn't match how these systems actually run. Stuffing conversation history, file outputs, and intermediate results into the context window isn't a workable substitute for real persistent storage. Once a long-running agent accumulates enough history, the context window itself becomes the bottleneck, and external persistent memory stops being optional.
Different isolation approaches handle persistence in noticeably different ways. Some platforms keep a sandbox's filesystem, memory, and running processes fully intact through standby, triggered after a fixed period of network inactivity, with separate volume storage guaranteeing persistence beyond that window. Others give agents a large NVMe-backed filesystem that persists across sessions without requiring an explicit snapshot step. Some let paused sandboxes persist indefinitely while capping active session duration, and still others retain filesystem snapshots for a set default period, memory snapshots for a shorter one, with newer beta features for attaching project-specific state to a pre-warmed sandbox.
Persistence of this kind carries its own isolation consequence: a disk that survives across turns is also a surface that survives across turns, and anything an agent writes in one turn can potentially be read or altered in the next if that storage isn't scoped tightly to a single tenant's identity. Isolating the VM layer well does nothing to protect a shared persistence layer sitting behind it. Per-tenant identity needs to propagate through signed credentials all the way to storage, to knowledge bases, and to whatever context the tool execution runs in. Treating the sandbox as something spun up fresh for every single call, with no persistent identity attached, doesn't hold up against how real agent products actually need to run. The architecture that does hold up looks more like a persistent, identity-scoped environment assigned per user, with isolation and persistence designed together rather than bolted on after the fact.
The economics of running isolated sandboxes at scale
MicroVM isolation carries a reputation for costing more than container-based alternatives, and the headline per-second compute rate is the wrong place to look for whether that reputation is deserved. Active-rate pricing across competing providers has largely converged: two independently operated, venture-funded companies arrived at the identical per-vCPU-hour rate in 2026, which suggests the market for this specific compute has settled on a price. If active compute pricing is roughly the same everywhere, it can't be the variable that explains real cost differences between providers.
Idle billing policy is where that difference actually lives. A benchmark run in August 2026 against identical agent workloads across providers found roughly a 2.4x spread in total cost, driven mostly by how each provider treats idle time rather than by the active compute rate. Metering based on active CPU usage came in lowest in that comparison, with providers billing more conventionally for idle sandbox time landing in the middle and higher tiers. A separate structural cost appears at moderate usage volumes, on the order of hundreds of runs per day across a month: a fixed plan fee can become the single largest line item in the bill, larger than the metered per-second compute charge itself, a cost a simple rate-sheet comparison hides entirely.
This idle-cost dynamic maps onto agent economics more directly than the per-second rate ever could, because most agents spend most of their time idle, waiting on a user, a tool response, or the next step in a reflection loop. A sandbox architecture that bills for storage during that idle stretch, rather than metering compute from the moment the sandbox exists, produces a fundamentally different cost curve at scale than one still charging for active CPU the entire time nothing is running. For teams weighing kernel-level isolation against the apparent savings of shared-kernel containers, the real economic question is how a provider bills for idle time. It's how a provider bills for the hours an agent spends doing nothing at all, because that is where most of the bill actually accumulates.
