Agent Sandboxing
How it works
The agent runs inside an isolation boundary โ a container, virtual machine, or microVM โ instead of directly on the host. System calls are intercepted and filtered (seccomp-bpf or a userspace application kernel such as gVisor), and resource visibility is limited by namespaces and cgroups. Filesystem access is narrowed to mounted, often read-only, working directories. Network traffic passes through an egress allowlist, blocking exfiltration and SSRF-style attacks. A policy-enforcement layer decides which tools and actions the agent may invoke, and sensitive operations can require human approval. After the task finishes, the environment is destroyed (ephemeral), clearing state and any contamination. For stronger multi-tenant isolation, microVMs (Firecracker) or gVisor are used instead of sharing the host kernel.
Problem solved
Autonomous AI agents take and execute actions without a human approving every step, so a model error or a successful prompt-injection attack can lead to deleted data, leaked secrets, an attack on the internal network, or unauthorized transactions. Agent Sandboxing limits the blast radius of such events by confining the agent to a minimal-privilege environment.
Components
A container, virtual machine, or microVM in which the agent runs. microVMs (Firecracker) and a userspace application kernel (gVisor) provide stronger isolation than a plain container sharing the host kernel.
Official
seccomp-bpf profiles that restrict the set of system calls available to the agent process, reducing the kernel attack surface.
Linux kernel mechanisms isolating the view of processes, network, mounts, and users (namespaces) and limiting CPU, memory, and time (cgroups).
An egress allowlist plus a layer deciding which tools and actions the agent may invoke; sensitive operations can require human approval.
Official
A non-persistent environment created for the duration of a task and destroyed afterwards, removing potential contamination and persistent state between runs.
Official
Implementation
A privileged container or excessive capabilities enables escape to the host.
Full egress access enables data exfiltration and SSRF attacks on internal services.
Containers sharing the host kernel are exposed to kernel vulnerabilities exploitable by hostile agent code.
Passing broad credentials or keys into the agent environment amplifies the impact of a successful attack.
Content from the web or files can hijack the agent and trigger harmful actions within its privileges.
Stronger isolation (microVMs, a fresh environment per task) increases latency and resource usage.
Evolution
The seccomp filter mode (seccomp-bpf) added in Linux 3.5 enables fine-grained restriction of a process's system calls โ a foundation for later sandbox isolation.
Google releases gVisor โ a userspace application kernel that intercepts system calls, reducing the host kernel attack surface for untrusted workloads.
AWS introduces Firecracker โ lightweight microVMs for secure, multi-tenant isolation of serverless and container workloads.
Sandboxes purpose-built for safely executing AI-agent-generated code and tools emerge (e.g. E2B).
Anthropic releases computer use for Claude, with documentation recommending a dedicated virtual machine or container with minimal privileges.
OpenAI ships the computer use tool, recommending an isolated browser or VM and an allow list of sites and actions.
Hyperparameters (configurable axes)
Choice between a container, a microVM, and a userspace kernel (gVisor). Affects isolation strength and overhead.
Scope of allowed network traffic (no network, allowlist, full access). Key to preventing exfiltration.
Scope and mode of mounts (read-only, narrowed working directories).
A seccomp profile defining the allowed system calls.
CPU, memory, and execution-time limits enforced via cgroups.
Whether the environment is ephemeral (destroyed after a task) or persistent across runs.