Sandbox Escape
How it works
A sandbox escape usually follows one of several patterns. (1) Kernel/hypervisor exploitation: containers share the host kernel, so a bug in the kernel or the runtime allows privilege escalation (e.g. CVE-2019-5736 in runc — overwriting the host runc binary by manipulating /proc/self/exe and gaining host root). For VMs the target is the hypervisor's emulated hardware (e.g. VENOM, CVE-2015-3456, an overflow in QEMU's floppy disk controller enabling code execution outside the guest). (2) Misconfiguration: running a container as root/--privileged, mounting the Docker socket or host paths, missing seccomp/AppArmor profiles, abusing kernel features (e.g. CVE-2022-0492 — abusing cgroups v1 release_agent to bypass namespace isolation). (3) Interpreter/application escape: bypassing language restrictions (access to system modules, reflection, native calls) or an engine bug (e.g. memory corruption in a browser's V8/JS as the first stage of a chain). (4) Side channels and microarchitecture: attacks like Spectre (CVE-2017-5753) exploit speculative execution and branch prediction to break isolation boundaries at the hardware level (e.g. reading memory outside the JS sandbox). In an AI-agent context the typical chain is: the model generates malicious or vulnerable code → the code runs in the sandbox → it exploits one of the above weaknesses → it gains access to the host, secrets (API keys, tokens), or the internal network.
Problem solved
Addresses the threat model of executing untrusted code: when a system must run code it does not control (LLM-generated code, a plugin, a user file, a browser ad), isolation is meant to contain the side effects. Sandbox Escape describes when and how that guarantee fails — understanding this class is a prerequisite for designing secure execution environments for AI agents, multi-tenant systems, and browsers.
Components
The mechanism separating sandboxed code from the host: kernel namespaces and cgroups, seccomp, AppArmor/SELinux, a hypervisor, or OS privilege separation.
Official
The interface where the sandbox contacts the host: syscalls, hypervisor-emulated hardware, IPC, shared file descriptors, the interpreter API. This is where the vulnerability typically lives.
The technical breaking mechanism: memory corruption, out-of-bounds read/write, abuse of a kernel feature, a race condition, or a side channel.
Official
The out-of-bounds resource the attack targets: host root, secrets (API keys, tokens), the host filesystem, the internal network, or other tenants.
Official
Implementation
A privileged container effectively removes isolation — it grants access to host devices and kernel capabilities.
Mounting /var/run/docker.sock or sensitive host paths into the sandbox gives full control over the host.
Without syscall filtering the kernel attack surface is maximal, easing exploitation of kernel bugs.
Standard containers share the host kernel, so a single kernel bug compromises the boundary. For untrusted code this is insufficient.
Unpatched runtimes (runc, QEMU, the kernel) contain publicly known escapes that can be exploited immediately.
Evolution
CVE-2015-3456: a bug in QEMU's floppy disk controller (used in Xen and KVM) lets a guest execute code outside the virtual machine.
CVE-2017-5753 (publicly disclosed in 2018): speculative execution and branch prediction as a side channel; showed that even in-browser JS isolation can be bypassed at the microarchitectural level.
CVE-2019-5736: overwriting the host runc binary by manipulating /proc/self/exe grants host root from inside a container.
CVE-2022-0492: abusing the cgroups v1 release_agent feature in the Linux kernel to escalate privileges and bypass namespace isolation.
With code interpreters in LLM assistants and autonomous agents, the code-execution sandbox (isolation via gVisor, Firecracker microVMs, dedicated services like E2B) becomes a critical security component — and escaping it a new, significant attack vector.