Robots Atlas>ROBOTS ATLAS
Safety

Sandbox Escape

ActivePublished
Key innovation
Frames the isolation boundary (the sandbox) itself as an active attack surface: code deliberately run inside an environment assumed “safe for executing untrusted code” can break that boundary and reach the host system.
Category
Safety
Abstraction level
Pattern
Operation level
Agent runtimeToolingDeployment
Use cases
AI agent code-execution sandboxes / LLM code interpretersMulti-tenant cloud systemsIsolation of untrusted code in browsers (JavaScript, WASM)CI/CD runners executing untrusted scriptsMalware analysis in isolated environmentsContainerization of production workloads

How it works

A sandbox escape usually follows one of several patterns. (1) Kernel/hypervisor exploitation: containers share the host kernel, so a bug in the kernel or the runtime allows privilege escalation (e.g. CVE-2019-5736 in runc — overwriting the host runc binary by manipulating /proc/self/exe and gaining host root). For VMs the target is the hypervisor's emulated hardware (e.g. VENOM, CVE-2015-3456, an overflow in QEMU's floppy disk controller enabling code execution outside the guest). (2) Misconfiguration: running a container as root/--privileged, mounting the Docker socket or host paths, missing seccomp/AppArmor profiles, abusing kernel features (e.g. CVE-2022-0492 — abusing cgroups v1 release_agent to bypass namespace isolation). (3) Interpreter/application escape: bypassing language restrictions (access to system modules, reflection, native calls) or an engine bug (e.g. memory corruption in a browser's V8/JS as the first stage of a chain). (4) Side channels and microarchitecture: attacks like Spectre (CVE-2017-5753) exploit speculative execution and branch prediction to break isolation boundaries at the hardware level (e.g. reading memory outside the JS sandbox). In an AI-agent context the typical chain is: the model generates malicious or vulnerable code → the code runs in the sandbox → it exploits one of the above weaknesses → it gains access to the host, secrets (API keys, tokens), or the internal network.

Problem solved

Addresses the threat model of executing untrusted code: when a system must run code it does not control (LLM-generated code, a plugin, a user file, a browser ad), isolation is meant to contain the side effects. Sandbox Escape describes when and how that guarantee fails — understanding this class is a prerequisite for designing secure execution environments for AI agents, multi-tenant systems, and browsers.

Components

Isolation boundaryThe security boundary the attack aims to break.

The mechanism separating sandboxed code from the host: kernel namespaces and cgroups, seccomp, AppArmor/SELinux, a hypervisor, or OS privilege separation.

Official

Attack surface / brokerThe exploit's entry point.

The interface where the sandbox contacts the host: syscalls, hypervisor-emulated hardware, IPC, shared file descriptors, the interpreter API. This is where the vulnerability typically lives.

Exploit primitiveConverts a weakness into execution control.

The technical breaking mechanism: memory corruption, out-of-bounds read/write, abuse of a kernel feature, a race condition, or a side channel.

Official

Host targetThe escape's ultimate objective.

The out-of-bounds resource the attack targets: host root, secrets (API keys, tokens), the host filesystem, the internal network, or other tenants.

Official

Implementation

Implementation pitfalls
Container run as root or --privilegedCritical

A privileged container effectively removes isolation — it grants access to host devices and kernel capabilities.

Fix:Run as a non-privileged user, drop capabilities, avoid --privileged, use user namespaces.
Mounting the Docker socket or host pathsCritical

Mounting /var/run/docker.sock or sensitive host paths into the sandbox gives full control over the host.

Fix:Never mount the Docker socket into untrusted workloads; minimize mounts and make them read-only.
Missing seccomp / AppArmor / SELinux profileHigh

Without syscall filtering the kernel attack surface is maximal, easing exploitation of kernel bugs.

Fix:Apply restrictive seccomp and MAC (AppArmor/SELinux) profiles; allow only necessary syscalls.
Shared kernel instead of hypervisor isolationHigh

Standard containers share the host kernel, so a single kernel bug compromises the boundary. For untrusted code this is insufficient.

Fix:For untrusted code (e.g. LLM-generated) use stronger isolation: gVisor, Firecracker microVMs, or full VMs.
Outdated runtime with known CVEsHigh

Unpatched runtimes (runc, QEMU, the kernel) contain publicly known escapes that can be exploited immediately.

Fix:Regularly patch and scan images and runtime components for CVEs.

Evolution

2015
VENOM — VM escape via QEMU FDC

CVE-2015-3456: a bug in QEMU's floppy disk controller (used in Xen and KVM) lets a guest execute code outside the virtual machine.

2018
Spectre — hardware-level isolation break
Inflection point

CVE-2017-5753 (publicly disclosed in 2018): speculative execution and branch prediction as a side channel; showed that even in-browser JS isolation can be bypassed at the microarchitectural level.

2019
runc / Docker — container escape to host root
Inflection point

CVE-2019-5736: overwriting the host runc binary by manipulating /proc/self/exe grants host root from inside a container.

2022
cgroups v1 release_agent — container escape

CVE-2022-0492: abusing the cgroups v1 release_agent feature in the Linux kernel to escalate privileges and bypass namespace isolation.

2023
AI-agent code sandboxes enter the mainstream
Inflection point

With code interpreters in LLM assistants and autonomous agents, the code-execution sandbox (isolation via gVisor, Firecracker microVMs, dedicated services like E2B) becomes a critical security component — and escaping it a new, significant attack vector.