Robots Atlas>ROBOTS ATLAS
Tooling

Code Execution

2022ActiveUpdated: 14 August 2026Published
Key innovation
Lets a model solve tasks by writing and running code in an isolated (sandbox) environment and using the execution result, instead of relying on text generation alone.
Category
Tooling
Abstraction level
Pattern
Operation level
InferenceAgent runtimeTooling
Use cases
Data analysis and computation over files (CSV, spreadsheets)Precise math and simulationsGenerating charts and visualizationsFile conversion and processingVerifying and running code by coding agentsSolving tasks via program-aided reasoning

How it works

The model is given an 'execute code' tool (usually via tool/function calling). When a task needs computation, it generates code (most often Python) that is sent to an isolated environment (a sandbox: a container/VM with CPU/RAM/time limits and network control). The environment runs the code, captures stdout, errors and artifacts (files, plots), and returns them to the model. The model interprets the result: if there was an error it fixes the code and retries (an iterative loop), and on success it uses the result in its answer. State (variables, uploaded files) can persist across calls within a single session.

Problem solved

LLMs are unreliable at precise computation (arithmetic, data processing, file manipulation) and lack a deterministic runtime. Code Execution solves this by delegating such tasks to an actual interpreter and returning a verified result.

Key mechanisms

Code generation by the model (usually Python)
An isolated execution environment (sandbox: container/VM, resource limits)
Exposure as a built-in tool within tool/function calling
Capturing stdout, errors and artifacts
An iterative fix loop (observe error -> correct)
Session state persistence (variables, files)

Strengths & limitations

Strengths
โœ“Deterministic, verifiable computation instead of guessing
โœ“Large accuracy gains in math and data analysis
โœ“Extends the model with real actions (files, charts)
โœ“Iterative self-correction based on execution errors
โœ“A foundation for coding agents and data assistants
Limitations
โœ—Security risk: running model-generated code (a sandbox is required)
โœ—Sandbox escapes, resource abuse, network access
โœ—Latency and cost of maintaining execution environments
โœ—Buggy or non-deterministic code, external dependencies
โœ—Environment limits (missing packages, time/memory caps)

Components

Code generatorDecision and generation

The model that produces code (usually Python) to accomplish the task.

Sandbox / execution runtimeSafe execution

An isolated container/VM with resource limits and network control that runs the model's code.

Official

Result & artifact channelFeedback

Captures stdout, errors and files/plots and returns them to the model.

Official

Implementation

Implementation pitfalls
Executing untrusted codeCritical

Code comes from the model and may be malicious or buggy; without isolation it risks compromising the host.

Fix:A strong sandbox (container/VM), no/allowlisted network, resource limits, no secrets in the environment.
Sandbox escape and resource abuseHigh

Infinite loops, memory exhaustion, attempts to break out of isolation.

Fix:Hard time/memory limits, kernel isolation, monitoring, ephemeral environments.
Dependencies and non-determinismMedium

Missing packages, differing library versions or randomness yield unstable results.

Fix:Pinned environments, preinstalled packages, fixed random seeds.
Data leakage via outputMedium

Large or sensitive outputs appended to context raise cost and risk.

Fix:Output size limits, redaction/summarization, artifact controls.

Evolution

Original paper ยท 2022 ยท ICML 2023 ยท Luyu Gao
PAL: Program-aided Language Models
Luyu Gao, Aman Madaan, i in. (et al.)
2022
PAL / Program of Thoughts โ€” program-aided reasoning
Inflection point

Models generate code whose execution yields the answer, improving math accuracy.

2023
ChatGPT Code Interpreter (Advanced Data Analysis)
Inflection point

Mass availability of code execution in an assistant: data analysis, files, charts.

2023
Open sandboxes (e.g. E2B)

Open environments emerge for safely running LLM-generated code.

2025
Code execution as a built-in tool (Anthropic, Gemini)

Code execution becomes a standard API tool across leading providers and a foundation for coding agents.

Hyperparameters (configurable axes)

Language / runtimeHigh

The language and runtime for code execution.

pythonMost common (data analysis, computation).
shell/bashFilesystem and CLI-tool operations.
Network accessCritical

Whether the sandbox has internet access.

disabledOff by default for security.
allowlistControlled access to selected domains.
Resource limitsHigh

CPU, memory and execution-time limits.

timeout, mem capProtection against infinite loops and abuse.

Computational complexity

Time complexity: O(kod). Space complexity: O(stan_sesji + artefakty).

Execution paradigm

Primary mode
Conditional

Code is run only when the task requires it (the model's decision).

Activation pattern
Input dependent

Parallelism

Parallelism level
Partially parallel

Many independent sandboxes/calls can run in parallel; a single generate->execute->fix loop is sequential.

Scope
Inference

Hardware requirements

Primary

Code (e.g. data analysis) runs on CPU in containers/VMs.

Good fit

A hardware-agnostic pattern; depends on the sandbox environment.