Claude Code sandbox vs undo: what each protects in an AI coding agent
A sandbox keeps an agent inside your project. Undo fixes what it breaks there. Here's how each works, where they fall short, and when you need an approval prompt instead.
Updated October 2026
A sandbox limits where an AI coding agent can write and what it can reach on the network. Undo puts back what it changed inside the area it was allowed to touch. You need both: the Claude Code sandbox and Codex's sandbox both leave your project folder writable, so a bad rm or rewrite inside it gets through.
Last checked: October 2026. Platform support and setting names come from each tool's current docs, linked below.
This page covers what agent sandboxes actually enforce, how Claude Code, Codex CLI, Gemini CLI and containers differ (including on Windows), what a sandbox can't protect, and a decision table for when you need a sandbox, undo or an approval prompt.
What a sandbox does for a coding agent
An agent runs shell commands as your user. Without isolation, a command can read your SSH keys, write anywhere your account can, and send data to any host. A sandbox puts an operating-system boundary around those commands, usually on two axes:
- Filesystem isolation: writes are limited to the project folder and a temp folder. Some sandboxes also restrict reads.
- Network isolation: outbound connections are blocked or routed through a proxy that checks each host against an allowlist.
Because the operating system enforces the boundary, it holds even if the model is tricked by a prompt injection. That's the main advantage over approval prompts, which depend on you reading every command.
How the Claude Code sandbox works
Claude Code's sandboxed Bash tool is off by default. Turn it on with /sandbox in a session, or "sandbox": {"enabled": true} in settings.
| Platform | Mechanism |
|---|---|
| macOS | Apple's built-in Seatbelt framework, no install needed |
| Linux | bubblewrap for filesystem isolation, socat to route traffic through the sandbox proxy |
| WSL2 | Same as Linux |
| WSL1 | Not supported ("Sandboxing requires WSL2") |
| Native Windows | Not supported: commands run unsandboxed |
It's built on Anthropic's open-source @anthropic-ai/sandbox-runtime. Default boundaries:
| Access | Default |
|---|---|
| Writes | The working directory, a per-user temp directory and directories you've added. Protected paths stay write-denied |
| Reads | Most of the machine, including ~/.ssh and ~/.aws/credentials |
| Network | No direct route out. A local proxy checks each host against your allowed domains, which start empty |
| Environment variables | Inherited, including any secrets in Claude Code's environment |
Three gaps to know about:
- It wraps shell commands only. Claude's own Read, Edit and Write tools, MCP servers, hooks and commands you type at the
!prompt run outside it. - There's an escape hatch. When a command fails in the sandbox, Claude can retry it with
dangerouslyDisableSandbox, which goes through the normal permission flow. Set"allowUnsandboxedCommands": falseto turn that off. - Sandboxed commands auto-approve by default. With
autoAllowBashIfSandboxedat its default oftrue, Bash commands that modify files inside the boundary run without a prompt, even in Manual mode.
To put the whole Claude Code process inside a boundary, Anthropic points to the sandbox runtime, a dev container, a custom container or a VM. Its comparison page says the Bash sandbox alone "is not sufficient for fully unattended runs."
Claude Code sandboxing on Windows
On native Windows, Claude Code's sandbox doesn't run. Anthropic's options for a Windows machine are:
- WSL2: install Claude Code inside a WSL2 distribution and run
/sandboxthere. On Ubuntu 24.04 and later you may need an AppArmor profile sobwrapcan create user namespaces; check withsysctl kernel.apparmor_restrict_unprivileged_userns. - A container or VM: a dev container in Docker Desktop, or a full VM.
Other agents differ here. Codex CLI has a native Windows sandbox (MXC, with older elevated and unelevated fallbacks) and uses its Linux sandbox under WSL2; WSL1 stopped being supported in Codex 0.115. Gemini CLI has a Windows Native Sandbox that uses icacls to set a Low Mandatory Level on files it needs to write, and those changes persist after the session.
Codex CLI sandbox
Codex turns its sandbox on by default. The docs describe three modes: read-only, workspace-write (the default for version-controlled folders) and danger-full-access.
- macOS: Seatbelt, running commands through
sandbox-execwith a profile for the selected mode. - Linux:
bwrapplusseccomp. - Windows: MXC natively, or the Linux sandbox in WSL2.
In workspace-write, network access is off unless you enable it, and .git, .agents and .codex inside the workspace are read-only, recursively. That .git rule is a real advantage: the agent can't rewrite your history or hooks. But everything else in the workspace is writable.
Gemini CLI sandbox
Gemini CLI's sandbox is off by default but turns on automatically with --approval-mode=yolo. You can pick the method: macOS Seatbelt (profiles from permissive-open to strict-proxied; the default confines writes to the project but allows broad reads and network), Docker or Podman, gVisor on Linux, or LXC (experimental). Enable it with -s, GEMINI_SANDBOX, or tools.sandbox in settings.json.
Docker and dev containers
A container moves the whole agent, including its file tools and MCP servers, inside the boundary. Anthropic ships a reference dev container with a non-root user and a default-deny init-firewall.sh. OpenAI ships a Codex secure devcontainer with bubblewrap and an allowlist firewall. Docker offers Docker Sandboxes, a microVM for running agents.
What a sandbox can't protect
Every sandbox above leaves one thing writable: the project you asked the agent to work on. Anthropic says so directly. "Any approach that mounts your project directory writable can still modify that code," and in a dev container, "Claude can still modify any file in the bind-mounted workspace, which appears directly on your host."
So inside the boundary, a sandbox does nothing about:
rm -rf src/orgit clean -fdxrun by mistake- a codemod, formatter or generator that rewrites hundreds of files badly
- an
npm installthat changespackage.jsonand the lockfile - uncommitted work the agent overwrote
And at the edges:
- Allowed hosts can still leak data. Anthropic warns that allowing broad domains like
github.com"can create paths for data exfiltration." Both Anthropic and OpenAI note that a malicious project inside a container can exfiltrate anything the container can read, including the agent's own credentials. - Allowed hosts can still be changed. If the sandbox allows your deploy target or database, a deploy or migration goes through.
Why undo is the complement
A sandbox shrinks the blast radius to your project. Undo covers the blast radius itself. They fail differently, which is why they work well together.
Most agent undo features only reverse the agent's own file edits. Claude Code's checkpointing "does not track files modified by Bash commands," and Codex CLI no longer has an undo command. The comparison is in AI coding agent checkpoints compared.
Darce approaches it from the undo side. Before every step that can change the project, it snapshots the working tree under private git refs, so /undo reverses files that shell commands created, changed or deleted, not just its own edits. It doesn't sandbox commands: they run as your user in the folder you started it from. Instead it scores every command, and in the default auto mode anything that reaches outside the machine (network, installs, deploys, migrations) or is destructive asks first. See Safety. The limits: undo history lasts for the session, gitignored files like node_modules aren't snapshotted, and shell effects are only covered in a git repo.
Darce has no OS-level sandbox, so on an untrusted repository a container is still the right call, with Darce or any other agent inside it.
Sandbox, undo or approval: a decision table
| Threat | Sandbox | Undo | Approval prompt |
|---|---|---|---|
| Agent deletes or rewrites files in the project | No, the project is writable | Yes | Only if you read the command |
Agent writes outside the project (~/.bashrc, ~/.ssh) | Yes | No, undo covers the project only | Yes |
rm -rf ~ or / | Yes | No | Yes |
Reads ~/.ssh or .env and sends it out | Only with read denies and network limits | No | Yes, if network access asks |
| Prompt injection from a web page or README | Yes, limits what it can reach | Yes, for local file damage | Weak: injected commands can look routine |
Malicious postinstall script | Yes, in a container | Only for files inside the project | Yes, if installs ask |
Deploy, migration, git push, publish | Only if the host isn't allowed | No | Yes, the only real defence |
| Bad refactor you notice ten steps later | No | Yes, step back through history | No |
| Untrusted repository | Yes, use a VM or container | Helps, not enough | Not enough |
Short version: a sandbox for things outside the project, undo for things inside it, approval for things outside the machine.
Recipe: run an agent in a dev container
The fastest path for Claude Code is the official Dev Container Feature. Save as .devcontainer/devcontainer.json:
{
"image": "mcr.microsoft.com/devcontainers/base:ubuntu",
"features": {
"ghcr.io/anthropics/devcontainer-features/claude-code:1.0": {}
}
}Then run Dev Containers: Rebuild Container in VS Code. For network egress limits, copy init-firewall.sh and the NET_ADMIN / NET_RAW runArgs from the reference container. Don't mount ~/.ssh or cloud credential files; use repository-scoped or short-lived tokens.
For any terminal agent, including Darce (Node.js 22+), a plain container works:
docker run --rm -it \
-v "$PWD":/work -w /work \
node:22 npx darce-cliThe agent can only write to /work, which is your project, and to the container's own disposable filesystem. Commit before you start, so git can restore anything inside /work that the agent breaks.
Recipe: read-only mounts
When the agent only needs to read something, mount it read-only with :ro. Docker enforces it, whatever the agent does:
docker run --rm -it \
-v "$PWD":/work -w /work \
-v "$HOME/reference-docs":/docs:ro \
node:22 bashFor a review-only task, mount the project itself read-only (-v "$PWD":/work:ro) and add --network none if the agent doesn't need to call a model API from inside. Most do, so in practice you'll pair a read-only mount with an egress allowlist instead.
In Claude Code's built-in sandbox, the equivalent is read denies:
{
"sandbox": {
"enabled": true,
"allowUnsandboxedCommands": false,
"filesystem": {
"denyRead": ["~/.ssh", "~/.aws", "~/**/.env"]
}
}
}Remember this only covers shell commands. Add matching Read deny rules under permissions for Claude's own file tools. More on permission rules and running without prompts in YOLO mode with a safety net.
Sources
- Claude Code: Configure the sandboxed Bash tool
- Claude Code: Choose a sandbox environment
- Claude Code: Development containers
- Claude Code: Checkpointing
- anthropics/sandbox-runtime
- anthropics/claude-code reference dev container
- Codex: Agent approvals and security
- openai/codex secure devcontainer
- Gemini CLI: Sandboxing
- Docker Sandboxes
- Darce: Safety
- Darce: Undo
Frequently asked questions
Does Claude Code run in a sandbox by default?
No. The sandboxed Bash tool is off by default. Turn it on with /sandbox or by setting sandbox.enabled to true in settings. It uses Seatbelt on macOS and bubblewrap on Linux and WSL2.
Does Claude Code sandboxing work on Windows?
Not natively. On native Windows, Claude Code runs commands unsandboxed. Run Claude Code inside WSL2 to use the sandbox, or use a container or VM. WSL1 isn't supported.
Does a sandbox stop an AI agent deleting my code?
No. Every common agent sandbox leaves the project folder writable, so the agent can still delete or rewrite files there. Commit first, and use an agent whose undo covers shell commands.
What sandbox does Codex CLI use?
Seatbelt via sandbox-exec on macOS, bwrap plus seccomp on Linux, and MXC on native Windows, with the Linux sandbox under WSL2. The default workspace-write mode keeps network off and .git read-only.
Is a dev container enough to run an agent with permissions skipped?
It protects your host, but Anthropic and OpenAI both warn that a malicious project can still exfiltrate anything inside the container, including the agent's credentials. Restrict network egress and don't mount host secrets.