Claude Code sandbox vs undo: what each protects in an AI coding agent

A sandbox keeps an agent inside your project. Undo fixes what it breaks there. Here's how each works, where they fall short, and when you need an approval prompt instead.

Updated October 2026

A sandbox limits where an AI coding agent can write and what it can reach on the network. Undo puts back what it changed inside the area it was allowed to touch. You need both: the Claude Code sandbox and Codex's sandbox both leave your project folder writable, so a bad rm or rewrite inside it gets through.

Last checked: October 2026. Platform support and setting names come from each tool's current docs, linked below.

This page covers what agent sandboxes actually enforce, how Claude Code, Codex CLI, Gemini CLI and containers differ (including on Windows), what a sandbox can't protect, and a decision table for when you need a sandbox, undo or an approval prompt.

What a sandbox does for a coding agent

An agent runs shell commands as your user. Without isolation, a command can read your SSH keys, write anywhere your account can, and send data to any host. A sandbox puts an operating-system boundary around those commands, usually on two axes:

  • Filesystem isolation: writes are limited to the project folder and a temp folder. Some sandboxes also restrict reads.
  • Network isolation: outbound connections are blocked or routed through a proxy that checks each host against an allowlist.

Because the operating system enforces the boundary, it holds even if the model is tricked by a prompt injection. That's the main advantage over approval prompts, which depend on you reading every command.

How the Claude Code sandbox works

Claude Code's sandboxed Bash tool is off by default. Turn it on with /sandbox in a session, or "sandbox": {"enabled": true} in settings.

PlatformMechanism
macOSApple's built-in Seatbelt framework, no install needed
Linuxbubblewrap for filesystem isolation, socat to route traffic through the sandbox proxy
WSL2Same as Linux
WSL1Not supported ("Sandboxing requires WSL2")
Native WindowsNot supported: commands run unsandboxed

It's built on Anthropic's open-source @anthropic-ai/sandbox-runtime. Default boundaries:

AccessDefault
WritesThe working directory, a per-user temp directory and directories you've added. Protected paths stay write-denied
ReadsMost of the machine, including ~/.ssh and ~/.aws/credentials
NetworkNo direct route out. A local proxy checks each host against your allowed domains, which start empty
Environment variablesInherited, including any secrets in Claude Code's environment

Three gaps to know about:

  1. It wraps shell commands only. Claude's own Read, Edit and Write tools, MCP servers, hooks and commands you type at the ! prompt run outside it.
  2. There's an escape hatch. When a command fails in the sandbox, Claude can retry it with dangerouslyDisableSandbox, which goes through the normal permission flow. Set "allowUnsandboxedCommands": false to turn that off.
  3. Sandboxed commands auto-approve by default. With autoAllowBashIfSandboxed at its default of true, Bash commands that modify files inside the boundary run without a prompt, even in Manual mode.

To put the whole Claude Code process inside a boundary, Anthropic points to the sandbox runtime, a dev container, a custom container or a VM. Its comparison page says the Bash sandbox alone "is not sufficient for fully unattended runs."

Claude Code sandboxing on Windows

On native Windows, Claude Code's sandbox doesn't run. Anthropic's options for a Windows machine are:

  • WSL2: install Claude Code inside a WSL2 distribution and run /sandbox there. On Ubuntu 24.04 and later you may need an AppArmor profile so bwrap can create user namespaces; check with sysctl kernel.apparmor_restrict_unprivileged_userns.
  • A container or VM: a dev container in Docker Desktop, or a full VM.

Other agents differ here. Codex CLI has a native Windows sandbox (MXC, with older elevated and unelevated fallbacks) and uses its Linux sandbox under WSL2; WSL1 stopped being supported in Codex 0.115. Gemini CLI has a Windows Native Sandbox that uses icacls to set a Low Mandatory Level on files it needs to write, and those changes persist after the session.

Codex CLI sandbox

Codex turns its sandbox on by default. The docs describe three modes: read-only, workspace-write (the default for version-controlled folders) and danger-full-access.

  • macOS: Seatbelt, running commands through sandbox-exec with a profile for the selected mode.
  • Linux: bwrap plus seccomp.
  • Windows: MXC natively, or the Linux sandbox in WSL2.

In workspace-write, network access is off unless you enable it, and .git, .agents and .codex inside the workspace are read-only, recursively. That .git rule is a real advantage: the agent can't rewrite your history or hooks. But everything else in the workspace is writable.

Gemini CLI sandbox

Gemini CLI's sandbox is off by default but turns on automatically with --approval-mode=yolo. You can pick the method: macOS Seatbelt (profiles from permissive-open to strict-proxied; the default confines writes to the project but allows broad reads and network), Docker or Podman, gVisor on Linux, or LXC (experimental). Enable it with -s, GEMINI_SANDBOX, or tools.sandbox in settings.json.

Docker and dev containers

A container moves the whole agent, including its file tools and MCP servers, inside the boundary. Anthropic ships a reference dev container with a non-root user and a default-deny init-firewall.sh. OpenAI ships a Codex secure devcontainer with bubblewrap and an allowlist firewall. Docker offers Docker Sandboxes, a microVM for running agents.

What a sandbox can't protect

Every sandbox above leaves one thing writable: the project you asked the agent to work on. Anthropic says so directly. "Any approach that mounts your project directory writable can still modify that code," and in a dev container, "Claude can still modify any file in the bind-mounted workspace, which appears directly on your host."

So inside the boundary, a sandbox does nothing about:

  • rm -rf src/ or git clean -fdx run by mistake
  • a codemod, formatter or generator that rewrites hundreds of files badly
  • an npm install that changes package.json and the lockfile
  • uncommitted work the agent overwrote

And at the edges:

  • Allowed hosts can still leak data. Anthropic warns that allowing broad domains like github.com "can create paths for data exfiltration." Both Anthropic and OpenAI note that a malicious project inside a container can exfiltrate anything the container can read, including the agent's own credentials.
  • Allowed hosts can still be changed. If the sandbox allows your deploy target or database, a deploy or migration goes through.

Why undo is the complement

A sandbox shrinks the blast radius to your project. Undo covers the blast radius itself. They fail differently, which is why they work well together.

Most agent undo features only reverse the agent's own file edits. Claude Code's checkpointing "does not track files modified by Bash commands," and Codex CLI no longer has an undo command. The comparison is in AI coding agent checkpoints compared.

Darce approaches it from the undo side. Before every step that can change the project, it snapshots the working tree under private git refs, so /undo reverses files that shell commands created, changed or deleted, not just its own edits. It doesn't sandbox commands: they run as your user in the folder you started it from. Instead it scores every command, and in the default auto mode anything that reaches outside the machine (network, installs, deploys, migrations) or is destructive asks first. See Safety. The limits: undo history lasts for the session, gitignored files like node_modules aren't snapshotted, and shell effects are only covered in a git repo.

Darce has no OS-level sandbox, so on an untrusted repository a container is still the right call, with Darce or any other agent inside it.

Sandbox, undo or approval: a decision table

ThreatSandboxUndoApproval prompt
Agent deletes or rewrites files in the projectNo, the project is writableYesOnly if you read the command
Agent writes outside the project (~/.bashrc, ~/.ssh)YesNo, undo covers the project onlyYes
rm -rf ~ or /YesNoYes
Reads ~/.ssh or .env and sends it outOnly with read denies and network limitsNoYes, if network access asks
Prompt injection from a web page or READMEYes, limits what it can reachYes, for local file damageWeak: injected commands can look routine
Malicious postinstall scriptYes, in a containerOnly for files inside the projectYes, if installs ask
Deploy, migration, git push, publishOnly if the host isn't allowedNoYes, the only real defence
Bad refactor you notice ten steps laterNoYes, step back through historyNo
Untrusted repositoryYes, use a VM or containerHelps, not enoughNot enough

Short version: a sandbox for things outside the project, undo for things inside it, approval for things outside the machine.

Recipe: run an agent in a dev container

The fastest path for Claude Code is the official Dev Container Feature. Save as .devcontainer/devcontainer.json:

{
  "image": "mcr.microsoft.com/devcontainers/base:ubuntu",
  "features": {
    "ghcr.io/anthropics/devcontainer-features/claude-code:1.0": {}
  }
}

Then run Dev Containers: Rebuild Container in VS Code. For network egress limits, copy init-firewall.sh and the NET_ADMIN / NET_RAW runArgs from the reference container. Don't mount ~/.ssh or cloud credential files; use repository-scoped or short-lived tokens.

For any terminal agent, including Darce (Node.js 22+), a plain container works:

docker run --rm -it \
  -v "$PWD":/work -w /work \
  node:22 npx darce-cli

The agent can only write to /work, which is your project, and to the container's own disposable filesystem. Commit before you start, so git can restore anything inside /work that the agent breaks.

Recipe: read-only mounts

When the agent only needs to read something, mount it read-only with :ro. Docker enforces it, whatever the agent does:

docker run --rm -it \
  -v "$PWD":/work -w /work \
  -v "$HOME/reference-docs":/docs:ro \
  node:22 bash

For a review-only task, mount the project itself read-only (-v "$PWD":/work:ro) and add --network none if the agent doesn't need to call a model API from inside. Most do, so in practice you'll pair a read-only mount with an egress allowlist instead.

In Claude Code's built-in sandbox, the equivalent is read denies:

{
  "sandbox": {
    "enabled": true,
    "allowUnsandboxedCommands": false,
    "filesystem": {
      "denyRead": ["~/.ssh", "~/.aws", "~/**/.env"]
    }
  }
}

Remember this only covers shell commands. Add matching Read deny rules under permissions for Claude's own file tools. More on permission rules and running without prompts in YOLO mode with a safety net.

Sources

Frequently asked questions

Does Claude Code run in a sandbox by default?

No. The sandboxed Bash tool is off by default. Turn it on with /sandbox or by setting sandbox.enabled to true in settings. It uses Seatbelt on macOS and bubblewrap on Linux and WSL2.

Does Claude Code sandboxing work on Windows?

Not natively. On native Windows, Claude Code runs commands unsandboxed. Run Claude Code inside WSL2 to use the sandbox, or use a container or VM. WSL1 isn't supported.

Does a sandbox stop an AI agent deleting my code?

No. Every common agent sandbox leaves the project folder writable, so the agent can still delete or rewrite files there. Commit first, and use an agent whose undo covers shell commands.

What sandbox does Codex CLI use?

Seatbelt via sandbox-exec on macOS, bwrap plus seccomp on Linux, and MXC on native Windows, with the Linux sandbox under WSL2. The default workspace-write mode keeps network off and .git read-only.

Is a dev container enough to run an agent with permissions skipped?

It protects your host, but Anthropic and OpenAI both warn that a malicious project can still exfiltrate anything inside the container, including the agent's credentials. Restrict network egress and don't mount host secrets.

Related guides