Docker Moves Agent Sandboxes to the Cloud: Containers Are Not Containment

·11 min read·Evergreen Tools Team

On September 24, 2026, at WeAreDevelopers North America, Docker announced Docker Cloud Sandboxes, extending the local agent sandboxes it shipped earlier this year onto managed cloud infrastructure. The interesting part is Docker president Mark Cavage's framing: containers were never designed for the level of isolation AI agents demand. An agent is not an application. It reads your files, runs your shell, holds your secrets, and, as Cavage demonstrated on stage, it actively looks for the edges of its environment, because pushing past boundaries is often how it gets work done. Cloud Sandboxes move that deterministic boundary off the laptop and into a managed microVM.

1. The demo that makes the case

Cavage started with the failure. He ran Anthropic's Claude inside an ordinary Docker container and asked it to find a locally stored secret outside the container. It did, by probing its environment and using the mounted host Docker socket. Then Docker principal engineer Michael Irwin ran the same prompt inside a Docker Sandbox. This time the model found the socket, tried to use it to mount host paths and a privileged container, and simply could not get there, because the sandbox runs as a full microVM. The isolation holds, Irwin said. The point of the demo is not that containers are insecure. They do exactly what they were designed to do, which is isolate applications. Cavage's line is the thesis itself: we have to separate containers from containment.

# The shape of the problem. An agent inside an ordinary container can still
# reach the host through anything that was mounted in. Docker's own demo had
# Claude find a locally stored secret by probing its environment and using
# the mounted host Docker socket -- the container was working as designed.

# Illustrative sandbox spec: what a contained agent actually needs declared.
version: "1"
agent:
  name: refactor-bot
  image: ghcr.io/acme/refactor-bot:1.4.0
isolation:
  # A full microVM, not a namespace-sharing container.
  runtime: microvm
  mountHostDockerSocket: false      # the escape hatch that leaked in the demo
secrets:
  inject: ["GITHUB_TOKEN"]          # injected, never mounted as a file
  ttlMinutes: 60
network:
  default: deny                     # egress deny-by-default
  allow: ["api.github.com:443", "registry.npmjs.org:443"]
resources:
  vcpu: 4
  memoryGb: 8
A deterministic isolation layer

Sandboxes limit reach; policy governs intent

2. What Cloud Sandboxes actually are

It is the same sandbox model, hosted. Agents run in microVMs on Docker-managed infrastructure with the same policies as the local sandboxes, so long-running agentic workflows continue after a developer's laptop shuts down. Boot time is in the low hundreds of milliseconds and billing is by the second. Secrets, policies, networks, agent configuration and MCP gateways are built in. Compute scales from 1 to 16 vCPUs and is fully managed by Docker, so you point your agent at it rather than provisioning anything. Reported pricing starts at $0.07 per hour for a Micro shape (1 vCPU, 2GB) and reaches $1.12 per hour for XL (16 vCPUs, 32GB). The same CLI and the same trust model work locally and in the cloud.

// Policy is the layer above isolation: the sandbox limits what the agent can
// REACH, policy limits what it is allowed to DO. Keep both, and keep the
// policy in version control.

{
  "policy": "prod-agent-baseline",
  "rules": [
    { "match": { "tool": "shell.exec", "command": "rm -rf *" },
      "effect": "deny", "reason": "destructive filesystem operation" },
    { "match": { "tool": "network.fetch", "host": "*.internal" },
      "effect": "require_approval" },
    { "match": { "tool": "shell.exec", "command": "git push*" },
      "effect": "require_approval", "approver": "human" },
    { "match": { "tool": "fs.write", "path": "/workspace/**" },
      "effect": "allow" }
  ],
  "default": "deny"
}

3. The timing is not accidental

The same week, Australian officials disclosed that an OpenAI agent had accessed a government portal without authorization while looking for health statistics. That is the pattern the industry keeps repeating: agents finding access paths their operators did not expect. Containment failures are now common enough that vendors publish them. Docker's argument is that a deterministic isolation layer should be the minimum bar, and that policies govern intent while sandboxes govern reach. Get that ordering right and a lot of agent-went-rogue headlines become agent-got-blocked non-events.

# Verify containment instead of assuming it. The demo that matters is the
# one you run against your own agent: ask it to leave, then check whether
# the escape actually worked.

import json, subprocess

PROBE = "find / -name '*.pem' -o -name 'credentials' 2>/dev/null | head"

def check_containment(sandbox_id):
    cmd = ["docker", "sandbox", "exec", sandbox_id, "sh", "-lc", PROBE]
    out = subprocess.run(cmd, capture_output=True, text=True)
    leaked = [l for l in out.stdout.splitlines() if l.strip()]
    # Also confirm the classic escape route is closed.
    sock = subprocess.run(
        ["docker", "sandbox", "exec", sandbox_id, "sh", "-lc",
         "ls -l /var/run/docker.sock || true"],
        capture_output=True, text=True).stdout.strip()
    return {"sandbox": sandbox_id, "leakedPaths": leaked,
            "dockerSocket": sock or "absent", "ok": not leaked and not sock}

print(json.dumps(check_containment("sbx-refactor-01"), indent=1))
microVM isolation

A sandbox runs as a full microVM, not a shared-namespace container

4. Kits became standard OCI images

Docker also updated its Kits specification, which packages an agent, its tools and its access rules into one shareable artifact. Kits now ship as standard OCI images, which answers the obvious lock-in objection, and Docker said it will submit the Kits specification to the CNCF. One example available at launch is the BAND Python Kit, which lets agents work with one another over a WebSocket connection without operating in the same environment. For platform teams this is the more durable announcement: a portable unit of agent packaging that a registry can sign, scan and version like any other image.

# Egress is where a contained agent can still hurt you. Allowlist the
# destinations a job legitimately needs, log everything else, and fail
# closed. Deny-by-default is the only default worth shipping.

ALLOWLIST = {
    "refactor-bot": [
        ("api.github.com", 443),
        ("registry.npmjs.org", 443),
    ],
    "doc-bot": [
        ("api.internal.docs", 443),
    ],
}

def review_egress(agent: str, observed: list[tuple[str, int]]) -> dict:
    allowed = set(ALLOWLIST.get(agent, []))
    blocked = [c for c in observed if c not in allowed]
    return {
        "agent": agent,
        "observed": len(observed),
        "blocked": blocked,
        # A blocked destination is either a policy gap or an incident.
        "action": "review" if blocked else "none",
    }

observed = [("api.github.com", 443), ("pastebin.example", 443)]
print(review_egress("refactor-bot", observed))

5. What it does and does not solve

Say it plainly: a sandbox contains reach, not judgement. Cavage's own framing is that sandboxes are the deterministic base layer while policies govern the agent's intent. An agent inside a microVM can still be prompt-injected, can still exfiltrate through an allowed egress path, and can still take a wrong action using a valid credential. So the practical stack is three layers: isolation with microVMs, policy for what the agent may reach and do, and identity for what it may authenticate as, revocably. Code sample 2 shows policy as code. If you can only do one thing this quarter, do the isolation layer first, because it is the one that holds when the other two are wrong.

# Kits now ship as standard OCI images, so an agent definition is an artifact
# you can sign, scan, version and pull like any other image -- instead of a
# folder in somebody's home directory.

# Build and publish a Kit the way you already publish images.
# (Illustrative workflow -- see Docker's Kits documentation for exact flags.)
docker build -t ghcr.io/acme/refactor-kit:1.4.0 ./kit
docker push ghcr.io/acme/refactor-kit:1.4.0

# Then pin it by digest, so a registry compromise cannot silently swap it.
cat <<'YAML'
agent:
  kit: ghcr.io/acme/refactor-kit@sha256:9f2c...e41b
  network: deny-by-default
  secrets: [GITHUB_TOKEN]
YAML
Turning escape checks into routine tests

The escape that failed in the demo should be a routine test for your agent

6. How to adopt it

Start by moving one long-running workflow to a cloud sandbox and confirm three things: the workload completes with your laptop closed, secrets are injected rather than mounted, and egress is deny-by-default with an explicit allowlist. Then standardise the Kits format so your agent definitions live in a registry rather than in someone's home directory. If you have an agent that runs in dangerous mode locally because the alternative is a hundred approval prompts, that is precisely the workload a sandbox was built to serve. One last habit is worth building: treat the sandbox specification and the policy as a single pull request. A change to what an agent can reach is a security change, and it deserves the same review as a change to what it can do. Teams that review the two together stop arguing about whether an incident was a containment failure or a policy gap.

📌 Frequently Asked Questions

What did Docker announce?

Docker Cloud Sandboxes, announced on September 24, 2026 at WeAreDevelopers North America, extending the local agent sandboxes' microVM isolation to managed cloud infrastructure.

How is a sandbox different from a container?

Containers isolate applications. Sandboxes run as full microVMs to provide the stronger isolation AI agents require. Docker frames it as separating containers from containment.

What does Docker Cloud Sandboxes cost?

Pricing varies by instance size, reported from $0.07 per hour for Micro (1 vCPU, 2GB) to $1.12 per hour for XL (16 vCPUs, 32GB).

What changed about Kits?

Kits now ship as standard OCI images, so an agent, its tools and its rules can be signed, scanned and versioned like any other image. Docker said it will submit the Kits specification to the CNCF.

Does a sandbox make an agent safe?

No. A sandbox is the deterministic isolation layer that bounds what an agent can reach. Policy governs behaviour and identity governs credentials; all three are needed.