Docker Cloud Sandboxes: Why Agents Need MicroVMs, Not Containers
Docker put AI coding agents in hosted microVMs instead of containers, because a container was never the isolation boundary this job needed.
Key takeaways
Docker put AI coding agents in hosted microVMs instead of containers, because a container was never the isolation boundary this job needed.
Docker launched Cloud Sandboxes at WeAreDevelopers North America on September 24, letting AI coding agents like Claude Code run in hosted microVMs instead of on a developer's laptop, and the detail that matters is not "cloud." It is that Docker decided a container was the wrong isolation primitive for this job in the first place, built a custom VMM to replace it, and is now selling that same boundary as a metered cloud service.
Containers were built for a cooperative tenant#
A container isolates a process from its neighbors using namespaces and cgroups, all sitting on one shared kernel. That model assumes the thing inside is roughly cooperative: it runs the software you packaged, it does not try to escape, and if it does something unexpected, the blast radius is bounded by whatever kernel vulnerability happens to be live that week. An AI coding agent breaks that assumption on purpose. Its job is to install packages it has never seen, run scripts an LLM just generated, and make shell decisions no human pre-approved. That is not a misbehaving tenant; it is a tenant whose normal operation looks exactly like an attack. A shared-kernel boundary around that workload is a convenience feature, not a security one, the same conclusion our AI agent security guide reaches from the credential side rather than the kernel side.
What a microVM actually buys you#
A microVM gives each sandbox its own kernel and, in Docker's case, its own Docker daemon, with separate network, workspace, and credential layers per sandbox. The boundary moves from "a kernel feature that separates processes" to "a hardware feature that separates virtual machines," using Intel VT-x or AMD-V through hypervisors Docker already had reason to trust: Hypervisor.framework on macOS, Windows Hypervisor Platform on Windows, and KVM on Linux. Docker built a dedicated VMM for this rather than reusing an existing one, specifically so the thing boots in low hundreds of milliseconds, the number that makes microVMs usable for a coding agent instead of just a batch job. A VM that takes ten seconds to start is fine for a CI runner and unusable for a loop where the agent spins up, tests an idea, and tears down a dozen times an hour.
The cloud part is the boring, correct part#
Cloud Sandboxes are not a different security model from Docker's local sandboxes, which is the actual selling point. Same CLI, same trust model, same Kits, just running on Docker-managed compute that scales from 1 to 16 vCPUs instead of whatever your laptop happens to have free. The practical win is continuity: an agent chewing through a long migration or a large test matrix does not stop when you close the lid, because it was never tied to your hardware's isolation in the first place. Moving a sandbox is a single command:
$ sbx --cloud create --name cloud-project --allow-network github.com:443 claude
$ sbx --cloud exec cloud-project git clone https://github.com/docker/welcome-to-docker.git /home/agent/workspace/project
$ sbx --cloud attach cloud-project
Note the --allow-network flag on creation. The interesting control here isn't the VM boundary at all, it's that network and credential access are declared per sandbox up front, which matters once you remember an agent inside even a perfectly isolated VM can still exfiltrate anything it's handed a token for.
Kits are the part that outlives this product#
A Kit packages an agent, its tools, and its access rules, including network and credential scope, as a standard OCI image rather than a Docker-specific format. That's a bigger decision than it sounds. A Kit builds with ordinary container tooling, publishes to Docker Hub like any other image, and isn't locked to Docker's runtime if someone else implements the spec. Docker has committed to submitting the Kits specification to the CNCF, the tell that they want "package an agent and its guardrails as one artifact" to become ecosystem infrastructure rather than a Docker feature, the way OCI itself stopped being a Docker-only concern years ago. If that submission lands, the access rules baked into a Kit become auditable and transferable, exactly the property the industry was missing when Plugin4Shell showed that a pinned reference nobody verifies is not a control at all.
The bill arrives by the second, and that changes how you use it#
Docker meters Cloud Sandbox compute by the second rather than by the hour, pay-as-you-go against whatever vCPU and memory configuration the sandbox is running, with no recurring subscription fee layered on top. That pricing shape nudges behavior in a useful direction: there is no reason to leave a 16-vCPU sandbox idle between agent turns when stopping it is free and resuming costs nothing but boot time. The tradeoff is state. A sandbox runs on a timer you set between 1 and 24 hours, so you need to have already pulled out whatever the agent produced before it stops, because the whole point of per-second billing is that you don't pay for a VM that just sits there holding your work.
The decision, concretely#
- Running an agent that only edits files in a repo you already trust? A container is still fine; you don't need a microVM to isolate a typing assistant.
- Running an agent that installs arbitrary dependencies or executes code it just generated? Use a sandboxed microVM, local or cloud, same as running AI CLI agents in a CI pipeline already assumes for the credential side of this problem.
- Need the agent to keep working after your laptop closes, or need to fan out dozens of parallel runs? Cloud Sandboxes is the right shape, because the isolation model does not change when you move off your hardware.
- Packaging an agent's environment for your team or for distribution? Build it as a Kit now rather than a bespoke Dockerfile, so you are not rewriting it if the CNCF submission actually standardizes the format.
The call we'd make#
Treat the microVM boundary as the default for anything that runs model-generated code unattended, and treat the cloud offering as a convenience on top of a security decision Docker had already made locally, not a new one. The part worth watching closely is the Kits spec's path through CNCF, because that is what determines whether "package the agent and its guardrails together" becomes something every vendor implements the same way, or stays another proprietary wrapper with better marketing. Until that lands, build Kits like you'd build any OCI image you intend to keep using after the vendor that invented them changes its mind.
Get the DevOps Troubleshooting Cheat Sheet
Subscribe and get our free one-page reference for the errors that eat an afternoon — CrashLoopBackOff, OOMKilled, Terraform state locks, and more — plus new guides as we publish them.
GKE Pod Snapshots: Cold Starts Drop 89%, If Your Nodes Match
Google's benchmarks show a 70B model restoring in 37 seconds instead of minutes. The catch is a hash and a hardware match that silently refuses to restore when either is off.
Storm-3068: A CI/CD Pipeline Is a Kubeconfig Exfiltration Machine
Microsoft's Storm-3068 report used zero malware to steal Kubernetes credentials, just a password reset, a pipeline edit, and permissions nobody had scoped down.
More from DevOps
Explore more articles in this category
GitLab's New Rate Limits: What to Fix Before Oct 19
GitLab is capping unauthenticated API calls at 60 an hour starting October 19, and the preview windows land before most teams will have noticed.
Storm-3068: A CI/CD Pipeline Is a Kubeconfig Exfiltration Machine
Microsoft's Storm-3068 report used zero malware to steal Kubernetes credentials, just a password reset, a pipeline edit, and permissions nobody had scoped down.
Zed Delta and the Pull Request: Obsolete, or Just Overloaded?
Zed says agents made pull requests obsolete and shipped Delta to replace them. The review unit needs rethinking, but review itself, and the gates around it, stay.
You might have missed
Evergreen posts worth revisiting.