Docker AI Without Desktop: Local GPU Inference on Linux
Skip Docker Desktop bloat on Linux servers. Configure native Docker Engine with GPU pass-through for local AI inference and free security scans.
Docker spent its stage time at the WeAreDevelopers conference spotlighting enterprise integrations: proprietary AI plugins, branded container catalogs, and commercial vulnerability scanning dashboards. For platform vendors and enterprise development teams tethered to Docker Desktop subscriptions, that product focus makes commercial sense.
For engineers running headless Linux servers, home lab racks, or small clusters, that conference lineup tells a very different story.
You do not need vendor-bundled desktop extensions to build a dependable local inference stack. You also do not need a paid Docker Hub subscription to audit container vulnerabilities. Stripping away the ecosystem marketing leaves you with Docker Engine, Compose v2, and lightweight open-source command-line utilities. Here is how the reality of running containers compares to the marketing hype, and how you can run high-throughput local inference and vulnerability scans on plain Linux.
Docker Desktop vs. Docker Engine: The Homelab Reality
Docker Desktop wraps a virtual machine, a desktop GUI, and an extension marketplace around standard container tooling. On macOS and Windows, that VM layer acts as a necessary translation bridge because those operating systems lack native Linux kernel primitives like cgroups and namespaces.
On a native Linux host, that architecture makes little sense. It introduces an unnecessary virtualization boundary, burns system memory, and runs persistent background helper daemons.
text +-----------------------------------------------------------+ | Docker Desktop (Linux/Mac/Win) | | [ Electron UI ] -> [ Background Helper VMs ] -> [ Daemons]| | Idle Footprint: ~1.8 GB to 3.5 GB RAM | +-----------------------------------------------------------+ vs +-----------------------------------------------------------+ | Docker Engine Community Edition (Linux Native) | | [ systemd dockerd ] -> [ containerd ] -> [ runc ] | | Idle Footprint: ~45 MB RAM | +-----------------------------------------------------------+
When you deploy local AI workloads, system memory and PCIe bus bandwidth are precious resources. An 8-billion parameter (8B) quantized model needs roughly 5.5 GB of VRAM just to sit in memory, plus several gigabytes of host RAM for prompt caching and context management. Wasting 2 GB of RAM on an Electron interface makes no sense on headless hardware.
| Feature | Docker Engine (CE 27.x) | Docker Desktop for Linux |
|---|---|---|
| Idle Host RAM Usage | ~45 MB | 1.8 GB – 3.5 GB |
| GPU Pass-through | Native via NVIDIA CDI | Virtualized bridge layer |
| Host OS Integration | Direct systemd service | Isolated VM / User namespace |
| License Model | Open Source (Apache 2.0) | Commercial / Subscription tiers |
| Security Auditing | Pluggable CLI (Trivy/Syft) | Docker Scout / Proprietary Hub UI |
| Best Used For | Homelabs, headless nodes, CI/CD | Mixed-OS developer laptops |
If your host runs Debian 12 or Ubuntu 24.04 on bare metal, skip the Desktop installer completely. Install native Docker Engine from Docker's official apt repository and configure hardware access directly.
GPU Access Without Desktop Extensions: NVIDIA CDI and Compose
Many vendor presentations present one-click AI setups inside Docker Desktop as the only convenient path to local machine learning. Under the hood, running local inference on native Linux requires only two things: your host's NVIDIA graphics drivers and the NVIDIA Container Toolkit configured with the Container Device Interface (CDI).
For years, passing an NVIDIA GPU into a container required modifying /etc/docker/daemon.json to inject a custom runtime (nvidia-container-runtime) and passing --gpus all on the command line. This setup was brittle. Upgrading Docker often broke the runtime path, and Docker Compose handled GPU reservations inconsistently across different versions.
CDI fixes this problem. Developed by the Container Orchestrated Device Workgroup, CDI uses a vendor-agnostic specification to describe host devices using a simple YAML file. Container runtimes like containerd and runc read this file directly. They configure device nodes (/dev/nvidia*), Unix capabilities, and driver libraries without needing custom runtime wrappers or altered daemon configs.
1. Generate the CDI Specification
First, install the nvidia-container-toolkit package using your Linux distribution's package manager. Confirm your NVIDIA kernel modules (nvidia and nvidia_uvm) are active on the host:
# Verify GPU visibility and driver version on the host
nvidia-smi
Next, generate the CDI specification file. This tool inspects your graphics hardware and writes device definitions to /etc/cdi/nvidia.yaml:
# Generate the CDI specification
sudo nvidia-ctk cdi generate --output=/etc/cdi/nvidia.yaml
# Confirm Docker and the Container Toolkit recognize the CDI devices
nvidia-ctk cdi list
You should see output listing devices such as nvidia.com/gpu=all alongside individual GPU indexes like nvidia.com/gpu=0. Docker can now reference your physical cards using these identifiers.
2. Compose Configuration for Local Inference
This production Compose file deploys an Ollama inference backend paired with the Open-WebUI interface. It reserves GPU access using the CDI driver and sets explicit hardware limits to prevent the container from overwhelming your server:
services:
ollama:
image: ollama/ollama:0.5.7
container_name: ollama-inference
restart: unless-stopped
environment:
# Keep the model in VRAM for 24 hours to prevent reload latency
- OLLAMA_KEEP_ALIVE=24h
# Limit parallel request processing to match card capacity
- OLLAMA_NUM_PARALLEL=2
volumes:
# Store model weights on persistent host storage
- /opt/ollama/data:/root/.ollama
deploy:
resources:
reservations:
devices:
- driver: cdi
device_ids:
- nvidia.com/gpu=all
limits:
cpus: '4.0'
memory: 14G
# Prevent inference buffers from writing to swap space
mem_swappiness: 0
webui:
image: ghcr.io/open-webui/open-webui:main
container_name: open-webui
restart: unless-stopped
ports:
- "3000:8080"
environment:
- OLLAMA_BASE_URL=http://ollama:11434
volumes:
- /opt/open-webui/data:/app/backend/data
depends_on:
- ollama
deploy:
resources:
limits:
cpus: '2.0'
memory: 2G
This CDI configuration bypasses legacy runtime hooks. It passes host device paths directly to the container, leaving zero virtualization overhead between the CUDA runtime and the physical hardware.
Managing Runaway Contexts: Docker Memory Limits
Local large language models introduce an operational failure mode that rarely occurs with traditional web applications: context bloat.
When you start an 8B model quantized at 4-bit precision, it sits comfortably inside roughly 5 GB of VRAM. Everything runs smoothly until a user pastes a massive log file or document into the chat interface. As the context window expands from 2,048 tokens to 32,768 tokens, the model's Key-Value (KV) cache grows significantly.
[ Context: 2,048 tokens ] -> KV Cache: ~250 MB -> Fits in VRAM
[ Context: 32,768 tokens ] -> KV Cache: ~4.2 GB -> Overflows VRAM into Host RAM
When an NVIDIA card runs out of VRAM, the inference runtime (such as llama.cpp or vLLM) spills memory allocations over to system RAM. If your system RAM fills up, the Linux kernel Out-Of-Memory (OOM) killer activates.
Without explicit container memory limits, the OOM killer selects targets using a heuristic algorithm (oom_badness). It often terminates the parent dockerd process, your SSH daemon, or database processes rather than the specific container worker that requested the extra memory.
To protect your system from crashing during long prompts, isolate your memory allocations directly in your Compose file:
deploy:
resources:
limits:
memory: 14G
reservations:
memory: 8G
Three configuration settings protect the host:
mem_swappiness: 0: By default, the Linux kernel moves idle memory pages to swap space on disk. If an LLM spills its weights or context layers onto a swap partition—even on a fast NVMe drive—generation speeds drop from 40 tokens per second to fractions of a token per second. This causes high I/O wait times that can freeze your storage controllers. Settingmem_swappiness: 0instructs the kernel to drop page caches instead of swapping container memory.- Hard Memory Ceilings (
limits.memory): Always set this limit below your host's total physical memory. On a server with 16 GB of RAM, cap the inference container at 12 GB or 14 GB. If a massive prompt exhausts that limit, the container's internal allocator fails or the container restarts on its own, leaving the host operating system completely stable. oom_score_adj: On headless setups, adjust the container's OOM score relative to other system processes:
# Add to your compose service definition
oom_score_adj: 500
The Linux kernel scores process survival from -1000 (never kill, used by system daemons) to 1000 (kill first). Standard processes sit at 0. Assigning a score of 500 ensures that if memory pressure spikes, the kernel terminates the inference engine before touching critical host services.
Open Container Vulnerability Scanning: Trivy and Syft
Vendor keynotes often emphasize automated cloud vulnerability scanners tied to paid subscription tiers. These tools require you to push your container images to an external registry before scanning them for known Common Vulnerabilities and Exposures (CVEs).
You can run thorough security scans locally on your server without sending private images to third-party services. Using open command-line tools like Trivy (maintained by Aqua Security) and Syft (maintained by Anchore), you can inspect images, review package versions, and build a complete Software Bill of Materials (SBOM) on your machine.
The workflow breaks down into two distinct operations:
- Syft catalogs every software artifact, binary, Python wheel, and OS package baked into the container layers, exporting them into structured formats like SPDX or CycloneDX.
- Trivy compares that package list against public vulnerability databases (including the NVD and distro security trackers), highlighting unpatched packages and sorting them by severity.
Here is an automated shell script that inspects any locally built image:
#!/usr/bin/env bash
set -euo pipefail
# Target image name and output location
IMAGE_TAG="local/custom-ai-worker:latest"
REPORT_DIR="/var/log/container-audits"
mkdir -p "${REPORT_DIR}"
echo "==> Step 1: Cataloging packages with Syft (SBOM)..."
# Generates a standard SPDX JSON document listing all installed components
syft packages "${IMAGE_TAG}" -o spdx-json > "${REPORT_DIR}/sbom.json"
echo "==> Step 2: Scanning image layers for vulnerabilities with Trivy..."
# Scans the image locally, skipping vendor noise and focusing on actionable fixes
trivy image \
--severity HIGH,CRITICAL \
--ignore-unfixed \
--exit-code 1 \
--format table \
--output "${REPORT_DIR}/scan-report.txt" \
"${IMAGE_TAG}" || EXIT_STATUS=$?
if [ "${EXIT_STATUS:-0}" -ne 0 ]; then
echo "[ALERT] Actionable security vulnerabilities found in ${IMAGE_TAG}."
echo "Check the report at: ${REPORT_DIR}/scan-report.txt"
exit 1
fi
echo "Scan complete: No unpatched HIGH or CRITICAL issues found."
This local process takes a few seconds to run, requires no account logins, avoids registry rate limits, and keeps internal container images private.
+-----------------------+ +-----------------------+
| Container Image | ---> | Syft: Generates SBOM |
| (Local Engine) | | (Packages & Deps) |
+-----------------------+ +-----------------------+
| |
v v
+------------------------------------------------------+
| Trivy: Cross-references packages with CVE databases |
| Outputs: Plain text tables or JSON vulnerability logs|
+------------------------------------------------------+
Common Pitfalls to Avoid
- Relying on deprecated
--gpusflags: Modern Docker deployments manage hardware access through Compose v2 device reservations and CDI. Avoid legacy wrapper scripts that require manual edits todaemon.jsonwith the oldnvidia-container-runtime. - Leaving model context limits unconfigured: Setting container memory limits solves only part of the problem. If your inference backend (such as Ollama, llama.cpp, or vLLM) lacks a designated
num_ctxceiling, a client can request an oversized token context window that exhausts both VRAM and host RAM simultaneously. Always set clear context limits in your service parameters or model modelfiles. - Storing models inside container layers: Never let inference weights write directly to the container's
overlay2storage layer. When weights are stored inside ephemeral container layers, updating an image deletes your downloaded models or fills your root partition. Mount an external directory on your host to store model weights cleanly. - Running desktop GPUs without persistent fan or power profiles: Desktop graphics cards (like the RTX 3080 or 4090) drop into aggressive low-power sleep states when containers stop. When a new inference request arrives, sudden power spikes can trigger voltage drops on budget power supplies. Set the persistence mode flag on the host using
nvidia-smi -pm 1to keep driver states stable across container runs. - Exposing unauthenticated model endpoints: By default, Ollama and similar runtimes do not enforce API authentication. Never expose port
11434directly to an open network. Keep the inference service within an internal Docker network, and route user traffic through a reverse proxy like Traefik, Caddy, or Nginx with proper authentication headers.
Frequently Asked Questions
How do I allocate GPU resources to Docker containers using Compose v2?
Add an explicit reservations block under deploy.resources in your Compose file. By pairing the NVIDIA Container Toolkit with the Container Device Interface (CDI), you specify the cdi driver and target the device ID nvidia.com/gpu=all or individual IDs (like nvidia.com/gpu=0). This configuration gives the container direct GPU access without legacy flags or graphical desktop interfaces.
Can I run container security scans without a commercial Docker Hub subscription?
Yes. Open-source CLI utilities like Trivy and Syft scan local container images directly on your host. They check image layers against public vulnerability feeds and export Software Bill of Materials (SBOM) reports in standard formats like SPDX and CycloneDX without requiring commercial subscriptions or cloud logins.
What is the difference between Docker Desktop partner extensions and running containers on Docker Engine?
Docker Desktop partner extensions are graphical add-ons running inside an Electron dashboard and a lightweight Linux virtual machine. Running native Docker Engine on Linux skips the hypervisor and desktop interfaces entirely. It uses roughly 45 MB of RAM at idle instead of several gigabytes, while communicating directly with host hardware like GPUs, storage controllers, and network devices.
How do I prevent out-of-memory (OOM) kernel panics when running local AI containers in Docker?
Combine hard memory limits (limits.memory in Compose) with a swappiness setting of 0 (mem_swappiness: 0). This stops the container from thrashing disk swap memory, prevents host memory exhaustion, and ensures the Linux kernel terminates only the runaway inference process instead of taking down the entire server.
Does NVIDIA CDI work with multi-GPU setups in Docker Compose?
Yes. You can target individual cards by their CDI device names. Run nvidia-ctk cdi list on the host to see your available devices. In your Compose file, replace nvidia.com/gpu=all with specific IDs such as nvidia.com/gpu=0 or nvidia.com/gpu=1 to pin separate containers to separate physical cards.
Related in this cluster
- /en/posts/cursor-xai-vs-local-llms-privacy-guide-for-homelabs/
- /en/posts/docker-cloud-sandboxes-secure-ai-agents-without-vendor-lock-in/
- /en/posts/docker-sandbox-kit-securing-ai-agent-container-isolation/
Ad space · not an Umbrel endorsement