Umbrel

Docker Compose: 5 Failure Modes That Pager You at 2AM

Stop 2am pager alerts. Learn the five Docker Compose failure modes—volume permissions, logs, resource limits, healthchecks, env vars—and how to fix them.

Audio Narration Listen to this article
00:00 / 00:00

Docker Compose: 5 Failure Modes That Pager You at 2AM

At 1:47 one Tuesday morning, my reverse proxy stopped resolving. The proxy hadn't died — the disk had. Twenty months of unrotated JSON logs from a single chatty container had quietly eaten 60 GB, and the whole stack collapsed under the weight. The fix was four lines of YAML. The diagnosis cost me an hour I didn't have, sitting in the dark with a laptop balanced on my knees.

Here's the thing: docker compose up is the easy part. What turns a clean compose file into a pager alert is everything the tutorial skips — permissions, logs, limits, startup order, and environment variables. Those five things are where self-hosted stacks actually break. This guide walks through each failure mode, with the YAML to prevent every one of them.

Failure mode 1: Permissions that fail silently

Bind a host folder into a container whose internal user is UID 0 or some random UID like 33 (www-data), and writes fail with Permission denied. That's the polite version. The nastier version: some apps catch the error, log it once, and keep running — looking perfectly healthy while writing absolutely nothing. Green dashboards, empty folders.

LinuxServer.io images solve this with PUID and PGID. Grab your host user's IDs with id -u and id -g, then set them in the service:

services: sonarr: image: lscr.io/linuxserver/sonarr:latest environment: - PUID=1000 - PGID=1000 - TZ=Europe/Lisbon volumes: - /srv/data/sonarr:/config

For images that don't support PUID/PGID, you've got two options: set user: "1000:1000" and chown the host folder once, or use a named volume. Named volumes inherit the ownership of the image's directory on first use; bind mounts don't — they take the host folder exactly as it is, warts and all.

Bind mount Named volume
Setup Path on the host Managed by Docker
Permissions Host UID applies as-is Inherits image directory ownership
Portability Breaks across machines Works anywhere
Backup rsync the folder docker run --rm -v vol:/data -v $(pwd):/backup alpine tar czf /backup/vol.tar.gz /data
Best for Config you edit by hand, LSIO apps Databases and state you never touch

Failure mode 2: Logs that eat your disk

The default json-file log driver is unbounded. Read that again: unbounded. One container with a debug-level logger can fill a disk in a weekend, and Docker will not lift a finger to stop it. Fix it daemon-wide in /etc/docker/daemon.json:

{
  "log-driver": "json-file",
  "log-opts": { "max-size": "10m", "max-file": "3" }
}

Then restart Docker with sudo systemctl restart docker. One catch worth knowing: this only affects newly created containers — existing ones keep their old settings, so anything already running needs a recreate to pick it up.

Or set limits per service in compose:

services:
  app:
    image: myapp:latest
    logging:
      driver: json-file
      options:
        max-size: "10m"
        max-file: "3"

That caps each container at 30 MB of logs — three files of 10 MB, rotated. To find the culprit before it bites you:

docker ps -q | xargs -I{} sh -c 'echo "{} $(docker logs {} 2>&1 | wc -c) bytes"'

Run that right now, on your current stack. I'll wait. If one of those numbers makes you wince, congratulations — you've just found your future 2am incident, while it's still cheap to fix.

Failure mode 3: One container starves the host

A single runaway container can consume all CPU and RAM, trigger the OOM killer, and take down your database, your reverse proxy, and your SSH session in one go. I've watched it happen — everything on the box degrades at once, and it's genuinely hard to tell which container is the arsonist from inside a dying SSH session.

Compose v2 honors deploy.resources.limits even with plain docker compose up, no swarm required:

services:
  app:
    image: myapp:latest
    deploy:
      resources:
        limits:
          cpus: "1.0"
          memory: "512M"

Set a limit that matches the app's real needs, not your wishful thinking. A database that needs 2 GB will fail with 512 MB — but it will fail loudly, on its own, instead of silently strangling the whole host. Loud failures are a gift. You want the database complaining at startup, not the kernel choosing victims at 3am.

Failure mode 4: depends_on doesn't mean "ready"

This one trips up almost everyone, and it's not obvious until it bites. depends_on only controls the order containers start. It does not wait for the database to accept connections. Your app starts, tries to connect, fails, and exits — or worse, retries forever in a crash loop while you stare at restart counts wondering what changed.

Fix it with a healthcheck and condition: service_healthy:

services:
  db:
    image: postgres:16-alpine
    healthcheck:
      test: ["CMD-SHELL", "pg_isready -U $${POSTGRES_USER} -d $${POSTGRES_DB}"]
      interval: 10s
      timeout: 5s
      retries: 5
      start_period: 30s

  app:
    image: myapp:latest
    depends_on:
      db:
        condition: service_healthy

Note the $$ — compose interpolates $ in the file itself, so $$ becomes a literal $ inside the container. Get this wrong and your healthcheck silently checks an empty variable, which is a fun hour of debugging you don't need. For web apps, use a health endpoint instead: test: ["CMD-SHELL", "curl -fsS http://localhost:3000/health || exit 1"]. Just remember that slim images often don't ship curl or wget — check before you commit to it.

Failure mode 5: Environment variables that leak or lie

Compose reads a .env file in the project directory for interpolation, and env_file: injects variables into the container. The two are easy to confuse, and the confusion has two classic failure modes: committing .env to git with real passwords in it, or relying on a variable that was never set and silently getting a default you didn't choose.

Keep secrets out of git:

echo ".env" >> .gitignore
chmod 600 .env

Use environment: for values you want visible in docker compose config, and env_file: for everything else. Then verify what compose actually resolves:

docker compose config

That prints the fully resolved configuration — the fastest way I know to catch a typo in a variable name or a missing .env entry. It's a thirty-second habit that has saved me from more than one "why is this empty in production" afternoon.

A stack that survives the night

Here's a complete example combining all five fixes. One deliberate choice worth pointing out: the database port is not published to the host. Containers on the same compose network reach each other by service name, so there's no reason to expose Postgres to the world — or even to your LAN.

services:
  db:
    image: postgres:16-alpine
    restart: unless-stopped
    volumes:
      - pgdata:/var/lib/postgresql/data
    environment:
      - POSTGRES_USER=${POSTGRES_USER}
      - POSTGRES_PASSWORD=${POSTGRES_PASSWORD}
      - POSTGRES_DB=${POSTGRES_DB}
    healthcheck:
      test: ["CMD-SHELL", "pg_isready -U $${POSTGRES_USER} -d $${POSTGRES_DB}"]
      interval: 10s
      timeout: 5s
      retries: 5
      start_period: 30s
    logging:
      driver: json-file
      options:
        max-size: "10m"
        max-file: "3"
    deploy:
      resources:
        limits:
          cpus: "1.0"
          memory: "512M"

  app:
    build: .
    restart: unless-stopped
    ports:
      - "8080:3000"
    depends_on:
      db:
        condition: service_healthy
    environment:
      - DATABASE_URL=postgres://${POSTGRES_USER}:${POSTGRES_PASSWORD}@db:5432/${POSTGRES_DB}
    logging:
      driver: json-file
      options:
        max-size: "10m"
        max-file: "3"
    deploy:
      resources:
        limits:
          cpus: "1.0"
          memory: "512M"

volumes:
  pgdata:

Debugging when it still breaks

Something will still break eventually — that's self-hosting. When a service won't start, work in this order:

docker compose ps -a
docker compose logs -f --tail=100 app
docker compose config

docker compose ps -a shows the actual state, including whether the healthcheck is passing. logs shows the crash. config shows what compose thinks it's running — which is sometimes very different from what you think it's running. If the container starts but the healthcheck fails, inspect it directly:

docker inspect --format='{{json .State.Health}}' app | jq

That dumps the healthcheck status and its last output as JSON, and jq makes it readable. Nine times out of ten, that failure output tells you exactly what's missing.

Three gotchas worth their weight

  • docker compose up -d won't recreate a container just because the image changed. It'll happily keep running the old one and tell you everything's fine. Use docker compose up -d --build or --force-recreate when you actually want a fresh container.
  • Healthcheck tools aren't guaranteed. curl and wget are missing from many slim images; use pg_isready, a tiny script, or a language-native check.
  • Port conflicts are the classic 2am surprise. Before you change anything, check what's already listening: sudo ss -tlnp | grep 8080.

FAQ

How do I set environment variables in docker compose?

Put them in a .env file next to your compose file for interpolation (${VAR}), or use env_file: to inject them into the container. For values you want visible in the resolved config, use environment:. And never commit .env to git — that's the one mistake on this list that follows you around.

How do I persist data with docker compose volumes?

Use a named volume for databases and state you never touch by hand — it's portable and Docker manages the ownership for you. Use a bind mount for config you want to edit or back up directly. Either way, check permissions first; it's the most common reason a "working" stack quietly isn't.

How do I limit CPU and memory in docker compose?

Use deploy.resources.limits with cpus and memory in the service definition. Compose v2 applies these limits with plain docker compose up, no swarm needed. Set them on anything that touches the network or the disk, honestly.

How do I debug a docker compose service that won't start?

Run docker compose ps -a to see the state, docker compose logs -f --tail=100 app for the error, and docker compose config to verify the resolved configuration. If the healthcheck is failing, inspect it with docker inspect. That order has never let me down.

The difference between a compose file that works and one that pages you at 2am is usually these five details. Add the healthchecks, cap the logs, set the limits, fix the permissions, and keep secrets out of git. None of it takes more than a few minutes, and all of it compounds. Your future self will thank you — from bed.