PID 1 in Containers: Signals, Zombies, and Why Your Container Won't Stop
Inside a container your app is PID 1, and PID 1 is special: the kernel won't apply default signal handlers to it and it must reap orphaned children. Why docker stop takes 10 seconds, why shell-form CMD swallows SIGTERM, zombies, and the fixes: exec form, tini/--init, signal handling.
A container starts one process, and inside its PID namespace that process is PID 1 — the same position init/systemd holds on a normal Linux system. (Linux namespaces explained) PID 1 has two special responsibilities that ordinary programs were never written for. Getting them wrong causes three classic symptoms:
docker stopalways takes 10 seconds, then the container is killed.- Deploys drop in-flight requests or corrupt work because the app never shut down gracefully.
- Zombie processes pile up until the container can't create new ones.
Special rule 1: no default signal actions for PID 1
Normally, if a process hasn't installed a handler for SIGTERM or SIGINT, the kernel applies the default action: terminate. For PID 1 in a PID namespace, the kernel ignores signals that have no handler installed (except SIGKILL and SIGSTOP sent from the parent namespace).
So a program that relies on default behaviour — many do — simply doesn't react to SIGTERM when it's PID 1. docker stop sends SIGTERM, waits the grace period (10 s by default), then sends SIGKILL. Your app never got to finish requests, close database connections or flush logs.
Node.js is a common case: without a SIGTERM listener, a Node process running as PID 1 ignores it.
Special rule 2: PID 1 reaps orphans
When a process exits, it becomes a zombie until its parent calls wait() to collect its exit status. If the parent has already died, the orphan is re-parented to PID 1, which is expected to reap it.
Your app isn't an init system. If it spawns children that spawn grandchildren (a shell script, headless Chrome, git, a language server, an AI agent running commands), orphaned grandchildren get re-parented to your app, which never reaps them. Zombies hold PID slots; with a cgroup pids.max limit, the container eventually fails with fork: Resource temporarily unavailable. (Container CPU and memory limits)
ps -eo pid,ppid,stat,cmd | awk '$3 ~ /Z/' # list zombies
Trap: shell-form CMD
CMD npm start # shell form
runs /bin/sh -c "npm start". Now sh is PID 1. sh doesn't forward SIGTERM to its child, so your app never hears it — even if it handles signals perfectly. And npm start adds another layer: npm as the parent of node.
Use the exec form, and run the runtime directly:
CMD ["node", "dist/server.js"]
Now node is PID 1 and receives signals directly. (Running via npm start is worth avoiding in production containers for this reason.) (Production Dockerfile for Node)
Trap: entrypoint scripts
Entrypoint scripts are useful for setup (waiting for a database, templating config). End them with exec so the app replaces the shell as PID 1:
#!/bin/sh
set -e
./bin/migrate-if-needed
exec node dist/server.js "$@"
Without exec, the shell stays PID 1 and swallows signals.
The fix for both rules: a tiny init
tini (and dumb-init) is a minimal init designed for containers. It runs as PID 1, forwards signals to your app, and reaps zombies.
Docker has it built in:
docker run --init myapp
# docker-compose.yml
services:
app:
image: myapp
init: true
or bake it into the image:
RUN apt-get update && apt-get install -y --no-install-recommends tini && rm -rf /var/lib/apt/lists/*
ENTRYPOINT ["/usr/bin/tini", "--"]
CMD ["node", "dist/server.js"]
In Kubernetes, there's no --init flag; bake tini into the image, or enable shareProcessNamespace (the pause container then reaps zombies, at the cost of sharing the PID namespace between containers in the pod).
The app still has to shut down properly
An init forwards SIGTERM; your app must do something with it: stop accepting new connections, finish in-flight requests, drain job workers, close the database pool, then exit.
const server = app.listen(port)
function shutdown(signal) {
console.log(`${signal} received, shutting down`)
server.close(async () => {
await db.end()
process.exit(0)
})
setTimeout(() => process.exit(1), 25_000).unref() // hard stop before the orchestrator's SIGKILL
}
process.on('SIGTERM', shutdown)
process.on('SIGINT', shutdown)
Make sure your internal timeout is shorter than the platform's grace period (docker stop -t, Compose stop_grace_period, Kubernetes terminationGracePeriodSeconds). Fail your readiness check as soon as shutdown begins so load balancers stop sending traffic. (Graceful shutdown in Node, health check endpoints, zero-downtime deploys)
STOPSIGNAL
Some programs expect a different signal for graceful shutdown (Nginx uses SIGQUIT for graceful stop; some apps use SIGINT). Set it in the image:
STOPSIGNAL SIGQUIT
Quick diagnosis
docker exec myapp ps -o pid,ppid,cmd # who is PID 1? sh? npm? node?
time docker stop myapp # ~10 s means SIGTERM was ignored
docker inspect -f '{{.State.ExitCode}}' myapp # 137 = killed by SIGKILL (or OOM)
The summary
- PID 1 gets no default signal handling and must reap orphaned processes.
- Shell-form
CMDand entrypoint scripts withoutexecput a shell at PID 1 that swallowsSIGTERM. - Use exec-form
CMD,execin scripts, and tini /--init/init: true. - Handle
SIGTERMin the app: stop accepting, drain, close, exit before the grace period ends. - A 10-second
docker stopand exit code 137 are the tell-tale signs.
EasySpawn runs your app and Claude Code's processes on a full VM with a real init system, so long-running agents, child processes and graceful restarts behave like they do on any Linux server. See how it works or join the waitlist.
Related: Graceful Shutdown in Node.js · Writing a Production Dockerfile · PM2 vs systemd · Docker Image Layers and OverlayFS
Keep reading
Docker Image Layers and OverlayFS: How Container Filesystems Really Work
A Docker image is a stack of read-only layers merged by OverlayFS, with a thin writable layer per container. How layers are built and cached, content addressing and digests, copy-up and whiteouts, why deleting files doesn't shrink images, and the performance traps for databases.
Container Networking Internals: veth, Bridges, NAT, and Embedded DNS
What happens when a container sends a packet: network namespaces, veth pairs, bridges, NAT for egress and published ports, why published ports bypass firewalls like ufw, Docker's embedded DNS, inter-container isolation, and debugging with nsenter and tcpdump.