Blog
5 min read

Docker Image Layers and OverlayFS: How Container Filesystems Really Work

A Docker image is a stack of read-only layers merged by OverlayFS, with a thin writable layer per container. How layers are built and cached, content addressing and digests, copy-up and whiteouts, why deleting files doesn't shrink images, and the performance traps for databases.

When a container starts in milliseconds from a 500 MB image without copying anything, that's layers and a union filesystem — normally OverlayFS via Docker's overlay2 storage driver. Understanding it explains build caching, image size, why rm in a later step doesn't help, and why databases shouldn't write to the container filesystem.

An image is a stack of layers

Each filesystem-changing instruction in a Dockerfile (RUN, COPY, ADD) produces a layer: a tarball of the files that instruction added, changed or deleted. Metadata instructions (ENV, CMD, EXPOSE) change the image config, not the filesystem.

docker history node:24-slim
docker image inspect myapp --format '{{json .RootFS.Layers}}'

Layers are content-addressed: identified by the SHA-256 digest of their contents. Two images built FROM the same base share those layers on disk and in registries — pulled and stored once. An image reference by digest (myapp@sha256:…) is immutable; a tag (myapp:latest) is a movable pointer.

OverlayFS: merging layers into one view

OverlayFS presents several directories as one:

  • lowerdir — one or more read-only layers (the image).
  • upperdir — a writable layer (the container's).
  • workdir — scratch space OverlayFS needs.
  • merged — the combined view the container sees as /.

You can do it by hand:

mkdir lower upper work merged
echo "from image" > lower/a.txt
sudo mount -t overlay overlay -o lowerdir=lower,upperdir=upper,workdir=work merged
cat merged/a.txt             # from image
echo "changed" > merged/a.txt
cat lower/a.txt              # still "from image"
cat upper/a.txt              # "changed" — copied up and modified

That's a container's filesystem: image layers as lowerdir (stacked), plus a fresh upperdir per container. Starting a container creates an empty directory, not a copy of the image.

On disk with overlay2, layers live under /var/lib/docker/overlay2/<id>/diff, with lower files recording the chain and merged mounted while the container runs. (Docker's newer containerd image store uses containerd snapshotters, typically still overlayfs underneath.)

Copy-up and whiteouts

  • Reading a file: served from the topmost layer that has it.
  • Modifying a file from a lower layer: OverlayFS first copies the whole file up into the writable layer, then modifies the copy. Changing one byte in a 2 GB file copies 2 GB.
  • Deleting a file from a lower layer: OverlayFS creates a whiteout (a special character device) in the upper layer that hides it. The original bytes remain in the lower layer.

Why deleting doesn't shrink images

Image layers work the same way. This Dockerfile:

RUN curl -o /tmp/big.tar.gz https://example.com/big.tar.gz
RUN tar -xzf /tmp/big.tar.gz -C /opt && rm /tmp/big.tar.gz

ships big.tar.gz forever in the first layer; the second layer only adds a whiteout. Do download, extract and delete in one RUN, or better, use a multi-stage build and copy only the result into the final image. The same applies to secrets: a key COPYed in one layer and deleted in the next is still in the image — use RUN --mount=type=secret instead. (Reduce Docker image size, production Dockerfile for Node)

The build cache follows the layers

BuildKit caches each step keyed on the instruction and its inputs (for COPY, the content of the copied files). If a step's cache key changes, every later step rebuilds. Hence the classic ordering:

COPY package.json package-lock.json ./
RUN npm ci                     # cached until dependencies change
COPY . .                       # source changes only invalidate from here
RUN npm run build

Cache mounts (RUN --mount=type=cache,target=/root/.npm npm ci) persist package caches across builds without putting them into layers.

Performance traps

Writing heavily to the container layer. Copy-up, whiteouts and the overlay indirection make write-heavy workloads slower, and the data vanishes with the container. Databases, uploads and caches belong in volumes (bind mounts or named volumes bypass OverlayFS entirely). (Docker volumes vs bind mounts)

Huge directories in lower layers. Operations that touch many files across layers — chown -R in a later RUN, recursive permission changes — copy up every file, duplicating it. Set ownership at COPY --chown= time instead.

node_modules and friends. Thousands of small files make layer extraction and copy-up slow; prune dev dependencies and use multi-stage builds.

Inode and open-file semantics. Copy-up changes a file's inode; a process holding the lower file open keeps seeing the old one. Rarely matters, occasionally baffling.

Rename of directories across layers may be emulated (redirect_dir) or fail with EXDEV; tools fall back to copying.

Inspecting what a container changed

docker diff mycontainer        # A = added, C = changed, D = deleted (relative to image)
docker container inspect mycontainer --format '{{.GraphDriver.Data.UpperDir}}'

dive is an excellent tool for exploring an image layer by layer and spotting wasted space.

Cleaning up

Unused layers accumulate with every build and pull:

docker system df
docker builder prune
docker image prune -a

(No space left on device)

The summary

  • Images are content-addressed stacks of read-only layers; containers add one writable layer.
  • OverlayFS merges them: reads fall through, writes copy up, deletes leave whiteouts.
  • Deleting in a later layer never shrinks an image — combine steps or use multi-stage builds.
  • Build cache invalidates from the first changed step; order Dockerfiles stable-first.
  • Write-heavy data belongs in volumes, not the container layer.

EasySpawn runs each server as its own VM with persistent disks, so your app's files, database and build caches live on real storage — not a throwaway container layer. See how it works or join the waitlist.

Related: What Is Docker? · Reduce Docker Image Size · Docker Volumes vs Bind Mounts · Linux Namespaces Explained

Keep reading