runc and OCI Runtimes: What Actually Starts Your Container
Below Docker and containerd sits a small program that creates the namespaces, cgroups and mounts and execs your process: the OCI runtime. The OCI image, runtime and distribution specs, a bundle and config.json, running runc by hand, alternatives (crun, youki, gVisor's runsc, Kata), and runc's notable escapes.
When you docker run, several layers pass the request down. At the bottom is a small, short-lived program whose only job is to turn a bundle on disk into a running, isolated process. That's the OCI runtime, and the reference implementation is runc.
docker CLI → dockerd → containerd → containerd-shim → runc → your process
The OCI specifications
The Open Container Initiative (Linux Foundation, 2015) standardised containers so tools interoperate:
- Image spec — how an image is packaged: a manifest, a config JSON, and layer tarballs addressed by digest. (Docker image layers)
- Runtime spec — how to run a container from a bundle: a root filesystem plus a
config.json. - Distribution spec — the HTTP API registries implement for pushing and pulling.
Because of these, an image built by Docker runs under Podman, containerd, CRI-O or Kubernetes, with runc, crun or Kata underneath.
The bundle and config.json
A bundle is a directory:
mybundle/
├── config.json
└── rootfs/ # the unpacked image filesystem
config.json describes everything about the container: the process (args, env, cwd, user), mounts, namespaces, cgroup resources, capabilities, seccomp profile, rlimits, hooks.
A trimmed example:
{
"ociVersion": "1.2.0",
"process": {
"user": { "uid": 1000, "gid": 1000 },
"args": ["node", "server.js"],
"cwd": "/app",
"capabilities": { "bounding": ["CAP_NET_BIND_SERVICE"], "effective": [], "permitted": [] },
"noNewPrivileges": true
},
"root": { "path": "rootfs", "readonly": true },
"linux": {
"namespaces": [{ "type": "pid" }, { "type": "mount" }, { "type": "network" },
{ "type": "ipc" }, { "type": "uts" }, { "type": "cgroup" }],
"resources": { "memory": { "limit": 536870912 }, "cpu": { "quota": 50000, "period": 100000 } },
"seccomp": { "defaultAction": "SCMP_ACT_ERRNO", "syscalls": [ ... ] }
}
}
Every isolation feature you configure in Docker ends up here. (Linux namespaces, Linux capabilities, Container CPU and memory limits)
Running runc by hand
mkdir -p mybundle/rootfs
docker export $(docker create alpine) | tar -C mybundle/rootfs -xf -
cd mybundle
runc spec # generates a default config.json
sudo runc run demo # creates and starts container "demo"
runc spec --rootless generates a config for unprivileged use with user namespaces. (Rootless containers)
What runc does, roughly:
- Reads
config.json. - Creates namespaces (via
clone/unshare) and a cgroup with the limits. - Sets up mounts,
pivot_roots intorootfs, masks sensitive/procpaths. - Drops capabilities, applies
no_new_privs, loads the seccomp filter, applies AppArmor/SELinux labels. execves the process — and runc exits (the shim stays as the parent).
The alternatives
| Runtime | What it is |
|---|---|
| runc | Reference implementation, Go, from Docker's libcontainer |
| crun | C implementation; faster start-up and lower memory; default in Podman on many distros |
| youki | Rust implementation |
| runsc (gVisor) | Runs the container against a user-space kernel that intercepts syscalls (Firecracker vs gVisor vs containers) |
| Kata | Runs the container inside a lightweight VM (Kata Containers) |
All accept the same bundle; swapping is a configuration change (--runtime in Docker, RuntimeClass in Kubernetes).
runc's notable vulnerabilities
Because runc runs with host privileges while setting up a container, bugs there are high impact:
- CVE-2019-5736 — a malicious container could overwrite the host's
runcbinary via/proc/self/exe, gaining root on the host the next time runc ran. Fixed by having runc execute from a sealed in-memory copy of itself. - CVE-2024-21626 ("Leaky Vessels") — a leaked file descriptor to the host filesystem let a container (or a malicious image's
WORKDIR) access host files. Fixed in runc 1.1.12.
Lessons: keep the runtime updated (it's often a separate package from Docker itself), run containers as non-root with user namespaces where possible, and use a stronger boundary for untrusted workloads. (Container hardening)
Debugging at this level
runc --root /run/containerd/runc/k8s.io list # containers managed via containerd
cat /run/containerd/io.containerd.runtime.v2.task/<ns>/<id>/config.json
Reading the generated config.json is the definitive answer to "what isolation does this container really have?"
EasySpawn runs containers inside a per-server virtual machine, so even a runtime escape like Leaky Vessels stays inside your own VM rather than reaching other tenants. See how it works or join the waitlist.
Related: containerd vs Docker · Linux Namespaces Explained · Docker Image Layers and OverlayFS · PID 1 in Containers
Keep reading
Container Networking Internals: veth, Bridges, NAT, and Embedded DNS
What happens when a container sends a packet: network namespaces, veth pairs, bridges, NAT for egress and published ports, why published ports bypass firewalls like ufw, Docker's embedded DNS, inter-container isolation, and debugging with nsenter and tcpdump.
How Container CPU and Memory Limits Actually Work
docker run --cpus 2 --memory 4g looks simple. Underneath, it's cgroup v2 files with behaviour that surprises people: CPU limits that throttle rather than slow, memory limits that count page cache, and tools inside the container that report the host's resources. How to read the real numbers.