Blog
4 min read

runc and OCI Runtimes: What Actually Starts Your Container

Below Docker and containerd sits a small program that creates the namespaces, cgroups and mounts and execs your process: the OCI runtime. The OCI image, runtime and distribution specs, a bundle and config.json, running runc by hand, alternatives (crun, youki, gVisor's runsc, Kata), and runc's notable escapes.

When you docker run, several layers pass the request down. At the bottom is a small, short-lived program whose only job is to turn a bundle on disk into a running, isolated process. That's the OCI runtime, and the reference implementation is runc.

docker CLI → dockerd → containerd → containerd-shim → runc → your process

(containerd vs Docker)

The OCI specifications

The Open Container Initiative (Linux Foundation, 2015) standardised containers so tools interoperate:

  • Image spec — how an image is packaged: a manifest, a config JSON, and layer tarballs addressed by digest. (Docker image layers)
  • Runtime spec — how to run a container from a bundle: a root filesystem plus a config.json.
  • Distribution spec — the HTTP API registries implement for pushing and pulling.

Because of these, an image built by Docker runs under Podman, containerd, CRI-O or Kubernetes, with runc, crun or Kata underneath.

The bundle and config.json

A bundle is a directory:

mybundle/
├── config.json
└── rootfs/        # the unpacked image filesystem

config.json describes everything about the container: the process (args, env, cwd, user), mounts, namespaces, cgroup resources, capabilities, seccomp profile, rlimits, hooks.

A trimmed example:

{
  "ociVersion": "1.2.0",
  "process": {
    "user": { "uid": 1000, "gid": 1000 },
    "args": ["node", "server.js"],
    "cwd": "/app",
    "capabilities": { "bounding": ["CAP_NET_BIND_SERVICE"], "effective": [], "permitted": [] },
    "noNewPrivileges": true
  },
  "root": { "path": "rootfs", "readonly": true },
  "linux": {
    "namespaces": [{ "type": "pid" }, { "type": "mount" }, { "type": "network" },
                   { "type": "ipc" }, { "type": "uts" }, { "type": "cgroup" }],
    "resources": { "memory": { "limit": 536870912 }, "cpu": { "quota": 50000, "period": 100000 } },
    "seccomp": { "defaultAction": "SCMP_ACT_ERRNO", "syscalls": [ ... ] }
  }
}

Every isolation feature you configure in Docker ends up here. (Linux namespaces, Linux capabilities, Container CPU and memory limits)

Running runc by hand

mkdir -p mybundle/rootfs
docker export $(docker create alpine) | tar -C mybundle/rootfs -xf -
cd mybundle
runc spec                      # generates a default config.json
sudo runc run demo             # creates and starts container "demo"

runc spec --rootless generates a config for unprivileged use with user namespaces. (Rootless containers)

What runc does, roughly:

  1. Reads config.json.
  2. Creates namespaces (via clone/unshare) and a cgroup with the limits.
  3. Sets up mounts, pivot_roots into rootfs, masks sensitive /proc paths.
  4. Drops capabilities, applies no_new_privs, loads the seccomp filter, applies AppArmor/SELinux labels.
  5. execves the process — and runc exits (the shim stays as the parent).

The alternatives

Runtime What it is
runc Reference implementation, Go, from Docker's libcontainer
crun C implementation; faster start-up and lower memory; default in Podman on many distros
youki Rust implementation
runsc (gVisor) Runs the container against a user-space kernel that intercepts syscalls (Firecracker vs gVisor vs containers)
Kata Runs the container inside a lightweight VM (Kata Containers)

All accept the same bundle; swapping is a configuration change (--runtime in Docker, RuntimeClass in Kubernetes).

runc's notable vulnerabilities

Because runc runs with host privileges while setting up a container, bugs there are high impact:

  • CVE-2019-5736 — a malicious container could overwrite the host's runc binary via /proc/self/exe, gaining root on the host the next time runc ran. Fixed by having runc execute from a sealed in-memory copy of itself.
  • CVE-2024-21626 ("Leaky Vessels") — a leaked file descriptor to the host filesystem let a container (or a malicious image's WORKDIR) access host files. Fixed in runc 1.1.12.

Lessons: keep the runtime updated (it's often a separate package from Docker itself), run containers as non-root with user namespaces where possible, and use a stronger boundary for untrusted workloads. (Container hardening)

Debugging at this level

runc --root /run/containerd/runc/k8s.io list   # containers managed via containerd
cat /run/containerd/io.containerd.runtime.v2.task/<ns>/<id>/config.json

Reading the generated config.json is the definitive answer to "what isolation does this container really have?"


EasySpawn runs containers inside a per-server virtual machine, so even a runtime escape like Leaky Vessels stays inside your own VM rather than reaching other tenants. See how it works or join the waitlist.

Related: containerd vs Docker · Linux Namespaces Explained · Docker Image Layers and OverlayFS · PID 1 in Containers

Keep reading