Blog
4 min read

Linux Capabilities: Splitting Root Into Pieces

Linux capabilities break root's power into about 40 separate privileges, so a process can have just the one it needs. The capability sets (permitted, effective, inheritable, bounding, ambient), file capabilities with setcap, Docker's default set and --cap-drop ALL, why CAP_SYS_ADMIN is 'the new root', and inspecting capabilities.

Traditionally Linux had two kinds of process: root (UID 0), which bypasses permission checks, and everyone else. That's a blunt instrument — a web server that only needs to bind port 443 doesn't need the power to load kernel modules.

Capabilities divide root's privileges into about 40 distinct units. A process can hold just the ones it needs; root is effectively "a process with all of them".

Some important capabilities

Capability Allows
CAP_NET_BIND_SERVICE Bind ports below 1024
CAP_NET_RAW Raw sockets (ping, packet crafting/spoofing)
CAP_NET_ADMIN Configure interfaces, routes, firewall rules
CAP_CHOWN, CAP_FOWNER Change file ownership; bypass owner checks
CAP_DAC_OVERRIDE, CAP_DAC_READ_SEARCH Bypass file read/write/execute permission checks
CAP_SETUID, CAP_SETGID Change user/group IDs
CAP_KILL Signal any process
CAP_SYS_PTRACE Trace/inspect other processes' memory
CAP_SYS_MODULE Load kernel modules
CAP_SYS_TIME Set the system clock
CAP_BPF, CAP_PERFMON Load BPF programs, performance monitoring (What is eBPF?)
CAP_SYS_ADMIN A huge grab-bag: mount, namespaces, many admin operations

CAP_SYS_ADMIN is often called "the new root" because so many operations were lumped into it; granting it to a container is close to granting root on the host. Newer capabilities like CAP_BPF and CAP_CHECKPOINT_RESTORE were split out of it precisely to avoid handing out CAP_SYS_ADMIN.

The capability sets

Each thread has several sets (visible in /proc/<pid>/status):

  • Permitted — the capabilities the thread may use.
  • Effective — the ones currently active for permission checks.
  • Inheritable — preserved across execve (in combination with file capabilities).
  • Bounding — an upper limit; capabilities removed here can never be regained by this process tree.
  • Ambient — preserved across execve of ordinary (non-privileged) programs, used to give a non-root service a capability without file caps.

Decode them:

grep Cap /proc/$$/status
capsh --decode=00000000a80425fb

File capabilities: setcap

Instead of making a binary setuid-root, grant it one capability:

sudo setcap 'cap_net_bind_service=+ep' /usr/local/bin/myserver
getcap /usr/local/bin/myserver

Now myserver can bind port 443 as an ordinary user. (Note: file capabilities are cleared when the binary is modified, and interpreters like node or python would grant the capability to every script they run — prefer other approaches for interpreted apps.)

With systemd, grant capabilities to a service without touching binaries:

[Service]
User=myapp
AmbientCapabilities=CAP_NET_BIND_SERVICE
CapabilityBoundingSet=CAP_NET_BIND_SERVICE
NoNewPrivileges=true

(systemd service file)

Capabilities in containers

A process running as root inside a Docker container doesn't get all capabilities. Docker's default set is 14:

CHOWN, DAC_OVERRIDE, FOWNER, FSETID, KILL, SETGID, SETUID, SETPCAP, NET_BIND_SERVICE, NET_RAW, SYS_CHROOT, MKNOD, AUDIT_WRITE, SETFCAP

Everything else — SYS_ADMIN, NET_ADMIN, SYS_MODULE, SYS_PTRACE — is dropped. That's a big part of why "root in a container" isn't "root on the host" (alongside namespaces, seccomp and AppArmor). (Container hardening)

Tighten further — most apps need none:

docker run --cap-drop ALL --cap-add NET_BIND_SERVICE --security-opt no-new-privileges myapp
# compose
services:
  app:
    cap_drop: [ALL]
    security_opt: ["no-new-privileges:true"]

Two commonly needless defaults worth dropping: NET_RAW (enables spoofing and some attacks on the container network) and MKNOD.

--privileged grants all capabilities, disables seccomp and AppArmor confinement, and exposes host devices. Treat a privileged container as root on the host.

In Kubernetes, the same lives in securityContext.capabilities (drop: ["ALL"]), and Pod Security Standards' "restricted" profile requires dropping all.

Capabilities and user namespaces

Inside a user namespace, a process can have full capabilities — but they apply only to resources owned by that namespace. Root in a rootless container can mount a tmpfs in its own mount namespace, but can't touch the host. That's what makes rootless containers possible — and why user namespaces expand kernel attack surface: lots of privileged kernel code becomes reachable by unprivileged users. (Rootless containers and user namespaces, Linux namespaces)

no_new_privs

PR_SET_NO_NEW_PRIVS (Docker's no-new-privileges, systemd's NoNewPrivileges) ensures execve can never add privileges — setuid binaries and file capabilities stop working. Cheap, and it closes a whole class of escalation paths. Unprivileged seccomp and Landlock require it. (Landlock)

Auditing

# capabilities of running containers
docker inspect --format '{{.Name}} {{.HostConfig.CapAdd}} {{.HostConfig.CapDrop}}' $(docker ps -q)
# processes with any effective capabilities
grep -l 'CapEff:\s*0*[1-9a-f]' /proc/[0-9]*/status

The principle is the same as everywhere else: start from nothing, add what's needed, document why. (Principle of least privilege)


EasySpawn runs servers as managed VMs without handing out root, and isolates each account at the hypervisor — so a capability you grant inside your workload never becomes a capability over someone else's. See how it works or join the waitlist.

Related: Hardening Containers With seccomp, AppArmor and User Namespaces · Rootless Containers and User Namespaces · Landlock · Docker vs Linux Users

Keep reading