Linux Capabilities: Splitting Root Into Pieces
Linux capabilities break root's power into about 40 separate privileges, so a process can have just the one it needs. The capability sets (permitted, effective, inheritable, bounding, ambient), file capabilities with setcap, Docker's default set and --cap-drop ALL, why CAP_SYS_ADMIN is 'the new root', and inspecting capabilities.
Traditionally Linux had two kinds of process: root (UID 0), which bypasses permission checks, and everyone else. That's a blunt instrument — a web server that only needs to bind port 443 doesn't need the power to load kernel modules.
Capabilities divide root's privileges into about 40 distinct units. A process can hold just the ones it needs; root is effectively "a process with all of them".
Some important capabilities
| Capability | Allows |
|---|---|
CAP_NET_BIND_SERVICE |
Bind ports below 1024 |
CAP_NET_RAW |
Raw sockets (ping, packet crafting/spoofing) |
CAP_NET_ADMIN |
Configure interfaces, routes, firewall rules |
CAP_CHOWN, CAP_FOWNER |
Change file ownership; bypass owner checks |
CAP_DAC_OVERRIDE, CAP_DAC_READ_SEARCH |
Bypass file read/write/execute permission checks |
CAP_SETUID, CAP_SETGID |
Change user/group IDs |
CAP_KILL |
Signal any process |
CAP_SYS_PTRACE |
Trace/inspect other processes' memory |
CAP_SYS_MODULE |
Load kernel modules |
CAP_SYS_TIME |
Set the system clock |
CAP_BPF, CAP_PERFMON |
Load BPF programs, performance monitoring (What is eBPF?) |
CAP_SYS_ADMIN |
A huge grab-bag: mount, namespaces, many admin operations |
CAP_SYS_ADMIN is often called "the new root" because so many operations were lumped into it; granting it to a container is close to granting root on the host. Newer capabilities like CAP_BPF and CAP_CHECKPOINT_RESTORE were split out of it precisely to avoid handing out CAP_SYS_ADMIN.
The capability sets
Each thread has several sets (visible in /proc/<pid>/status):
- Permitted — the capabilities the thread may use.
- Effective — the ones currently active for permission checks.
- Inheritable — preserved across
execve(in combination with file capabilities). - Bounding — an upper limit; capabilities removed here can never be regained by this process tree.
- Ambient — preserved across
execveof ordinary (non-privileged) programs, used to give a non-root service a capability without file caps.
Decode them:
grep Cap /proc/$$/status
capsh --decode=00000000a80425fb
File capabilities: setcap
Instead of making a binary setuid-root, grant it one capability:
sudo setcap 'cap_net_bind_service=+ep' /usr/local/bin/myserver
getcap /usr/local/bin/myserver
Now myserver can bind port 443 as an ordinary user. (Note: file capabilities are cleared when the binary is modified, and interpreters like node or python would grant the capability to every script they run — prefer other approaches for interpreted apps.)
With systemd, grant capabilities to a service without touching binaries:
[Service]
User=myapp
AmbientCapabilities=CAP_NET_BIND_SERVICE
CapabilityBoundingSet=CAP_NET_BIND_SERVICE
NoNewPrivileges=true
Capabilities in containers
A process running as root inside a Docker container doesn't get all capabilities. Docker's default set is 14:
CHOWN, DAC_OVERRIDE, FOWNER, FSETID, KILL, SETGID, SETUID, SETPCAP, NET_BIND_SERVICE, NET_RAW, SYS_CHROOT, MKNOD, AUDIT_WRITE, SETFCAP
Everything else — SYS_ADMIN, NET_ADMIN, SYS_MODULE, SYS_PTRACE — is dropped. That's a big part of why "root in a container" isn't "root on the host" (alongside namespaces, seccomp and AppArmor). (Container hardening)
Tighten further — most apps need none:
docker run --cap-drop ALL --cap-add NET_BIND_SERVICE --security-opt no-new-privileges myapp
# compose
services:
app:
cap_drop: [ALL]
security_opt: ["no-new-privileges:true"]
Two commonly needless defaults worth dropping: NET_RAW (enables spoofing and some attacks on the container network) and MKNOD.
--privileged grants all capabilities, disables seccomp and AppArmor confinement, and exposes host devices. Treat a privileged container as root on the host.
In Kubernetes, the same lives in securityContext.capabilities (drop: ["ALL"]), and Pod Security Standards' "restricted" profile requires dropping all.
Capabilities and user namespaces
Inside a user namespace, a process can have full capabilities — but they apply only to resources owned by that namespace. Root in a rootless container can mount a tmpfs in its own mount namespace, but can't touch the host. That's what makes rootless containers possible — and why user namespaces expand kernel attack surface: lots of privileged kernel code becomes reachable by unprivileged users. (Rootless containers and user namespaces, Linux namespaces)
no_new_privs
PR_SET_NO_NEW_PRIVS (Docker's no-new-privileges, systemd's NoNewPrivileges) ensures execve can never add privileges — setuid binaries and file capabilities stop working. Cheap, and it closes a whole class of escalation paths. Unprivileged seccomp and Landlock require it. (Landlock)
Auditing
# capabilities of running containers
docker inspect --format '{{.Name}} {{.HostConfig.CapAdd}} {{.HostConfig.CapDrop}}' $(docker ps -q)
# processes with any effective capabilities
grep -l 'CapEff:\s*0*[1-9a-f]' /proc/[0-9]*/status
The principle is the same as everywhere else: start from nothing, add what's needed, document why. (Principle of least privilege)
EasySpawn runs servers as managed VMs without handing out root, and isolates each account at the hypervisor — so a capability you grant inside your workload never becomes a capability over someone else's. See how it works or join the waitlist.
Related: Hardening Containers With seccomp, AppArmor and User Namespaces · Rootless Containers and User Namespaces · Landlock · Docker vs Linux Users
Keep reading
WebAssembly as a Sandbox: Running Untrusted Code With Wasmtime and WASI
WebAssembly's design makes it a strong in-process sandbox: linear memory, no ambient authority, and capability-based access through WASI. How the isolation works, limiting CPU with fuel and epochs, memory limits, the component model, real uses for plugins and user code, and where it falls short.
Landlock: Unprivileged Sandboxing Built Into the Linux Kernel
Landlock lets an unprivileged process restrict its own filesystem and network access, permanently, for itself and its children. How rulesets, access rights and ABI versions work, a minimal C/Rust-style walkthrough, how it compares with seccomp, namespaces and AppArmor, and why coding agents use it.