Blog
6 min read

io_uring and Security: Why Sandboxes Turn It Off

io_uring made Linux I/O dramatically faster — and became one of the kernel's biggest sources of privilege-escalation bugs. How io_uring works, why it sidesteps seccomp filtering, which platforms disable it, the io_uring_disabled sysctl, restriction APIs, and what to do for containers and code sandboxes.

io_uring, added in Linux 5.1, is the most significant change to how programs do I/O on Linux in decades. It's also one of the reasons sandbox builders have grown wary of the kernel's attack surface: for several years it was a leading source of kernel privilege-escalation exploits, and most hardened environments now switch it off for untrusted code.

What io_uring does

Classic I/O is one system call per operation: read, write, send, accept — each a transition into the kernel and back. At millions of operations per second, the transitions dominate.

io_uring replaces that with two ring buffers shared between user space and the kernel:

 user space                        kernel
┌───────────────────┐   SQEs    ┌──────────────┐
│ Submission Queue  │ ───────▶ │  io_uring    │
│ Completion Queue  │ ◀─────── │  workers     │
└───────────────────┘   CQEs    └──────────────┘

The program writes submission queue entries (read this fd, accept on that socket, open this file) into shared memory, calls io_uring_enter once to submit a batch — or with SQPOLL, not at all, because a kernel thread polls the ring — and reads results from the completion queue. Fewer syscalls, true async I/O for files and sockets, and very high throughput. Databases, proxies and runtimes (including some Node/Bun/Deno internals and Rust's async ecosystem) use it.

Why it's a security problem

1. A huge, fast-moving attack surface

io_uring grew from a few operations to well over fifty — files, sockets, timeouts, openat, statx, splice, buffer registration, networking zero-copy and more — with complex asynchronous lifetime handling. Asynchronous code with shared memory and reference counting is exactly where use-after-free bugs live.

In 2023 Google reported that around 60% of the kernel exploits submitted to its kCTF vulnerability reward programme in 2022 targeted io_uring, and announced it was disabling io_uring in ChromeOS, for Android apps, and on its production servers. Many of the bugs have been fixed, and the subsystem has matured, but the history shaped defaults across the industry.

2. It sidesteps syscall filtering

Container and sandbox security leans heavily on seccomp: a filter that allows or denies system calls by number and arguments. (Container hardening with seccomp and AppArmor)

With io_uring, operations aren't system calls. A program that is forbidden from calling openat or connect directly can submit IORING_OP_OPENAT or IORING_OP_CONNECT through the ring. seccomp only sees io_uring_enter — it can't inspect what's in the submission queue. So a seccomp policy that carefully blocks networking, for example, is bypassed unless it also blocks io_uring entirely.

LSMs fare better: file and socket operations performed via io_uring still pass through most of the same in-kernel permission hooks (and io_uring has gained its own LSM hooks, for credential overrides, SQPOLL threads and uring_cmd), so AppArmor, SELinux and Landlock checks on what is accessed still apply. (Landlock on Linux) But syscall-based auditing and filtering lose visibility.

3. Work happens in kernel threads

Some requests are executed by io-wq worker threads or the SQPOLL thread on behalf of the process. Early on this caused confusing credential and accounting behaviour; most of it has been fixed, but it complicates reasoning about which security context performed an action, and tracing tools that hook syscalls miss it. (What is eBPF)

Who disables it

  • Docker / Moby and containerd — the default seccomp profiles block io_uring_setup, io_uring_enter and io_uring_register in current releases. A container calling them gets EPERM, and well-behaved libraries fall back to classic I/O.
  • Google — ChromeOS, Android apps (via seccomp in the app sandbox), and Google production.
  • gVisor — implements only a small subset, and disables it by default. (Firecracker vs gVisor vs containers)
  • Many hardened distributions and Kubernetes security baselines that use RuntimeDefault seccomp.

Controls you have

The io_uring_disabled sysctl (Linux 6.6+)

sysctl kernel.io_uring_disabled
# 0 = available to all (default)
# 1 = only processes with CAP_SYS_ADMIN, or in the group set by kernel.io_uring_group
# 2 = disabled for everyone

sudo sysctl -w kernel.io_uring_disabled=2
echo 'kernel.io_uring_disabled = 2' | sudo tee /etc/sysctl.d/90-io-uring.conf

Mode 1 with a dedicated group lets you allow a specific trusted service (say, a database) while denying everything else.

seccomp

Block the three io_uring syscalls in any profile that runs untrusted code. If you write your own seccomp policy, also make sure the default action for unknown syscalls is to deny — new syscalls appear regularly.

Restricting a ring

A process that sets up a ring can lock it down before handing it to less-trusted code: IORING_REGISTER_RESTRICTIONS limits which opcodes and register operations the ring accepts, and IORING_SETUP_R_DISABLED lets you create the ring disabled, apply restrictions, then enable it. Useful for programs that want io_uring internally but sandbox a component; it's not a substitute for disabling it for untrusted processes.

Kernel version

If you must allow io_uring, run a recent, well-maintained kernel. Many exploited bugs lived in older long-term branches where backports lagged.

What this means for AI agents and code sandboxes

If you run code you didn't write — AI-generated scripts, user plugins, CI jobs from forks — the threat model is "this code will try to escape". (Run untrusted AI-generated code safely)

  1. Container sandboxes: keep the runtime's default seccomp profile on (--security-opt seccomp=unconfined turns io_uring back on, along with much else). Don't run privileged containers. (Linux capabilities)
  2. Process sandboxes (bubblewrap, Landlock, Claude Code's sandbox modes): check that the seccomp layer blocks io_uring if you rely on syscall filtering for network isolation. (Bubblewrap sandbox)
  3. VM sandboxes (Firecracker, Kata, Cloud Hypervisor): a guest kernel bug is contained by the VM boundary — the guest can exploit its own kernel, but still faces the hypervisor. You can leave io_uring on inside the guest if workloads benefit, and keep it off on the host. (Kata Containers)

The broader lesson: the kernel's attack surface is the sandbox's attack surface. Every feature reachable from untrusted code — io_uring, unprivileged user namespaces, eBPF, obscure socket families — is a potential escape route, so the safest sandboxes expose as few of them as possible. (Linux namespaces explained)

Will it be re-enabled?

io_uring today is far more mature than in 2021–2022, and its performance benefits are real. Expect it to stay on for trusted services on hosts you control and off for untrusted code for the foreseeable future — the same trade-off as unprivileged eBPF.


On EasySpawn, each server is its own VM, so the code Claude Code writes and runs is isolated by a hypervisor boundary, not just by a seccomp profile on a shared kernel. See how it works or join the waitlist.

Related: Container Hardening: seccomp and AppArmor · Run Untrusted AI-Generated Code Safely · Landlock on Linux · Firecracker vs gVisor vs Containers

Keep reading