Kata Containers: Container Ergonomics With VM Isolation
Kata Containers runs each pod or container inside a lightweight virtual machine with its own kernel, while keeping the OCI and Kubernetes interfaces. How the runtime, guest agent and virtio-fs fit together, hypervisor choices, the performance costs, what it protects against, and when it's worth it.
Containers are fast and convenient, but every container on a host shares one kernel. A kernel vulnerability reachable from inside one container can compromise all of them. Virtual machines give each workload its own kernel, but traditionally at the cost of slow boots and clunky tooling. (Containers vs virtual machines)
Kata Containers splits the difference: you use containers exactly as before — OCI images, docker/nerdctl, Kubernetes pods — but each pod (or container) actually runs inside a lightweight VM with its own guest kernel.
It's an Open Infrastructure Foundation project, descended from Intel Clear Containers and Hyper runV.
Architecture
kubelet / containerd
│ (CRI, runtime class "kata")
containerd-shim-kata-v2 ─── on the host
│ launches
Hypervisor (QEMU / Cloud Hypervisor / Firecracker / Dragonball)
│
Guest VM: minimal kernel + kata-agent
│
Containers of the pod (namespaces/cgroups inside the VM)
- Shim (runtime v2) — implements the containerd shim API, so containerd talks to Kata like to runc. (containerd vs Docker, runc and OCI runtimes)
- Hypervisor — boots a small VM per pod using KVM. (How KVM works)
- Guest kernel — a trimmed kernel tuned for fast boot.
- kata-agent — runs inside the guest, receives requests over vsock (create container, exec, signals) and sets up namespaces and cgroups inside the VM.
- Storage — the container root filesystem is shared from the host via virtio-fs (or block devices for some snapshotters).
- Networking — the pod's network namespace on the host is bridged into the VM via a tap device.
Hypervisor choices
| Hypervisor | Trade-off |
|---|---|
| QEMU | Most features and device support; larger |
| Cloud Hypervisor | Rust, modern, good balance; hot-plug support |
| Firecracker | Minimal and fast, fewer features (e.g. no virtio-fs, so block-based rootfs) (Firecracker snapshots) |
| Dragonball | Built-in VMM in the Rust runtime, tuned for containers |
Using it in Kubernetes
Install Kata on the nodes (e.g. via kata-deploy), then select it per pod with a RuntimeClass:
apiVersion: node.k8s.io/v1
kind: RuntimeClass
metadata:
name: kata
handler: kata
---
apiVersion: v1
kind: Pod
metadata:
name: untrusted-job
spec:
runtimeClassName: kata
containers:
- name: job
image: ghcr.io/example/job:1.0
Trusted system pods keep using runc; untrusted workloads get a VM. Nodes must support hardware virtualisation — bare metal or cloud instances with nested virtualisation.
What it protects against
- Kernel exploits from inside the container — the attacker reaches the guest kernel, not the host's. They'd then need a hypervisor or virtio device escape, a much smaller and better-defended surface.
- Noisy-neighbour kernel resource interference (e.g. shared kernel data structures) is reduced.
It doesn't protect against:
- Misconfiguration — mounting the Docker socket or host paths into a Kata pod still hands them over.
- Network access you allow — egress policy is still your job. (Egress control for AI agents)
- Side channels at the hardware level (mitigated by the hypervisor and CPU features, not eliminated).
For the strongest isolation of genuinely hostile code, Kata can also run confidential containers (CoCo) on TEEs like AMD SEV-SNP and Intel TDX, where even the host can't read guest memory.
Costs
- Start-up — typically sub-second to a few seconds, versus tens of milliseconds for runc.
- Memory overhead — each pod carries a guest kernel and agent, tens of MB or more.
- I/O — virtio-fs and virtualised networking add latency compared with native; heavy-I/O workloads notice.
- Compatibility — some features behave differently: host networking, privileged containers, certain device plugins,
hostPathsemantics, and memory accounting (pod memory must be sized for the VM). - Operational — another runtime to upgrade and debug; guest kernel and agent versions to manage.
When it's worth it
- Multi-tenant platforms running customers' code on shared nodes.
- CI runners executing untrusted pull requests.
- AI agents and code execution sandboxes where arbitrary model-generated code runs. (Run AI-generated code safely)
- Regulated workloads needing a VM boundary but wanting Kubernetes.
When it isn't: single-tenant apps you wrote and trust, where runc plus hardening is plenty. (Container hardening)
Kata vs gVisor vs Firecracker directly
| Kata | gVisor | Firecracker (direct) | |
|---|---|---|---|
| Boundary | Hardware VM per pod | User-space kernel intercepting syscalls | Hardware microVM |
| Compatibility | High (real Linux kernel in guest) | Good, some syscalls unsupported | You build the rest |
| Kubernetes integration | Native (RuntimeClass) | Native (runsc) | Via Kata or custom |
| Overhead | VM boot + memory | Syscall interception cost | Very small VMM |
(Firecracker vs gVisor vs containers)
EasySpawn takes the VM-per-tenant idea all the way: every server is its own virtual machine with a dedicated kernel, so containers and agents inside it share nothing with anyone else's. See how it works or join the waitlist.
Related: Firecracker vs gVisor vs Containers · How KVM Virtualization Works · Docker vs Linux Users for Multi-Tenant Isolation · runc and OCI Runtimes
Keep reading
What Is eBPF? Safe Programs Inside the Linux Kernel, Explained
eBPF lets you run small, verified programs inside the Linux kernel at hook points — syscalls, network packets, function entry — without kernel modules. How it works (verifier, JIT, maps, hooks), what it powers, bpftrace one-liners, and its limits and security implications.
vsock Explained: Talking Between a VM and Its Host Without a Network
vsock is a socket family for communication between a virtual machine and its host — no IP addresses, no network interface, no firewall rules. How AF_VSOCK addressing with CIDs and ports works, vhost-vsock in QEMU, Firecracker's Unix-socket bridge, how sandboxes use it for guest agents, and the security model.