How KVM Virtualization Works: The Isolation Behind Cloud Servers
Most cloud VMs run on KVM. How hardware-assisted virtualization works — VT-x/AMD-V, the VMM, virtio, two-level page tables — what a VM boundary protects against compared with containers, the remaining risks (side channels, noisy neighbours, steal time), and what shared vs dedicated vCPUs mean.
When you rent a cloud server, you're almost certainly getting a virtual machine running on KVM — the Kernel-based Virtual Machine, built into Linux. It underpins the major clouds' modern instance types and most VPS providers. Understanding how it works explains both why VMs are a strong isolation boundary and where that boundary still leaks.
The pieces
KVM: the kernel half
KVM is a Linux kernel module (since 2007) that turns the Linux kernel into a hypervisor. It exposes /dev/kvm, through which a user-space program can create virtual machines and virtual CPUs. KVM handles the privileged, performance-critical parts: running guest code on the real CPU, managing guest memory, and handling exits back to the host.
The VMM: the user-space half
A virtual machine monitor (VMM) runs as an ordinary process on the host and uses KVM to run a guest. It sets up memory, emulates or provides devices (disks, network cards), and handles the events KVM passes up. QEMU is the classic VMM; Firecracker and Cloud Hypervisor are minimal, security-focused VMMs designed for the cloud. (Firecracker vs gVisor vs containers compares their trade-offs.)
From the host's point of view, each VM is a process — which means host tools like cgroups can limit a VM's CPU and memory just like any other process. (How container CPU and memory limits work.)
How the guest runs at near-native speed
Hardware-assisted CPU virtualization
Modern CPUs have virtualization extensions — Intel VT-x and AMD-V. They add a special mode for guest code. The guest kernel runs its own instructions directly on the CPU, at full speed, believing it's in control. When it does something the hypervisor must handle — touching certain hardware, a privileged instruction, an I/O port — the CPU performs a VM exit, handing control to KVM. KVM deals with it (or passes it to the VMM) and resumes the guest.
Performance tuning in virtualisation is largely about reducing VM exits.
Memory: two levels of page tables
The guest kernel manages its own page tables, mapping guest-virtual to "guest-physical" addresses. But guest-physical memory is really just memory belonging to the VMM process on the host. CPUs handle this with a second level of translation in hardware — EPT (Intel) or NPT (AMD) — mapping guest-physical to host-physical addresses. The guest can't even express an address outside its own memory.
Devices: virtio
Emulating real hardware (a specific network card, a SATA controller) is slow and complex. Instead, guests use virtio — paravirtualised devices designed for VMs, where the guest knows it's virtualised and exchanges data with the host through shared-memory queues. Far fewer VM exits, much less emulated code.
What the VM boundary protects
Each VM runs its own kernel. This is the crucial difference from containers, which all share the host's kernel. (Containers vs virtual machines.)
- A guest process can't see host processes or other VMs' processes, files or memory.
- A kernel exploit inside the guest compromises the guest kernel — the attacker still has to break out of the VM.
- The attack surface exposed to the guest is the KVM interface plus the VMM's device implementations — much narrower than the hundreds of Linux system calls a container can reach.
That's why multi-tenant platforms that run untrusted code use VMs (or micro-VMs) as the boundary between customers rather than containers alone. (How to run AI-generated code safely.)
What it doesn't fully protect against
VM escapes
Bugs in the VMM's device emulation or in KVM itself can allow a guest to break out. They're rare and highly valued, which is why minimal VMMs (less device code) and running the VMM process with sandboxing (seccomp, separate users, jails) are standard hardening. (Hardening containers with seccomp and AppArmor covers the same tools.)
CPU side channels
VMs share physical CPU cores, caches and memory buses. Attacks like Spectre and L1TF showed that one guest could infer data from another through timing effects. Mitigations — microcode updates, kernel patches, flushing caches on context switches, and not scheduling different tenants' vCPUs on sibling hyperthreads of the same core — are now standard on major clouds, at some performance cost.
Noisy neighbours
Isolation of data isn't isolation of performance. On shared hardware, another tenant's heavy workload competes for CPU cache, memory bandwidth, disk and network.
Shared vs dedicated vCPUs, and steal time
A vCPU is usually one hardware thread. Providers sell two kinds:
- Shared vCPUs — your vCPUs are scheduled on host cores alongside other tenants' vCPUs. Cheaper, and fine for typical web workloads with bursty CPU use.
- Dedicated vCPUs — host cores reserved for you. More consistent performance for sustained CPU-heavy work.
On shared vCPUs, watch steal time — the st column in top, or %steal in vmstat/sar. It's the percentage of time your vCPU was ready to run but the hypervisor ran something else. A few percent is normal; sustained high steal means you're contending with neighbours and the CPU you're paying for isn't fully available.
vmstat 1 5 # watch the "st" column
Where VMs sit among isolation options
| Shared kernel? | Typical boundary | Startup | Density | |
|---|---|---|---|---|
| Containers | Yes | Namespaces + cgroups + seccomp | Milliseconds | Highest |
| gVisor | User-space kernel | Syscall interception | Fast | High |
| Micro-VMs (Firecracker) | No | KVM, minimal VMM | ~100+ ms | High |
| Full VMs (KVM + QEMU) | No | KVM, full VMM | Seconds | Lower |
For a long-lived server per customer, a full VM's startup time doesn't matter, and its separate kernel gives the strongest standard boundary.
The summary
- KVM turns Linux into a hypervisor; a user-space VMM (QEMU, Firecracker) runs each VM as a process.
- VT-x/AMD-V run guest code natively; EPT/NPT isolate memory in hardware; virtio keeps I/O efficient.
- Each VM has its own kernel — a much narrower attack surface than containers sharing one.
- Remaining risks: VMM bugs, CPU side channels, and noisy neighbours (watch steal time).
EasySpawn runs every server as its own virtual machine — its own kernel, CPU, memory and disk — not a container sharing a host with other customers. Other tenants can't see your processes, files or network traffic. See how it works or join the waitlist.
Related: Docker vs Linux Users for Multi-Tenant Isolation · Rootless Containers and User Namespaces · What Is a VPS? · Container Networking Internals
Keep reading
What Is eBPF? Safe Programs Inside the Linux Kernel, Explained
eBPF lets you run small, verified programs inside the Linux kernel at hook points — syscalls, network packets, function entry — without kernel modules. How it works (verifier, JIT, maps, hooks), what it powers, bpftrace one-liners, and its limits and security implications.
Linux Namespaces Explained: The Kernel Feature Containers Are Made Of
A container is a process with its own namespaces. What each of the eight Linux namespaces isolates — mount, PID, network, UTS, IPC, user, cgroup, time — how to build a container-like process by hand with unshare, inspect one with nsenter and /proc, and what namespaces do not protect against.