Blog
6 min read

Firecracker Snapshots: Restoring MicroVMs in Milliseconds

Firecracker can pause a microVM, write its memory and device state to disk, and restore it later — or many times over. How the snapshot API works, full vs diff snapshots, lazy memory loading with userfaultfd, and the hard part: making clones unique (randomness, network identity, clocks) and safe.

Firecracker boots a minimal Linux microVM in a fraction of a second. That's fast, but it isn't fast enough for everything: a sandbox that needs Python, a language server and a warmed-up runtime may still take seconds to become useful after the kernel is up. (Firecracker vs gVisor vs containers)

Snapshots skip all of that. You boot once, warm everything up, pause the VM, and save its full state. Later you restore it — memory, CPU registers, devices — and execution continues exactly where it stopped, typically in a few to a few tens of milliseconds. Serverless platforms and AI code sandboxes lean on this heavily. (Run untrusted AI-generated code safely)

What a snapshot contains

A Firecracker snapshot is two files plus your disks:

  • Microvm state file — vCPU registers, KVM state, device emulation state (virtio queues, serial, etc.). Small.
  • Guest memory file — the full contents of guest RAM. As big as the VM's memory size, or smaller if sparse.
  • Block devices — not included. The root filesystem and any data disks are referenced by path, and you are responsible for having them in a consistent state at restore time.

That last point catches people: if the guest had dirty data in its page cache, it's in the memory file; if you restore with a disk that has since changed, the guest's view and the disk disagree. Use read-only root filesystems plus a fresh, copy-on-write writable layer per restore.

The API flow

Firecracker is driven over a Unix-socket REST API. Creating a snapshot:

# 1. Pause the VM
curl --unix-socket /tmp/fc.sock -X PATCH http://localhost/vm \
  -d '{"state": "Paused"}'

# 2. Write the snapshot
curl --unix-socket /tmp/fc.sock -X PUT http://localhost/snapshot/create \
  -d '{"snapshot_type": "Full",
       "snapshot_path": "/snap/vm.state",
       "mem_file_path": "/snap/vm.mem"}'

Restoring, in a fresh Firecracker process that hasn't been configured with a kernel or drives:

curl --unix-socket /tmp/fc2.sock -X PUT http://localhost/snapshot/load \
  -d '{"snapshot_path": "/snap/vm.state",
       "mem_backend": {"backend_type": "File", "backend_path": "/snap/vm.mem"},
       "resume_vm": true}'

Exact field names have shifted between releases, so check the API spec for the version you run. Snapshots are only guaranteed to load on the same Firecracker version and a compatible host CPU — pin CPU templates if you move snapshots between machines.

Full vs diff snapshots

  • Full writes all guest memory every time.
  • Diff writes only pages dirtied since the last snapshot. It requires dirty-page tracking (enabled when the VM is configured or loaded), which adds some runtime overhead. A diff generally isn't loadable on its own: you merge it onto its base to get a full memory file. Diff snapshots are still marked developer preview in Firecracker's docs.

Diff snapshots can make periodic checkpointing of long-running VMs cheap. For the "golden image, restore many times" pattern, a single full snapshot is all you need.

Fast restore: don't read all of memory up front

Loading a 2 GB memory file before resuming would defeat the point. Two approaches avoid it:

  • File backend with mmap — Firecracker maps the memory file privately (copy-on-write). Pages are faulted in from the page cache on first access, so hot snapshots that are already cached restore almost instantly, and many clones share the same physical pages until they write.
  • UFFD backend — Firecracker registers guest memory with userfaultfd and hands the file descriptor to a page-fault handler process you write. Each first access to a page traps to your handler, which can fetch it from local disk, a remote store, or a predicted working set. This is how platforms stream snapshots from object storage and prefetch the pages a workload is known to need.

The cost is moved, not removed: page faults after resume are slower than local RAM access, so the first request after restore carries some latency. Recording the access pattern of a warm-up run and prefetching those pages is the usual fix.

The hard part: clones aren't unique

Restoring one snapshot into ten VMs gives you ten identical machines — identical in ways that break things.

Randomness

Every clone resumes with the same kernel entropy pool and the same userspace PRNG state. Two clones can generate the same UUIDs, TLS session keys or nonces. That is a real security bug, not a theoretical one.

Mitigations:

  • Firecracker always exposes a VM Generation ID device and updates the ID on every snapshot resume; Linux guests with a VMGenID driver reseed the kernel RNG when it changes. Make sure your guest kernel has the driver, and don't snapshot during very early boot — Firecracker warns the guest may not handle the notification yet. VMGenID only fixes the kernel entropy pool; everything else in memory is still duplicated.
  • Userspace libraries that cache randomness (OpenSSL's DRBG, language runtimes' seeded PRNGs) won't notice by themselves. Snapshot before they initialise, or explicitly reseed them after restore via an in-guest agent.
  • Never snapshot after generating long-lived secrets (SSH host keys, API tokens) that each clone should own.

Network identity

Clones come back with the same MAC address, IP and open TCP connections. Connections are dead on restore anyway (the peer moved on), so:

  • Run each clone in its own network namespace on the host with an identical tap device name and guest IP, and NAT outward. The guest never needs to know it's a clone. (Container networking internals)
  • Close or don't open connections before snapshotting; reconnect after resume.

Clocks

The guest's wall clock resumes at the moment the snapshot was taken. Have the guest agent step the clock (or run a quick NTP sync) after restore, or TLS certificate checks and token expiries go wrong.

Ephemeral state

Temp files, PIDs, hostnames and caches written before the snapshot are shared by every clone. Keep the snapshot point as early and generic as possible: runtime loaded, dependencies imported, nothing user-specific.

Security notes

  • The memory file contains everything in RAM — environment variables, keys, user data. Treat snapshots as secrets: encrypt at rest, restrict who can read them, and never snapshot a VM after a user's data has been in it unless that snapshot belongs to that user only.
  • Restoring many clones from one memory file shares host pages copy-on-write, which is efficient but means side-channel considerations similar to memory deduplication. Multi-tenant platforms typically don't share a golden snapshot across trust boundaries without thinking about this. (How KVM works)
  • The jailer still applies: run restored Firecracker processes in the same chroot, cgroup and seccomp sandbox as freshly booted ones. (Container hardening with seccomp and AppArmor)

When snapshots are worth it

Worth it: sandboxes started on demand per request or per user, serverless functions with heavy initialisation, CI runners with large pre-warmed toolchains.

Not worth it: long-lived servers that boot once a week. A normal boot is simpler, and you avoid every uniqueness problem above.


EasySpawn runs your apps on long-lived, dedicated VMs rather than snapshot clones, so there's no shared memory image between customers to worry about. See how it works or join the waitlist.

Related: Firecracker vs gVisor vs Containers · Kata Containers · How KVM Works · vsock Explained

Keep reading