The Postgres Write-Ahead Log (WAL) Explained
Every change in Postgres is written to the write-ahead log before the data files. How WAL gives durability and crash recovery, LSNs and segments, checkpoints and full-page writes, synchronous_commit, wal_level, archiving and replication, and the replication-slot trap that fills disks.
PostgreSQL's durability rests on one rule: log the change before you make it. Every modification is first recorded in the write-ahead log (WAL), flushed to disk, and only then is the transaction considered committed. The actual table and index files are updated later, lazily.
Understanding WAL explains crash recovery, replication, point-in-time recovery, checkpoint I/O spikes, and one of the most common ways Postgres servers run out of disk.
Why log first?
Writing data pages in place is random I/O — an update touches a heap page, several index pages, maybe a TOAST page, scattered across files. Doing that synchronously on every commit would be slow, and a crash halfway through would leave files inconsistent.
Instead, Postgres:
- modifies pages in memory (shared buffers),
- appends a compact WAL record describing the change — sequential I/O,
- on commit, flushes WAL up to that point (
fsync), - writes the dirty data pages to disk later (background writer, checkpointer).
If the server crashes, on restart Postgres replays WAL from the last checkpoint, redoing changes that hadn't reached the data files. Committed transactions survive; uncommitted ones are ignored thanks to MVCC visibility. (Postgres MVCC explained, Database transactions)
LSNs and segments
WAL is a single logical stream addressed by LSN (log sequence number), a byte position like 3/A1B2C3D0. It's stored in segment files (16 MB by default) in pg_wal/.
SELECT pg_current_wal_lsn();
SELECT pg_walfile_name(pg_current_wal_lsn());
SELECT pg_size_pretty(pg_wal_lsn_diff(pg_current_wal_lsn(), '0/0')); -- total WAL generated
Every data page records the LSN of its last change, so recovery knows which records to apply.
Checkpoints
A checkpoint flushes all dirty pages to disk and writes a checkpoint record. After it, older WAL isn't needed for crash recovery (and can be recycled, unless kept for archiving or replication).
Controlled by:
checkpoint_timeout(default 5 min) andmax_wal_size(default 1 GB) — whichever comes first.checkpoint_completion_target(default 0.9) — spread the writes over most of the interval to avoid I/O spikes.
Frequent checkpoints mean less replay after a crash but more I/O. A log line like "checkpoints are occurring too frequently" means max_wal_size is too small for your write load.
Full-page writes
After each checkpoint, the first modification of any page writes the whole page into WAL (full_page_writes = on). That protects against torn pages — a crash mid-way through writing an 8 KB page to a filesystem with smaller atomic writes. It's also why WAL volume spikes right after checkpoints, and why wal_compression (lz4/zstd) can help a lot.
synchronous_commit: trading durability for latency
SET synchronous_commit = off; -- per transaction or session
With off, commit returns before WAL is flushed. A crash can lose the last fraction of a second of commits — but the database stays consistent (unlike fsync = off, which risks corruption and should never be used in production). Useful for high-volume, low-value writes like analytics events.
With synchronous replication, values like remote_apply make commit wait until a standby has applied the change.
wal_level
| Level | Contains | Enables |
|---|---|---|
minimal |
Just enough for crash recovery | Some bulk-load optimisations |
replica (default) |
+ info for archiving and physical replication | Standbys, PITR |
logical |
+ info to decode row changes | Logical replication, CDC (Change data capture) |
Archiving and point-in-time recovery
Copy every completed segment somewhere safe (archive_command / archive_library, or tools like pgBackRest and WAL-G). A base backup plus the continuous WAL archive lets you restore to any moment — say, one second before someone ran DELETE without a WHERE. (Postgres point-in-time recovery, The 3-2-1 backup rule)
Replication
Streaming replication ships WAL to standbys, which replay it continuously — the same mechanism as crash recovery, running forever. (Postgres read replicas)
The replication slot trap
A replication slot makes the primary retain WAL until the consumer (a standby or a CDC tool) confirms it. If the consumer goes away — a deleted replica, a stopped Debezium connector — WAL accumulates in pg_wal without limit until the disk fills and Postgres stops. (No space left on device)
Check for it:
SELECT slot_name, active,
pg_size_pretty(pg_wal_lsn_diff(pg_current_wal_lsn(), restart_lsn)) AS retained
FROM pg_replication_slots;
Safeguards: set max_slot_wal_keep_size to cap retention (the slot is invalidated instead of the disk filling), alert on retained WAL, and drop slots you no longer need. A failing archive_command causes the same buildup.
Monitoring WAL
SELECT * FROM pg_stat_wal; -- WAL records, bytes, full-page images
SELECT * FROM pg_stat_archiver; -- archiving successes/failures
SELECT * FROM pg_stat_checkpointer; -- checkpoint counts and timing (PG17+)
Sudden WAL growth usually traces to bulk updates, index rebuilds, or many full-page images after frequent checkpoints.
EasySpawn runs Postgres on your server with daily backups and sensible WAL and checkpoint defaults, and Claude Code can read pg_stat_wal and the slot views with you when disk usage climbs. See how it works or join the waitlist.
Related: Postgres Point-in-Time Recovery · Postgres MVCC Explained · Postgres Read Replicas · Change Data Capture in Postgres
Keep reading
Postgres TOAST: How Large Values Are Stored (and Why It Affects Performance)
Postgres pages are 8 KB, so large values are compressed and moved out of line into a TOAST table. The ~2 KB threshold, storage strategies (PLAIN, MAIN, EXTERNAL, EXTENDED), pglz vs lz4 compression, measuring TOAST size, and the performance traps with SELECT *, big JSONB documents and frequent updates.
Postgres Timeouts: statement_timeout, lock_timeout and Friends
Postgres ships with every timeout disabled, so a stuck query or forgotten transaction can hold locks forever. What statement_timeout, lock_timeout, idle_in_transaction_session_timeout, transaction_timeout and idle_session_timeout do, the lock-queue outage they prevent, safe values, and where to set them.