Distributed Locks: Redis, Postgres and Why Fencing Tokens Matter
A lock across multiple processes or servers is harder than it looks: processes pause, leases expire and two holders can both believe they own the lock. Redis SET NX PX done right, the Redlock debate, Postgres advisory locks, fencing tokens, and when you don't need a distributed lock at all.
Inside one process, a mutex is simple: the runtime guarantees only one thread holds it. Once you have multiple processes or servers — two app instances, a cron job on each of three machines, a fleet of queue workers — you need a lock that lives somewhere they can all see. That's a distributed lock, and the hard part isn't acquiring it. It's what happens when the holder stops responding while still believing it holds the lock.
What you're using the lock for
Martin Kleppmann's distinction is the most useful starting point:
- Efficiency — avoid doing the same work twice (sending the daily digest email once, not three times). If the lock occasionally fails and work happens twice, it's annoying but not harmful.
- Correctness — prevent two processes from doing something that corrupts state (two writers updating the same file, double-spending a balance). If the lock fails, you lose data.
For efficiency locks, a simple Redis lock is fine. For correctness locks, the lock alone is never enough — you need fencing, or a different design.
The naive Redis lock and its bugs
// ❌ Two bugs
if (await redis.get('lock:report') === null) {
await redis.set('lock:report', 'me')
await doWork()
await redis.del('lock:report')
}
- Race between check and set — two processes both see
nulland both set. - No expiry — if the holder crashes, the lock is held forever.
The correct single-instance Redis lock
import { randomUUID } from 'node:crypto'
async function withLock(redis, key, ttlMs, fn) {
const token = randomUUID()
// Atomic: set only if absent, with an expiry
const ok = await redis.set(key, token, 'PX', ttlMs, 'NX')
if (!ok) return false
try {
await fn()
} finally {
// Release only if we still own it (compare-and-delete, atomically)
await redis.eval(
`if redis.call('get', KEYS[1]) == ARGV[1] then
return redis.call('del', KEYS[1])
else return 0 end`,
1, key, token)
}
return true
}
SET ... NX PXacquires atomically with a lease.- The random token ensures you don't delete someone else's lock: if your lease expired and another process acquired it, a blind
DELwould release their lock. - The Lua script makes the check-and-delete atomic. (Redis: when you need it)
For long jobs, extend the lease periodically with another compare-and-set script (PEXPIRE only if the token still matches), and abort the work if extension fails.
The fundamental problem: leases and pauses
Every lock with an expiry has this failure mode:
Client A: acquire lock (30 s lease) ───── GC pause / VM freeze / network stall 40 s ──── write to storage ✗
Client B: lease expires → acquire lock ── write to storage ✓
Client A doesn't know its lease expired. It wakes up, still inside its critical section, and performs a write — concurrently with B. Long garbage-collection pauses, a VM being live-migrated, swapping, a slow disk, or a network partition can all cause it. Checking "do I still hold the lock?" right before writing just shrinks the window; the pause can happen between the check and the write.
No lock service can fix this on its own, because the problem is in the client.
Fencing tokens
The fix is to make the resource reject stale holders. Each time the lock is granted, the lock service hands out a monotonically increasing number — a fencing token. The client passes it with every write, and the storage rejects any write with a token lower than the highest it has seen:
A acquires → token 33 (pauses)
B acquires → token 34 writes with 34 ✓ (storage records 34)
A wakes writes with 33 ✗ rejected: 33 < 34
In a database, that's an ordinary conditional update:
UPDATE documents
SET body = $1, fence = $2
WHERE id = $3 AND fence < $2; -- 0 rows updated → you lost the lock
Redis's basic lock has no monotonic token (the random value isn't ordered). You can generate one with INCR alongside acquisition, or use a system that provides them: ZooKeeper's znode versions/zxid, etcd's revision numbers, or a database sequence.
If the protected resource can't check a token — an external API, an email send — you can't get correctness from a lock. Use idempotency instead. (Idempotency keys)
Redlock: the debate
Redlock acquires the same lock on a majority of N independent Redis nodes (typically 5) to survive a node failure. It's the algorithm in many Redis lock libraries.
The criticism (Kleppmann, 2016) is that Redlock depends on timing assumptions — bounded clock drift, bounded pauses, bounded network delay — and provides no fencing tokens, so it's unsafe for correctness-critical locks; and for efficiency-only locks, a single Redis instance is simpler and good enough. Redis's author responded defending the timing model. The practical consensus:
- Efficiency lock → single-instance Redis
SET NX PXis fine. - Correctness lock → use fencing tokens with a store that supports them, or a consensus system (etcd, ZooKeeper, Consul) — or avoid the lock entirely.
Postgres as a lock service
If you already have Postgres, you often don't need Redis for locking at all. (Postgres advisory locks)
Advisory locks are application-defined locks keyed by a number:
-- Session-level: held until released or the connection closes
SELECT pg_try_advisory_lock(hashtext('daily-report'));
-- ... work ...
SELECT pg_advisory_unlock(hashtext('daily-report'));
-- Transaction-level: released automatically at COMMIT/ROLLBACK
BEGIN;
SELECT pg_try_advisory_xact_lock(hashtext('daily-report'));
Advantages: no lease to tune — the lock is tied to the connection, so a crashed holder's lock is released when the connection drops; and with the transaction-level variant, the lock and your data changes commit together. Caveats: session-level locks and transaction-pooling poolers like PgBouncer don't mix, and a paused holder with a live connection keeps the lock (safe, but can stall others). (Postgres connection pooling)
Row locks are often even better, because they lock the actual thing you're changing:
BEGIN;
SELECT * FROM accounts WHERE id = 42 FOR UPDATE; -- others wait
UPDATE accounts SET balance = balance - 100 WHERE id = 42;
COMMIT;
Here the database is both the lock service and the resource, so there's no fencing gap. (Optimistic vs pessimistic locking)
Patterns that avoid distributed locks
Before reaching for a lock, check whether the problem has a lock-free shape:
- Unique constraints — "only one active subscription per user" is a partial unique index, not a lock.
- Conditional updates / optimistic concurrency —
UPDATE ... WHERE version = $expected; retry on zero rows. - Job queues with
SKIP LOCKED— each job is claimed by exactly one worker without a global lock. (Postgres SKIP LOCKED job queue) - Single-writer partitioning — route all work for a key to one consumer (Kafka partitions, BullMQ group keys). (BullMQ tutorial)
- Idempotent operations — if doing it twice is harmless, you don't need to prevent it.
Checklist
- Decide: efficiency or correctness?
- Always use an expiry or connection-bound lock — never one that survives a crash forever.
- Release only your own lock (token compare).
- For correctness, add fencing tokens checked by the resource, or use database row locks.
- Set the lease well above the normal job duration, extend it for long jobs, and abort on failed extension.
- Log acquisition failures and lease expiries; they're the signals that your timing assumptions are wrong.
On EasySpawn you can run Postgres and Redis on your own server, so advisory locks and SKIP LOCKED queues work without extra infrastructure — and Claude Code can help you find the races before they double-charge a customer. See how it works or join the waitlist.
Related: Postgres Advisory Locks · Optimistic vs Pessimistic Locking · Idempotency Keys · Postgres SKIP LOCKED Job Queue
Keep reading
Postgres HOT Updates and Fillfactor: Cheaper UPDATEs
Every Postgres UPDATE writes a new row version — and normally new entries in every index. Heap-only tuple (HOT) updates skip the index work. When HOT applies, why indexed columns and full pages prevent it, tuning fillfactor, measuring the HOT ratio, and designing hot tables to stay HOT.
Change Data Capture in Postgres: Streaming Every Row Change
Change data capture streams every insert, update and delete out of Postgres as it happens. How logical decoding and replication slots work, publications and pgoutput, REPLICA IDENTITY, Debezium and lighter alternatives, the outbox pattern, and the operational traps — retained WAL, schema changes and failover.