Node.js Event Loop Lag: Finding What Blocks Your Server
When one request's CPU work blocks the event loop, every other request waits. How to measure event loop delay and utilisation with perf_hooks, the usual culprits (big JSON, sync crypto, regex backtracking, sync fs), profiling with --cpu-prof, the libuv threadpool, and fixes from chunking to worker threads.
Node.js serves thousands of concurrent requests on one thread. That works because most of a request's time is spent waiting — for the database, an API, the disk — and while one request waits, the event loop runs others.
The flip side: while JavaScript is running, nothing else happens. A 300 ms synchronous operation in one request adds up to 300 ms of latency to every request in flight. Health checks time out, WebSocket heartbeats miss, and p99 latency explodes while average CPU looks fine. (Why is my website slow)
This is event loop lag (or event loop delay/blocking), and it's the most common performance problem in Node services that isn't a slow query.
Measuring it
Event loop delay
perf_hooks samples how late timers fire — a direct measure of blocking:
import { monitorEventLoopDelay } from 'node:perf_hooks'
const h = monitorEventLoopDelay({ resolution: 20 })
h.enable()
setInterval(() => {
const ms = (ns) => (ns / 1e6).toFixed(1)
console.log(`loop delay p50=${ms(h.percentile(50))} p99=${ms(h.percentile(99))} max=${ms(h.max)}ms`)
h.reset()
}, 10_000).unref()
Rough interpretation for a web server: p99 under ~20 ms is healthy; sustained p99 above ~100 ms means users are feeling it; spikes into seconds mean something is badly blocking.
Event loop utilisation (ELU)
ELU is the fraction of time the loop was busy running code rather than idle waiting for I/O:
import { performance } from 'node:perf_hooks'
let last = performance.eventLoopUtilization()
setInterval(() => {
const now = performance.eventLoopUtilization()
const elu = performance.eventLoopUtilization(now, last).utilization
last = now
console.log(`ELU ${(elu * 100).toFixed(0)}%`)
}, 10_000).unref()
ELU near 100% means the process is CPU-saturated on its main thread — more requests will just queue. It's a better autoscaling and load-shedding signal for Node than OS CPU%, which also counts GC and threadpool work. Libraries like @fastify/under-pressure use these metrics to return 503 before the process falls over. (Load testing your app)
Export both to your metrics system and alert on them. (Error monitoring for beginners)
The usual culprits
Huge JSON. JSON.parse and JSON.stringify are synchronous. Serialising a 50 MB API response or parsing a large webhook body blocks for hundreds of milliseconds. Paginate, stream (stream-json, NDJSON), or don't load it all at once.
Synchronous crypto. bcrypt.hashSync, crypto.pbkdf2Sync, scryptSync are deliberately slow — 100 ms+ per call. Under a login spike, they serialise the whole server. Use the async versions, which run on the libuv threadpool.
Regex catastrophic backtracking. A pattern like /^(a+)+$/ against a crafted string takes exponential time — a ReDoS. Validate input lengths, avoid nested quantifiers, and lint for unsafe patterns.
Sync filesystem calls. fs.readFileSync in a request handler, often hidden in a template loader or config read. Fine at startup, harmful per request.
Big loops over big arrays. Sorting, filtering or reducing hundreds of thousands of items in memory — often data that should have been filtered or aggregated in the database. (Database indexes)
Server-side rendering of large pages, image processing in JS, PDF generation, Markdown/syntax highlighting of large documents.
Garbage collection. Big heaps and high allocation rates cause GC pauses that show up as loop delay too. If lag correlates with heap size, see Node memory leaks.
Finding the blocking code
CPU profile
node --cpu-prof --cpu-prof-dir=./profiles server.js
# run load against it, then stop the process (SIGINT)
Open the .cpuprofile in Chrome DevTools (Performance panel → load profile) or speedscope. Look for wide, flat bars at the top of the stack: functions with high self time in a single long task.
For a running production process, start with --inspect and attach DevTools, or send SIGUSR1 to enable the inspector on demand (only bind it to localhost and tunnel over SSH — the inspector gives full code execution).
Blocked-at stack traces
Packages like blocked-at use async hooks to print the stack of the code that started a long synchronous section. Heavy overhead — use in staging, not production.
Logging long requests
Cheap and effective: log any request whose handler takes more than N ms of wall time along with its route and payload size. Blocking operations show up as specific routes with specific inputs.
The libuv threadpool
Not everything async is free. fs operations, dns.lookup, async crypto (pbkdf2, scrypt, randomBytes) and zlib run on libuv's threadpool, which has 4 threads by default. Five concurrent bcrypt hashes and one is waiting — and so is every fs.readFile behind it.
UV_THREADPOOL_SIZE=16 node server.js # must be set before the pool starts
Symptoms of threadpool saturation look different from loop lag: loop delay is low, but file reads and DNS lookups are slow. Note that dns.lookup (used by http by default) goes through the threadpool, so a slow resolver can starve your file I/O.
Fixes, from cheapest to most involved
- Don't do the work. Paginate, filter in SQL, cache results, precompute at write time. (Redis: when you need it)
- Use async APIs for crypto, compression and files.
- Break up the work so other events get a turn:
async function processInChunks(items, fn, size = 1000) {
for (let i = 0; i < items.length; i += size) {
items.slice(i, i + size).forEach(fn)
await new Promise(setImmediate) // yield to the event loop
}
}
- Move it to a worker thread. CPU-heavy, self-contained tasks (image resizing, parsing, PDF generation) belong off the main thread. Use a pool like
piscinarather than spawning a worker per request:
import Piscina from 'piscina'
const pool = new Piscina({ filename: new URL('./resize.worker.js', import.meta.url).href })
app.post('/resize', async (req, res) => {
const out = await pool.run({ buffer: req.body })
res.type('png').send(out)
})
- Move it to a job queue if the user doesn't need the result in the response. (BullMQ tutorial)
- Run more processes (cluster mode, multiple containers) — this adds capacity but doesn't fix a single 2-second block; that request still blocks its process.
Protecting the server under load
Even with fixes, set guard rails:
- Body size limits so nobody can send you a 200 MB JSON payload. (413 Request Entity Too Large)
- Timeouts on handlers and upstream calls.
- Load shedding on ELU or loop delay: return
503withRetry-Afterwhen overloaded, rather than queuing requests that will time out anyway. - Health checks that don't share the bottleneck — a liveness probe that fails because the loop is busy causes restarts that make an overload worse. (Graceful shutdown in Node.js)
EasySpawn runs your Node service on a VM where you can add CPU profiles and worker pools freely — and Claude Code can read a .cpuprofile with you to find the function blocking the loop. See how it works or join the waitlist.
Related: Node Memory Leaks · Graceful Shutdown in Node.js · Load Testing Your App · BullMQ Tutorial
Keep reading
WebAssembly as a Sandbox: Running Untrusted Code With Wasmtime and WASI
WebAssembly's design makes it a strong in-process sandbox: linear memory, no ambient authority, and capability-based access through WASI. How the isolation works, limiting CPU with fuel and epochs, memory limits, the component model, real uses for plugins and user code, and where it falls short.
V8 Isolates: How Edge Platforms Run Thousands of Tenants in One Process
A V8 isolate is an independent JavaScript heap and execution context inside one process. How isolates make edge runtimes start in milliseconds and pack thousands of tenants together, what isolation they provide, Spectre and the extra layers platforms add, limits, and isolates vs containers vs VMs.