Finding Memory Leaks in Node.js: Heap Snapshots, Retainers, and Common Culprits
Your Node.js server's memory climbs until it crashes. How to confirm a leak, read process.memoryUsage, capture heap snapshots in production safely, compare them in Chrome DevTools, follow retainer chains, and fix the usual causes: unbounded caches, listeners, timers and closures.
A Node.js service starts at 150 MB, sits at 600 MB by evening, and crashes overnight with JavaScript heap out of memory — or gets killed with exit code 137. Restarting "fixes" it until tomorrow. That's a memory leak: something keeps references to objects that should have been garbage-collected.
Here's a systematic way to find it.
First: confirm it's a leak
Memory that rises and then plateaus is usually fine — caches warming, the V8 heap growing to a comfortable size. A leak looks like a sawtooth that trends upward: garbage collection reclaims some memory each cycle, but the floor keeps rising with traffic.
Log memory periodically:
setInterval(() => {
const m = process.memoryUsage();
console.log(JSON.stringify({
rss_mb: Math.round(m.rss / 1e6),
heap_used_mb: Math.round(m.heapUsed / 1e6),
heap_total_mb: Math.round(m.heapTotal / 1e6),
external_mb: Math.round(m.external / 1e6),
array_buffers_mb: Math.round(m.arrayBuffers / 1e6),
}));
}, 60_000).unref();
What the numbers tell you:
heapUsedrising → a JavaScript object leak. Heap snapshots will find it.rssrising butheapUsedflat → memory outside the JS heap:Buffers (external/arrayBuffers), native addons, or memory fragmentation. Heap snapshots won't show it directly; look at Buffer-heavy code (streams, file handling, image processing) and native modules.
Correlate growth with traffic: does memory rise per request, per WebSocket connection, per job?
Reproduce it locally if you can
Run the app with the inspector and generate load against the suspected endpoint. (Load testing your app.)
node --inspect dist/server.js
Open chrome://inspect in Chrome, click inspect under your process, and go to the Memory tab.
The three-snapshot technique
- Warm up the app (a few hundred requests), then take snapshot 1.
- Send a batch of requests — say 1,000 — then snapshot 2.
- Send another 1,000, then snapshot 3.
DevTools runs garbage collection before each snapshot, so what remains is genuinely retained.
In snapshot 3, switch the view to Comparison against snapshot 2 and sort by # Delta or Size Delta. Objects whose count grows by roughly the number of requests you sent are your leak candidates — for example, 1,000 new IncomingMessage objects, or 1,000 new closures from one function.
Follow the retainers
Select a leaking object and look at the Retainers panel at the bottom. It shows the chain of references keeping the object alive, from a GC root down:
(GC root) → global → requestCache (Map) → entry → { req, user, ... }
The first thing in that chain that shouldn't be holding on is your bug — here, a global Map used as a cache. Ignore entries in parentheses like (system) and focus on your own variable and function names. Naming functions (rather than anonymous arrows) makes this far easier to read.
Capturing snapshots in production
Sometimes the leak only happens with real traffic. Options:
- On demand by signal: start Node with
--heapsnapshot-signal=SIGUSR2, thenkill -USR2 <pid>writes a snapshot to the working directory. - Programmatically:
require("node:v8").writeHeapSnapshot()from an admin-only endpoint. - Automatically before a crash:
--heapsnapshot-near-heap-limit=1writes a snapshot when the heap approaches its limit.
Be careful:
- Taking a snapshot pauses the process and can need memory comparable to the heap itself — do it on one instance taken out of the load balancer, or when traffic is low.
- Snapshots contain everything in memory, including user data, tokens and secrets. Treat the files as sensitive and delete them afterwards.
The usual culprits
Unbounded caches
const cache = new Map();
app.get("/user/:id", async (req, res) => {
if (!cache.has(req.params.id)) cache.set(req.params.id, await loadUser(req.params.id));
res.json(cache.get(req.params.id));
});
Every distinct ID stays forever. Use an LRU cache with a maximum size and TTL, or an external cache. (What is caching?)
Event listeners added repeatedly
app.get("/stream", (req, res) => {
emitter.on("update", (data) => res.write(data)); // never removed
});
Each request adds a listener that holds res forever. Remove it on close: req.on("close", () => emitter.off("update", handler)). Node's "MaxListenersExceededWarning" is often the first clue.
Timers that are never cleared
A setInterval created per request or per connection, never cleared, keeps its closure — and everything it references — alive.
Closures capturing large objects
A small callback stored long-term (in a queue, a map of pending operations) that closes over a large request or response object keeps the whole thing alive. Capture only what you need.
Per-request data in module-level state
Arrays used as logs, metrics buckets or "recent requests" that only ever grow.
Pending promises that never settle
Promises waiting on something that never happens (a lost callback, a request with no timeout) accumulate with their closures. Add timeouts to outbound calls.
Libraries and native modules
Sometimes the leak is in a dependency — an old version of a client library or a native addon. If retainers point into node_modules, check the library's issue tracker and upgrade.
Prevent regressions
- Track
heapUsedandrssas metrics with an alert on sustained growth. - Use bounded data structures by default.
- Run a soak test (steady load for an hour) before major releases.
- Set a sane
--max-old-space-sizeand a memory limit with restart as a safety net — not a fix.
The summary
- A leak shows as a rising floor in memory, not just high memory.
heapUsedrising → JS objects;rssrising alone → Buffers or native memory.- Take three heap snapshots around load, compare, and follow retainer chains.
- Usual causes: unbounded caches, unremoved listeners, uncleared timers, closures, ever-growing arrays.
- Snapshots pause the process and contain sensitive data — handle with care.
EasySpawn gives Claude Code your real server, where it can load-test your app, capture heap snapshots and compare them — and you can watch memory on server sizes up to 16 GB. See how it works or join the waitlist.
Related: How Container CPU and Memory Limits Actually Work · Graceful Shutdown in Node.js · Structured Logging · Error Monitoring for Beginners
Keep reading
What Is eBPF? Safe Programs Inside the Linux Kernel, Explained
eBPF lets you run small, verified programs inside the Linux kernel at hook points — syscalls, network packets, function entry — without kernel modules. How it works (verifier, JIT, maps, hooks), what it powers, bpftrace one-liners, and its limits and security implications.
The TLS 1.3 Handshake Explained: What Happens Before the First Byte
What actually happens when a browser connects over HTTPS: ClientHello, key shares, the server's certificate and signature, Finished messages, one round trip instead of two, 0-RTT resumption and its replay risk, SNI and ECH, certificate chain validation, and how to inspect it all with openssl.