Blog
6 min read

Linux Pressure Stall Information (PSI): Measuring Resource Contention

CPU percentage and free memory don't tell you whether your workload is actually being slowed down. Pressure stall information does: the share of time tasks were stalled waiting for CPU, memory or I/O. Reading /proc/pressure, some vs full, per-cgroup pressure, triggers, and using PSI for alerts and OOM handling.

Is this server overloaded? The usual metrics answer a different question:

  • CPU at 100% might be fine — a batch job soaking up idle cycles at low priority — or it might mean every request is queueing.
  • Free memory near zero is normal on Linux; the page cache uses whatever is free. It says nothing about whether processes are waiting on reclaim.
  • Load average mixes CPU waiting with uninterruptible I/O waits and doesn't scale meaningfully across core counts.

Pressure stall information (PSI), added in Linux 4.20, measures what you actually care about: how much time work was delayed because a resource was contended.

Reading /proc/pressure

$ cat /proc/pressure/cpu
some avg10=12.40 avg60=8.91 avg300=4.02 total=183273912
full avg10=0.00 avg60=0.00 avg300=0.00 total=0

$ cat /proc/pressure/memory
some avg10=0.55 avg60=1.20 avg300=0.40 total=9123311
full avg10=0.21 avg60=0.48 avg300=0.15 total=3810029

$ cat /proc/pressure/io
some avg10=3.10 avg60=2.75 avg300=1.98 total=88123401
full avg10=1.02 avg60=0.91 avg300=0.66 total=29012034
  • avg10, avg60, avg300 — the percentage of wall-clock time stalled, averaged over 10 s, 60 s and 5 min.
  • total — cumulative stall time in microseconds. Take deltas of this for your own windows and for graphs.

"some" vs "full"

  • some — at least one task was stalled on this resource. Work was being delayed, but something else may have been making progress.
  • full — all non-idle tasks were stalled at the same time. The CPU was doing nothing useful; this is lost productivity.

For CPU at the system level, full is not meaningful (if every task is waiting for CPU, someone is running on it), so it's reported but usually zero; in cgroups it's meaningful. For memory and I/O, full is the alarm signal: memory full avg10=20 means for 2 of the last 10 seconds, nothing ran because everything was waiting on reclaim or swap-in.

What each resource's pressure means

CPU pressure — runnable tasks waiting for a CPU. Sustained some above ~10–20% on a latency-sensitive service means requests are queueing behind each other. In a container, CPU pressure also includes throttling by a CPU quota, which ordinary CPU% doesn't show. (Container CPU and memory limits)

Memory pressure — tasks stalled on memory reclaim, waiting for pages to be swapped back in, or refaulting pages of the working set that the kernel evicted from page cache. This is the early warning before the OOM killer: pressure climbs first, then the kernel kills. (Linux OOM killer)

I/O pressure — tasks waiting on block I/O. Includes waits caused by memory pressure (swapping, re-reading evicted files), so memory and I/O pressure often rise together.

Kernels 6.1+ built with CONFIG_IRQ_TIME_ACCOUNTING also report IRQ pressure (/proc/pressure/irq): time stolen from running tasks by interrupt handling. It has only a full line, and the file is missing on kernels built without that option.

Per-cgroup pressure

The real power is that cgroup v2 tracks pressure per group. Every cgroup directory has the same files:

cat /sys/fs/cgroup/system.slice/postgresql.service/memory.pressure
cat /sys/fs/cgroup/system.slice/docker-<id>.scope/cpu.pressure
cat /sys/fs/cgroup/user.slice/io.pressure

So you can answer "which service is suffering" and "is my container's slowness caused by its own limits or by the host". A container with high cpu.pressure while the host's CPU pressure is low is being throttled by its own quota, not by its neighbours. (Linux namespaces explained)

systemd exposes the same data: systemctl status shows it on newer versions, and systemd-cgtop lists cgroups by resource use.

Using PSI for alerts

Pressure makes better alert conditions than utilisation:

Alert Why
memory full avg60 > 5 Workload is losing real time to memory shortage — act before the OOM killer does
io full avg60 > 10 Disk is the bottleneck for everything
cpu some avg60 > 25 (per service cgroup) Requests are queueing for CPU; scale up or find the hot loop

Node exporter exposes PSI as node_pressure_*_seconds_total counters; cAdvisor exposes container pressure on recent versions. Use rate() over the counters rather than the kernel's pre-averaged values.

Triggers: get notified, don't poll

Programs can ask the kernel to notify them when pressure crosses a threshold, by writing a trigger to a pressure file and polling it:

// Notify when "some" memory stall exceeds 150 ms within any 1 s window
int fd = open("/proc/pressure/memory", O_RDWR | O_NONBLOCK);
const char *trig = "some 150000 1000000";
write(fd, trig, strlen(trig) + 1);

struct pollfd p = { .fd = fd, .events = POLLPRI };
while (poll(&p, 1, -1) > 0) {
    if (p.revents & POLLPRI) { /* shed load, drop caches, alert */ }
}

The format is <some|full> <stall threshold µs> <window µs>. This is how systemd-oomd and Android's low-memory killer react to memory pressure in near real time — killing or throttling a cgroup when its pressure stays high, before the system thrashes into the ground. Creating triggers on system-wide files has historically required privileges; per-cgroup files are usable by whoever owns the cgroup.

Using PSI in applications

PSI is also a load-shedding signal:

  • A job runner can pause pulling work when cpu.pressure for its cgroup is high, instead of piling more onto a saturated machine.
  • A cache can shrink itself when memory.pressure rises.
  • An autoscaler can scale on pressure rather than CPU%, which avoids scaling up for a harmless low-priority batch job.

Read the file directly; it's cheap:

import { readFileSync } from 'node:fs'

function pressure(resource = 'cpu', kind = 'some') {
  const line = readFileSync(`/proc/pressure/${resource}`, 'utf8')
    .split('\n').find(l => l.startsWith(kind))
  return Number(/avg10=([\d.]+)/.exec(line)[1])
}

Inside a container, /proc/pressure usually shows the host values; read your own cgroup's cpu.pressure under /sys/fs/cgroup/ for numbers that reflect your limits.

Availability

  • Kernel 4.20+ with CONFIG_PSI=y. Some distributions build it in but disable it by default; enable with the psi=1 kernel boot parameter. If /proc/pressure doesn't exist, that's why.
  • Per-cgroup files require cgroup v2.
  • Overhead is small enough that it's on by default in most major distributions and in large production fleets.

A quick diagnosis routine

When a server "feels slow":

grep -H . /proc/pressure/*      # which resource is contended at all?
  1. CPU some high → find the busy cgroup (systemd-cgtop), then the process (top), then the code.
  2. Memory some/full high → check for swap activity (vmstat 1), working set vs RAM, and leaks. (Node memory leaks)
  3. I/O full high → iostat -x 1 for the device, then iotop or per-cgroup io.stat for the culprit. Often a database checkpoint, a backup job, or swapping. (Postgres memory tuning)
  4. All low → the slowness isn't local resource contention: look at the network, an upstream API, or locks inside your app. (Why is my website slow)

Each EasySpawn server is its own VM rather than a slice of a shared container host, so cgroup pressure maps cleanly to your own services — and Claude Code can read pressure and cgroup stats with you to find which service is starving. See pricing or join the waitlist.

Related: The Linux OOM Killer · Container CPU and Memory Limits · What Is eBPF? · Postgres Memory Tuning

Keep reading