Advanced guides.
Under the hood: container isolation, cgroups, sandboxing untrusted code, and the architecture behind multi-tenant platforms.
20 posts
The Transactional Outbox Pattern: Reliable Events Without Dual Writes
Writing to your database and publishing an event can't be made atomic, so one eventually happens without the other. How the transactional outbox fixes it: polling relays vs CDC, ordering, at-least-once delivery, idempotent consumers with an inbox, cleanup, and monitoring.
Securing MCP Servers: Threats and Controls for Tool-Connected Agents
An MCP server turns a model's text into real actions against real systems. The threat model — tool poisoning, prompt injection via tool output, confused deputies, token passthrough, DNS rebinding on local servers, over-broad scopes — and the controls for building and deploying MCP servers safely.
How to Run AI-Generated Code Safely
AI-generated code is usually well-intentioned and occasionally destructive, and the packages it installs are a supply-chain risk of their own. A practical, layered approach — what the code can see, reach, consume, and outlive — with a hardened Docker command you can use today.
Rootless Containers and User Namespaces: What They Actually Protect
Root in a container is root on the host unless something remaps it. How user namespaces work, subuid/subgid ranges, Docker userns-remap vs rootless mode vs Podman, Kubernetes hostUsers: false, the file-ownership and networking costs, and where rootless fits.
Prompt Injection in Coding Agents: A Threat Model
A coding agent with a shell, credentials, and network access reads text written by strangers all day. A threat model — sources, capabilities, sinks — why detection-based defences fail, and the architectural controls that actually bound the damage.
Postgres as a Job Queue: FOR UPDATE SKIP LOCKED Done Properly
You may not need Redis or a broker for background jobs. How SKIP LOCKED makes Postgres a safe concurrent queue: claim/lease/ack, visibility timeouts and crash recovery, retries with backoff, LISTEN/NOTIFY, transactional enqueue, indexing, bloat, and its limits.
Postgres Point-in-Time Recovery: WAL Archiving, Base Backups, and Restore Drills
A nightly dump can lose a day of data. Point-in-time recovery restores to the second before the bad migration. How WAL archiving and base backups combine, the settings that matter, recovery targets and timelines, pgBackRest and WAL-G, and restore drills.
Postgres Migrations on Large Tables Without Downtime
The migration that took 40 ms in staging locked production for minutes. Postgres lock levels and the lock queue, lock_timeout with retries, which ALTER TABLE operations rewrite, CREATE INDEX CONCURRENTLY, NOT VALID constraints, safe NOT NULL, and batched backfills.
Postgres Major Version Upgrades: pg_upgrade, Logical Replication, and Minimal Downtime
Major versions change the on-disk format, so upgrading PostgreSQL isn't a package update. Dump/restore vs pg_upgrade (copy, link, clone) vs logical replication cutover; extension and collation pitfalls; sequences and DDL gaps in logical replication; statistics after upgrade; and a rehearsed runbook.
npm Supply Chain Security: Install Scripts, Release Cooldowns, Provenance, and Trusted Publishing
Compromised maintainer accounts and self-propagating worms made npm installs an attack surface. The threat model, disabling install scripts, release cooldowns in npm, pnpm, Yarn, and Bun, lockfile discipline, provenance, trusted publishing, and isolating installs.
Multi-Tenant SaaS on Postgres: Shared Schema + RLS vs Schema-per-Tenant vs Database-per-Tenant
The tenancy model is the hardest SaaS decision to reverse. Shared schema with RLS vs schema-per-tenant vs database-per-tenant — isolation, migrations, pooling, per-tenant restore — plus the owner-bypass, pooling, and foreign-key traps that silently break row-level security.
Firecracker vs gVisor vs Containers: Choosing Isolation for Untrusted Code
Containers share a kernel; gVisor intercepts it; Firecracker gives each workload its own. How the three isolation models actually work, what each costs in performance and compatibility, and how to match the boundary to the threat — for AI agents, multi-tenant platforms, and code execution.
Evals for Coding Agents: Measuring Whether Your Agent Setup Actually Works
Changing a CLAUDE.md, model, skill, or MCP server changes agent behaviour, usually untested. How to build an eval suite for coding-agent workflows: task selection, hermetic environments, graders, pass@k vs pass^k, cost and trajectory metrics, and running headless in CI.
Container Networking Internals: veth, Bridges, NAT, and Embedded DNS
What happens when a container sends a packet: network namespaces, veth pairs, bridges, NAT for egress and published ports, why published ports bypass firewalls like ufw, Docker's embedded DNS, inter-container isolation, and debugging with nsenter and tcpdump.
Hardening Containers With Capabilities, seccomp, AppArmor, and User Namespaces
A default container shares the host kernel and starts with more privilege than most workloads need. A layer-by-layer guide: dropping capabilities, no-new-privileges, seccomp, AppArmor and SELinux, read-only filesystems, and user namespaces — and how to verify each.
How Container CPU and Memory Limits Actually Work
docker run --cpus 2 --memory 4g looks simple. Underneath, it's cgroup v2 files with behaviour that surprises people: CPU limits that throttle rather than slow, memory limits that count page cache, and tools inside the container that report the host's resources. How to read the real numbers.
Preventing Cache Stampedes: Coalescing, Locks, Early Expiration, and Stale Serving
When a hot cache key expires, hundreds of requests miss at once and all hit the database. How stampedes happen, and the fixes: singleflight coalescing, distributed locks, probabilistic early recomputation (XFetch), stale-while-revalidate, TTL jitter, and negative caching.
Egress Control for AI Agents: Designing an Allowlist Proxy
Restricting where an agent can send data is the most reliable defence against exfiltration, and easy to get subtly wrong. Network-layer enforcement, SNI vs TLS interception, DNS as a covert channel, allowlisted domains as leak paths, credential brokering, and testing.
The Sandbox Is the Wrong Abstraction for AI Coding Agents
The industry settled on ephemeral sandboxes for AI agents — isolated, disposable, destroyed after each task. That's exactly right for running untrusted code and exactly wrong for building software. Here's the distinction that matters.
Docker vs Linux Users for Multi-Tenant Workspace Isolation
Separate Linux users look like a cheap way to isolate tenants until you try to enforce a CPU limit. A walkthrough of why containers win for multi-tenant development workspaces — and how to verify the limits are real.