Architecture
How the pieces of an app fit together — frontends and backends, APIs, databases, queues and caches — and how to choose between designs.
45 posts · page 2 of 2
Monorepo or Separate Repos? A Practical Guide for Small Teams
Should your frontend, backend, and shared code live in one repository or several? The real trade-offs — atomic changes, tooling cost, deploy independence, access control — how monorepo tooling like workspaces and Turborepo helps, and why AI agents tip the balance.
Monolith vs Microservices: Why Small Teams Should Start With a Monolith
Microservices solve organisational problems most small teams don't have and add distributed-systems problems they can't afford. What each costs, the modular monolith as a middle path, the signals that justify splitting a service out, and how AI coding agents change the maths.
Implementing Rate Limiting: Algorithms, Redis, and Response Headers
Fixed window, sliding window, token bucket, and GCRA — how each behaves at the edges, how to implement them atomically in Redis or Postgres, choosing keys behind proxies, fail-open vs fail-closed, headers clients can use, and layered limits for login, APIs, and AI endpoints.
HTTP Caching Headers: Cache-Control, ETags, and Getting It Right
Most caching bugs are header bugs. How Cache-Control directives actually behave (max-age, s-maxage, no-cache vs no-store, private, immutable, stale-while-revalidate), how ETags and 304s work, Vary, and a practical header policy for static assets, HTML, APIs, and personalised pages.
Handling Webhooks Reliably: Signatures, Idempotency, and Retries
Webhooks arrive late, twice, out of order, or from someone pretending to be Stripe. How to verify signatures against the raw body, acknowledge fast and process in the background, make handlers idempotent, cope with ordering, and test the whole thing locally.
GraphQL vs REST: What's the Difference?
REST gives you many endpoints that each return a fixed shape; GraphQL gives you one endpoint where the client asks for exactly the fields it wants. How each works, over-fetching and under-fetching, the costs GraphQL adds, and why most small apps should start with REST.
Frontend vs Backend: What's the Difference?
Every app has a part that runs in your browser and a part that runs on a server. Knowing which is which explains why secret keys leak, why some apps need a server and others don't, and what your AI tool actually built. A plain-English guide with a restaurant analogy that actually holds up.
Firecracker vs gVisor vs Containers: Choosing Isolation for Untrusted Code
Containers share a kernel; gVisor intercepts it; Firecracker gives each workload its own. How the three isolation models actually work, what each costs in performance and compatibility, and how to match the boundary to the threat — for AI agents, multi-tenant platforms, and code execution.
Feature Flags for Small Teams: Ship Code Without Shipping Features
Feature flags separate deploying code from releasing it: merge unfinished work safely, try features yourself first, roll out gradually, and switch things off without a redeploy. A simple implementation, when to use a service, and how to stop flags becoming clutter.
How to Design Your First Database (Without a Computer Science Degree)
Before you ask an AI tool to 'build the database', spend fifteen minutes on paper. How to find your tables, choose columns and types, connect tables with foreign keys, handle one-to-many and many-to-many relationships, and avoid the mistakes that are painful to fix later.
Database Transactions Explained: ACID, Isolation Levels, and Race Conditions
Transactions make several changes succeed or fail together — but they don't automatically prevent race conditions. ACID in practice, PostgreSQL's isolation levels, lost updates and write skew, SELECT FOR UPDATE, serializable retries, and the transaction mistakes that cause outages.
Database Indexes: Why Your App Got Slow and How to Fix It
The app was fast with 100 rows and crawls with 100,000. The fix is usually an index. How indexes work, how to find the slow queries, how to read EXPLAIN ANALYZE, which columns to index (including the foreign keys ORMs forget), and what indexes cost.
CSRF Explained: Cross-Site Request Forgery and How Modern Apps Prevent It
CSRF tricks a logged-in user's browser into making a request they didn't intend. How the attack works, what SameSite cookies do and don't cover, CSRF tokens, Origin and Fetch Metadata checks, framework defaults, and why token-in-header APIs are different.
Content Security Policy: A Practical Guide to CSP Headers
A Content Security Policy tells the browser which scripts, styles, and connections your page may use, turning many XSS bugs into blocked requests. The directives that matter, nonce-based strict CSP, Report-Only rollout, Next.js specifics, and mistakes that make CSP useless.
How Container CPU and Memory Limits Actually Work
docker run --cpus 2 --memory 4g looks simple. Underneath, it's cgroup v2 files with behaviour that surprises people: CPU limits that throttle rather than slow, memory limits that count page cache, and tools inside the container that report the host's resources. How to read the real numbers.
What Is Caching? A Beginner's Guide to Making Apps Faster
Caching means keeping a copy of something so you don't have to fetch or compute it again. The caches between your user and your database — browser, CDN, server, database — what each is good for, why 'hard refresh' fixes things, and the one hard problem: stale data.
Preventing Cache Stampedes: Coalescing, Locks, Early Expiration, and Stale Serving
When a hot cache key expires, hundreds of requests miss at once and all hit the database. How stampedes happen, and the fixes: singleflight coalescing, distributed locks, probabilistic early recomputation (XFetch), stale-while-revalidate, TTL jitter, and negative caching.
Your App Needs Background Jobs. Here's the Simplest Way to Add Them.
Sending emails, processing uploads, calling slow AI models, and nightly cleanups don't belong in the middle of a web request. What background jobs and scheduled tasks are, the simplest reliable way to add them — often your existing Postgres database — and the mistakes that make jobs fail silently.
API Pagination: Offset vs Cursor (Keyset) and When Each Breaks
Returning every row works until it doesn't. How offset pagination works and why it gets slow and skips rows, how cursor (keyset) pagination fixes both, how to encode cursors and index for them, tie-breakers, total counts, and a response shape clients can rely on.
The Sandbox Is the Wrong Abstraction for AI Coding Agents
The industry settled on ephemeral sandboxes for AI agents — isolated, disposable, destroyed after each task. That's exactly right for running untrusted code and exactly wrong for building software. Here's the distinction that matters.
Why AI Coding Agents Need Persistent Workspaces
Stateless sandboxes make AI agents repeat themselves, lose context, and guess at results they could have measured. Persistent workspaces fix the feedback loop — here's the mechanism, and what it costs to build.