Horizontal vs Vertical Scaling: How Apps Handle More Users
Vertical scaling means a bigger server; horizontal scaling means more servers. How each works, what load balancers do, why databases are the hard part, what has to change in your app to scale out, and why most small apps should scale up first.
Your app is getting slower as more people use it. There are two ways to give it more capacity: make the server bigger (vertical scaling, "scaling up") or add more servers (horizontal scaling, "scaling out"). Each has its place — and the order you try them in matters.
Vertical scaling: a bigger machine
Move to a server with more CPU, more memory, faster disks.
Pros:
- Simple. No code changes. Often a resize button and a restart.
- No new moving parts — still one machine, one database, one place to look when something breaks.
- Fast. Everything stays on one machine, talking over
localhost.
Cons:
- There's a ceiling. Eventually you hit the biggest machine available (though that ceiling is very high today).
- Brief downtime while resizing, on many providers.
- One machine is a single point of failure. If it goes down, everything goes down.
Horizontal scaling: more machines
Run several copies of your app on several servers, with a load balancer in front.
Pros:
- Almost unlimited capacity — add machines as needed.
- Resilience. One server dies; the others keep serving.
- Gradual deploys — update servers one at a time.
Cons:
- Your app has to be ready for it (see below).
- More moving parts: a load balancer, several servers, shared storage, more to monitor.
- The database doesn't scale out this easily.
What a load balancer does
A load balancer sits in front of your app servers and spreads incoming requests between them — round-robin, or to whichever server is least busy. It runs health checks and stops sending traffic to servers that aren't responding. (Health check endpoints.) Reverse proxies like Nginx, Caddy and Traefik can act as load balancers, and every cloud offers a managed one.
What has to change to scale out
With several copies of your app, any request can land on any copy. So your app must be stateless — nothing important can live only in one server's memory or disk:
- Sessions: store them in a database or Redis, or use signed tokens — not in a single server's memory. (Session vs JWT.)
- Uploaded files: store them in object storage, not the local disk. (Where should user uploads go?.)
- Caches: a shared cache like Redis, or accept per-server caches. (Redis: when you need it.)
- Background jobs and cron: make sure scheduled jobs run once, not once per server. (Background jobs.)
- WebSockets: connections stick to one server; you need a way to broadcast across servers. (WebSockets explained.)
Building this way from the start costs little and keeps your options open.
The database: the hard part
App servers are easy to copy. Databases aren't — every copy needs the same data. Options, roughly in order:
- Scale the database vertically. A bigger database server goes a very long way.
- Optimise queries and add indexes. Usually the biggest win of all. (Database indexes.)
- Connection pooling, so many app servers don't exhaust the database's connections. (Postgres connection pooling.)
- Caching frequently read data.
- Read replicas — copies that handle read queries, while one primary handles writes.
- Partitioning and sharding — splitting data across tables or servers. Complex; most apps never need it. (Postgres table partitioning.)
Which should you do first?
Scale up first. For most apps it's cheaper in time and complexity, and a single modern server handles far more than people expect — often thousands of concurrent users for a typical web app with a well-indexed database.
Scale out when:
- you need redundancy — downtime from a single machine failing is unacceptable,
- you've hit the practical limit of one machine,
- your workload is naturally spiky and parallel,
- or you're separating different jobs onto different machines (web servers, background workers, a database server) — a kind of scaling out that's often the natural first step.
Before either: find out what's actually slow. Often it's one bad query, not a lack of servers. (Why is my website slow? and load testing your app.)
The summary
- Vertical: a bigger server. Simple, fast, but one machine and a ceiling.
- Horizontal: more servers behind a load balancer. Resilient, but your app must be stateless.
- Databases are the hard part; optimise and scale them up before splitting them.
- Most small apps should scale up first, and design statelessly so scaling out stays possible.
EasySpawn makes both simple: resize a server when you need more power — Small, Medium or Large — or add more servers to split production from experiments or give a busy app its own machine. See pricing or join the waitlist.
Related: Do You Need Kubernetes? · Monolith vs Microservices · What Is Latency? · What Is a VPS? · What Is a Load Balancer?
Keep reading
What Is a Load Balancer? Spreading Traffic Across Servers
A load balancer sits in front of several copies of your app and spreads requests between them, skipping any that are unhealthy. How it works, the common algorithms, sticky sessions, health checks, and why a small app probably doesn't need one yet.
What Is Latency? Why Distance Makes Your App Feel Slow
Latency is the delay before data arrives; bandwidth is how much can arrive at once. What latency is, why physical distance sets a floor on it, why round trips multiply it, latency vs bandwidth vs throughput, and practical ways to make an app feel faster for faraway users.