Blog
4 min read

Message Queues Explained: When You Need One and Which to Use

A message queue lets one part of your system hand work to another without waiting. Queues vs pub/sub vs event streams, delivery guarantees, dead-letter queues, and an honest comparison of Postgres, Redis, RabbitMQ, SQS and Kafka for a small team.

A message queue sits between the part of your system that creates work and the part that does it. The producer drops a message ("send welcome email to user 42") into the queue and moves on; a consumer (worker) picks it up and processes it, now or later.

That decoupling buys you three things:

  • Responsiveness — the web request returns immediately instead of waiting for slow work.
  • Resilience — if the email provider is down, messages wait and retry instead of failing the signup.
  • Smoothing — a burst of 10,000 jobs is processed at a steady rate instead of overwhelming everything.

Three patterns that get called "queues"

Pattern Each message goes to Message after processing Example use
Work queue One of many workers Deleted Send emails, resize images, run AI jobs
Pub/sub Every subscriber Gone once delivered "Order placed" → email, analytics, inventory
Event stream / log Every consumer group, reading at its own position Kept for a retention period Event sourcing, analytics pipelines, replay

Most small apps need a work queue. (Background jobs)

Delivery guarantees

  • At-most-once — may lose messages, never duplicates. Rarely what you want.
  • At-least-once — never loses, may deliver twice (e.g. a worker crashes after doing the work but before acknowledging it). The common default.
  • Exactly-once — mostly a marketing term end-to-end; in practice you get it by combining at-least-once delivery with idempotent consumers.

So: design every consumer so that processing the same message twice is harmless. Check "already sent?" before sending; use unique constraints; use idempotency keys with external APIs. (Idempotency keys)

Features you'll want

  • Acknowledgements — a message is only removed once the worker confirms success.
  • Retries with backoff — wait longer between each attempt.
  • Dead-letter queue (DLQ) — after N failures, park the message somewhere you can inspect it instead of retrying forever.
  • Visibility timeout / lease — if a worker dies mid-job, the message becomes available again.
  • Delayed messages — "run this in 24 hours."
  • Ordering — usually not guaranteed across workers; if order matters, partition by key.

The options, honestly

Postgres (FOR UPDATE SKIP LOCKED) — a jobs table in the database you already have. Transactional with your data (the job is created in the same transaction as the order), easy to inspect with SQL, no new infrastructure. Libraries: pg-boss, Graphile Worker (Node), Procrastinate (Python), and others. Handles far more throughput than most small apps need. (Postgres as a job queue)

Redis-based (BullMQ, Sidekiq, RQ, Celery with Redis) — fast, mature, great dashboards and features (delays, rate limits, priorities). Requires running Redis and configuring it not to lose data on restart. (Redis: when you need it)

RabbitMQ — a dedicated message broker with flexible routing (direct, topic, fanout exchanges). Strong for complex routing and pub/sub between services. Another system to operate.

Amazon SQS / Google Pub/Sub / cloud queues — fully managed, scale without thought, pay per message. Natural if you're already on that cloud.

Kafka (or Redpanda) — a distributed event log built for very high throughput, retention and replay. Excellent for data pipelines and event-driven architectures at scale; heavy to operate and overkill as a job queue for a small app.

Postgres Redis/BullMQ RabbitMQ SQS Kafka
New infrastructure None Redis Broker Managed Cluster
Transactional with app data Yes No No No No
Pub/sub Basic (LISTEN/NOTIFY) Yes Yes With SNS Yes
Replay history If you keep rows No No No Yes
Best for Small/medium apps Feature-rich job queues Routing between services Cloud-native apps High-volume streams

The transactional trap

If your code saves an order to Postgres and publishes a message to a separate broker, one can succeed while the other fails — an order with no confirmation email, or an email for an order that rolled back. Fixes: use a Postgres-backed queue (same transaction), or the outbox pattern. (Transactional outbox pattern)

Do you need one?

You need some queue as soon as you have work that's slow, flaky, or scheduled. You probably don't need a dedicated broker yet. Start with a Postgres-backed queue; move to Redis or a broker when you hit a specific limit — throughput, fan-out to many services, or replay.

The summary

  • Queues decouple producing work from doing it: faster responses, retries, smoothing.
  • Work queue (one consumer) vs pub/sub (all subscribers) vs event log (replayable).
  • Assume at-least-once delivery; make consumers idempotent; use a DLQ.
  • Small team? Start with Postgres; add Redis, RabbitMQ, SQS or Kafka for specific needs.

EasySpawn servers run your app, its background workers and Postgres together, so a database-backed queue needs no extra infrastructure — and workers keep running when you close your laptop. See how it works or join the waitlist.

Related: Your App Needs Background Jobs · Postgres as a Job Queue · Transactional Outbox Pattern · Monolith vs Microservices

Keep reading