Blog
4 min read

Server-Sent Events vs WebSockets: Which for Real-Time Features and AI Streaming?

SSE streams updates from server to browser over plain HTTP; WebSockets open a two-way channel. How each works, code for both, why AI chat responses use SSE-style streaming, proxy buffering and connection-limit gotchas, and a decision guide including plain polling.

Live notifications, progress bars, dashboards that update themselves, and AI chat replies that appear word by word all need the server to push data to the browser. The three common ways are polling, Server-Sent Events (SSE) and WebSockets. Choosing well saves a lot of complexity.

Polling: just ask again

The browser asks every few seconds: "anything new?"

setInterval(async () => {
  const res = await fetch("/api/notifications");
  render(await res.json());
}, 10_000);

Simple, works everywhere, scales like normal requests. Fine when a few seconds' delay is acceptable — order status, a job queue, a dashboard. Don't dismiss it: it's often the right answer.

Server-Sent Events: a one-way stream over HTTP

SSE keeps one HTTP response open and the server writes events into it as they happen. The format is plain text:

event: progress
data: {"percent": 40}

data: {"percent": 80}

Each event is one or more data: lines followed by a blank line.

Server (Node.js / Express):

app.get("/api/events", (req, res) => {
  res.set({
    "Content-Type": "text/event-stream",
    "Cache-Control": "no-cache",
    Connection: "keep-alive",
  });
  res.flushHeaders();

  const timer = setInterval(() => {
    res.write(`data: ${JSON.stringify({ time: Date.now() })}\n\n`);
  }, 1000);

  req.on("close", () => clearInterval(timer));
});

Browser:

const source = new EventSource("/api/events");
source.onmessage = (e) => console.log(JSON.parse(e.data));

Strengths:

  • Plain HTTP — works through most proxies, load balancers and CDNs; uses your normal cookies and auth.
  • Automatic reconnection built into EventSource, which resends the last event ID it saw (Last-Event-ID) so the server can resume.
  • Simple to implement on any backend.

Limits:

  • One direction: server → browser. The browser sends data with normal requests.
  • EventSource only does GET and can't set custom headers. If you need a POST body (as with AI chat), read the stream with fetch instead (below).
  • Text only (send JSON).
  • Connection limits on HTTP/1.1: browsers allow about six connections per domain, so many open tabs with SSE can exhaust them. HTTP/2 removes this problem in practice.

WebSockets: a two-way channel

A WebSocket starts as an HTTP request, then upgrades into a persistent, two-way connection. Either side can send messages at any time. (What are WebSockets? covers the basics.)

const ws = new WebSocket("wss://example.com/ws");
ws.onmessage = (e) => console.log(e.data);
ws.send(JSON.stringify({ type: "typing", room: 42 }));

Strengths: true bidirectional, low-latency messaging; binary data; ideal for chat, multiplayer games, collaborative editing, live cursors.

Costs:

  • A separate protocol — your proxy and load balancer must support the upgrade; authentication and reconnection are yours to implement (libraries help).
  • Stateful connections — scaling to several servers means a way to broadcast between them (Redis pub/sub, for example). (Horizontal vs vertical scaling.)
  • Serverless platforms often don't support long-lived WebSocket connections without a separate service.

How AI chat streaming works

When an AI reply appears word by word, it's almost always SSE-style streaming over HTTP. The Claude and OpenAI APIs stream responses as server-sent events. In your app, the browser usually sends a POST with the conversation, and your backend streams text back in the same response:

const res = await fetch("/api/chat", {
  method: "POST",
  headers: { "Content-Type": "application/json" },
  body: JSON.stringify({ messages }),
});

const reader = res.body.getReader();
const decoder = new TextDecoder();
for (;;) {
  const { done, value } = await reader.read();
  if (done) break;
  appendToReply(decoder.decode(value, { stream: true }));
}

That's a one-way stream per request — no WebSocket needed. How to add an AI chatbot to your app has the full backend.

Gotchas that break streaming in production

These cause the classic "streaming works locally, arrives all at once in production":

  • Proxy buffering. Nginx and some proxies buffer responses until they're complete. For SSE routes, disable it — in Nginx, proxy_buffering off; or send an X-Accel-Buffering: no header from your app. (Reverse proxies explained.)
  • Compression can buffer streamed output; exclude text/event-stream from gzip middleware.
  • Timeouts. Proxies and platforms close idle connections (often after 60 seconds). Send a comment line (: ping) every 15–30 seconds as a heartbeat.
  • Serverless function limits cap how long a response can stay open.
  • CDN caching — make sure stream endpoints are never cached (Cache-Control: no-cache).

Decision guide

You need Use
Updates every few seconds are fine Polling
Server → browser updates: notifications, progress, live feeds, AI responses SSE (or a streamed fetch)
Frequent messages both ways: chat rooms, games, collaboration WebSockets
You're on serverless and need real-time A managed real-time service, or polling

Start with the simplest option that meets the requirement. Many apps that reach for WebSockets only ever push from the server.

The summary

  • Polling: simplest, slight delay, scales like normal requests.
  • SSE: one-way server push over HTTP, auto-reconnect, great for notifications and AI streaming.
  • WebSockets: two-way, low-latency, more infrastructure to run.
  • Disable proxy buffering, send heartbeats, and never cache streaming endpoints.

EasySpawn runs your app as a long-lived server process rather than time-limited functions, so streaming responses, SSE and WebSockets just work — and Claude Code can test them against the real deployment. See how it works or join the waitlist.

Related: What Is Serverless? · HTTP Caching Headers · Structured Output From LLMs · Graceful Shutdown in Node.js · Streaming LLM Responses to the Browser

Keep reading