Server-Sent Events vs WebSockets: Which for Real-Time Features and AI Streaming?
SSE streams updates from server to browser over plain HTTP; WebSockets open a two-way channel. How each works, code for both, why AI chat responses use SSE-style streaming, proxy buffering and connection-limit gotchas, and a decision guide including plain polling.
Live notifications, progress bars, dashboards that update themselves, and AI chat replies that appear word by word all need the server to push data to the browser. The three common ways are polling, Server-Sent Events (SSE) and WebSockets. Choosing well saves a lot of complexity.
Polling: just ask again
The browser asks every few seconds: "anything new?"
setInterval(async () => {
const res = await fetch("/api/notifications");
render(await res.json());
}, 10_000);
Simple, works everywhere, scales like normal requests. Fine when a few seconds' delay is acceptable — order status, a job queue, a dashboard. Don't dismiss it: it's often the right answer.
Server-Sent Events: a one-way stream over HTTP
SSE keeps one HTTP response open and the server writes events into it as they happen. The format is plain text:
event: progress
data: {"percent": 40}
data: {"percent": 80}
Each event is one or more data: lines followed by a blank line.
Server (Node.js / Express):
app.get("/api/events", (req, res) => {
res.set({
"Content-Type": "text/event-stream",
"Cache-Control": "no-cache",
Connection: "keep-alive",
});
res.flushHeaders();
const timer = setInterval(() => {
res.write(`data: ${JSON.stringify({ time: Date.now() })}\n\n`);
}, 1000);
req.on("close", () => clearInterval(timer));
});
Browser:
const source = new EventSource("/api/events");
source.onmessage = (e) => console.log(JSON.parse(e.data));
Strengths:
- Plain HTTP — works through most proxies, load balancers and CDNs; uses your normal cookies and auth.
- Automatic reconnection built into
EventSource, which resends the last event ID it saw (Last-Event-ID) so the server can resume. - Simple to implement on any backend.
Limits:
- One direction: server → browser. The browser sends data with normal requests.
EventSourceonly does GET and can't set custom headers. If you need a POST body (as with AI chat), read the stream withfetchinstead (below).- Text only (send JSON).
- Connection limits on HTTP/1.1: browsers allow about six connections per domain, so many open tabs with SSE can exhaust them. HTTP/2 removes this problem in practice.
WebSockets: a two-way channel
A WebSocket starts as an HTTP request, then upgrades into a persistent, two-way connection. Either side can send messages at any time. (What are WebSockets? covers the basics.)
const ws = new WebSocket("wss://example.com/ws");
ws.onmessage = (e) => console.log(e.data);
ws.send(JSON.stringify({ type: "typing", room: 42 }));
Strengths: true bidirectional, low-latency messaging; binary data; ideal for chat, multiplayer games, collaborative editing, live cursors.
Costs:
- A separate protocol — your proxy and load balancer must support the upgrade; authentication and reconnection are yours to implement (libraries help).
- Stateful connections — scaling to several servers means a way to broadcast between them (Redis pub/sub, for example). (Horizontal vs vertical scaling.)
- Serverless platforms often don't support long-lived WebSocket connections without a separate service.
How AI chat streaming works
When an AI reply appears word by word, it's almost always SSE-style streaming over HTTP. The Claude and OpenAI APIs stream responses as server-sent events. In your app, the browser usually sends a POST with the conversation, and your backend streams text back in the same response:
const res = await fetch("/api/chat", {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({ messages }),
});
const reader = res.body.getReader();
const decoder = new TextDecoder();
for (;;) {
const { done, value } = await reader.read();
if (done) break;
appendToReply(decoder.decode(value, { stream: true }));
}
That's a one-way stream per request — no WebSocket needed. How to add an AI chatbot to your app has the full backend.
Gotchas that break streaming in production
These cause the classic "streaming works locally, arrives all at once in production":
- Proxy buffering. Nginx and some proxies buffer responses until they're complete. For SSE routes, disable it — in Nginx,
proxy_buffering off;or send anX-Accel-Buffering: noheader from your app. (Reverse proxies explained.) - Compression can buffer streamed output; exclude
text/event-streamfrom gzip middleware. - Timeouts. Proxies and platforms close idle connections (often after 60 seconds). Send a comment line (
: ping) every 15–30 seconds as a heartbeat. - Serverless function limits cap how long a response can stay open.
- CDN caching — make sure stream endpoints are never cached (
Cache-Control: no-cache).
Decision guide
| You need | Use |
|---|---|
| Updates every few seconds are fine | Polling |
| Server → browser updates: notifications, progress, live feeds, AI responses | SSE (or a streamed fetch) |
| Frequent messages both ways: chat rooms, games, collaboration | WebSockets |
| You're on serverless and need real-time | A managed real-time service, or polling |
Start with the simplest option that meets the requirement. Many apps that reach for WebSockets only ever push from the server.
The summary
- Polling: simplest, slight delay, scales like normal requests.
- SSE: one-way server push over HTTP, auto-reconnect, great for notifications and AI streaming.
- WebSockets: two-way, low-latency, more infrastructure to run.
- Disable proxy buffering, send heartbeats, and never cache streaming endpoints.
EasySpawn runs your app as a long-lived server process rather than time-limited functions, so streaming responses, SSE and WebSockets just work — and Claude Code can test them against the real deployment. See how it works or join the waitlist.
Related: What Is Serverless? · HTTP Caching Headers · Structured Output From LLMs · Graceful Shutdown in Node.js · Streaming LLM Responses to the Browser
Keep reading
UUID vs Auto-Increment IDs: Which Primary Key Should You Use?
Sequential integers or UUIDs for your primary keys? The real trade-offs — size, index performance, guessability, merging data, leaking business metrics — why UUIDv7 changes the answer, Postgres 18's uuidv7(), and the common hybrid of internal IDs plus public IDs.
Structured Output From LLMs: Getting Reliable JSON Every Time
How to get a language model to return JSON your code can trust: why "respond in JSON" isn't enough, schema-constrained structured outputs with Claude and Zod, strict tool use, validation, handling refusals and truncation, and designing schemas models fill well.