Blog
4 min read

Streaming LLM Responses to the Browser: SSE, Fetch Streams, and Gotchas

Streaming makes AI features feel fast: words appear as they're generated instead of after a ten-second wait. How to stream from the Claude API on your server, forward it to the browser, read it in React, and fix the proxies and timeouts that buffer or cut off streams.

A model might take ten seconds to write a full answer. If your UI waits for all of it, users stare at a spinner. If you stream, the first words appear in under a second and the rest flow in — the same total time, but it feels instant.

Streaming has three legs: model → your server, your server → browser, and browser → screen. Each has its own details.

Leg 1: Model → your server

Never call the model API from the browser — your key would be public. (Hide API keys) Your server calls it with streaming on.

With the Anthropic TypeScript SDK:

import Anthropic from '@anthropic-ai/sdk'
const anthropic = new Anthropic()

const stream = anthropic.messages.stream({
  model: 'claude-sonnet-5-5',
  max_tokens: 1024,
  messages: [{ role: 'user', content: question }],
})

stream.on('text', (delta) => {
  // a few characters or words at a time
})

const final = await stream.finalMessage()   // full message, usage, stop reason

Under the hood, the API sends server-sent events: message_start, a series of content_block_delta events with text, and message_stop. The SDK assembles them for you. OpenAI's SDK has an equivalent streaming mode. (OpenAI API vs Claude API)

Leg 2: Your server → browser

The simplest robust approach: return a streaming HTTP response of plain text chunks. In a Next.js route handler (or any runtime with web streams):

// app/api/chat/route.ts
export async function POST(req: Request) {
  const { question } = await req.json()
  // authenticate the user and rate-limit here!

  const stream = anthropic.messages.stream({
    model: 'claude-sonnet-5-5',
    max_tokens: 1024,
    messages: [{ role: 'user', content: question }],
  })

  const encoder = new TextEncoder()
  const body = new ReadableStream({
    start(controller) {
      stream.on('text', (t) => controller.enqueue(encoder.encode(t)))
      stream.on('end', () => controller.close())
      stream.on('error', (e) => controller.error(e))
    },
    cancel() {
      stream.abort()   // user navigated away — stop paying for tokens
    },
  })

  return new Response(body, {
    headers: {
      'Content-Type': 'text/plain; charset=utf-8',
      'Cache-Control': 'no-cache, no-transform',
      'X-Accel-Buffering': 'no',
    },
  })
}

If you need structured events (text, tool calls, "done," errors), use SSE format (text/event-stream with data: …\n\n lines) instead of raw text. (Server-sent events vs WebSockets) Libraries like the Vercel AI SDK wrap all of this if you'd rather not hand-roll it.

Leg 3: Browser → screen

EventSource only supports GET, so for a POST with a body, read the response stream directly:

async function ask(question: string) {
  setAnswer('')
  const controller = new AbortController()
  abortRef.current = controller

  const res = await fetch('/api/chat', {
    method: 'POST',
    headers: { 'Content-Type': 'application/json' },
    body: JSON.stringify({ question }),
    signal: controller.signal,
  })
  if (!res.ok || !res.body) throw new Error(`Chat failed: ${res.status}`)

  const reader = res.body.pipeThrough(new TextDecoderStream()).getReader()
  while (true) {
    const { value, done } = await reader.read()
    if (done) break
    setAnswer(prev => prev + value)   // functional update — see useState
  }
}

Give users a Stop button that calls abortRef.current.abort(); with the cancel() handler above, that also stops the model call. (useState explained)

The gotchas

Proxies buffer the stream. Nginx, some CDNs and compression middleware collect the whole response before sending it — so the user sees nothing, then everything. Fixes: X-Accel-Buffering: no (Nginx), proxy_buffering off; in the location block, Cache-Control: no-transform, and excluding the route from response compression. (What is Nginx?)

Timeouts cut it off. Serverless functions have execution limits; proxies have idle timeouts (Cloudflare's is around 100 seconds). Long generations can be cut mid-sentence. Raise limits where you can, send periodic events, or run long jobs in the background and stream progress.

Rendering Markdown mid-stream produces flicker as half-finished syntax (**bol) appears. Use a Markdown renderer that tolerates partial input, or render plain text while streaming and Markdown at the end. Always sanitise rendered output. (XSS explained)

Errors after the first byte. Once streaming starts, the status code is already 200. Send errors as an in-band event (with SSE) so the UI can show them.

Re-render cost. Updating React state on every tiny chunk is fine for most chats; for very fast streams, batch updates with requestAnimationFrame.

Cost control. Streaming doesn't change token prices, but abandoned streams that keep running do. Abort on disconnect, set max_tokens, authenticate and rate-limit the endpoint. (Stop bots running up your AI bill)

The summary

  • Stream from the model on your server, never from the browser.
  • Return a ReadableStream (plain text or SSE) with no-buffering headers.
  • Read it in the browser with res.body.getReader(); support Stop via AbortController.
  • Watch for proxy buffering, timeouts, partial Markdown and in-stream errors.

EasySpawn runs your backend as an always-on server rather than short-lived functions, so long AI streams aren't cut off by execution limits. See how it works or join the waitlist.

Related: How to Add an AI Chatbot to Your App · Server-Sent Events vs WebSockets · Prompt Caching Explained · Structured Output From LLMs

Keep reading