Exponential Backoff and Jitter: Retrying Failed Requests Properly
When an API call fails with a timeout, 429 or 503, retrying immediately makes things worse. Exponential backoff waits longer after each failure, and jitter spreads retries out. Which errors to retry, Retry-After, a TypeScript implementation, idempotency, and retrying LLM API calls.
Network calls fail: timeouts, dropped connections, 429 Too Many Requests, 503 Service Unavailable, AI providers reporting they're overloaded. Many of these failures are temporary, so retrying is right. Retrying badly — immediately, forever, all at once — turns a blip into an outage.
Exponential backoff
Wait longer after each failed attempt, doubling each time:
attempt 1 fails → wait 1s
attempt 2 fails → wait 2s
attempt 3 fails → wait 4s
attempt 4 fails → wait 8s
give up after N attempts
The struggling service gets breathing room, and you don't burn through rate limits retrying a request that can't succeed yet.
Jitter: don't retry in lockstep
If a service hiccups and a thousand clients all fail at the same moment, plain backoff makes them all retry at 1s, then 2s, then 4s — a synchronised stampede each time. Jitter adds randomness so retries spread out:
// "full jitter": random delay between 0 and the backoff cap
const delay = Math.random() * Math.min(maxDelay, base * 2 ** attempt)
Full jitter is a simple, widely recommended default. (Preventing cache stampedes covers the same herd problem elsewhere.)
Which errors to retry
| Retry | Don't retry |
|---|---|
| Network errors, timeouts, connection resets | 400 Bad Request — the request is wrong |
429 Too Many Requests |
401 / 403 — credentials or permissions (401 vs 403) |
500, 502, 503, 504 |
404 Not Found |
Provider "overloaded" errors (Anthropic's 529) |
422 validation errors |
Retrying a 400 just fails slower. (HTTP status codes)
Respect Retry-After
Many APIs send a Retry-After header with 429 and 503 responses — seconds to wait (or a date). If present, use it instead of your own calculation. Rate-limit headers like x-ratelimit-reset tell you even more. (What is rate limiting?)
An implementation
const RETRYABLE = new Set([429, 500, 502, 503, 504, 529])
export async function fetchWithRetry(url: string, init: RequestInit = {}, maxAttempts = 5) {
const base = 500 // ms
const maxDelay = 30_000
for (let attempt = 0; ; attempt++) {
try {
const res = await fetch(url, { ...init, signal: AbortSignal.timeout(15_000) })
if (res.ok || !RETRYABLE.has(res.status) || attempt === maxAttempts - 1) return res
const retryAfter = Number(res.headers.get('retry-after'))
const delay = retryAfter > 0
? retryAfter * 1000
: Math.random() * Math.min(maxDelay, base * 2 ** attempt)
await new Promise(r => setTimeout(r, delay))
} catch (err) {
if (attempt === maxAttempts - 1) throw err // network error or timeout
await new Promise(r => setTimeout(r, Math.random() * Math.min(maxDelay, base * 2 ** attempt)))
}
}
}
Key details: a timeout per attempt, a maximum number of attempts, and a maximum delay.
Only retry what's safe to repeat
Retrying a GET is harmless. Retrying "charge the card" might charge twice if the first attempt actually succeeded but the response was lost. For operations that change things:
- make them idempotent — send an idempotency key so the server recognises the repeat and returns the original result, (Idempotency keys)
- or don't retry automatically.
Stripe, and many payment and email APIs, support idempotency keys for exactly this.
Retrying LLM API calls
AI APIs see rate limits and overload more than most. Good news: the official SDKs (Anthropic, OpenAI) already retry with exponential backoff on connection errors, 408, 409, 429 and 5xx — by default twice. Configure rather than reimplement:
const client = new Anthropic({ maxRetries: 4, timeout: 60_000 })
Also consider:
- Falling back to another model or provider after repeated overloads. (What is OpenRouter?)
- Queueing non-urgent work and processing it at a steady rate instead of bursting. (Background jobs, BullMQ)
- Streaming long responses so a slow generation doesn't hit a proxy timeout. (Streaming LLM responses)
Beyond retries: circuit breakers
If a dependency is clearly down, every request still pays for several retries. A circuit breaker notices repeated failures and fails fast for a while, then tries again. Worth adding when one dependency's outage would otherwise slow your whole app.
Checklist
- Retry only retryable errors.
- Exponential backoff with jitter, a cap, and a max attempts.
- Honour
Retry-After. - Timeouts on every attempt.
- Idempotency for anything that changes state.
- Log retries, so you notice when a dependency is degrading.
EasySpawn runs your app with background workers and Redis or Postgres queues on the same server, so retries and slow jobs happen off the request path. See how it works or join the waitlist.
Related: Idempotency Keys · What Is Rate Limiting? · Handling Webhooks Reliably · 504 Gateway Timeout
Keep reading
How to Build an AI Agent: The Loop, the Tools, and the Guardrails
An AI agent is a model in a loop: it decides on an action, a tool runs it, the result goes back, repeat. Build one from scratch with the Claude API's tool use, then see when to reach for the Claude Agent SDK or other frameworks — plus stopping conditions, memory, cost control and safety.
Structured Output From LLMs: Getting Reliable JSON Every Time
How to get a language model to return JSON your code can trust: why "respond in JSON" isn't enough, schema-constrained structured outputs with Claude and Zod, strict tool use, validation, handling refusals and truncation, and designing schemas models fill well.