Blog
5 min read

Claude API Pricing Explained: What Your AI Feature Will Actually Cost

How Claude API billing works: per-token input and output prices for each model, prompt caching, the 50% Batch discount, and worked examples for a chatbot and a summariser — so you can estimate your bill before you ship.

If you're adding AI to your app — a chatbot, a summariser, a "write this for me" button — you'll pay for each call to the model. The pricing is simple once you see how it's measured, and estimating your bill before launch takes five minutes.

Prices below are Anthropic's first-party API rates as published in October 2026. Check Anthropic's pricing page before you budget.

You pay per token

A token is a chunk of text — on average about three-quarters of an English word. (More in what are tokens in AI?.) Every API call is billed on two counts:

  • Input tokens — everything you send: your system prompt, the conversation so far, any documents, the user's message.
  • Output tokens — everything the model writes back.

Output costs five times as much as input on every current model.

Current prices

Per million tokens:

Model Input Output Cache read
Claude Fable 5.1 $10.00 $50.00 $0.25
Claude Opus 5.5 $4.00 $20.00 $0.20
Claude Sonnet 5.5 $2.00 $10.00 $0.20
Claude Haiku 4.5 $1.00 $5.00 $0.10

"Per million" sounds abstract. In practice: one million input tokens is roughly 750,000 words — several novels.

Not sure which model? See Claude Opus vs Sonnet vs Haiku.

Worked example 1: a support chatbot

Say each conversation turn sends a 1,000-token system prompt, 1,500 tokens of history and a 100-token question, and gets a 300-token answer. On Sonnet 5.5:

  • Input: 2,600 tokens × $2 / 1,000,000 = $0.0052
  • Output: 300 tokens × $10 / 1,000,000 = $0.0030
  • Total: about $0.008 per turn.

At 10,000 turns a month, that's around $80. On Haiku 4.5, it would be about $40.

Note how history grows: by the tenth turn of a long conversation you're re-sending all nine earlier turns. Long conversations cost more per message than short ones.

Worked example 2: summarising documents

Summarise 2,000 support tickets of ~800 tokens each into 150-token summaries, on Haiku 4.5:

  • Input: 1.6M tokens × $1 = $1.60
  • Output: 300K tokens × $5 = $1.50
  • Total: about $3.10 — or about $1.55 with the Batch API (below).

Three ways to pay less

1. Prompt caching

If every request starts with the same long prefix — a big system prompt, a product catalogue, a set of documents — you can cache it. The first request writes the cache (at 1.25× the input price for a five-minute cache, or 2× for a one-hour cache). Later requests read it at roughly a tenth of the input price or less. For a chatbot with a long fixed system prompt, caching can cut input cost dramatically. See prompt caching explained.

2. The Batch API — 50% off

If the result doesn't need to come back immediately — nightly reports, bulk tagging, back-filling summaries — send requests through the Message Batches API. They're processed asynchronously (usually well within 24 hours) at half price.

3. Pick the smallest model that does the job well

Haiku costs a quarter of Opus. For classification and extraction it's often just as good. Test with 20 real examples before deciding.

Other things that cost money

  • Images and PDFs are converted to tokens; a page or a screenshot costs roughly what a page of text does.
  • Thinking. Current models can reason before answering; those reasoning tokens are billed as output. Lower the effort setting for simple tasks.
  • Server tools such as web search have their own per-use fees.
  • Long contexts. The newest models accept up to a million tokens per request — which is powerful and also a way to spend a lot per call.

Estimate before you ship

  1. Write down the typical input and output size of one call (use the API's token-counting endpoint for real numbers).
  2. Multiply by your expected calls per month.
  3. Add 50% for retries, long conversations and growth.
  4. Set a monthly spend limit in the Anthropic Console so a bug or a bot can't run up an unlimited bill.

And protect the endpoint that calls the model: an unprotected AI route is an open invitation to abuse. See stop bots running up your AI bill.

API pricing vs Claude subscriptions

The API is for software calling Claude. Claude Pro and Max are subscriptions for a person using the Claude apps and Claude Code. You can't use a Pro subscription to power your app's chatbot. See Claude Pro vs Max vs API key.

The summary

  • You pay per input token and per output token; output costs 5× input.
  • Sonnet 5.5 is $2 / $10 per million tokens; Haiku 4.5 is $1 / $5.
  • Caching cuts the cost of repeated prefixes; the Batch API halves the price of non-urgent work.
  • Estimate per call, multiply by volume, add headroom, and set a spend limit.

EasySpawn runs your app's backend on its own server with API keys stored as server-side environment variables, so the code that calls Claude — and the rate limits that protect it — never touch the browser. See how it works or join the waitlist.

Related: What Is an API Key? · Hide API Keys in an AI-Built App · Structured Output From LLMs · How Much Does It Cost to Run an App? · OpenAI API vs Claude API

Keep reading