As of October 8, 2026, Claude API prices run from $0.10 per million input tokens on Claude Haiku 5.5 to $50 per million output tokens on Claude Fable 5.1. The four current models cost, input / output per 1M tokens: Haiku 5.5 $0.10 / $0.50 for prompts up to 100,000 tokens ($0.50 / $2.50 above that), Sonnet 5.5 $2 / $10, Opus 5.5 $4 / $20, and Fable 5.1 $10 / $50. The Batch API halves every rate, and a cache hit costs 2.5–10% of the normal input price, depending on the model.
To turn tokens into dollars, divide the token count by 1,000,000 and multiply by the rate. For example, 200,000 input tokens on Sonnet 5.5 cost 0.2 × $2 = $0.40. All prices below come from Anthropic's pricing page and are in USD.
Claude API prices per 1M tokens as of October 2026
Anthropic lists four models as its current lineup. Opus 5.5 launched on September 22, 2026, Sonnet 5.5 on September 28 and Haiku 5.5 on October 7.
| Model (per 1M tokens) | Input | Output | 5-min cache write | 1-hour cache write | Cache hit | Batch input / output |
|---|---|---|---|---|---|---|
| Claude Fable 5.1 | $10 | $50 | $12.50 | $20 | $0.25 | $5 / $25 |
| Claude Opus 5.5 | $4 | $20 | $5 | $8 | $0.20 | $2 / $10 |
| Claude Sonnet 5.5 | $2 | $10 | $2.50 | $4 | $0.10 | $1 / $5 |
| Claude Haiku 5.5, prompt ≤100,000 tokens | $0.10 | $0.50 | $0.125 | $0.20 | $0.01 | $0.05 / $0.25 |
| Claude Haiku 5.5, prompt >100,000 tokens | $0.50 | $2.50 | $0.625 | $1 | $0.05 | $0.25 / $1.25 |
Three things in this table change real bills more than the headline numbers:
- Haiku 5.5 has two price rows. Once a single prompt goes past 100,000 tokens, every category for that request moves to the higher row, which is 5x the lower one.
- Output costs 5x input on every model. Long answers and thinking tokens (billed as output) usually dominate the cost of short prompts.
- The other three models have no long-context surcharge. Fable 5.1, Opus 5.5 and Sonnet 5.5 bill a 900,000-token request at the same per-token rate as a 9,000-token one. The old "2x input above 200K tokens" rule applied to Sonnet 4 and 4.5-era models and is not current.
Prices on Amazon Bedrock and Google Cloud are set by those platforms. Their regional or multi-region endpoints add 10% over the global endpoint for Claude 4.5 and later models.
Older Claude models still on the price list, and which are retired
Anthropic's price table also lists older versions. Most are still callable, but several rows belong to models that no longer answer on the Claude API, according to the model deprecations page.
| Model | Input / output per 1M | Cache hit | Status on October 8, 2026 |
|---|---|---|---|
| Claude Mythos 5.1 | $10 / $50 | $0.25 | Active |
| Claude Fable 5, Mythos 5 | $10 / $50 | $1 | Active |
| Claude Opus 5, 4.8, 4.7, 4.6 | $5 / $25 | $0.50 | Active |
| Claude Opus 4.5 | $5 / $25 | $0.50 | Active, retirement not before November 24, 2026 |
| Claude Sonnet 5 | $2 / $10 | $0.20 | Active |
| Claude Sonnet 4.6 | $3 / $15 | $0.30 | Active |
| Claude Sonnet 4.5 | $3 / $15 | $0.30 | Deprecated, retires November 30, 2026 (replacement: Sonnet 5.5) |
| Claude Haiku 4.5 | $1 / $5 | $0.10 | Active, no deprecation date yet |
Ignore the rows for Opus 4.1, Opus 4, Sonnet 4 and Haiku 3.5 if you see them on the pricing page. Those models, along with Haiku 3 and Sonnet 3.7, were retired between February and August 2026, so their prices no longer apply to anything you can call.
Within each family, the newest version costs the same or less per token than the one it replaces. Opus 5.5 is cheaper than Opus 5 in every column, and Sonnet 5.5 and Fable 5.1 keep their predecessors' input and output prices while cutting cache hits in half or more.
How much do 1,000,000, 200,000 and 5,000 Claude tokens cost?
Multiply the token count by the per-million rate and divide by 1,000,000. Input and output are priced separately, so the table splits them.
| Model | 1M input | 1M output | 200K input | 200K output | 5K input | 5K output |
|---|---|---|---|---|---|---|
| Fable 5.1 | $10 | $50 | $2 | $10 | $0.05 | $0.25 |
| Opus 5.5 | $4 | $20 | $0.80 | $4 | $0.02 | $0.10 |
| Sonnet 5.5 | $2 | $10 | $0.40 | $2 | $0.01 | $0.05 |
| Haiku 5.5, prompts ≤100K | $0.10 | $0.50 | $0.02 | $0.10 | $0.0005 | $0.0025 |
| Haiku 5.5, prompt >100K | $0.50 | $2.50 | $0.10 | $0.50 | $0.0025 | $0.0125 |
| Haiku 4.5 | $1 | $5 | $0.20 | $1 | $0.005 | $0.025 |
On Haiku 5.5, the row depends on how the tokens arrive. Two hundred thousand input tokens spread over several prompts under 100K cost $0.02. The same 200,000 tokens sent as one prompt cost $0.10, because that request crosses the threshold.
How much text is a million tokens? Anthropic's rule of thumb is 1 token ≈ 4 characters ≈ 0.75 English words, so 1M tokens is roughly 750,000 words. Claude 4.7 and later models, which include all four current ones, use a newer tokenizer that produces about 30% more tokens for the same text. Code, other languages and structured data vary further, so count tokens on your own text with the target model before you budget.
Tokens to dollars: the formula and three worked examples
The cost of one request is:
cost = (input tokens × input rate + output tokens × output rate + cache-write tokens × write rate + cache-read tokens × hit rate) ÷ 1,000,000
The examples below use list prices from the first table and assume the same token counts on every model. Real token counts differ by model because of the tokenizer, and output length depends on your prompts.
One chat request: 3,000 input + 800 output tokens
A typical chatbot turn with some conversation history:
| Model | Cost per request | Cost per 100,000 requests |
|---|---|---|
| Fable 5.1 | $0.070 | $7,000 |
| Opus 5.5 | $0.028 | $2,800 |
| Sonnet 5.5 | $0.014 | $1,400 |
| Haiku 4.5 | $0.007 | $700 |
| Haiku 5.5 | $0.0007 | $70 |
For Sonnet 5.5: (3,000 × $2 + 800 × $10) ÷ 1,000,000 = $0.014.
A monthly workload: 50M input + 10M output tokens
| Model | Standard API | Batch API |
|---|---|---|
| Fable 5.1 | $1,000 | $500 |
| Opus 5.5 | $400 | $200 |
| Sonnet 5.5 | $200 | $100 |
| Haiku 4.5 | $100 | $50 |
| Haiku 5.5 (all prompts ≤100K) | $10 | $5 |
For Opus 5.5: 50 × $4 + 10 × $20 = $400. The Batch column applies only to work that can wait for asynchronous results.
A cached RAG day: 40K-token prefix, 10,000 requests
Inputs: each request sends the same 40,000-token system prompt and document set, plus 2,000 fresh input tokens, and gets 600 output tokens back. There are 10,000 requests a day. The 5-minute cache stays warm, so the prefix is written once every five minutes (288 writes a day) and read on the other requests.
| Model | Without caching | With caching | Saving |
|---|---|---|---|
| Fable 5.1 | $4,500/day | $741/day | 84% |
| Opus 5.5 | $1,800/day | $335/day | 81% |
| Sonnet 5.5 | $900/day | $168/day | 81% |
| Haiku 5.5 | $45/day | $10.32/day | 77% |
On Sonnet 5.5, the uncached day is 10,000 × (42,000 × $2 + 600 × $10) ÷ 1M = $900. With caching, fresh input and output cost $100, the 400M cached reads cost $40, and the 288 writes add about $28. The 42,000-token prompt stays under Haiku 5.5's 100K threshold, so the lower Haiku row applies throughout.

Prompt caching and Batch API: how much each discount saves
Prompt caching charges extra once to store a repeated prefix, then much less each time it is reused:
- A 5-minute cache write costs 1.25x the input price and pays for itself after one cache read.
- A 1-hour cache write costs 2x the input price and pays for itself after two reads.
- A cache hit costs 0.1x input on most models, 0.05x on Opus 5.5 and Sonnet 5.5, and 0.025x on Fable 5.1 and Mythos 5.1.
You can turn caching on with one top-level cache_control field or set explicit breakpoints. Haiku 5.5 caches prompts from 512 tokens. For code, see the Claude API prompt caching guide.
The Batch API takes 50% off both input and output for requests you submit as a batch and collect later. It stacks with caching and data residency. Fast mode can't run in Batch, and Managed Agents have no batch mode.
The two discounts solve different problems. Caching cuts the cost of a long prefix you resend often, even in real time. Batch cuts everything by half but only for jobs that can wait, such as nightly classification, evaluations or bulk summaries.
What raises a Claude bill above the list price
| Factor | Effect on cost | Applies to |
|---|---|---|
| Prompt over 100,000 tokens | Every rate for that request is 5x (input $0.50, output $2.50) | Haiku 5.5 only |
| Newer tokenizer | About 30% more tokens for the same text, varying by content | Claude 4.7 and later, including all current models |
| Extended or adaptive thinking | Thinking tokens are billed as output | All models with thinking on |
US-only inference (inference_geo: "us") | 1.1x on every token category; Sonnet 5.5 becomes $2.20 / $11 | Claude 4.6 and later |
| Fast mode (research preview) | Opus 5.5 at $8 / $40; Opus 5 and Opus 4.8 at $10 / $50 | Claude API only, no Batch |
| Web search | $10 per 1,000 searches, plus the tokens it adds | Any model using the tool |
| Code execution | Free when used with web search or web fetch (20260209 versions or later); otherwise 1,550 free container-hours per organization per month, then $0.05 per hour | Organization-wide |
| Managed Agents | Token cost at model rates plus $0.08 per running session-hour | Agent sessions |
| Tool definitions | The tool-use system prompt adds 286 input tokens on Opus 5.5, Sonnet 5.5 and Haiku 5.5 (406 on Haiku 5.5 when a tool is forced) | Every tool-use request |
Haiku 5.5's threshold is the sharpest jump in the table. A request with 100,000 input and 2,000 output tokens costs $0.011. Add one input token and the whole request moves to the higher row: 100,001 input and 2,000 output tokens cost $0.055.

The tokenizer change matters most when you compare old and new models. A prompt that counts as 10,000 tokens on Haiku 4.5 comes to roughly 13,000 on Haiku 5.5. That's $0.010 versus about $0.0013, still far cheaper but not the full 10x the list prices suggest. Moving from Sonnet 4.5 to Sonnet 5.5 works the same way: $0.030 for 10,000 input tokens becomes about $0.026 for 13,000.
On Haiku 5.5, thinking from earlier turns stays in the conversation by default, so it is billed again as input on later turns. For how each vendor bills reasoning, see Are Reasoning Tokens Billed? OpenAI, Claude, and Gemini Costs.
Web fetch has no extra fee beyond its tokens. Claude Platform on AWS and Microsoft Foundry bill in Claude Consumption Units at $0.01 each, at the same USD rates.
Which Claude model is cheapest for your workload
Pick the cheapest model whose output you can accept, then apply caching and Batch on top. On price alone:
- High-volume short prompts (classification, extraction, routing, chat under 100K tokens): Haiku 5.5 at $0.10 / $0.50 is the cheapest Claude option by a wide margin. GPT-6 Luna lists the same $0.10 / $0.50 and keeps it up to 272K input tokens. Claude Haiku 5.5 vs GPT-6 Luna Pricing: Same Until 100K, Then 5x works through where each one wins.
- Long single prompts over 100K tokens: Haiku 5.5 jumps to $0.50 / $2.50. That is still a quarter of Sonnet 5.5's price, so test whether Haiku's answers hold up on your documents before paying for Sonnet.
- General production work: Sonnet 5.5 at $2 / $10. If you still run Sonnet 4.5, you have to move before November 30, 2026, and Sonnet 5.5 lists at a lower price.
- The hardest reasoning and coding tasks: Opus 5.5 at $4 / $20 undercuts every older Opus. Details are in Claude Opus 5.5: API Price, Model ID, and Migration Checks.
- Fable 5.1 at $10 / $50 costs 2.5x Opus 5.5. Reserve it for tasks where you have seen the cheaper models fall short. See Claude Fable 5.1: API Pricing, Changes, and Migration Guide.
Capability differences between these models are covered in Claude Sonnet vs Opus vs Haiku vs Fable: Which Model to Use.
For reference, other vendors' list prices as of October 8, 2026: GPT-6.1 Sol $2 / $10, GPT-6 Astra $10 / $50, Gemini 3.1 Pro Preview $2 / $12 up to 200K tokens, and Gemini 2.5 Flash-Lite $0.10 / $0.40. The Gemini vs OpenAI vs Claude cost guide compares them on full workloads.
Checking the real cost: usage fields and a cost snippet
Every Messages API response includes a usage object with four token counts:
input_tokens: input after the last cache breakpoint, billed at the input ratecache_creation_input_tokens: tokens written to the cache, billed at the write ratecache_read_input_tokens: tokens served from the cache, billed at the hit rateoutput_tokens: the response, including thinking, billed at the output rate
Total input is the sum of the first three. This Python snippet prices one request on Sonnet 5.5, assuming 5-minute cache writes. It needs pip install anthropic and an ANTHROPIC_API_KEY environment variable.
import anthropic
# Claude Sonnet 5.5, USD per 1M tokens, as of October 8, 2026
RATES = {"input": 2.00, "output": 10.00, "cache_write": 2.50, "cache_read": 0.10}
client = anthropic.Anthropic()
msg = client.messages.create(
model="claude-sonnet-5-5",
max_tokens=800,
messages=[{"role": "user", "content": "Summarize our refund policy in five bullets."}],
)
u = msg.usage
cost = (
u.input_tokens * RATES["input"]
+ u.output_tokens * RATES["output"]
+ (u.cache_creation_input_tokens or 0) * RATES["cache_write"]
+ (u.cache_read_input_tokens or 0) * RATES["cache_read"]
) / 1_000_000
print(f"{u.input_tokens} in, {u.output_tokens} out: ${cost:.6f}")To estimate before sending, use the token counting endpoint with the model you plan to run. Its count reflects that model's tokenizer.
Billing, free credits and buying Claude API tokens
You don't buy Claude tokens in packs. The API is pay-as-you-go: Anthropic bills usage monthly in USD, by card or invoice, at the rates above.
- Free credits: new accounts receive a small amount of free credits. Anthropic doesn't publish the amount, and there is no ongoing free API tier.
- Volume discounts: negotiated with Anthropic case by case. No public discount schedule exists.
- Rate limits: set by your usage tier, which is separate from price. See Claude API Quota Tiers and Limits Explained.
- Subscriptions: Claude Pro and Max are monthly plans for the Claude apps and Claude Code, not API token billing. Claude Code Pricing: Pro vs Max and When the Upgrade Pays Off compares the two routes.
- Regions: Anthropic's supported countries list does not include Russia, mainland China or Hong Kong.
Two cost questions have no public answer yet: whether cached tokens count toward Haiku 5.5's 100,000-token threshold, and the per-model 5.5 prices on Bedrock and Google Cloud. Until they're documented, read the usage fields and your console's cost report after a test run.
Claude token cost questions
How much do Claude API tokens cost per million?
Between $0.10 and $10 per million input tokens and $0.50 to $50 per million output tokens across the current models, as of October 8, 2026. Haiku 5.5 is cheapest at $0.10 / $0.50 for prompts up to 100K tokens, and Fable 5.1 is the most expensive at $10 / $50. Batch halves those numbers.
How much is 200,000 tokens on Claude?
As input: $0.40 on Sonnet 5.5, $0.80 on Opus 5.5, $2 on Fable 5.1, and $0.02 on Haiku 5.5 if split into prompts under 100K ($0.10 as a single prompt). As output, multiply each by five.
Why did my token count go up after switching to a newer Claude model?
Claude 4.7 and later use a newer tokenizer that turns the same text into about 30% more tokens. Compare costs in dollars per request, not price per token.
Does Claude charge more for long prompts?
Only Haiku 5.5 does, at 5x above 100,000 tokens per prompt. Fable 5.1, Opus 5.5 and Sonnet 5.5 bill a full 1M-token context at their standard rates.



