As of October 8, 2026, Claude Haiku 5.5 and GPT-6 Luna have the same list price on their own APIs: $0.10 per 1M input tokens, $0.01 per 1M cached input tokens, $0.125 per 1M cache-write tokens and $0.50 per 1M output tokens (Anthropic pricing, OpenAI pricing). The match only holds while the Haiku prompt stays at or under 100,000 tokens. One token more and the whole Haiku request, output included, moves to rates five times higher: $0.50 input and $2.50 output. Luna keeps its base rates up to 272,000 input tokens.
Per request, that gives three bands:
- Prompts up to 100,000 tokens: identical cost. A 2,000-token classification prompt with a 200-token answer costs $0.0003 on either model.
- Prompts from 100,001 to 272,000 tokens: Luna is exactly 5x cheaper.
- Prompts above 272,000 tokens: Luna's rates rise too, and the gap narrows to about 2.5x.
Equal rates still don't guarantee an equal bill, because the two models spend different numbers of tokens on the same job. In Artificial Analysis's benchmark run, Haiku 5.5 at medium effort wrote about three times as many output tokens per task as Luna at medium (33k vs 11k), cost $0.05 per task against $0.02, and scored 34 against 30. Compared at similar scores instead of similar effort labels, the per-task costs come out close.
The short version: Luna is the safe default for long-context requests and for agent sessions you can't keep under 100,000 tokens. For short, high-volume work, the price sheets tie, and the cheaper model is whichever reaches your quality bar with fewer output tokens.
Claude Haiku 5.5 vs GPT-6 Luna API prices as of October 8, 2026
Both models were priced for the same job: Anthropic positions Haiku 5.5 "for high-volume, latency-sensitive tasks such as classification, extraction, and routing" (model overview), and OpenAI calls Luna its "most efficient model for focused, high-volume tasks" (model page). Standard first-party rates, USD per 1M tokens:
| Rate | Haiku 5.5, prompt ≤100,000 | Haiku 5.5, prompt >100,000 | GPT-6 Luna, input ≤272K | GPT-6 Luna, input >272K |
|---|---|---|---|---|
| Input | $0.10 | $0.50 | $0.10 | $0.20 |
| Cached input (cache hit) | $0.01 | $0.05 | $0.01 | $0.02 |
| Cache write (Haiku: 5-minute) | $0.125 | $0.625 | $0.125 | $0.25 |
| Cache write, 1-hour | $0.20 | $1.00 | no separate rate listed | no separate rate listed |
| Output (includes thinking) | $0.50 | $2.50 | $0.50 | $0.75 |
How each threshold works:
- Haiku 5.5: Anthropic's pricing page says a prompt "of over 100,000 tokens pays higher prices." A prompt of exactly 100,000 tokens stays in the low tier. The table labels every row, output included, by prompt length, so a request with a 100,001-token prompt pays $2.50 per 1M output tokens as well. Everything you send counts as input: system prompt, messages, tool results, images, documents and tool definitions (context windows).
- GPT-6 Luna: requests over 272K input tokens pay 2x input and cache rates and 1.5x output "for the entire request." The jump is smaller and starts 172,000 tokens later. GPT-6 Luna vs. Sol pricing: calculate the cost of your workload covers that rule in more detail.
- Cached tokens: neither vendor's pricing documentation says whether cache reads and writes count toward its pricing threshold. Anthropic does say all three input counts (uncached, cache read, cache write) count toward the context window. The calculations below assume cached tokens count toward the threshold, which is the safer assumption for a budget.
Modifiers that change the numbers above:
- Batch: 50% off on both. Haiku 5.5 Batch is $0.05/$0.25 under 100K and $0.25/$1.25 over it. Luna's Batch and Flex tiers are $0.05/$0.25 under 272K.
- Faster tiers: Luna has a Fast tier at 2x Standard ($0.20/$1.00). Haiku 5.5 has no Fast mode (that's Opus-only) and doesn't support Priority Tier.
- Data residency: US-only inference on Haiku 5.5 (
inference_geo: "us") multiplies every rate by 1.1. - Other clouds: Haiku 5.5 also runs on Amazon Bedrock, Google Cloud and Microsoft Foundry. Bedrock and Google Cloud set their own prices; every figure here is a first-party Claude API or OpenAI API price.
Cost per request by prompt size: equal to 100K tokens, then 5x
Luna is cheaper per request only once the prompt passes 100,000 tokens, assuming both models see the same token counts. The formula for one request without caching:
cost = (prompt tokens × input rate + output tokens × output rate) / 1,000,000
with the rates picked by prompt size from the table above. For a 150,000-token prompt and a 2,000-token answer, Haiku costs 150,000 × $0.50 + 2,000 × $2.50 = $80,000 per million requests, or $0.080 each. Luna costs 150,000 × $0.10 + 2,000 × $0.50 = $16,000 per million, or $0.016 each.
| Prompt tokens | Output tokens | Claude Haiku 5.5 | GPT-6 Luna | Haiku ÷ Luna |
|---|---|---|---|---|
| 2,000 | 200 | $0.0003 | $0.0003 | 1x |
| 20,000 | 1,000 | $0.0025 | $0.0025 | 1x |
| 100,000 | 2,000 | $0.0110 | $0.0110 | 1x |
| 100,001 | 2,000 | $0.0550 | $0.0110 | 5x |
| 150,000 | 2,000 | $0.0800 | $0.0160 | 5x |
| 272,000 | 2,000 | $0.1410 | $0.0282 | 5x |
| 300,000 | 5,000 | $0.1625 | $0.06375 | 2.55x |
Standard prices as of October 8, 2026, no caching, same token count on both models.
At volume, the first row is $300 per million requests on either model. The 150,000-token row is $80,000 per million on Haiku and $16,000 on Luna. Batch halves both ($0.040 vs $0.008 per request for the 150,000-token case) but keeps the 5x ratio.
The line is a cliff, not a slope. Adding one token to a 100,000-token prompt multiplies the Haiku request cost by five. Requests that sit near the line, such as RAG answers with a variable number of retrieved chunks or prompts that carry long tool definitions, can flip between tiers from one call to the next. That's why the share of your traffic above 100,000 tokens matters more than the average prompt size.
Why the same task can cost more on Haiku 5.5 below 100K tokens
The table above assumes both models read and write the same number of tokens. In practice they don't, for three reasons.
Haiku 5.5 thinks longer at the same effort setting
Thinking tokens are billed as output on both models (OpenAI's reasoning guide says the same for Luna's reasoning). Haiku 5.5 thinks adaptively by default at medium effort, and Luna's default reasoning effort is also medium, but the shared label doesn't mean the same amount of work. In Artificial Analysis's runs (Intelligence Index v4.3.2, 10 evaluations), output tokens per task were:
| Effort label | Haiku 5.5 output tokens per task | GPT-6 Luna output tokens per task |
|---|---|---|
| Low | 17k | 2k |
| Medium | 33k (20k of it reasoning) | 11k (6k reasoning) |
| High | 55k (37k reasoning) | 20k (13k reasoning) |
At medium, the output alone costs 33,000 × $0.50 / 1M = $0.0165 per task on Haiku against $0.0055 on Luna, before any input. For how reasoning tokens show up in usage fields and invoices, see Are Reasoning Tokens Billed? OpenAI, Claude, and Gemini Costs.
Kept thinking blocks make the next prompt bigger
Haiku 5.5 keeps previous thinking blocks in the conversation by default, and the kept blocks are billed as input on later requests (context windows). Haiku 4.5 dropped them. In a multi-turn conversation, each Haiku turn therefore carries the reasoning of earlier turns forward, which pushes the prompt toward 100,000 tokens sooner. Context editing can clear old thinking blocks if you don't need them.
Tokenizers count the same text differently
Haiku 5.5 uses Anthropic's newer tokenizer, which counts the same text as about 30% more tokens than Haiku 4.5 did (what's new in Haiku 5.5); Simon Willison measured about 1.25x on his own prompt. No public comparison exists between Haiku 5.5 and Luna token counts for identical text. A document that is 95,000 tokens for one model may not be 95,000 for the other, so count your real prompts on both before assuming the same bill.
Agent loops: a 20-turn run crosses Haiku's 100K line at turn 16
Agent loops and long chats are where the 100,000-token tier bites, because the prompt grows every turn. A worked example with these assumptions:
- 20 turns. The first prompt is 12,000 tokens, and each turn adds 6,000 (tool results plus the previous answer), so turn 20 sends 126,000 tokens.
- 1,500 output tokens per turn. Standard prices. The same token counts on both models.
- With caching: each turn reads the previous prompt from cache and writes the new 6,000 tokens at the 5-minute rate. Cached tokens count toward the threshold.
Turns 1–15 stay at or under 96,000 tokens. Turns 16–20 run from 102,000 to 126,000 and land in Haiku's high tier.
| One 20-turn run | Claude Haiku 5.5 | GPT-6 Luna | Haiku ÷ Luna |
|---|---|---|---|
| No caching | $0.396 | $0.153 | 2.59x |
| Prompt caching | $0.0949 | $0.0433 | 2.19x |
| Caching, context reset to 40,000 before turn 16 | $0.0441 | $0.0441 | 1x |

With caching, the last five turns are a quarter of the run but $0.0645 of Haiku's $0.0949, or 68% of the bill. The third row shows the fix: if the agent compacts its context so no turn exceeds 100,000 tokens, both models charge the same rates for every turn and the bills match. Compaction doesn't make Haiku cheaper than Luna; it removes the Haiku penalty. The cost of producing the summary isn't included.
Real agents will drift from this model. Haiku's kept thinking blocks make its context grow faster than the fixed 6,000 tokens per turn assumed here, and output per turn differs between the models as shown above. On launch day, users on Reddit reported multi-turn jobs that cost many times more on Haiku than on Luna, with commenters pointing at the 100K tier; the runs' settings weren't published.
Cost per task at matched quality: within 1.7x on Artificial Analysis
Per-token prices say little about what one finished task costs. Artificial Analysis ran both models through its Intelligence Index v4.3.2 (10 evaluations) at each effort setting, one day after Haiku 5.5's launch (low, medium, high, max):
| Effort label | Haiku 5.5 index | Haiku 5.5 cost per task | GPT-6 Luna index | GPT-6 Luna cost per task |
|---|---|---|---|---|
| Low | 29 | $0.02 | 22 | $0.0045 |
| Medium | 34 | $0.05 | 30 | $0.02 |
| High | 38 | $0.08 | 33 | $0.03 |
| Max | 43 | $0.21 | 38 | $0.07 |
The max row comes from Artificial Analysis's default comparison page and is the least firm figure in the table; it may be revised.
At the same label, Haiku costs 2.5–4.4x more per task and scores 4–7 points higher. Lining up similar scores instead gives a different picture:
| Similar index score | Claude Haiku 5.5 | GPT-6 Luna |
|---|---|---|
| About 30 | Low: 29 at $0.02 | Medium: 30 at $0.02 |
| About 33–34 | Medium: 34 at $0.05 | High: 33 at $0.03 |
| 38 | High: 38 at $0.08 | Max: 38 at $0.07 |
At comparable quality the gap shrinks to between nothing and about 1.7x, and Haiku gets there faster where both latencies are shown. At an index of 38, Haiku at high effort had a 28.01-second time to first token against 111.31 seconds for Luna at max. Haiku at medium (13.40 seconds) also beat Luna at high (21.08 seconds).
Keep the limits in view. This is one evaluator's harness, measured a day after launch, and the figures may be revised. Haiku's low- and high-effort cost figures carry an asterisk on those pages, so read its footnote before relying on them. The index includes a long-context test (AA-LCR), and the pages don't say whether any Haiku requests in it crossed 100,000 tokens. An index score is not your task: use these pairs to decide which settings to test, not which model to buy.
The same pattern shows up at the flagship level; GPT-6.1 Sol vs Claude Opus 5.5: Half the Price, 1/6 the Task Cost walks through it for the larger models.
Which model is cheaper for your workload: a decision rule
Sort your traffic by prompt size first, then by how many tokens each model needs to pass your quality check.

- Short, high-volume requests under 100,000 tokens (classification, extraction, routing, support replies): the price sheets are identical, so output tokens per task decide. Start each model at its lowest setting that passes: Haiku at
loweffort or with thinking disabled, Luna atnoneorlow. Whichever passes with fewer output tokens is cheaper. - Prompts that regularly land between 100,001 and 272,000 tokens (long RAG context, large documents): choose Luna. Each Haiku request costs 5x as much, so to cost less per accepted answer Haiku would need five times Luna's pass rate. That only happens if Luna succeeds less than 20% of the time.
- Multi-turn agents and long chats: Luna by default. Haiku matches Luna's price only if every turn's prompt stays at or under 100,000 tokens, through compaction, clearing old thinking blocks or trimming tool output.
- Prompts over 272,000 tokens: Luna, at roughly 2.5x cheaper per request in the example above.
- Latency-bound work with a quality bar above Luna's medium effort: Haiku may be worth its per-task premium, since it reached matched scores with a shorter time to first token in Artificial Analysis's runs.
Put together: Luna is cheaper for any traffic that crosses 100,000 prompt tokens. Below that line, neither model is cheaper on paper, and your own output-token and pass-rate numbers decide. One test changes the answer in either direction: run a few hundred real requests on both at matched quality and compare the cost per accepted result.
Settings that lower the bill on Haiku 5.5 and GPT-6 Luna
| Cost lever | Claude Haiku 5.5 | GPT-6 Luna |
|---|---|---|
| Reasoning control | effort, default medium; thinking: {"type": "disabled"} turns thinking off at high effort or below | Reasoning effort none, low, medium (default), high, xhigh, max |
| Long-prompt surcharge | All rates 5x for prompts over 100,000 tokens | 2x input and cache, 1.5x output over 272K input tokens |
| Earlier reasoning in context | Kept by default and billed as input; context editing can clear it | — |
| Prompt caching | Cache hit $0.01; 5-minute write $0.125; 1-hour write $0.20; minimum 512 tokens | Cached input $0.01; cache write $0.125 |
| Asynchronous discount | Batch API, 50% off | Batch and Flex, 50% off |
| Paying for speed | No Fast mode, no Priority Tier | Fast tier at 2x Standard |
Anthropic's migration guide recommends a lower effort level over turning thinking off: at lower effort, the model thinks less and can skip thinking entirely on simple requests. Simon Willison's single test shows how wide the spread is: the same prompt cost 0.0936 cents at low effort and 3.3826 cents at max effort.
Measure before switching: token counts to log and a cost function
Five numbers settle the comparison for your own traffic:
- Prompt size distribution, counted with
modelset toclaude-haiku-5-5. Anthropic says to recount rather than reuse Haiku 4.5 counts. Note the share of requests above 100,000 and above 272,000 tokens. - Output tokens per task, thinking included, at each effort setting you consider.
- Pass rate on your own check (exact match, schema validity, a reviewer's accept/reject) at each setting.
- Prompt growth per turn in agents, and the turn where Haiku crosses 100,000 tokens.
- Cache reads and writes. On the Claude API, the usage block splits input into
input_tokens,cache_read_input_tokensandcache_creation_input_tokens, plusoutput_tokens. OpenAI's Responses API reportsinput_tokensandoutput_tokens, with the reasoning share underoutput_tokens_details.reasoning_tokens; it is already insideoutput_tokens, so don't add it twice.
Feed the logged counts into a cost function that applies each vendor's threshold. This Python snippet uses the Standard prices above and the 5-minute cache-write rate; multiply by 0.5 for Batch:
# Standard list prices, USD per 1M tokens, as of October 8, 2026
HAIKU_LOW = {"input": 0.10, "cache_write": 0.125, "cache_read": 0.01, "output": 0.50}
HAIKU_HIGH = {k: v * 5 for k, v in HAIKU_LOW.items()} # prompt over 100,000 tokens
LUNA_LOW = {"input": 0.10, "cache_write": 0.125, "cache_read": 0.01, "output": 0.50}
LUNA_HIGH = {"input": 0.20, "cache_write": 0.25, "cache_read": 0.02, "output": 0.75} # input over 272,000
def request_cost(model: str, uncached: int, cache_write: int, cache_read: int, output: int) -> float:
"""Cost of one request in USD. `output` includes thinking/reasoning tokens."""
prompt = uncached + cache_write + cache_read # assumption: cached tokens count toward the threshold
if model == "claude-haiku-5-5":
rates = HAIKU_HIGH if prompt > 100_000 else HAIKU_LOW
else: # "gpt-6-luna"
rates = LUNA_HIGH if prompt > 272_000 else LUNA_LOW
total = (uncached * rates["input"] + cache_write * rates["cache_write"]
+ cache_read * rates["cache_read"] + output * rates["output"])
return round(total / 1_000_000, 6)
print(request_cost("claude-haiku-5-5", 150_000, 0, 0, 2_000)) # 0.08
print(request_cost("gpt-6-luna", 150_000, 0, 0, 2_000)) # 0.016Sum request_cost over a day of logged requests per model, divide by the number of accepted results, and you have the figure that matters. If you don't have a Claude key yet, How to Get a Claude API Key Without Buying Someone Else’s Secret covers the legitimate routes.
Coming from Haiku 4.5: what the price cut changes
Haiku 4.5 lists at $1 input and $5 output per 1M tokens, so Haiku 5.5 is 90% cheaper per token under 100,000 and 50% cheaper above it. Anthropic says that after the tokenizer change it costs around 75% less to run on average, an average across its traffic rather than a promise for yours.
A worked example: a prompt that was 10,000 tokens on Haiku 4.5 becomes about 13,000 on Haiku 5.5. With a 500-token answer, the request goes from $0.0125 to $0.00155, about 88% less. The new threshold sits lower than it looks: 100,000 Haiku 5.5 tokens is roughly 76,900 tokens as Haiku 4.5 counted them, so check prompts that used to run above that size.
Switching also changes API behavior that can break existing calls:
- Thinking is on by default at
mediumeffort, and thinking tokens count towardmax_tokens. Calls that ran without thinking on Haiku 4.5 will produce more output unless you lower the effort. - A non-default
temperature,top_portop_k, or an assistant prefill, returns a 400 error. - Safety classifiers can return
stop_reason: "refusal", with no server-side fallback.
Haiku 4.5 remains active, with retirement scheduled no sooner than October 15, 2026, according to Anthropic's model deprecations page. If Haiku 5.5 at high effort still misses your quality bar, the next Claude step up is covered in Claude Opus 5.5: API Price, Model ID, and Migration Checks.
Claude Haiku 5.5 vs GPT-6 Luna FAQ
Is GPT-6 Luna cheaper than Claude Haiku 5.5?
Per token, only for prompts over 100,000 tokens: there Luna is 5x cheaper up to 272,000 input tokens and about 2.5x cheaper beyond. At or under 100,000 tokens, the list prices are identical as of October 8, 2026. Per task, Luna was cheaper at every matching effort label in Artificial Analysis's runs, but close to Haiku at matched scores.
Is GPT-6 Luna better than Claude Haiku 5.5?
Not on Artificial Analysis's index: Haiku 5.5 scored higher at every effort label (34 vs 30 at medium, 38 vs 33 at high). Press coverage of Anthropic's launch charts also shows Haiku 5.5 ahead of Luna on GDPval-AA and Terminal-Bench 4.0, but those results are vendor-reported. Luna can roughly match Haiku's low-, medium- and high-effort scores by running one or two effort levels higher, at a similar or lower cost per task. Nothing on Luna matched Haiku at max effort (43 vs 38).
How expensive is Claude Haiku 5.5?
On the Claude API, $0.10 per 1M input tokens and $0.50 per 1M output tokens for prompts up to 100,000 tokens; $0.50 and $2.50 for prompts over 100,000. Cache hits cost $0.01 (or $0.05 in the high tier), and Batch takes 50% off. A 2,000-token request with a 200-token answer costs $0.0003.
Do cached tokens count toward Haiku 5.5's 100,000-token threshold?
Anthropic's pricing, prompt caching and context window pages don't say. Cached tokens do count toward the context window. Budget as if the full prompt, cached part included, decides the tier, and check the billed rate on a long cached request in your usage data.



