Skip to content

Claude Sonnet vs Opus vs Haiku vs Fable: Which Model to Use

Haiku 4.5 is the cheapest and fastest Claude model, Fable 5.1 the most capable. Where to start depends on whether you spend plan limits or pay per API token.

A
•••14 min read•AI Model Comparison
Four stacks on a platform sized by Claude API input price per million tokens: Haiku 4.5 $1, Sonnet 5 $2, Opus 5.5 $4 and Fable 5.1 $10

As of September 24, 2026, the current Claude lineup is Haiku 4.5, Sonnet 5, Opus 5.5 and Fable 5.1. They run from fastest and cheapest (Haiku) to most capable, slowest and most expensive (Fable). The model to start with depends on how you pay for Claude:

  • Claude app on Free, Pro or Max: start with Sonnet 5. It uses a moderate share of your usage limit and handles coding, writing and analysis. Use Haiku 4.5 for quick answers and summaries, Opus 5.5 when Sonnet struggles, and Fable 5.1 for long, accuracy-critical work (included on Max, pay-as-you-go on Pro).
  • Claude Code: the account default is now Opus 5.5 on Pro and Max, but Anthropic still recommends Sonnet for most coding because Opus uses several times more quota per turn. /model opusplan plans with Opus and executes with Sonnet.
  • Claude API: Anthropic's docs say to start with Opus 5.5 for most workloads, or with Haiku 4.5 when cost and latency come first. Move to Fable 5.1 only when Opus 5.5 at xhigh or max effort still falls short on your tests.

On every surface, try a different effort level before you switch models. The sections below show the numbers behind each choice and the signals that tell you to move up or down.

Haiku, Sonnet, Opus and Fable at a glance

All figures are from Anthropic's models overview, API pricing page and plan pricing page as of September 24, 2026. API prices are standard first-party rates in US dollars per million tokens.

Haiku 4.5Sonnet 5Opus 5.5Fable 5.1
Anthropic's positioningFastest model with near-frontier intelligenceBest combination of speed and intelligenceLong-running agentic coding and knowledge workDemanding reasoning and long-horizon agentic work
Relative speedFastestFastModerateSlower
API input / output$1 / $5$2 / $10$4 / $20$10 / $50
API cache read$0.10$0.20$0.20$0.25
Context window200K tokens1M tokens1M tokens1M tokens
Max output (standard API)64K tokens128K tokens128K tokens128K tokens
Default effort on the APINot supportedhighmediumhigh
Reliable knowledge cutoffFeb 2025Jan 2026Jun 2026Jun 2026
Free planYesYesNoNo
Pro planYesYesYesUsage credits only
Max planYesYesYesUp to 50% of weekly limits
API model IDclaude-haiku-4-5-20251001claude-sonnet-5claude-opus-5-5claude-fable-5-1

A few details in that table change real decisions:

  • Sonnet 5 costs $2 / $10, and that price is now permanent. It launched as introductory pricing, and a rise to $3 / $15 was scheduled for September 1, 2026. Anthropic's pricing page now says that increase "will not occur." If you see $3 / $15 quoted for Sonnet 5, it refers to the canceled increase.
  • Sonnet 5 and Opus 5.5 charge the same $0.20 for cache reads. A cache read is a part of the prompt that Claude already stored from an earlier request, such as a long system prompt or a codebase an agent keeps re-reading. In work dominated by cache reads, the gap between Opus 5.5 and Sonnet 5 shrinks (see the cost examples below).
  • Haiku 4.5 is the only model with a 200K context window. If a single request needs more than that, Haiku is out.
  • Default effort differs by model. Opus 5.5 runs at medium unless you set it, while Sonnet 5 and Fable 5.1 run at high. Comparing models at their defaults therefore mixes a model difference with an effort difference.
  • Equal text is not equal tokens. Anthropic says Claude 4.7 and later models use a tokenizer that produces about 30% more tokens for the same text. Haiku 4.5 was released before that change, so the same prompt is likely to count as fewer tokens on Haiku than on the other three. Per-token price comparisons therefore understate Haiku's advantage slightly.

Opus 5.5 was released on September 22, 2026. Older models such as Opus 5, Fable 5 and Sonnet 4.6 are still available as legacy models. If you are comparing "Sonnet 5 vs Opus 5," note that Opus 5.5 is cheaper than Opus 5 ($4 / $20 against $5 / $25); the Opus 5.5 pricing and migration guide covers the switch. Anthropic's Opus 5.5 announcement also says Sonnet 5.5 and Haiku 5.5 will follow "in the coming weeks." Haiku 4.5 has no deprecation notice as of September 24, 2026, and its retirement commitment is "not sooner than October 15, 2026." That is a floor, not a scheduled shutdown date.

Which Claude model to start with

Anthropic gives two different starting points, and both are correct for their audience. A subscriber spends a shared usage limit, so the lighter model that still does the job is the better default. An API developer pays per token and can test quality against their own evaluation set, so starting strong and then cutting cost is a reasonable plan.

In the Claude app: Sonnet 5 first

Anthropic's Claude Academy tutorial ranks the models by how much of your rate limit they use: Haiku is the lightest, Sonnet moderate, Opus heavy, and Fable the heaviest. It calls Sonnet "your versatile default" and says, "If you're not sure which model to pick, start here." Its own examples map everyday jobs to models:

TaskModel the tutorial suggests
Summarizing articles, quick answers, simple extractionHaiku
Debugging code, writing, multi-step analysisSonnet
Deep research you steer as you go; problems Sonnet struggled withOpus
Long, accuracy-critical tasks; building a working project from a rough idea; problems Opus struggled withFable

The tutorial's Opus section still names Opus 5. The Opus in the model picker today is Opus 5.5.

Anthropic publishes only the order of limit use, not a multiplier. There is no official figure for how many Sonnet messages equal one Opus or Fable message. Usage depends on conversation length, the model, the features you use and the effort level.

In Claude Code: the default is Opus 5.5, the everyday choice is Sonnet

Since Claude Code v2.1.280, the default model setting resolves to Opus 5.5 on Pro, Max, Team, Enterprise and the Anthropic API, according to the Claude Code model configuration docs. Before that version, Pro users started on Sonnet 5.

That default is not advice to keep Opus on for everything. Anthropic's Claude Code help page on models and limits says "Sonnet is the right choice for the large majority of coding work" and that "Opus costs several times more per turn than Sonnet, and Sonnet more than Haiku." It suggests Opus when you are genuinely stuck or the change is wide, and Haiku for renames, log lines and boilerplate.

The /model aliases on the Anthropic API:

AliasResolves to
sonnetSonnet 5
opusOpus 5.5 (Claude Code v2.1.280 or later)
haikuHaiku 4.5
fableFable 5.1
bestFable where your plan allows it, otherwise Opus
opusplanOpus in plan mode, then Sonnet for execution

Switching models mid-session keeps the conversation. Fable is never the account default, so you have to pick it. If it does not appear in /model, see Fable 5.1 Not Available in Claude Code? Diagnose It by Symptom.

On the API: Opus 5.5 for hard work, Haiku for volume

Anthropic's model selection guide says, "If you're unsure which model to use, start with Claude Opus 5.5 for most workloads." It describes two ways to begin:

  • Efficiency-first: build with Haiku 4.5 and upgrade only where you find a capability gap. This suits prototyping, tight latency budgets, cost-sensitive products and high-volume, straightforward tasks.
  • Capability-first: build with Opus 5.5, then cut cost by lowering effort or moving parts of the work to a cheaper model. If your evals at xhigh or max effort still fall short on demanding reasoning or long-horizon agent work, move to Fable 5.1.

Sonnet 5 sits between them. Anthropic lists it for everyday coding, agent and enterprise workloads such as code generation, data analysis, content creation and tool use.

You also do not have to pick one model per product. Anthropic describes two common mixes: an executor model that escalates hard decisions to a stronger advisor, and an orchestrator that hands bulk work to cheaper workers. Claude Code's opusplan is a built-in example of the first pattern.

Which plan gives you which model

PlanPrice (US)Models you can pick
Free$0Haiku, Sonnet
Pro$20/month, or $17/month billed annuallyHaiku, Sonnet, Opus; Fable only through usage credits
Max 5x / Max 20xFrom $100/monthAll four; Fable can use up to 50% of your weekly limits

Source: claude.com/pricing and Anthropic's help article on Fable models by plan, as of September 24, 2026. Prices exclude tax.

Every plan has a rolling five-hour session limit, and paid plans add a weekly limit. Claude on the web, desktop, mobile and Claude Code all draw from the same pool, so heavy Opus use in Claude Code leaves less for chat.

The Fable rules trip people up:

  • Pro: Fable 5.1 is not part of your plan's limits. It runs on usage credits, which are billed at standard API rates, from the first request.
  • Max: Fable 5.1 can use up to half of your weekly limit at no extra cost. That half comes out of your regular weekly limit and is used up faster than with other models. It is not extra capacity.

For the details, see Claude Fable 5.1 Usage Limits: Weekly Allowance, Extra Usage, and Model Switching and Claude Pro vs. Max Usage Limits. If you are deciding which plan to buy, Claude Cowork Pricing: What You Actually Pay on Each Plan compares the plans.

One more behavior can look like a model switch you did not ask for. For a narrow set of biology and cybersecurity requests, Opus 5.5 and Fable conversations can fall back to an older Opus model. It is a safety measure and does not mean your plan lost access.

When to move up a model, and when to move down

Effort is the setting that controls how much Claude thinks and how thorough it is. On the API it has five levels: low, medium, high, xhigh and max. The Claude app model picker has its own effort control. Higher effort usually helps on hard problems and uses more tokens, which means a bigger bill or more of your limit. Effort affects every output token, including thinking and tool calls. Sonnet 5, Opus 5.5 and Fable 5.1 all support effort up to xhigh and max (effort docs); Haiku 4.5 does not support effort.

Anthropic's guidance puts it plainly: "Tuning effort is often a better lever than switching models." A practical sequence:

  1. Define a check before you compare. Use a test suite that must pass, fields that must be extracted correctly, or a task you already know the answer to. Without a check, you are comparing impressions.
  2. Move up when the current setup keeps failing that check. Raise effort on the same model first. On Opus 5.5 the default is medium, so high and xhigh are the obvious next steps. Switch to the next model only if the failures persist.
  3. Move down when a cheaper model passes the same check. The Academy tutorial suggests running the same known task on Haiku and then Sonnet in new chats and comparing where the answers differ, not how long they are. On the API, run the same eval set at a lower effort or on the cheaper model.
  4. Split the work if only part of it is hard. Plan with Opus and execute with Sonnet, or route simple subtasks to Haiku.

The Opus 5.5 effort guide shows how to set each level in Claude Code and the API and how to measure the cost.

Opus 5.5 or Fable 5.1 for the hardest work

This is the escalation question with the least obvious answer. Fable 5.1 costs 2.5 times as much as Opus 5.5 per input and output token. Anthropic's Opus 5.5 announcement claims that Opus 5.5 "performs at the level of Claude Fable 5.1 on most work." It adds that in Anthropic's own use, the gap between the two is narrower than benchmark scores suggest. In one internal test, both models translated HAProxy from C to Rust and passed nearly all of its regression tests. Opus 5.5 took 9.5 hours against 12 for Fable 5.1, at 51% less cost. These are Anthropic's claims and a single internal test, not an independent measurement.

Fable 5.1 still has a clear role in Anthropic's docs. It covers agent sessions that run for hours and research carried through to a finished document, spreadsheet or deck. It is also the next step when Opus 5.5 at xhigh or max still falls short. The way to find out is a matched trial on your own hard tasks:

Five-step matched trial of Claude Opus 5.5 and Fable 5.1: pick representative hard tasks, keep inputs, tools, permissions, time limits and acceptance checks equal, run Opus 5.5 at medium effort and retry unresolved cases at higher effort, send the remaining cases to Fable 5.1 with the same setup, then compare accepted outcomes and total cost

Judge the trial by cost per accepted task, not cost per call:

cost per accepted task = (token costs + tool charges + retries + review and repair time) ÷ tasks that passed the check

Count every attempt, including timeouts and unusable answers. Fable 5.1 earns a place in your routing only if it passes checks that Opus 5.5 misses, or if it needs so few retries that it costs less per accepted result despite its higher rates. For Fable's full specs, see the Fable 5.1 pricing and migration guide.

What the same work costs on the API

Per-token prices only tell you part of the story, because your token mix decides how far apart the models land. The two examples below use the standard rates from the table above and assume each model uses exactly the same number of tokens. They are illustrations, not measurements.

Example A: a typical request. 100,000 fresh input tokens + 20,000 output tokens + 200,000 cache-read tokens, with no cache writes or tool charges. Tokens are counted in millions, so 100,000 tokens is 0.1:

  • Haiku 4.5: 0.1 × $1 + 0.02 × $5 + 0.2 × $0.10 = $0.22
  • Sonnet 5: 0.1 × $2 + 0.02 × $10 + 0.2 × $0.20 = $0.44
  • Opus 5.5: 0.1 × $4 + 0.02 × $20 + 0.2 × $0.20 = $0.84
  • Fable 5.1: 0.1 × $10 + 0.02 × $50 + 0.2 × $0.25 = $2.05

Example B: a cache-heavy agent loop. 20 turns that each re-read the same 100,000-token cached prefix (2,000,000 cache-read tokens in total, with every request under Haiku's 200K window), plus 50,000 new input tokens and 10,000 output tokens:

  • Haiku 4.5: 2 × $0.10 + 0.05 × $1 + 0.01 × $5 = $0.30
  • Sonnet 5: 2 × $0.20 + 0.05 × $2 + 0.01 × $10 = $0.60
  • Opus 5.5: 2 × $0.20 + 0.05 × $4 + 0.01 × $20 = $0.80
  • Fable 5.1: 2 × $0.25 + 0.05 × $10 + 0.01 × $50 = $1.50
Bar chart of two equal-token API cost examples at standard rates: a typical request costs $0.22 on Haiku 4.5, $0.44 on Sonnet 5, $0.84 on Opus 5.5 and $2.05 on Fable 5.1, while a cache-heavy 20-turn loop costs $0.30, $0.60, $0.80 and $1.50

What the two examples show:

  • Haiku 4.5 costs half as much as Sonnet 5 in both. The tokenizer difference would widen that gap slightly for the same text.
  • Opus 5.5 costs about 1.9 times as much as Sonnet 5 in Example A, but only 1.33 times as much in Example B. Both models read cache at $0.20, so the more of your work is cache reads, the closer Opus 5.5 gets to Sonnet 5.
  • Fable 5.1 costs about 2.4 times as much as Opus 5.5 in Example A and about 1.9 times in Example B. At equal token counts it is the most expensive model in every category.

Example B leaves out the first cache write. Writing 100,000 tokens to the five-minute cache adds $0.125 on Haiku, $0.25 on Sonnet, $0.50 on Opus 5.5 and $1.25 on Fable 5.1.

Real tasks do not use equal tokens. Higher effort produces more thinking tokens. A stronger model may finish in fewer turns, and output lengths differ from model to model. The examples tell you where the rate card puts each model. Only your own workload tells you the cost of a finished task.

Other pricing options change the numbers but not the order. The Batch API halves input and output prices: $0.50 / $2.50 for Haiku 4.5, $1 / $5 for Sonnet 5, $2 / $10 for Opus 5.5 and $5 / $25 for Fable 5.1. Opus 5.5 also has a Fast mode research preview at $8 / $40, which should not be mixed into a standard-rate comparison. Sonnet 5, Opus 5.5 and Fable 5.1 bill their full 1M-token context at standard rates, with no long-context surcharge. If you are also weighing OpenAI models, Claude Opus 5.5 vs GPT-5.6 Sol runs the same kind of comparison across vendors.

Quick answers

What is the difference between Claude Sonnet and Fable?

Sonnet 5 is the fast, mid-priced model for everyday coding, writing and analysis. It is on every plan, including Free, and costs $2 / $10 per million tokens on the API. Fable 5.1 is Anthropic's most capable widely released model, built for long, demanding reasoning and agent work. It is slower, costs $10 / $50 on the API, and is not on Free. On Pro it runs on usage credits, and on Max it can use up to half of your weekly limit.

Which Claude model is the best?

Fable 5.1 is Anthropic's highest-capability model, but it is rarely the best first choice. Anthropic says Opus 5.5 performs at Fable's level on most work at a lower price. For everyday tasks, Sonnet 5 usually gives the best balance of quality, speed and limit use. The best model is the cheapest one that passes your check.

When should you use Fable instead of Opus?

Use Fable 5.1 when Opus 5.5 at a high effort level (xhigh or max) still fails a task that matters. Typical cases are agent sessions that run for hours, or research that has to end in a finished document. Confirm the choice with a matched trial on your own tasks rather than benchmark scores.

Is Haiku cheaper and faster than Sonnet?

Yes. Haiku 4.5 costs $1 / $5 per million tokens against $2 / $10 for Sonnet 5, and Anthropic lists it as the fastest model in the lineup. Anthropic does not publish a speed multiplier. Haiku's limits are a 200K-token context window, 64K max output and no effort control.