You cannot call Gemini 4 Argon yet. Google announced it on September 30, 2026, but the rollout is limited to "a set of trusted cyber defenders" in its Fairwind Program, and the model does not appear in the Gemini API model list or price list. Claude Opus 5.5 and Claude Sonnet 5.5 are available to every Claude API customer. So as of October 1, 2026, the three-way comparison is a two-way decision plus a plan for later:
- Well-scoped work with a clear pass/fail check (code generation, data analysis, tool-using agents on bounded tasks): start on Claude Sonnet 5.5 at $2 input / $10 output per million tokens.
- Long, open-ended work (multihour autonomous coding, large refactors, computer use, vision-heavy pipelines): use Claude Opus 5.5 at $4 / $20.
- Gemini 4 Argon: put it on your re-test list. Its announced introductory price matches Sonnet 5.5 and its later price matches Opus 5.5, and it leads the one independent index that scores all three. None of that helps until Google opens access.
| Gemini 4 Argon | Claude Opus 5.5 | Claude Sonnet 5.5 | |
|---|---|---|---|
| Callable by developers today | No (Fairwind Program only) | Yes | Yes |
| API model ID | Not published | claude-opus-5-5 | claude-sonnet-5-5 |
| Input / output per 1M tokens | $2 / $10 introductory, then $4 / $20 | $4 / $20 | $2 / $10 |
| Cached input per 1M tokens | $0.10 at the introductory price | $0.20 cache read | $0.20 cache read |
| Context window | Not published | 1M tokens | 1M tokens |
| Maximum output | 1M tokens | 128K tokens | 128K tokens |
| Vals Index v2.1 (September 30, 2026) | 68.90% | 66.97% | 67.04% |
Prices are USD list prices on each vendor's own API, before tax.
What you can call today: Opus 5.5 and Sonnet 5.5, not Argon
Google's Gemini 4 Argon announcement says the model is "rolling out to a set of trusted cyber defenders through our Fairwind Program" and will reach "developers, enterprises, and consumers as soon as possible," starting with paid API customers and Google AI Ultra subscribers. It gives no date, no model ID, no context window, and no rate limits. The Gemini API models page and pricing page do not list Argon. The name is Gemini 4 Argon; "Gemini 4 Pro" came from pre-launch leak posts and is not what Google shipped.
If you need a Gemini model in production this week, the newest one on those pages is Gemini 3.8 Flash (gemini-3.8-flash) at $0.75 input / $3.75 output per million tokens through December 31, 2026. It is a different class of model from Argon, so treat it as a separate decision; the Gemini 3.8 Flash API guide covers its model ID, pricing, and limits.
On the Claude side, Opus 5.5 (released September 22, 2026) and Sonnet 5.5 (released September 28, 2026) are both generally available on the Claude API, Amazon Bedrock (anthropic.claude-opus-5-5, anthropic.claude-sonnet-5-5), Google Cloud, Microsoft Foundry, and Claude Platform on AWS, according to Anthropic's models overview. Whether a given cloud region has them enabled depends on your account. Both take text and images in and return text, with a 1M-token context window, a 128K synchronous output limit, and a June 2026 knowledge cutoff.
Argon starts at Sonnet 5.5's price, then moves to Opus 5.5's
Argon's price sheet maps onto Anthropic's almost exactly. Google's announcement sets an introductory price of $2 per million input tokens and $10 per million output tokens, which is Sonnet 5.5's list price. A footnote says $4 and $20 will apply after the introductory period, which is Opus 5.5's list price. Google has not said when the introductory period ends.
| Per 1M tokens (USD) | Gemini 4 Argon, introductory | Gemini 4 Argon, later | Claude Opus 5.5 | Claude Sonnet 5.5 |
|---|---|---|---|---|
| Input | $2 | $4 | $4 | $2 |
| Output | $10 | $20 | $20 | $10 |
| Cached input / cache read | $0.10 (95% off input) | Not stated | $0.20 | $0.20 |
| 5-minute cache write | Not published | Not published | $5 | $2.50 |
| Batch input / output | Not published | Not published | $2 / $10 | $1 / $5 |
The Claude figures come from Anthropic's pricing table. Two details there matter for the comparison. Cache reads cost the same $0.20 on both Claude models, so the 2× gap between Sonnet and Opus applies only to fresh input and output. And the full 1M-token context is billed at the standard rate on both. Opus 5.5 also has a fast mode in research preview at $8 / $40, Claude API only; Sonnet 5.5 has none. Bedrock and Google Cloud set their own prices.
For Argon, batch pricing, tool fees, any context-length tiers, and Vertex AI pricing are unpublished. Nobody outside the Fairwind cohort can pay the introductory price today.
Recompute the cost for your own task
The formula is the same for all three: tokens in millions times the per-million rate, summed across input, cached input, and output.
Example A: 100,000 input tokens and 20,000 output tokens, no caching.
| Model | Calculation | Cost |
|---|---|---|
| Claude Sonnet 5.5 | 0.1 × $2 + 0.02 × $10 | $0.40 |
| Claude Opus 5.5 | 0.1 × $4 + 0.02 × $20 | $0.80 |
| Gemini 4 Argon, introductory | 0.1 × $2 + 0.02 × $10 | $0.40 |
| Gemini 4 Argon, later price | 0.1 × $4 + 0.02 × $20 | $0.80 |
Example B: one turn of a cache-heavy agent, with 200,000 cached input tokens, 20,000 fresh input tokens, and 5,000 output tokens.
| Model | Calculation | Cost |
|---|---|---|
| Claude Sonnet 5.5 | 0.2 × $0.20 + 0.02 × $2 + 0.005 × $10 | $0.13 |
| Claude Opus 5.5 | 0.2 × $0.20 + 0.02 × $4 + 0.005 × $20 | $0.22 |
| Gemini 4 Argon, introductory | 0.2 × $0.10 + 0.02 × $2 + 0.005 × $10 | $0.11 |
In Example B, Sonnet 5.5 costs 41% less than Opus 5.5, not 50%, because the cached portion is priced identically. The more of your prompt is cached, the smaller the saving from dropping to Sonnet. Argon's introductory row comes out two cents under Sonnet purely from the cheaper cached input.

Both examples assume all three models consume and produce the same number of tokens, which they will not. Tokenizers differ, and effort and thinking settings change how many tokens a model spends on the same task. Opus 5.5 defaults to medium effort and Sonnet 5.5 to high. The examples also leave out cache writes, tools, and tax, and the Argon rows stay hypothetical until the model is callable. To get your real cost per task, take the token counts from your own API responses and plug them into the formula. For the billing rules behind thinking tokens, see how reasoning tokens are billed on OpenAI, Claude, and Gemini.
Who measured each benchmark number
Three sets of numbers are in circulation, and each one comes from a different party with a different setup. All of the figures below are published by a vendor or a leaderboard, and each is labeled with who ran it.
Google's table compares Argon with Opus 5.5 only
Google's model page puts Argon next to Opus 5.5 on 18 rows. Sonnet 5.5 is not in the table.
| Benchmark (%) | Gemini 4 Argon | Claude Opus 5.5 |
|---|---|---|
| Vals Index | 68.9 | 67.0 |
| AutomationBench | 51.3 | 42.5 |
| Vals Finance Agent v2 | 65.4 | 58.6 |
| Harvey Legal Agent Benchmark | 19.6 | 3.8 |
| DeepSWE v1.1 | 77.9 | 74.2 |
| Vibe Code Bench | 91.9 | 90.3 |
| LABBench 2 | 88.8 | 73.1 |
| RiemannBench | 76.0 | 69.6 |
| GraphWalks, up to 128K | 99.7 | 90.6 |
| GraphWalks, 256K–1M | 84.2 | 66.8 |
| Agent's Last Exam | 39.5 | 38.2 |
| Chartography | 71.6 | 66.3 |
| LVBench | 91.7 | 83.7 |
| CWE-bench v1 | 68.0 | 67.0 |
| FrontierSWE v2 | 55.0 | 62.3 |
| Terminal-Bench 4.0 | 57.4 | 66.4 |
| PostTrainBench | 45.3 | 49.3 |
| Terminal-Bench Science 0.1 | 57.6 | 63.3 |
Argon is ahead on 14 rows and Opus 5.5 on the last four, which cover software engineering, terminal work, and ML engineering. What the table can support is narrower than the row count suggests. Google chose the benchmarks. Its evaluation methodology says Argon ran at the highest thinking setting, while most Opus 5.5 figures are Anthropic's self-reported numbers or public leaderboard entries at maximum or best-available reasoning. Google computed several rows itself for every model: PostTrainBench, LABBench 2, GraphWalks, and LVBench. On LVBench the models did not even see the same input, with Gemini sampled at one frame per second and Opus 5.5 given 600 frames "due to API limitations."
Read it as a list of where to look, not as a ranking. The largest gaps (legal agent tasks, long-context graph traversal, lab-bench science) are the ones worth reproducing on your own data once you have access. The rows Opus 5.5 wins tell you where Argon is least likely to displace it.
Vals Index is the only scale with all three
The Vals Index is run by Vals AI, not by either vendor. It weights finance, coding, legal, and tax tasks by their share of GDP. Version 2.1, updated September 30, 2026, shows:
| Model | Vals Index score | Cost per test | Time per test |
|---|---|---|---|
| Gemini 4 Argon | 68.90% | $15.68 | 46m 33s |
| Claude Sonnet 5.5 | 67.04% | $21.34 | 1h 18m |
| Claude Opus 5.5 | 66.97% | $32.14 | 1h 19m |
Two things stand out. Sonnet 5.5 and Opus 5.5 are 0.07 points apart, which Vals describes as effectively tied, and Sonnet gets there at about two-thirds of Opus's cost per test. That ratio is itself a lesson in cost per task: Sonnet's token price is half of Opus's, yet its cost per test is 66% of Opus's, because the two models do not use the same number of tokens on the same work.
Argon leads both by just under two points, at a lower cost per test and in roughly 60% of the time. The page does not say which Argon price that $15.68 is based on. If it is the introductory price and token use stays the same, doubling it gives about $31.36 at the later price, level with Opus 5.5.
Leaderboard rows also move with reruns, and a 0.07-point gap between Sonnet 5.5 and Opus 5.5 can flip on the next one, so quote these scores with their date and treat a gap under two points as unsettled.

Anthropic's numbers compare Sonnet 5.5 with Opus 5.5
Anthropic's Sonnet 5.5 announcement is the only place the two Claude models are measured side by side by their maker. Higher is better on every row.
| Benchmark | Claude Sonnet 5.5 | Claude Opus 5.5 |
|---|---|---|
| Terminal-Bench 4.0 | 70.6% | 66.4% |
| FrontierCode 1.1 | 46.2% | 54.4% |
| CursorBench 4.0 | 55.5% | 57.8% |
| GDPval-AA v2.1 | 1844 | 1846 |
| AA-Briefcase v1.1 | 1811 | 1822 |
| Humanity's Last Exam (with tools) | 64.5% | 67.7% |
| OSWorld 2.1 | 80.1% | 81.8% |
| Chartography (no tools) | 61.6% | 64.4% |
Sonnet 5.5 leads on Terminal-Bench 4.0, where Opus's 66.4% is its best result (at xhigh effort). Opus 5.5 leads everywhere else, by about two to three percentage points on most rows and by eight on FrontierCode 1.1. That FrontierCode gap narrows when Sonnet is run at xhigh (52.1%) instead of max (46.2%). Anthropic's own summary matches the table: Sonnet 5.5 is for "well-scoped everyday tasks," while on complex, open-ended work Opus 5.5 "remains clearly stronger."
Two cautions on reading across tables. Anthropic lists Opus 5.5's Chartography at 64.4% and Google's table lists it at 66.3%. Those are different runs under different settings, so compare numbers only within one table. And Anthropic's claims that Sonnet 5.5 "generates outputs 30%+ faster" and "costs up to 30% less per task" are measured against Sonnet 5, not against Opus 5.5 or Gemini.
There is no vendor-published head-to-head between Argon and Sonnet 5.5. The Vals Index is the only place the two share a scale.
Sonnet 5.5 vs Opus 5.5: which Claude model to use now
The evidence above points to a simple rule: choose by how open-ended the task is, then confirm with your own token counts.
| Your workload | Start with | Why |
|---|---|---|
| Code generation, data analysis, content, tool-using agents on bounded tasks | Claude Sonnet 5.5 | Tied with Opus 5.5 on the Vals Index at about two-thirds of the cost per test; half the token price; "Fast" latency label versus "Moderate" |
| Multihour autonomous coding agents, large-scale refactoring, complex systems engineering | Claude Opus 5.5 | Anthropic says it "remains clearly stronger" on complex, open-ended work; eight-point lead on FrontierCode 1.1 |
| Computer use and vision-heavy workflows | Claude Opus 5.5 | Anthropic's model selection guide assigns these to Opus; small leads on OSWorld 2.1 and Chartography |
| Latency-sensitive tool loops | Claude Sonnet 5.5 | It can skip up-front thinking with thinking: {"type": "between_tools"} at low, medium, or high effort; Opus 5.5's adaptive thinking is always on |
| Cache-heavy agents where most input is a cached prefix | Either; measure first | Cache reads cost $0.20 on both, so the saving from Sonnet shrinks (41% in Example B) |
Anthropic's selection guide says "most workloads start with Claude Opus 5.5," and that is a reasonable default if quality on hard tasks matters more to you than the bill. The cost-first path runs the other way. Start on Sonnet 5.5, keep a set of your own tasks with a pass/fail check, and move a workload to Opus 5.5 only when Sonnet fails that check. Before switching in either direction, try changing effort. Anthropic notes that tuning effort "is often a better lever than switching models," and the guide to choosing and setting effort on Opus 5.5 shows how to measure what each level costs.
If you are moving from an older Claude model, both 5.5 models reject forced tool use (tool_choice set to any or tool returns a 400) and neither accepts thinking: {"type": "disabled"}. The full list is in the Claude Opus 5.5 migration checks, and Sonnet 5.5's differences are in Anthropic's what's new page. For where Haiku and Fable fit, see which Claude model to use.
When Argon opens, what is worth re-testing
Waiting for Argon is not a reason to delay a Claude decision, because there is no date to wait for. It is worth a re-test when access opens if one of these describes you:
- You run Sonnet 5.5 and price is the constraint. At the introductory price Argon costs the same per token, with cached input at half of Claude's cache-read rate. Run your pass/fail set on both and compare actual token counts, not rate cards.
- You run Opus 5.5 for long-context, legal, finance, or science work. Those are the rows with Google's largest claimed leads. Check them on your own documents.
- You need very long outputs. Argon's announced output limit is 1M tokens, against 128K on both Claude models.
It is less likely to pay off if your workload looks like the rows Opus 5.5 wins in Google's own table: terminal-driven agents and long software engineering tasks.
Before you move anything, confirm four things that are unpublished as of October 1, 2026: the model ID, the context window, the rate limits on your tier, and the date the introductory price ends. A workload that is cheaper at $2 / $10 may not be at $4 / $20.
FAQ
Can I use Gemini 4 Argon through the API right now?
No. As of October 1, 2026, access is limited to cyber defenders in Google's Fairwind Program. Google says paid API customers and Google AI Ultra subscribers will be first when it widens, with no date given. The newest Gemini model on the public API pages is Gemini 3.8 Flash.
Is Gemini 4 Argon cheaper than Claude Opus 5.5?
At the announced introductory price of $2 / $10 per million tokens it is half of Opus 5.5's $4 / $20 and equal to Sonnet 5.5. After the introductory period Google's footnote sets $4 / $20, the same as Opus 5.5. The end date is not published, and equal token prices do not guarantee equal cost per task.
Is Sonnet 5.5 better than Opus 5.5?
On one benchmark in Anthropic's table, Terminal-Bench 4.0, yes (70.6% versus 66.4%). On the Vals Index the two are effectively tied. On Anthropic's other seven published benchmarks Opus 5.5 is ahead, mostly by small margins. Sonnet 5.5 is the better value on well-scoped tasks; Opus 5.5 is the stronger model on complex, open-ended ones.
Does Gemini 4 Argon beat Claude Opus 5.5?
In Google's own table it leads on 14 of 18 benchmarks and trails on FrontierSWE v2, Terminal-Bench 4.0, PostTrainBench, and Terminal-Bench Science 0.1. On the independent Vals Index it leads by about two points. The Google table is vendor-selected and mixes self-computed and self-reported results, so it shows where Argon may be stronger, not by how much on your workload.
What is Gemini 4 Argon's context window?
Google has not published it. The announcement states a 1M-token output limit, and Google's long-context evaluation uses inputs up to 1M tokens, but neither is a stated context window.



