As of October 1, 2026, GPT-6.1 Sol costs $2 per 1M input tokens and $10 per 1M output tokens, exactly half of Claude Opus 5.5's $4 and $20. Independent benchmarks put the real gap much wider: Artificial Analysis measures Opus 5.5 at 6.4× Sol's cost per task at medium effort and 8.3× at max effort. The same benchmarks score Opus 5.5 higher, by 3 index points at medium and 6 at max.
That gives a workable starting position:
- Try GPT-6.1 Sol first for high-volume coding, agent, and document work where requests stay at or under 272,000 input tokens and a failed attempt is cheap to catch and retry.
- Keep Opus 5.5 for the tasks where each failed attempt costs a person real review time, for max-effort work on hard problems, and for prompts that routinely run past 272K input tokens, where Sol's price advantage shrinks to a few percent.
- Decide with your own pass rates, not the leaderboard. The rule further down turns two pass rates and a review cost into an answer.
The figures below come from vendor documentation and third-party benchmarks, and each one carries its source and settings.
Which model to try first, by workload
| Your workload | Start with | Why | What would change it |
|---|---|---|---|
| High-volume agent or coding tasks, checked by tests | GPT-6.1 Sol | 6.4× lower measured cost per task at medium effort | Sol's pass rate on your tasks falls far below Opus 5.5's |
| Hard tasks run at max effort, reviewed by a senior engineer | Opus 5.5 | Higher scores at max effort in both independent indexes and in OpenAI's own charts | Review is cheap, or the score gap doesn't show up on your tasks |
| Prompts over 272K input tokens | Either | At equal tokens Sol is only 1–7% cheaper there | Output length: Sol wrote about a third as many tokens per task in Artificial Analysis's runs |
| Cache-heavy loops under 272K | GPT-6.1 Sol | Cache reads at $0.10 vs. $0.20 per 1M | Your cache hit rate differs between the two APIs |
| Security-adjacent work | Test both | Opus 5.5 can refuse or hand off to another model | You don't enable fallback, or refusals are rare on your prompts |
| Already on GPT-6 Sol | GPT-6.1 Sol | Same token price, cheaper cache reads | You call it with reasoning effort none |
What "GPT 6.1" refers to
"GPT 6.1" is one model: GPT-6.1 Sol, API ID gpt-6.1-sol. OpenAI introduced it at DevDay at the end of September 2026 as an upgrade to GPT-6 Sol. There is no GPT-6.1 Astra or GPT-6.1 Luna, and OpenAI says GPT-6 Astra remains its most capable model. Claude Opus 5.5 (claude-opus-5-5) was released on September 22, 2026.
| GPT-6.1 Sol | Claude Opus 5.5 | |
|---|---|---|
| API model ID | gpt-6.1-sol | claude-opus-5-5 |
| Context window | 1,050,000 tokens (922,000 max input) | 1M tokens |
| Max output | 128,000 tokens | 128K tokens (300K on Message Batches with a beta header) |
| Knowledge cutoff | April 30, 2026 | June 2026 (reliable cutoff) |
| Reasoning control | Reasoning effort low, medium (default), high, xhigh, max | Effort, default medium; adaptive thinking is always on |
| Input / output per 1M tokens | $2 / $10 | $4 / $20 |
| Where you can use it | OpenAI API; ChatGPT Work and Codex on Plus, Pro, Business, Enterprise, and Edu | Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry; Claude Pro and Max |
Sources: OpenAI's model page and Anthropic's Opus 5.5 overview, as of October 1, 2026.
One availability detail matters if you mostly chat: OpenAI's announcement says GPT-6.1 Sol is not yet offered in Chat, only in ChatGPT Work and Codex.
Price: half per token, until a request passes 272K input tokens
All prices are USD per 1 million tokens on each vendor's direct API at the standard tier, before tax and regional surcharges.
| Line item | GPT-6.1 Sol | Opus 5.5 |
|---|---|---|
| Input | $2 | $4 |
| Output | $10 | $20 |
| Cache read | $0.10 | $0.20 |
| Cache write | $2.50 | $5 (5-minute) / $8 (1-hour) |
| Batch input / output | $1 / $5 | $2 / $10 |
| Fast mode input / output | $4 / $20 | $8 / $40 (research preview, Claude API only) |
| Requests over 272K input tokens | $4 input, $0.20 cache read, $5 cache write, $15 output | No change |
Sources: OpenAI API pricing and Claude API pricing. Both vendors add about 10% for region-pinned processing. Bedrock and Google Cloud set their own Opus 5.5 prices.
OpenAI's pricing page carries a "promotional pricing" note with a November 21, 2026 date. That note names GPT-5.6 Sol, not GPT-6.1 Sol. OpenAI's announcement also mentions a GPT-6.1 Sol Ultrafast variant, but the pricing page has no row for it.
The same request on both models
Take a request with 100,000 input tokens and 10,000 output tokens, no caching, no tools:
- GPT-6.1 Sol: 0.1 × $2 + 0.01 × $10 = $0.30
- Opus 5.5: 0.1 × $4 + 0.01 × $20 = $0.60
A cache-heavy agent turn with 200,000 cached tokens read, 20,000 fresh input tokens, and 5,000 output tokens:
- GPT-6.1 Sol: 0.2 × $0.10 + 0.02 × $2 + 0.005 × $10 = $0.11
- Opus 5.5: 0.2 × $0.20 + 0.02 × $4 + 0.005 × $20 = $0.22
The cache example leaves out the earlier cache write and assumes the same hit rate on both sides. The two vendors bill writes differently, which OpenAI vs Claude prompt caching cost works through.
Both examples assume equal token counts, which real workloads won't give you. The two models write different amounts for the same task, as the next section shows, and they tokenize text differently. Anthropic says Claude 4.7 and later models produce about 30% more tokens for the same text than its previous tokenizer did, and no published figure compares Claude's tokenizer with OpenAI's. A 2:1 price per token is not a 2:1 price per document. The usage fields in each API response are the only reliable count.
What happens at 272K
OpenAI prices any request with more than 272,000 input tokens at 2× the input and cache rates and 1.5× the output rate, for the whole request. Anthropic bills the full 1M context at the standard rate. The examples are uncached.
| Request | GPT-6.1 Sol | Opus 5.5 | Sol's saving |
|---|---|---|---|
| 272,000 in + 20,000 out | 0.272 × $2 + 0.02 × $10 = $0.744 | 0.272 × $4 + 0.02 × $20 = $1.488 | 50% |
| 273,000 in + 20,000 out | 0.273 × $4 + 0.02 × $15 = $1.392 | 0.273 × $4 + 0.02 × $20 = $1.492 | 6.7% |
| 300,000 in + 20,000 out | 0.3 × $4 + 0.02 × $15 = $1.50 | 0.3 × $4 + 0.02 × $20 = $1.60 | 6.3% |
| 900,000 in + 10,000 out | 0.9 × $4 + 0.01 × $15 = $3.75 | 0.9 × $4 + 0.01 × $20 = $3.80 | 1.3% |
Exactly 272,000 input tokens stays at the short rate. One thousand more tokens raises the Sol request from $0.744 to $1.392, an 87% jump, while the Opus request rises by $0.004. Above the line the input rates are identical at $4, and only Sol's $15 output rate keeps it ahead.
Two practical consequences follow. If your prompts hover near 272K on Sol, trimming them under the line is worth more than any other optimization. If they sit well above it, token price stops being a reason to switch, and output length and quality decide.
Why a 2× price gap becomes 6–8× per task
Artificial Analysis runs both models through its Intelligence Index and reports what each run cost. Two matched pairs are published: medium effort and max effort.
| Artificial Analysis, as of October 1, 2026 | GPT-6.1 Sol (Medium) | Opus 5.5 (Medium) | GPT-6.1 Sol (Max) | Opus 5.5 (Max) |
|---|---|---|---|---|
| Intelligence Index | 48 | 51 | 52 | 58 |
| Cost per task | $0.21 | $1.34 | $0.72 | $5.98 |
| Cost to run the index | $361 | $1,627 | $1,082 | $8,708 |
| Output tokens per task | 8k | 26k | 38k | 119k |
| Reasoning tokens per task | 3k | 12k | 25k | 84k |

The Opus 5.5 rows are labeled "Adaptive Reasoning, Default Fallback"; the fallback section explains what that label changes.
Opus 5.5's cost per task is $1.34 ÷ $0.21 ≈ 6.4× Sol's at medium and $5.98 ÷ $0.72 ≈ 8.3× at max. Measured as cost to run the whole index, the ratios are 4.5× and 8.0×. These are two different metrics, so quote one or the other, not a blend.
The extra multiple comes mostly from volume. Opus 5.5 wrote about 3.3× as many output tokens per task at medium (26k vs. 8k) and 3.1× at max (119k vs. 38k), and three to four times as many reasoning tokens. Three times the output at twice the output price is already about six times the output bill.
To estimate your own difference, replace the price ratio with a token ratio you measure:
Bill ratio ≈ 2 × (Opus 5.5 billed tokens per task ÷ Sol billed tokens per task), for requests at or under 272K input.
Run the same 20 to 50 representative tasks on both, sum usage from the responses, and multiply each side by its rates. If Opus 5.5 uses 1.5× the tokens on your tasks, expect roughly a 3× bill; at 3×, roughly 6×. Effort is the main lever on the Opus side, and choosing and measuring an Opus 5.5 effort level covers how to set it.
Which one is smarter depends on effort and on who measured
Independent indexes: Opus 5.5 scores higher and costs far more
On Artificial Analysis, Opus 5.5 leads by 3 points at medium effort (51 vs. 48) and 6 at max (58 vs. 52). Sol at max effort scores 52, one point above Opus 5.5 at medium, for $0.72 per task against $1.34.
The Vals Index v2.1, last updated September 30, 2026, shows the same shape:
| Vals Index v2.1 | Score | Cost per test | Latency per test |
|---|---|---|---|
| Claude Opus 5.5 | 66.97% (±0.9) | $32.14 | 1h 19m |
| GPT-6.1 Sol | 61.15% | $3.24 | 43m 10s |
Opus 5.5 scores 5.8 points higher at about 9.9× the cost per test ($32.14 ÷ $3.24). On the same index Claude Sonnet 5.5 scores 67.04% at $21.34 per test, level with Opus 5.5 within the margin, which is worth knowing if you are choosing inside the Claude family.
Both indexes are one evaluator's task suite in launch week, and the values can be revised.
OpenAI's launch charts: the direction flips with effort
OpenAI's announcement compares GPT-6.1 Sol with Opus 5.5 on three benchmarks. Its text states the cost figures and the 2.2-point AutomationBench gap; the scores below are read from the charts in that post, as transcribed by Kingy AI.
| OpenAI launch charts | Effort | GPT-6.1 Sol | Opus 5.5 | Higher score |
|---|---|---|---|---|
| GDP.pdf | Medium | 30.0% at $0.34 per task | 25.6% at $0.80 | Sol |
| GDP.pdf | High | 32.0% at $0.35 | 28.8% at $0.83 | Sol |
| AutomationBench 1.0.6 | Medium | 31.7% at $0.19 | 29.5% at $0.65 | Sol |
| AutomationBench 1.0.6 | Max | 36.1% at $0.30 | 42.5% at $1.44 | Opus 5.5 |
| Terminal-Bench Science 0.1 | Max | 57.0% at $5.47 | 63.3% at $23.21 | Opus 5.5 |
At medium effort Sol scores higher and costs less. At max effort, in the vendor's own charts, Opus 5.5 scores 6.4 points higher on AutomationBench and 6.3 higher on Terminal-Bench Science, at about 4.8× and 4.2× the cost per task.
Read these with OpenAI's footnote in mind: it ran the GPT evaluations itself and took competitor results from public reports. The benchmarks and settings are the vendor's choice. Anthropic's launch numbers for Opus 5.5 (Terminal-Bench 4.0 at 66.4%, FrontierCode v1.1 at 54.4%) predate GPT-6.1 Sol and have no row for it, so the two vendors' tables share no benchmark version with both models on it.
Early user reports disagree as well. One startup founder posted that Sol beat Opus 5.5 on internal evals and became the team's default, while other widely shared posts call Opus 5.5 clearly better. Neither is evidence about your workload.
The fallback footnote on Opus 5.5 rows
Opus 5.5's safety classifiers can decline a request. The API returns HTTP 200 with stop_reason: "refusal" rather than an error. If the caller has turned on server-side fallback (fallbacks: "default" with the beta header server-side-fallback-2026-07-01), the API retries on the model Anthropic recommends for that refusal category, and the response names the model that answered. Anthropic's documentation shows Claude Opus 4.8 in its example, and its launch post says most cybersecurity tasks route there.
Three things follow for a comparison:
- Scores are partly another model's. "Default Fallback" on Artificial Analysis and the Vals note that Opus 5.5 scores "include some component tasks served by a fallback model after a provider refusal" both mean the row measures Opus 5.5 plus its fallback. OpenAI's GDP.pdf claim is likewise against Opus 5.5 "with fallback."
- Fallback is not automatic. You configure it. Without it, the refusal stands and your code has to handle it.
- Each attempt is billed at the rate of the model that ran it. A refusal that produced no output is billed only for a few categories (
bio,frontier_llm, andreasoning_extractionas of September 2026), but it still counts against rate limits.
How often this triggers depends on the work. Kingy AI relays Vals counts of 44 fallback-served tasks out of 2,281 overall (1.9%), but 22 of 198 on Terminal-Bench 4.0. If your prompts touch security tooling, measure the refusal rate on your own tasks before trusting any published Opus 5.5 score. OpenAI's refusal behavior for GPT-6.1 Sol isn't documented in comparable detail in the launch materials.
Cost per accepted result: a rule to run on your own numbers
Token cost is only part of what a result costs. Someone or something has to check each attempt, and failed attempts get retried. If one attempt costs c in tokens, costs R in review, and passes with probability p, then:
Expected cost per accepted result = (c + R) ÷ p
Opus 5.5 is the cheaper choice only when:
p(Opus) ÷ p(Sol) > (c(Opus) + R) ÷ (c(Sol) + R)
R and both pass rates are your inputs. The examples use Artificial Analysis's medium-effort cost per task ($0.21 and $1.34) as stand-ins for c; substitute your measured cost per task when you have it.
Case 1: automated checks, R = $0. The right side is 1.34 ÷ 0.21 ≈ 6.4. Even an Opus 5.5 that never fails wins only if Sol passes fewer than 1 ÷ 6.38 ≈ 15.7% of attempts. Checking both sides of that line: Sol at 15% costs $0.21 ÷ 0.15 = $1.40 per accepted result, more than a perfect Opus at $1.34. Sol at 20% costs $1.05, less than $1.34. When tests do the reviewing, Sol wins unless it almost never works.
Case 2: human review, R = $5 per attempt (for example, five minutes at $60 an hour). The right side is (1.34 + 5) ÷ (0.21 + 5) = 6.34 ÷ 5.21 ≈ 1.217. If Sol passes 70% of attempts, Opus 5.5 has to pass more than 0.70 × 1.217 ≈ 85.2%.
| Sol passes 70% | Opus 5.5 pass rate | Sol cost per accepted result | Opus 5.5 cost per accepted result | Cheaper |
|---|---|---|---|---|
| Above the threshold | 90% | $5.21 ÷ 0.70 = $7.44 | $6.34 ÷ 0.90 = $7.04 | Opus 5.5 |
| Below the threshold | 80% | $5.21 ÷ 0.70 = $7.44 | $6.34 ÷ 0.80 = $7.93 | Sol |

The pattern is the useful part. The more expensive each review is, the less the token price matters, and the smaller the quality edge Opus 5.5 needs to pay for itself. With $5 of review per attempt, a 15-point pass-rate lead is enough; with free review, nothing realistic is.
The rule assumes you retry until something passes and that reviewing a Sol attempt takes as long as reviewing an Opus attempt. If Opus 5.5's longer output takes longer to read, raise its R.
Speed: Opus 5.5 streams faster, Sol answers sooner
Artificial Analysis measures Opus 5.5 at 72 output tokens per second against Sol's 59 at medium effort, and 92 against 65 at max. Time to the first answer token runs the other way: 5.3 seconds for Sol and 22.2 for Opus 5.5 at medium, and about 282 seconds against 703 at max.
Faster streaming doesn't mean a task finishes sooner when the model writes three times as many tokens. Vals reports end-to-end latency of 43 minutes per test for Sol and 1 hour 19 minutes for Opus 5.5. These numbers move with effort, load, and route.
The $20 plan: no published number settles it
OpenAI's ChatGPT and Codex pricing page lists 15–160 local messages per five-hour window for GPT-6.1 Sol on Plus, against 15–150 for GPT-6 Sol and 5–45 for GPT-6 Astra. It adds that the number "depends on the model used, size and complexity of your tasks," so the range is not a guarantee. In Codex credits, GPT-6.1 Sol costs 50 per 1M input tokens, 2.5 cached, and 250 output.
Anthropic says Opus 5.5 is included in Claude Pro and Max and that five-hour limits were raised with the release, but it publishes no numbers. The two $20 plans can't be compared by a figure, and dividing the monthly fee by API prices doesn't produce one.
Early forum reports contradict each other: one r/codex commenter got about five small tasks per window from Sol at medium, while an r/OpenAI commenter described Sol burning a full window on one bug as Opus 5.5 finished six tasks on half of its own. Those are single sessions on unknown tasks.
What you can decide on:
- If you mainly use the chat interface, GPT-6.1 Sol isn't there yet. It lives in ChatGPT Work and Codex.
- If you already hold one plan, count how many of your usual tasks finish inside a five-hour window for a week before paying for the other.
- If you hit limits daily, the API prices above, not the plan, are the comparison that applies to you.
Moving existing code to GPT-6.1 Sol or Opus 5.5
From GPT-6 Sol to GPT-6.1 Sol
Input and output stay at $2 and $10, the 272K rule is the same, and cache reads drop from $0.20 to $0.10 per 1M. One thing can break: gpt-6.1-sol doesn't support reasoning effort none or minimal, and GPT-6 Sol accepts none. Calls that turned reasoning off need a new setting; low is the lowest available. Those calls will now reason where they didn't before, so recheck their cost and latency.
OpenAI's own comparisons of the two (DeepSWE v1.1, OSWorld 2.0) favor 6.1, but they are vendor benchmarks. Rerun your regression set before switching the default. Pricing across the rest of the family is covered in GPT-6 Luna vs. Sol pricing.
From an OpenAI model to Opus 5.5
Changing the model string isn't enough. Per Anthropic's migration notes:
- Thinking can't be disabled.
thinking: {"type": "disabled"}or a manual budget returns a 400. tool_choiceaccepts onlyautoandnone. Forcinganyor a named tool returns a 400.- Text between tool calls arrives in thinking blocks, which are empty at the default display setting. Agents that log or parse that text need a change.
- At the same effort setting it tends to think more per turn than Opus 5, most of all at
xhighandmax. - Handle
stop_reason: "refusal"explicitly, or enable fallback.
The full list of checks is in Claude Opus 5.5: API price, model ID, and migration checks. If you are comparing against the older, same-priced OpenAI model instead, see Claude Opus 5.5 vs GPT-5.6 Sol.
Quick answers
Is GPT-6.1 Sol better than Opus 5.5?
On the two independent indexes, no: Opus 5.5 scores 3 to 6 points higher on Artificial Analysis and 5.8 points higher on Vals. On OpenAI's launch benchmarks Sol scores higher at medium effort and Opus 5.5 at max effort. Sol is far cheaper per task in every one of those comparisons.
How much cheaper is GPT-6.1 Sol?
Half per token for requests at or under 272K input tokens, and 1–7% at equal token counts above that. Per completed benchmark task, Artificial Analysis measures Opus 5.5 at 6.4× (medium) to 8.3× (max) Sol's cost, and Vals at 9.9× per test, because Opus 5.5 writes about three times as many tokens.
Is there a GPT-6.1 Astra or Luna?
No. GPT-6.1 exists only as Sol. GPT-6 Astra and GPT-6 Luna remain on their original versions.
What is still unknown about GPT-6.1 Sol and Opus 5.5?
No neutral party has published a head-to-head on an identical harness beyond Artificial Analysis and Vals. OpenAI hasn't said whether GPT-6.1 Sol will come to Chat, hasn't listed an Ultrafast price, and neither vendor has said how long current prices hold. Anthropic hasn't published Opus 5.5 allowances for its subscription plans.



