Claude Opus 5.5 now has the highest score on the Artificial Analysis Intelligence Index: 58, against 53 for GPT-6 Astra, the flagship of OpenAI's new GPT-6 family.12 Opus 5.5 is also 60% cheaper per token than Astra. Neither fact settles Claude Opus 5.5 vs GPT-6. In that same independent test suite, Astra completes the average task for USD 3.26 and Opus 5.5 for USD 5.98. The reason is that Opus 5.5 at maximum effort writes about four times as many output tokens.12

Anthropic released Opus 5.5 on 22 September 2026.3 OpenAI released GPT-6 Astra earlier in September, and Artificial Analysis published its benchmarks on 9 September. OpenAI added the cheaper GPT-6 Sol and GPT-6 Luna on 22 September.42 This comparison is part of our coverage of language models. It draws on four kinds of evidence: the vendors' launch posts, their API documentation and price lists, independent measurements from Artificial Analysis, and the places where those sources contradict each other.

Versions covered: claude-opus-5-5, gpt-6-astra and gpt-6-sol. Prices and specifications were checked on 26 September 2026.

The short answer

At a glance

Claude Opus 5.5GPT-6 AstraGPT-6 Sol
Released22 September 2026September 202622 September 2026
Best suited toLong-running agentic coding and knowledge workThe hardest end-to-end work, computer use, scienceCost-sensitive coding and agent workflows
Main strength (evidence)Highest Artificial Analysis Intelligence Index score (58); leads GDPval-AA v2.1 at 1,846 Elo1About 27k output tokens per Artificial Analysis task, a third of Claude Fable 5.1's for the same score; 64.6% on Terminal-Bench-Science 0.125Artificial Analysis Coding Agent Index of 57 at USD 2.99 per task6
Main limitation (evidence)About 119k output tokens per task at maximum effort1Highest list price, and higher rates above 272K tokens; in the API, tasks stop if misalignment monitoring intervenes75About 100 Elo lower than GPT-5.6 Sol on GDPval-AA v2.1 in Artificial Analysis testing6
Price (USD per million tokens, input / output)4 / 2010 / 502 / 10
Context window / max output1M / 128K tokens1.05M / 128K tokens1.05M / 128K tokens
Knowledge cutoff (as stated by vendor)June 202630 April 202620 April 2026
Data residency and privacyZero data retention; US-only inference at 1.1x; regional endpoints on Bedrock and Google Cloud at +10%8Zero data retention for eligible customers; EU data residency at +10%, Standard processing only7As Astra

How we compared

Why the benchmark tables disagree

Other comparisons skip this part. Anthropic's and OpenAI's launch posts both include tables with each other's models, and on several rows they report different numbers for the same model on the same benchmark.

BenchmarkModelAnthropic reportsOpenAI reportsArtificial Analysis measured
Terminal-Bench 4.0Claude Opus 5.566.4% (xhigh effort)–59.6%
Terminal-Bench 4.0GPT-6 Astra57.9% (citing OpenAI)57.9%59%
FrontierCode 1.1 MainClaude Opus 548.0%53.4%–
FrontierCode 1.1 MainClaude Fable 5.150.3%50.9%–
AutomationBenchGPT-5.6 Sol28.8% (Zapier leaderboard)18.1%–
Humanity's Last Exam (with tools)Claude Fable 5.165.6%65.0%–
OSWorld 2.0Claude Opus 5.581.8% (partial score)––
OSWorld 2.0, offline setGPT-6 Astra–72.6% (partial score)–

Sources: Anthropic,3 OpenAI,5 Artificial Analysis.12 A dash means the source did not report the figure.

None of this points to bad faith. Both vendors disclose the causes in their footnotes. Effort settings differ: Anthropic's Terminal-Bench figure for Opus 5.5 is at xhigh, and it says the figures represent "each model's highest score". Harnesses and trial counts differ too. Safeguards change results: Anthropic ran Opus 5.5 with production safeguards on, and when they intervened, cybersecurity tasks were completed by Opus 4.8, which it says "likely reduces" Opus 5.5's scores.3 OpenAI notes that its OSWorld figures for Claude use the official settings, not the modified tasks and grading in Anthropic's Fable 5.1 system card.5 So the two OSWorld rows above measure different things and cannot be compared.

Anthropic's own reported standard error on Terminal-Bench 4.0 is plus or minus 2.6 points for Opus 5.5.3 On independent measurement, the two models are level on that benchmark.

Coding

Claude Opus 5.5

Anthropic reports 54.4% on FrontierCode 1.1 Main, which grades whether changes are ready to merge, and 57.8% on CursorBench 4.0.3 Its main claim is about efficiency. At its default effort, Anthropic says Opus 5.5 beats GPT-6 Astra on FrontierCode at roughly 20% of the cost per task, and matches Astra on Terminal-Bench 4.0 at about 40% of the cost. It also reports that an early tester audited and fixed a 200,000-line codebase in under three hours, where Opus 5 took more than 20.3 These are vendor claims, and the tester is unnamed.

GPT-6 Astra

OpenAI reports 57.9% on Terminal-Bench 4.0, 74.1% on DeepSWE v1.1 and 53.3% on FrontierCode 1.1 Main.5 Artificial Analysis measured Astra at 62 on its Coding Agent Index, running in OpenAI's Codex harness. That ties Claude Fable 5.1 in Claude Code, at USD 7.09 per task, about 40% cheaper than Fable 5.1 at maximum effort.2 For long sessions, OpenAI has added an experimental Codex feature that keeps notes across context windows instead of repeatedly compacting them into one summary.5

Knowledge work and business workflows

Claude Opus 5.5

This is Opus 5.5's clearest lead. Artificial Analysis measured 1,846 Elo on GDPval-AA v2.1, a test of real work tasks across 44 occupations. That is 111 points above Claude Fable 5.1 and 138 above Opus 5. It also measured 1,822 Elo on AA-Briefcase v1.1, a private evaluation of multi-week projects with thousands of source files, 143 points above Fable 5.1.1 Anthropic reports that in an internal research test, 16 of 18 Opus 5.5 reports cleared a bar where any invented figure or quote counted as a failure. Neither Fable 5.1 nor Opus 5 cleared it in any attempt.3

GPT-6 Astra

Anthropic's table puts Astra at 1,542 Elo on GDPval-AA v2.1.3 Artificial Analysis observed that Astra dropped about 45 Elo points on this benchmark relative to GPT-5.6 Sol. It also found Astra used far fewer turns per task: 24, against 60 for Claude Fable 5.1 and Claude Opus 5.2 On Zapier's AutomationBench, the two flagships are close: Astra 41.4% and Opus 5.5 40.0%. Zapier ran Opus 5.5 without fallback models, so every safeguard intervention counted as a failure.3

Computer use, science and maths

GPT-6 Astra

OpenAI reports 59.3% on Agents' Last Exam, 92.7% on ScreenSpot-Pro without tools and 72.6% on the OSWorld 2.0 offline set. In science, it reports 96.0% on GPQA Diamond, 97.6% on FrontierMath Tier 4 and 64.6% on Terminal-Bench-Science 0.1.5 OpenAI also says Astra helped establish a new bound on gaps between prime numbers, and it has published the proofs.5

Claude Opus 5.5

Anthropic reports 81.8% on OSWorld 2.0 (partial score, on its own settings), 58.7% on Terminal-Bench-Science 0.1 and 89.0% on the Chartography chart-reading test with tools.3 We found no GPQA Diamond or FrontierMath figure for Opus 5.5 in Anthropic's launch post.

Pricing

USD per million tokensClaude Opus 5.5GPT-6 AstraGPT-6 SolNotes
Input4.0010.002.00
Cached input (read)0.201.000.20Opus 5.5 discounts cache reads by 95%; OpenAI by 90%
Cache write5.00 (5-minute) or 8.00 (1-hour)12.502.50
Output20.0050.0010.00
Input above 272K tokens4.0020.004.00Anthropic charges one rate across the full 1M window
Output above 272K tokens20.0075.0015.00
Batch (input / output)2 / 105 / 251 / 5
Fast mode (input / output)8 / 4020 / 1004 / 20Opus 5.5: up to 2.5x speed, Claude API only. Astra: up to 2x speed.

Sources: Anthropic pricing documentation and Opus 5.5 announcement;83 OpenAI pricing documentation.7 Checked 26 September 2026. Standard, global processing.

Price per token is not cost per task

Artificial Analysis publishes what it costs to run each model on its Intelligence Index, including reasoning tokens. At maximum effort, Opus 5.5 costs USD 5.98 per task, Astra USD 3.26 and Sol USD 1.06.126 Opus 5.5 is cheaper per token than Astra but more expensive per task at maximum effort. It uses about 119k output tokens per task against Astra's 27k.1

Maximum effort is only one setting. Artificial Analysis found that four of Opus 5.5's five effort levels sit on its cost-performance frontier, costing less than, or outperforming, every other model that scores above 50.1 Anthropic says that at its default medium effort, Opus 5.5 beats Astra at maximum effort on GDPval-AA for about a fifth of the cost per task.3 Choosing the effort level has more effect on cost than choosing between these two models.

Two worked examples

An agent turn. 180K tokens of cached context, 20K new input tokens and 8K output tokens, a typical shape for a coding agent partway through a task:

  • Claude Opus 5.5: USD 0.036 + 0.080 + 0.160 = USD 0.28
  • GPT-6 Astra: USD 0.180 + 0.200 + 0.400 = USD 0.78
  • GPT-6 Sol: USD 0.036 + 0.040 + 0.080 = USD 0.16

A long document. 500K uncached input tokens and 4K output tokens, for example a contract bundle or a large log file:

  • Claude Opus 5.5: USD 2.00 + 0.08 = USD 2.08
  • GPT-6 Astra (long-context rates): USD 10.00 + 0.30 = USD 10.30
  • GPT-6 Sol (long-context rates): USD 2.00 + 0.06 = USD 2.06

Safeguards and refusals in production

Both models ship with safeguards that can interrupt legitimate work, and they fail in different ways.

Claude Opus 5.5 uses cybersecurity, biology and anti-distillation safeguards similar to Claude Fable 5.1's. Most cybersecurity tasks are routed to Opus 4.8, although routine bug-finding and fixing stay on Opus 5.5.3 A declined request returns HTTP 200 with stop_reason: "refusal" and a stop_details object naming the policy area. Anthropic recommends configuring a fallback model, either server-side or in your own code.9

GPT-6 Astra is the first model OpenAI has designated at the Critical cybersecurity threshold under its Preparedness Framework.10 It refuses advanced tasks such as writing proof-of-concept exploits. OpenAI plans wider defensive access through its Daybreak programme.5 Astra runs with production misalignment monitoring. In ChatGPT and Codex, a paused task may ask the user to review the action. In the API, the task stops.5 Enterprise administrators must also enable Astra for their workspace, because access was off by default at launch.5

API changes that break existing code

Claude Opus 5.5 has four breaking changes relative to Opus 5:9

  1. Thinking cannot be disabled. Sending thinking: {"type": "disabled"} or a manual token budget returns a 400 error. Depth is controlled with the effort parameter instead.
  2. Forced tool use is not supported. A tool_choice of any or a named tool returns a 400 error. Use auto with strict tool use, or structured outputs. If your agents reach tools through the Model Context Protocol, check how your client sets tool_choice.
  3. Thinking blocks are tied to the model and the conversation. Opus 5.5 reads thinking blocks from Opus 5 and earlier, but not from Fable or Mythos models. For newer accounts, editing anything before a thinking block and replaying it returns an error.
  4. The computer_20251124 tool is rejected on the Claude API and Google Cloud. Use the computer_toolset_20260801 toolset instead.

Two behaviour changes need no code but will change your costs. The default effort is now medium, where Opus 5 defaulted to high. The model also tends to think more per turn at any given effort level. Anthropic recommends re-running your effort calibration instead of carrying settings over.9

GPT-6 changes less at the API level. Astra supports reasoning effort from low to max with no none option, while Sol and Luna also offer none.11 Two caching changes are worth using. Reasoning effort can now be changed between responses without breaking the prompt cache, and OpenAI gives cache discounts for shared prefixes reused within a 30-minute window.12

When to choose Claude Opus 5.5

  • Multi-hour agent runs over large codebases, where cache reads dominate the bill
  • Research, analysis and document production, where the independent evidence is strongest
  • Requests that regularly exceed 272K tokens
  • Teams willing to tune effort per task, since cost per task depends on it

Whichever model you choose, put the provider behind a thin internal interface so that switching is a configuration change. Three different models have held or shared the top of the Artificial Analysis Intelligence Index since 1 September.13 If you are designing that layer into a product, it is part of Revere Group's SaaS development work.

When to choose GPT-6 Astra, or GPT-6 Sol

  • Astra for computer and browser automation, scientific and mathematical work, and short reasoning-heavy tasks where its low output token count keeps cost per task down
  • Astra if your organisation is standardised on Azure OpenAI or already runs Codex
  • Sol for cost-sensitive coding agents. Artificial Analysis found it improved on GPT-5.6 Sol's Coding Agent Index score at half the cost per task, with a much lower hallucination rate.6

Alternatives worth considering

  • Claude Fable 5.1 (USD 10 / 50). Anthropic recommends it when Opus 5.5 at higher effort still falls short.14
  • Claude Sonnet 5 (USD 2 / 10). The same price as GPT-6 Sol, and the natural Anthropic comparison for it. Anthropic says Claude Sonnet 5.5 and Haiku 5.5 will follow Opus 5.5 "in the coming weeks".3
  • GPT-6 Luna (USD 0.10 / 0.50). For high-volume classification, extraction and routing.

Limitations and open questions

  • No in-house testing. This comparison relies on vendor and independent benchmarks. Your prompts, tools and data may reverse any of the conclusions above.
  • Headline figures are at maximum effort. Most production traffic runs lower, where the cost and quality trade-offs differ.
  • Speed is unmeasured. Artificial Analysis had not published an output speed for Opus 5.5 when we checked. It measured Astra at about 61 tokens per second.15 Anthropic says Opus 5.5 generates output more than 30% faster than Opus 5.3
  • Both families are still rolling out. Astra was reaching ChatGPT plans and cloud platforms in stages. OpenAI's promotional price for GPT-5.6 Sol runs at least until 21 November 2026, which affects the older comparison point.7

Frequently asked questions

Is Claude Opus 5.5 better than GPT-6?
It depends on the work. On the independent Artificial Analysis Intelligence Index, Claude Opus 5.5 scores 58 and GPT-6 Astra 53. Opus 5.5 leads clearly on long-horizon knowledge work such as GDPval-AA. Astra leads on the science benchmark both vendors report, Terminal-Bench-Science 0.1. On agentic coding the two are level in Artificial Analysis testing of Terminal-Bench 4.0, at 59.6% and 59%.
How much does Claude Opus 5.5 cost compared with GPT-6?
Claude Opus 5.5 costs USD 4 per million input tokens and USD 20 per million output tokens, with cache reads at USD 0.20. GPT-6 Astra costs USD 10 and USD 50, with cache reads at USD 1.00, rising to USD 20 and USD 75 for long-context requests above 272K tokens. GPT-6 Sol costs USD 2 and USD 10. Prices were checked on 26 September 2026.
What is the difference between GPT-6 Astra, Sol and Luna?
Astra is OpenAI's flagship for the hardest end-to-end work. Sol is built for complex coding and agentic workflows at a fifth of Astra's token price. Luna is the low-cost model for focused, high-volume tasks. All three have a 1.05 million token context window and 128K maximum output.
What changed in the Claude Opus 5.5 API?
Four breaking changes: thinking can no longer be disabled, forced tool use returns an error, thinking blocks are tied to the model and conversation that produced them, and the older computer_20251124 tool is not accepted on the Claude API and Google Cloud. The default effort level also dropped from high to medium.
Which has the larger context window, Opus 5.5 or GPT-6?
The GPT-6 models accept 1.05 million tokens and Claude Opus 5.5 accepts 1 million. The pricing differs more than the size: Anthropic charges the standard rate across the whole window, while OpenAI applies higher long-context rates to requests above 272K tokens.
Can Opus 5.5 and GPT-6 be used with EU data residency?
OpenAI offers EU data residency for the GPT-6 models at a 10% price uplift, with Standard processing only. Anthropic's own API documents a US-only inference option at 1.1 times the price. For in-region processing, Claude is available through regional endpoints on Amazon Bedrock and Google Cloud at a 10% premium. Both vendors offer zero data retention to eligible customers.

Changelog

  • 26 September 2026: First published. Prices and specifications checked on this date.

Sources

  1. Artificial Analysis, "Claude Opus 5.5 takes the top spot on the Artificial Analysis Intelligence Index", 22 September 2026. https://artificialanalysis.ai/articles/claude-opus-5-5 ↩ ↩2 ↩3 ↩4 ↩5 ↩6 ↩7 ↩8 ↩9

  2. Artificial Analysis, "Benchmarking GPT-6 Astra", 9 September 2026. https://artificialanalysis.ai/articles/benchmarking-gpt-6-astra ↩ ↩2 ↩3 ↩4 ↩5 ↩6 ↩7 ↩8

  3. Anthropic, "Claude Opus 5.5", 22 September 2026. https://www.anthropic.com/claude-opus-5-5 ↩ ↩2 ↩3 ↩4 ↩5 ↩6 ↩7 ↩8 ↩9 ↩10 ↩11 ↩12 ↩13 ↩14 ↩15 ↩16

  4. OpenAI, "Introducing GPT-6 Sol and Luna", 22 September 2026. https://openai.com/index/introducing-gpt-6-sol-and-luna/ ↩

  5. OpenAI, "GPT-6 Astra: A new generation of intelligence", September 2026, updated 22 September 2026. https://openai.com/index/gpt-6-astra/ ↩ ↩2 ↩3 ↩4 ↩5 ↩6 ↩7 ↩8 ↩9 ↩10 ↩11

  6. Artificial Analysis, "GPT-6 Sol and Luna push the cost efficiency frontier", 22 September 2026. https://artificialanalysis.ai/articles/gpt-6-sol-and-luna-push-the-cost-efficiency-frontier ↩ ↩2 ↩3 ↩4

  7. OpenAI, "Pricing", OpenAI API documentation, accessed 26 September 2026. https://developers.openai.com/api/docs/pricing ↩ ↩2 ↩3 ↩4

  8. Anthropic, "Pricing", Claude Platform documentation, accessed 26 September 2026. https://platform.claude.com/docs/en/about-claude/pricing ↩ ↩2 ↩3

  9. Anthropic, "What's new in Claude Opus 5.5", Claude Platform documentation, accessed 26 September 2026. https://platform.claude.com/docs/en/models/opus-5-5/whats-new-opus-5-5 ↩ ↩2 ↩3

  10. OpenAI, "Path to Astra: critical capabilities and frontier safeguards", 1 September 2026. https://openai.com/index/path-to-astra/ ↩

  11. OpenAI, "Models", OpenAI API documentation, accessed 26 September 2026. https://developers.openai.com/api/docs/models ↩

  12. OpenAI, "Better prompt caching for GPT-6", 22 September 2026. https://openai.com/index/better-prompt-caching-for-gpt-6/ ↩

  13. Artificial Analysis, articles index, accessed 26 September 2026: "Claude Fable 5.1 tops the Artificial Analysis Intelligence Index" (1 September), "Benchmarking GPT-6 Astra" (9 September) and "Claude Opus 5.5 takes the top spot" (22 September). https://artificialanalysis.ai/articles ↩

  14. Anthropic, "Models overview", Claude Platform documentation, accessed 26 September 2026. https://platform.claude.com/docs/en/models/overview ↩

  15. Artificial Analysis, "GPT-6 Astra (max): Intelligence, Performance and Price Analysis", accessed 26 September 2026. https://artificialanalysis.ai/models/gpt-6-astra ↩