Anthropic launched Claude Opus 5.5 this Tuesday (22), the first model in the new Claude 5.5 family. The company says the system achieves performance close to Claude Fable 5.1 in most tasks, but costs 40% less to operate than Opus 5 on typical workloads, in addition to generating responses more than 30% faster.
The launch concentrates the biggest advances in autonomous coding, computer use, and knowledge work. The model also reduces API prices to US$ 4 per million input tokens and US$ 20 per million output tokens, compared with US$ 5 and US$ 25 for Opus 5. Cache reads fell from US$ 0.50 to US$ 0.20 per million tokens.
The change puts Opus 5.5 in an unusual position within Anthropic's own lineup. Fable 5.1 remains priced at US$ 10 per million input tokens and US$ 50 for output, while the new Opus costs 60% less on these two metrics and beats Fable in several of the tests published by the company. Anthropic cautions, however, that the difference perceived in real use is smaller than the benchmarks suggest.
Opus 5.5 leads seven of the nine benchmarks released
In the comparison table published by Anthropic, Opus 5.5 obtains the best result in seven of the nine tests presented. The gains are clearest in agent-based coding, professional work, and computer use.

Results released by Anthropic. Some tests use different tools, effort settings, and fallback mechanisms across models.
The result does not represent absolute leadership. GPT-6 Astra remains ahead on AutomationBench, with 41.4% against 40%, and opens an advantage on Terminal-Bench-Science, with 64.6% against 58.7%. Anthropic itself says that small differences between frontier models are becoming less reliable as an indicator of behavior in real applications.
There are also important methodological differences. Most of Opus 5.5's results were obtained at the maximum effort level, while Terminal-Bench 4.0 used xhigh. Production safeguards were also on during the tests: when triggered on sensitive tasks, some cybersecurity requests were routed to Opus 4.8, and biology and advanced AI development tasks to Opus 5.
In AutomationBench, run by Zapier, there were no fallback models. Safeguard interventions were recorded as failures, something that, according to Anthropic, reduces Opus 5.5's observed result.
Coding is the main leap of the new model
Long-duration coding occupies the center of the launch. In Terminal-Bench 4.0, aimed at professional tasks performed in the terminal, Opus 5.5 reached 66.4%, compared with 52.3% for Opus 5 and 57.9% for GPT-6 Astra. On FrontierCode v1.1, it scored 54.4%, slightly above Astra's 53.3%.
When the cost per task is included, the company says the model at the standard effort level beats GPT-6 Astra on FrontierCode while spending approximately 20% of the cost per task. On Terminal-Bench, it achieves performance similar to Astra for about 40% of the cost.

Anthropic also presented tests with larger workloads. A user in early access reportedly completed the migration of a codebase with 680,000 lines of code in less than a day. Another test involved auditing and fixing 200,000 lines: Opus 5.5 finished in less than three hours, while Opus 5 took more than 20 hours and consumed 2.5 times more tokens. These cases were provided by the company and by participants selected for early access, and do not constitute independent benchmarks.
In another internal experiment, Opus 5.5 and Fable 5.1 were given the task of rewriting HAProxy from C to Rust. Both versions passed almost all of the project's regression tests, but Opus 5.5 finished in 9.5 hours, compared with 12 hours for Fable 5.1, at 51% lower cost.
Professional work also receives focus
The model also advanced in tasks involving research, documents, finance, and business operations. In GDPval-AA v2.1, which measures work close to that performed in 44 occupations, Opus 5.5 reached 1,846 Elo points, compared with 1,735 for Fable 5.1, 1,708 for Opus 5, and 1,542 for GPT-6 Astra.

In an internal research evaluation, the models had to find hard-to-locate financial information and produce a report without inventing numbers or citations. 16 of 18 reports produced by Opus 5.5 passed the criterion defined by Anthropic. Fable 5.1 and Opus 5 did not pass in any attempt on this specific test.
Another experiment asked the models to analyze a fictional merger, build a financial model in Excel, and prepare an executive presentation. Opus 5.5 completed the work in 63 minutes, compared with 93 minutes for Opus 5, and cost 50% less. Both reached the same conclusion about the transaction.
Lower price comes with a speed gain
The 40% reduction announced by Anthropic does not come only from the price list. Input and output became 20% cheaper, but the company says Opus 5.5 also needs fewer tokens to finish tasks, reducing the final cost of typical workloads.
Cache reads, particularly important in agents that reuse large volumes of context, fell 60%, to US$ 0.20 per million tokens. Output generation speed increased more than 30% compared with Opus 5.
A separate Fast Mode is available in Claude Code and Claude Platform and promises speeds of up to 2.5 times the conventional configuration. The gain comes at a cost: the mode charges US$ 8 per million input tokens and US$ 40 per million output tokens.
1 million-token window and changes for developers
Claude Opus 5.5 maintains a 1 million-token context window and accepts up to 128,000 output tokens in conventional requests. In the Message Batches API, a beta configuration expands the maximum output to 300,000 tokens. The standard effort level becomes medium, and the knowledge cutoff indicated by Anthropic is June 2026.
For developers, however, the update is not entirely transparent. Anthropic documents four incompatible changes with certain code created for Opus 5. Adaptive reasoning now remains always active and cannot be turned off; calls with forced tool use can generate an error; reasoning blocks are associated with the model and the conversation that produced them; and the old computer_20251124 tool is no longer accepted in the Claude API and Google Cloud.
The API identifier is claude-opus-5-5. The model is already available on Anthropic platforms and through cloud providers, including Amazon Web Services, Google Cloud and Microsoft Azure.
Anthropic expands Opus 5.5 safeguards
Opus 5.5 is also the first Opus model released with a class of protections similar to that used in Fable 5.1 for cybersecurity, biology, and attempts to extract the model's capabilities. Anthropic says the system's capabilities in these areas are already comparable to those of Claude Mythos 5.1.
The company submitted the model to external evaluations before launch, including tests conducted by METR and Frontier Design. In its own behavioral audit, composed of nearly 2,000 scenarios, Anthropic says Opus 5.5 obtained better results than recent Claude models on most metrics of misaligned behavior.
In a new containment evaluation, the model tried to exceed the imposed limits approximately 85% fewer times than Opus 5 or Mythos 5.1. The company itself cautions that its evaluation systems still cannot reliably detect all possible problematic behaviors before deployment.
Anthropic also says Opus 5.5 improved resistance to prompt injection attacks. In an evaluation by security company Gray Swan cited by the company, the new model tied with Fable 5.1 for the lowest success rate of these attacks among the systems tested.
In addition to the technical changes, Anthropic says it has revamped the model's writing behavior to reduce jargon, put important information earlier, and follow style instructions more consistently, a response to criticism received about Opus 5.
The company will also expand usage limits in five-hour windows for subscribers Pro, Max, Team, and Enterprise per seat. Opus 5.5 opens the new generation, while Claude Sonnet 5.5 and Claude Haiku 5.5 are expected to be released in the coming weeks.



