Quick Facts

  • Claude Opus 5.5 scores 66.4% on Terminal-Bench 4.0, beating GPT-6 Astra at 57.9% and Fable 5.1 at 55.8%
  • API pricing drops to $4 per million input tokens and $20 per million output tokens, with cache-read costs falling 60% to $0.20 per million tokens
  • Early testers report Opus 5.5 completed a 680,000-line code migration in under a day and a 200,000-line codebase audit in under three hours

Anthropic released Claude Opus 5.5 on September 22, 2026, targeting enterprise teams running agentic coding, computer use, and knowledge work workflows. The model is the first in the Claude 5.5 family, with Sonnet 5.5 and Haiku 5.5 to follow.

Opus 5.5 beats Fable 5.1 across all nine benchmarks Anthropic tested, at roughly 40% of the cost. Fable 5.1 and Mythos 5.1 remain available as separate frontier options.

Benchmark Results

On Terminal-Bench 4.0, which measures autonomous multi-step engineering tasks in a real command-line environment, Opus 5.5 scores 66.4% at extra-high effort. That beats OpenAI's GPT-6 Astra at 57.9%, Fable 5.1 at 55.8%, Opus 5 at 52.3%, and GPT-5.6 Sol at 37.3%.

On FrontierCode v1.1, which evaluates whether an agent's pull requests would be accepted into production codebases, Opus 5.5 scores 54.6% at default effort. That edges GPT-6 Astra's top score of 53.3% at about one-fifth the per-task cost.

On GDPval-AA v2.1, a knowledge work benchmark spanning 44 occupations, Opus 5.5 scores 1,846 against Fable 5.1's 1,735. On Humanity's Last Exam, it reaches 61.4%, up from Fable 5.1's previous best of 59.1%.

Pricing Structure

Anthropic prices Opus 5.5 at $4 per million input tokens and $20 per million output tokens, down 20% from Opus 5. Cache-write costs fall from $6.25 to $5 per million tokens. Cache-read costs drop from $0.50 to $0.20 per million tokens, a 60% reduction.

Because cache reads account for the majority of costs in agentic and coding workflows, and because the model uses fewer tokens per task, Anthropic says typical workloads run about 40% cheaper overall. Output generation is also more than 30% faster than Opus 5.

Anthropic says Opus 5.5 at default effort outperforms Opus 5 at maximum effort for roughly one-fifth the cost, and matches GPT-6 Astra's performance at about 40% of the price.

Early Customer Results

Anthropic says Opus 5.5 completed a 680,000-line code migration in under a day during early testing. It finished an internal HAProxy conversion from C to Rust in 9.5 hours, compared with 12 hours for Fable 5.1. A 200,000-line codebase audit took less than three hours. Opus 5 needed more than 20 hours for the same task and consumed 2.5 times as many tokens.

Deepak Singh, VP of Agentic AI at Kiro, said Opus 5.5 solved more tasks than Opus 5 while making 40% fewer calls and consuming half the tokens.

Carl Bennett, CIO at Deloitte Consulting, said Opus 5.5 caught 72% of known bugs in code reviews at its lowest effort setting, compared with 56% for Opus 5 at high effort, with fewer false positives and less output volume.

LexisNexis Legal and Professional said early evaluations showed the model consistently identified relevant legal citations and structured answers around central legal frameworks, capabilities the company looks for in tools built into its Lexis+ platform.

What This Means for Buyers

For software teams running long agentic pipelines, the cache-read price drop is the most significant change. Those costs accumulate quickly in multi-step workflows, and cutting them by 60% directly lowers the per-task bill without requiring any code changes.

The benchmark results suggest enterprises that adopted Fable 5.1 or GPT-6 Astra for coding agents have a clear reason to retest with Opus 5.5. The combination of higher task completion rates, lower token consumption per task, and reduced API pricing makes a cost and performance case that procurement teams can quantify.

Read more: Anthropic releases Claude Opus 5.5, beating Fable 5.1 on key agentic benchmarks at 60% cheaper API price

This article was written by an AI agent. Spotted an error? Send a correction and we will fix it.