Quick Facts

  • GPT-5.5 scores 82.7% on Terminal-Bench 2.0, beating Claude Mythos Preview’s 82% and Claude Opus 4.7’s 69.4%
  • API pricing doubles to $5/$30 per million tokens, with Pro version at $30/$180 per million tokens
  • Model achieves state-of-the-art performance across 14 benchmarks versus 4 for Claude Opus 4.7

OpenAI released GPT-5.5 on April 23, marking the company’s first fully retrained base model since GPT-4.5. The model narrowly defeated Anthropic’s Claude Mythos Preview on Terminal-Bench 2.0, scoring 82.7% accuracy compared to Claude’s 82%.

GPT-5.5 significantly outperformed other models on the benchmark. Claude Opus 4.7 scored 69.4% while Google Gemini 3.1 Pro reached 68.5%. Terminal-Bench 2.0 evaluates 89 tasks in computer terminal environments designed to test multi-step reasoning and tool use.

The model dominates across multiple benchmarks. GPT-5.5 achieves state-of-the-art performance on 14 benchmarks compared to 4 for Claude Opus 4.7 and 2 for Google Gemini 3.1 Pro. On FrontierMath Tier 4, GPT-5.5 scored 35.4% versus Opus 4.7’s 22.9%.

“What is really special about this model is how much more it can do with less guidance,” said Greg Brockman, OpenAI co-founder and president. He described GPT-5.5 as “a new class of intelligence” and “a big step towards more agentic and intuitive computing.”

OpenAI doubled API pricing with the release. Standard GPT-5.5 costs $5 per million input tokens and $30 per million output tokens. The Pro version commands $30 per million input tokens and $180 per million output tokens.

Despite higher costs, the company argues the model delivers better efficiency. According to Artificial Analysis, GPT-5.5 medium matches Claude Opus 4.7 max performance at one quarter of the cost.

The release comes one week after Anthropic launched Opus 4.7, intensifying competition in agentic AI. Claude Mythos Preview remains limited to approved companies through Project Glasswing due to safety concerns around its cybersecurity capabilities.

GPT-5.5 shows a concerning trade-off in accuracy versus reliability. While achieving 57% accuracy on the AA-Omniscience benchmark, the model has an 86% hallucination rate compared to Opus 4.7’s 36%.

The model is available to Plus, Pro, Business, and Enterprise users in ChatGPT and Codex. API deployments require additional safety measures as OpenAI works with partners on security requirements for scale deployment.

Read more: OpenAI’s GPT-5.5 is here, and it’s no potato: narrowly beats Anthropic’s Claude Mythos Preview on Terminal-Bench 2.0

This article was written by an AI agent. Spotted an error? Send a correction and we will fix it.