Quick Facts
- OpenAI’s Ultrafast mode generates up to 750 output tokens per second, compared to roughly 53 tokens per second in Standard processing.
- The tier runs on Cerebras’ Wafer-Scale Engine hardware and is currently in limited preview via invitation only through the OpenAI API.
- OpenAI has not published pricing or a general availability date for Ultrafast.
OpenAI on Aug. 13 announced Ultrafast, a new service tier that runs its GPT-5.6 Sol model up to 14 times faster than standard processing. The tier generates up to 750 output tokens per second and is launching first in the OpenAI API.
Ultrafast is not a new model. It runs the existing GPT-5.6 Sol, OpenAI’s flagship reasoning and agentic model, which reached general availability on July 9, 2026. The speed gain comes entirely from the underlying hardware.
How It Works
Ultrafast runs on Cerebras’ Wafer-Scale Engine, which keeps model weights on-chip in 44 GB of SRAM per wafer-sized chip. Standard GPU-based inference must move weights between on-chip memory and off-chip storage, creating a memory-bandwidth bottleneck. Cerebras eliminates that bottleneck entirely.
OpenAI already offers a Fast mode, announced in late July 2026, that runs GPT-5.6 Sol at up to 2.5 times Standard speed at twice the Standard price using GPU infrastructure. Ultrafast operates at a different scale, reaching speeds OpenAI says are 5 times faster than Claude Opus 4.8 in Fast mode and 11 times faster than Claude Fable 5, based on output speeds reported by Artificial Analysis.
The Cerebras Deal Behind It
Ultrafast is the most visible output of a sweeping infrastructure agreement OpenAI signed with Cerebras in January 2026. The original contract, valued at more than $10 billion, secured 750 megawatts of computing power through 2028. By Cerebras’ first earnings report as a public company, that figure had grown to more than $20 billion.
Cerebras completed its IPO on May 14, 2026, pricing shares at $185. The company reported first-quarter 2026 core revenue between $191 million and $193 million, a 92% year-over-year jump driven largely by the OpenAI deal. As of Cerebras’ Q2 2026 earnings release on Aug. 12, the company had 600 megawatts of data center capacity under contract.
The deal includes unusual terms. OpenAI received warrants for roughly 10% of Cerebras and provided about $1 billion in working capital to help Cerebras build the data centers OpenAI will then pay to use.
Target Workflows
OpenAI is already using Ultrafast internally for incident response. Engineers feed the system logs, code changes, and diagnostic reports while an outage is in progress, using the model to identify causes and help prepare fixes in real time. OpenAI also applies it to research workflows involving knowledge searches, data queries, and information synthesis.
The company points to customer service, financial market analysis, and e-commerce as priority deployment targets for external customers. Jane Street is among the early preview customers. John Crepezzi, who works on AI assistants at Jane Street, said the speed gain changes how developers work alongside the models and makes more focused workflows practical.
Access and Pricing
Ultrafast is currently available by invitation only. OpenAI has opened a signup form for businesses that want to be notified when access expands. No price has been published and no model ID string has been assigned.
Sachin Katti, VP of Compute Strategy and GPT-Infra at OpenAI, said the company is starting with a small group of customers to learn where speed creates meaningful value before expanding the service.
Read more: OpenAI introduces ‘Ultrafast,’ a new mode that makes GPT-5.6 Sol work at 14x the speed
This article was written by an AI agent. Spotted an error? Send a correction and we will fix it.
