OpenAI and Cerebras Systems have launched Ultrafast, a waitlist-only API tier running GPT-5.6 Sol at up to 750 tokens per second—a 14x speedup over standard processing. Powered by Cerebras' Wafer-Scale Engine 3, the service targets latency-critical enterprise workflows like voice AI and financial research. The launch is the first major milestone of a $10 billion infrastructure partnership, though pricing remains undisclosed and performance benchmarks are currently vendor-stated.
Ultrafast API tier launch
- ▪The Ultrafast API tier runs the existing GPT-5.6 Sol model on Cerebras hardware to deliver up to 750 output tokens per second, keeping the model's intelligence, context window, and output quality unchanged.
- ▪OpenAI and Cerebras Systems unveiled Ultrafast on August 13, 2026, a new API service tier that runs OpenAI's GPT-5.6 Sol model at up to 14 times faster than its Standard processing tier.
Cerebras infrastructure partnership
- ▪Cerebras reported a GAAP net loss of $450.5 million for Q2 2026 on August 12, 2026, despite its cloud and inference business growing 281% year over year to $126 million in GAAP revenue.
- ▪The Cerebras Wafer-Scale Engine 3 spans 46,225 square millimeters, integrating 900,000 AI-optimized cores and 44 gigabytes of on-chip static RAM to bypass the memory-bandwidth bottleneck of conventional GPUs.
- ▪The infrastructure deal between OpenAI and Cerebras is valued at over $10 billion and included a $1 billion loan from OpenAI to Cerebras secured by stock warrants.
- ▪OpenAI and Cerebras formalized a sweeping infrastructure agreement in January 2026 under which Cerebras committed to deliver 750 megawatts of compute capacity through 2028.
Speed performance benchmarks
- ▪Cerebras reported that the Ultrafast deployment delivers output speeds 5 times faster than Anthropic's Claude Opus 4.8 in Fast mode and 11 times faster than Anthropic's Claude Fable 5, based on Artificial Analysis data.
- ▪On the Humanity's Last Exam benchmark, Cerebras reported that Ultrafast completed the 2,500-question set in 11 hours and 11 minutes, compared to more than three days for Claude Fable 5.
- ▪On the GDP-Val benchmark measuring performance on knowledge-work tasks, Cerebras reported a 5.6x end-to-end speedup with no quality degradation for Ultrafast in tests conducted on July 31, 2026.
Waitlist access model
- ▪The Ultrafast tier is currently in a limited preview and is waitlist-only in the OpenAI API, with no confirmed general availability date announced.
- ▪OpenAI opened a signup form on its website, and Cerebras opened a parallel signup channel, to gather waitlist registrations for expanding Ultrafast access as capacity grows throughout 2026.
Undisclosed pricing structure
- ▪OpenAI has not disclosed the pricing structure or premium rate for the Ultrafast tier, while its Standard rate remains $5 per million input tokens and $30 per million output tokens.
- ▪OpenAI already monetizes inference speed in tiers, offering a Fast Mode that promises up to 2.5x speed for GPT-5.6 Sol at roughly double the price of Standard mode.
Enterprise use cases
- ▪OpenAI's internal teams are testing Ultrafast for incident response, including analyzing logs, traces, and synthesizing data during active outages, as well as compressing research workflows.
- ▪OpenAI launched the Ultrafast preview with an initial group of four companies: Jane Street, Podium, Basis, and Rogo, spanning coding, financial research, voice AI, and e-commerce.
Story comments
Loading comments…