OpenAI and Cerebras Launch 'Ultrafast' Mode for GPT-5.6 Sol
A new API tier leveraging Wafer-Scale Engine architecture pushes output speeds to 750 tokens per second.
OpenAI and Cerebras have introduced "Ultrafast mode," a high-performance service tier for the OpenAI API designed to drastically reduce inference latency. The new tier allows the GPT-5.6 Sol model to reach output speeds of up to 750 tokens per second without sacrificing model quality.
The performance leap is powered by Cerebras' Wafer-Scale Engine (WSE) architecture. Unlike traditional GPU clusters, which typically stream frontier models at 40 to 120 tokens per second, the WSE utilizes 44 GB of on-chip SRAM. This design eliminates the memory bandwidth bottlenecks common in GPU-based systems by keeping model weights on-chip. Currently, the service is available in a limited preview for a select group of customers.
The Performance Gap
The speed advantages of the Ultrafast tier are evident in both direct comparisons and complex benchmarks. According to Cerebras, GPT-5.6 Sol on Ultrafast mode runs 11 times faster than Fable 5 and five times faster than Opus 4.8 when running in Fast mode.
These gains translate to significant time savings for large-scale tasks. In a benchmark consisting of 2,500 PhD-level questions known as "Humanity's Last Exam," GPT-5.6 Sol Ultrafast completed the set in 11 hours and 11 minutes. In contrast, Claude Fable 5 required 78 hours and 27 minutes to finish the same task. This model is part of a broader three-tier GPT-5.6 family that includes Sol, Terra, and Luna.
Implications for Real-Time AI
For years, AI developers have navigated a strict tradeoff between intelligence and speed, as larger, more capable models typically introduce higher latency. By removing this bottleneck, OpenAI and Cerebras are enabling the transition toward real-time agentic workflows.
This shift is critical for time-sensitive industrial applications. High-speed inference allows AI agents to operate on the "critical path" for urgent operations, such as root-causing production outages in real-time or responding to high-stakes cyberattacks. In these scenarios, the minutes typically spent waiting for a model response are unacceptable. Jeffrey Wang, an OpenAI researcher, noted that tasks that previously took several minutes now finish before a user has the opportunity to context-switch.
The Path Forward
As the industry moves toward more autonomous agents, the ability for AI to keep pace with human thought and collaboration becomes a primary competitive advantage. Rohan Varma, a product lead at OpenAI, stated that the Ultrafast mode enables AI that can keep up with how users think, code, and collaborate.
While the current rollout is limited to a preview group, the success of the WSE integration suggests a shift in how frontier models are deployed. The industry will now be watching to see if this architecture can be scaled to the larger Terra and Luna models within the GPT-5.6 family, and how other providers respond to the new benchmark for inference speed.