750 Tokens a Second, in Limited Preview

OpenAI and Cerebras announced Ultrafast mode for GPT-5.6 Sol on August 13, 2026, generating up to 750 output tokens per second, up to 14 times faster than the model's standard processing, according to OpenAI's own announcement of the preview. The tier is live now only for a select group of customers while OpenAI evaluates how the added speed changes real-world products; other businesses can join a waitlist by submitting their workload details. OpenAI researcher Jeffrey Wang described the practical effect in the company's own materials: 'It now finishes for me before I even have the opportunity to context-switch. Makes me way more productive.'

OpenAI points to voice interfaces, customer support, commerce, developer agents, financial research, and security incident response as the intended use cases, and says its own developers have used the preview to analyze logs and traces during incidents and to compress research cycles that previously ran overnight into multiple iterations during a single workday. GPT-5.6 Sol itself launched in June 2026 alongside the balanced Terra and speed-focused Luna variants, and became broadly available across ChatGPT, Codex, and the API in July.

The Chip Doing the Work Is Not Nvidia's

The speed gain comes from Cerebras' Wafer-Scale Engine, which keeps an entire model's weights resident on-chip across 44 gigabytes of SRAM built into each wafer-sized processor, according to Cerebras' own blog post on the partnership. That removes the repeated transfers between memory and compute that create latency on conventional GPU clusters, letting tokens flow uninterrupted through pipelined model layers spread across multiple wafers - a design built to scale as models grow larger, per Cerebras' own technical description.

Cerebras has spent years positioning this architecture as a faster, if less flexible, alternative to Nvidia's GPU clusters for inference workloads, but its highest-profile public partnerships until now have leaned toward smaller or mid-tier labs rather than the industry's single largest buyer of Nvidia capacity. OpenAI's own infrastructure commitments to Nvidia run into the hundreds of billions of dollars, a relationship this outlet has covered in detail, which makes its choice to route even a limited preview tier through a competing chip architecture a notable data point rather than routine vendor diversification.

What the Speed Actually Buys

On GDP-Val, a benchmark built around economically valuable work tasks, OpenAI reports a 5.6 times end-to-end speedup with no measured quality loss. On Humanity's Last Exam, a 2,500-question PhD-level test, GPT-5.6 Sol Ultrafast completed the full set in 11 hours 11 minutes against 78 hours 27 minutes for Claude Fable 5, according to Cerebras' own published comparison - roughly seven times faster on that specific benchmark, alongside Cerebras' separate throughput claims of 11 times over Claude Fable 5 and 5 times over Claude Opus 4.8's own fast mode.

None of this makes GPT-5.6 Sol a smarter model. It is the same model, answering the same way, arriving faster - which matters specifically for workloads where latency itself is the bottleneck rather than reasoning quality: a support agent a customer will not wait on, a security response that needs to happen inside an active incident, or a financial model that has to run several scenarios before a market closes.

A Hedge Against Nvidia, Even for OpenAI

The practical lesson here is not that Cerebras beat Nvidia on a benchmark. It is that the company most financially entangled with Nvidia's AI buildout still had a commercial reason to run its most latency-critical product tier on a different chip architecture, because the physics of keeping weights on-chip beats the physics of moving them across a GPU cluster for this specific job. Any European business building a long-term AI infrastructure strategy around a single compute vendor, however dominant that vendor looks today, should treat this as evidence that even the biggest customers hedge when a specific workload demands it.

This is a limited preview with no published pricing, not a general product decision enterprises need to act on immediately. The useful move is not to switch vendors on the strength of one announcement, but to note it as a data point: track which of your own workloads are genuinely latency-bound rather than reasoning-bound, because those are the ones where a second infrastructure option, not just a second cloud contract, may eventually be worth the added complexity.