Artificial intelligenceAugust 14, 2026· via The Decoder

OpenAI’s GPT-5.6 Sol hits 750 tokens/sec with Ultrafast mode

OpenAI’s GPT-5.6 Sol hits 750 tokens/sec with Ultrafast mode

Image : The Decoder

OpenAI just turned up the speed dial on GPT-5.6 Sol, launching an “Ultrafast” inference mode that clocks in at up to 750 output tokens per second. The move is powered by Cerebras hardware, the fruit of the companies’ $10 billion partnership, and it slots into a new three-tier pricing structure alongside existing “Standard” and “Fast” modes. For developers and businesses that live or die by latency, the step change is hard to ignore.

Speed as a product

Ultrafast isn’t just a technical footnote—it’s a deliberate product tier. OpenAI is now selling inference speed the way airlines sell seat classes: Standard for everyday throughput, Fast for a moderate bump, and Ultrafast for when every millisecond counts. The pricing isn’t disclosed yet, but the implication is clear: faster models cost more, and customers will pay for the privilege. Cerebras chips, with their wafer-scale silicon, are what make the jump possible by crunching tokens at an order of magnitude beyond conventional GPUs.

Who benefits—and who pays

Ultrafast mode is squarely aimed at latency-sensitive workloads: real-time chat assistants, live translation pipelines, or high-frequency coding copilots where sub-second responses are table stakes. Startups running at scale can carve out a competitive edge by shaving precious milliseconds off each interaction. Yet the three-tier pricing also risks deepening the divide between well-funded teams that can afford Ultrafast and everyone else still on Standard. OpenAI’s decision to tie hardware and pricing together signals a future where raw speed becomes a premium feature rather than a baseline expectation.

Why it matters

OpenAI is turning inference speed into a marketable commodity, which has two concrete consequences. First, it pressures rivals to match or beat the 750 tokens/sec bar if they want to stay relevant in high-throughput scenarios. Second, it entrenches a pay-to-play model where latency-sensitive applications face a clear cost trade-off between performance and budget. For developers, the calculus is now simpler: if your use case demands blistering speed, Ultrafast provides the runway—but only if your wallet can keep pace.


Source: The Decoder. AI-assisted editorial synthesis — TechnoExpress.

Read the original source on The Decoder →

← Back to home