OpenAI Ships Ultrafast Preview for GPT-5.6 Sol, Promising 750 Tokens Per Second
The new mode, powered by Cerebras chips, targets enterprise workflows where latency matters more than deliberation - but the rollout stays gated for now.

Speed as a Product Category
OpenAI announced Thursday that its flagship model, GPT-5.6 Sol, now supports an Ultrafast mode capable of generating up to 750 output tokens per second - a fourteen-fold increase over the model's standard processing pace. The feature enters preview with a narrow customer cohort, and OpenAI says broader availability will track capacity expansion.
At DailyTechWire, we've tracked the inference-speed arms race across foundation-model labs for the past eighteen months. What began as a latency footnote in technical documentation has matured into a headline product differentiator. Ultrafast represents OpenAI's most explicit bet yet that some enterprise buyers will pay a premium for raw throughput, even when the underlying model already ranks among the most capable on the market.
The Cerebras Connection
Ultrafast runs on silicon from Cerebras, the Sunnyvale chipmaker whose wafer-scale engines have carved out a niche in high-speed AI inference. OpenAI framed the partnership as enabling "more useful work per second," language that underscores a conceptual pivot: historically, deploying a frontier model meant accepting slower responses in exchange for deeper reasoning. Ultrafast inverts that trade-off for workloads where latency trumps contemplation.
The model emits tokens - discrete chunks of text that large language models produce during generation - at a rate designed to feel instantaneous in interactive applications. Seven hundred fifty tokens per second translates to roughly one hundred fifty words per second under typical tokenization schemes, fast enough to render multi-paragraph answers in under two seconds.
Target Workloads and the Real-Time Threshold
OpenAI named incident response, customer support, financial-market analysis, and e-commerce operations as ideal use cases. Each shares a common constraint: humans or downstream systems expect answers measured in milliseconds, not minutes. A security operations center triaging alerts, a chatbot deflecting tier-one support tickets, or an algorithmic trading desk parsing earnings transcripts all value speed as a first-order requirement.
The company acknowledged that until now, achieving real-time performance typically forced customers to drop down to smaller or task-specific models. Ultrafast aims to collapse that hierarchy, letting organizations keep the reasoning ceiling of GPT-5.6 Sol while meeting latency budgets previously reserved for lightweight alternatives.
That framing matters. Across the enterprise AI stack we follow in Seoul, Singapore, and San Francisco, buyers have complained that flagship models deliver impressive benchmark scores yet remain too slow for production at scale. Ultrafast is OpenAI's direct answer to that friction, though the gated preview suggests the company is still stress-testing infrastructure before opening the floodgates.
Competitive Context and the Anthropic Benchmark
Anthropic offers a fast mode for its Claude family, though OpenAI's quoted throughput appears meaningfully higher. The gap reflects both architectural choices and chip partnerships: Cerebras wafer-scale engines prioritize parallelism and on-chip memory bandwidth, design decisions that favor throughput over energy efficiency or cost per token.
We expect other labs to respond. The foundation-model market has entered a phase where differentiation hinges less on benchmark leaderboards and more on operational variables - latency, cost, deployment flexibility, and compliance tooling. Speed is now table stakes for any model targeting high-volume enterprise workflows, and Ultrafast raises the bar.
Capacity Constraints and the Preview Gate
OpenAI limited initial access to a small customer group and tied broader rollout to capacity growth. That language signals supply-side constraints, likely a function of Cerebras chip availability or data-center build-out timelines. Wafer-scale engines remain expensive and scarce relative to conventional GPU clusters, and OpenAI's partnership does not yet appear to have reached hyperscale volume.
The phased launch also lets OpenAI tune pricing. Ultrafast will almost certainly command a premium over standard inference, but the company has not published rate cards. Enterprise customers we've spoken with across the region report that pricing opacity remains a persistent pain point in foundation-model procurement, and Ultrafast's preview phase likely doubles as a price-discovery exercise.
Implications for the Inference Layer
Ultrafast underscores a broader trend: the inference layer is fragmenting into speed tiers. Foundation-model providers now offer not just different model sizes but different inference profiles for the same model - standard, fast, and in OpenAI's case, ultrafast. This mirrors the evolution of cloud compute, where customers choose instance types optimized for memory, CPU, or network depending on workload.
For enterprises, the calculus becomes more complex. A chatbot handling routine queries might run on Ultrafast to minimize user wait time, while a research assistant drafting patent applications might use standard inference to preserve reasoning depth. Managing that segmentation - routing requests to the right speed tier - will fall to orchestration layers and prompt routers, categories that have attracted meaningful venture capital in the past year.
What Remains Unsaid
OpenAI's announcement offered no detail on how Ultrafast affects output quality. Faster inference can introduce trade-offs: more aggressive sampling strategies, shorter context windows, or reduced internal reasoning steps. The company's framing - "more useful work per second" - sidesteps the question of whether the work itself changes.
We also lack visibility into energy consumption and cost per token. Cerebras chips deliver speed but at unknown expense, both financial and environmental. If Ultrafast proves prohibitively expensive or power-hungry, its addressable market shrinks to a narrow set of latency-critical, high-margin applications.
Finally, the preview gate raises questions about OpenAI's infrastructure strategy. The company has historically relied on Microsoft Azure and NVIDIA GPUs for the bulk of its compute. Cerebras represents a diversification bet, but one that comes with execution risk if the partnership cannot scale to meet demand from OpenAI's millions of users.
Looking Ahead
Ultrafast is a feature, not a model, but it signals where OpenAI believes competitive pressure will intensify. As foundation models converge in capability - GPT, Claude, and Gemini all cluster near the top of most benchmarks - operational attributes become the wedge. Speed, cost, and deployment flexibility matter more than an extra percentage point on MMLU.
For now, Ultrafast remains a limited preview, and its ultimate impact will depend on pricing, availability, and whether enterprises find the speed gains worth the likely cost premium. But the direction is clear: the inference wars are no longer just about intelligence. They are about clock speed.


