OpenAI unveiled Ultrafast mode for GPT-5.6 Sol on August 13, 2026, claiming it runs 14 times faster than standard processing and delivers up to 750 output tokens per second. The feature runs on chips from Cerebras and is currently in preview for a small group of customers.
Key takeaways
- Announced August 13, 2026, Ultrafast mode for GPT-5.6 Sol
- Claimed speed: 14x versus standard processing
- Throughput up to 750 output tokens per second
- Hardware backing comes from Cerebras chips
- Availability: preview for a small customer group, expanding as capacity grows
Speed instead of a smaller model
OpenAI positions Ultrafast as a way to get near real-time responses without dropping to a weaker model. The mode runs on the full GPT-5.6 Sol, not a trimmed-down version.
Until now, getting real-time speed typically meant choosing a smaller or more specialized model. Ultrafast points to progress in a new direction: more useful work per second.
OpenAI statement cited by TechCrunch.
The claimed 14x speedup and up to 750 tokens per second are numbers that matter most where response time counts. OpenAI points to incident response, customer service, financial market analysis and e-commerce as the main use cases.
The role of Cerebras and the competition
The mode is backed by chips from Cerebras, specialized in very fast inference?inference: The phase where a trained model generates outputs from input, as opposed to training.. That sets Ultrafast apart from the standard infrastructure most OpenAI models run on. For many uses what matters is not average response time but the smoothness of the stream, at 750 tokens per second a long answer appears almost at once rather than building up in front of the user. Fast modes are not new to the market, Anthropic offers a fast mode for its Claude models, though OpenAI claims higher speed.
Why it matters
Generation speed is becoming a separate axis of competition in AI, alongside quality and price. For uses such as customer service or incident response, latency can matter more than the last few benchmark points, the user is waiting for an answer in real time. Building the mode on specialized Cerebras hardware shows that further performance jumps increasingly come not from the model itself but from the hardware layer beneath it. If OpenAI keeps full model quality at this speed, it will raise the bar for what counts as real-time work.
What's next
- Access is set to expand as available compute grows, per OpenAI, the mode is currently in preview for a small group of customers
- No price was given for Ultrafast, pricing terms will be decisive for adoption in large-scale use





