OpenAI released a preview of Ultrafast mode for the GPT-5.6 Sol model, capable of generating up to 750 output tokens per second.
The acceleration is provided by Cerebras, with whom OpenAI signed a US$ 10 billion partnership earlier this year. Initially, the service will be available only through OpenAI's API and to select customers, with gradual expansion as capacity increases.
According to OpenAI, the mode combines the speed of smaller models with the full capability of a large reasoning model, which the company calls "more useful work per second."
Use cases
OpenAI says it already uses the mode internally for incident response, allowing it to analyze logs, code changes, and reports during an outage. In finance, the model can assess market signals and flag suspicious transactions. In customer support, complex queries can be resolved in real time, and in e-commerce, it can answer questions, check inventory, and personalize recommendations before the buyer gives up.
In research, experiments that previously ran overnight as batch tasks can become interactive sessions, making it possible to test an idea, review results, and adjust the approach without interrupting the workflow.
Speed as a price differentiator
OpenAI already offers speed tiers on its API. Fast Mode promises up to 2.5 times the speed with lower latency for GPT-5.6 Sol, for about twice the price. Ultrafast adds a third, faster and possibly more expensive tier.
The logic resembles that of cloud providers such as AWS, which charge more for the same service at higher performance levels. With this, OpenAI can directly benefit from revenue generated by faster inference.



