OpenAI and Google announced on Thursday (13) new artificial intelligence offerings that turn speed into a pricing item. OpenAI presented the Ultrafast tier, in preview for a small group of customers, and Google launched the Gemini 3.7 Flash model.
In the announcement, OpenAI said the Ultrafast tier runs the GPT-5.6 Sol model up to 14 times faster than the standard tier, with up to 750 output tokens per second. The model itself is the same; the response comes out faster because it runs on Cerebras chips. The first customers include Jane Street, Podium, Basis and Rogo, which are testing the technology for programming, financial research, customer service, voice and commerce.
“Until now, getting real-time speed usually meant choosing a smaller or more specialized model. Ultrafast points to progress in a new direction: more useful work per second,” the announcement says.
Gemini 3.7 Flash costs half until the end of the year
On the same day, Google launched Gemini 3.7 Flash. The model costs US$0.75 per million input tokens and US$3.75 per million output tokens until December 31, half the price of the previous Flash. As of January 1, 2027, the price doubles and becomes the same as Gemini 3.6 Flash charged. In the statement, Google said Gemini 3.7 Flash is the “smartest work model so far for coding and agents.”
Benchmarking company Artificial Analysis measured Gemini 3.7 Flash's output at around 340 tokens per second, nearly three times the speed of GPT-5.6 Terra and GLM-5.2, and put the model on the frontier between intelligence and time per task.
AI pricing splits into fast and slow lanes
The launches point to the same shift: for two years, AI pricing has depended basically on model capability and usage volume. Speed becomes a third resource for which companies begin paying separately.
Not every AI task needs to be fast. A bank checking whether a transaction is fraudulent needs to decide in a fraction of a second. According to the report “Where Payment Decisions Happen: How Issuer Data Is Powering the Next Era of Commerce,” from PYMNTS Intelligence, AI fraud detection has already saved at least US$5 million for 42% of card issuers. Slow fraud checking can let a fraudulent charge pass before the system detects it.
In the Ultrafast announcement, OpenAI cited voice applications, customer service, commerce, programming and financial research as tasks for the new tier, plus incident response, which includes fraud and security threats that require quick decisions. A company that runs an AI task overnight to organize old documents, on the other hand, has no one waiting on the other end. That work can run on cheaper, slower computers, with no real cost to the business. It is the same AI capability, at two different prices, depending only on whether someone is waiting for the answer.
The price split resembles buying internet: companies pay more for a fast, guaranteed connection when time matters and use a cheaper, slower connection in other cases. Cerebras, responsible for the Ultrafast tier's chips, used the same argument in its Thursday announcement, comparing the launch to earlier shifts, such as the transition from dial-up internet to broadband.
If that comparison holds, companies will split their AI spending into two lanes: fast, expensive AI for tasks where delay costs money, such as live customer service or real-time fraud checking; slow, cheaper AI for tasks where no one notices the wait, such as overnight reports or document batches. Thursday's launches suggest OpenAI and Google expect that split to become normal.



