Google and OpenAI announced new artificial intelligence models focused on speed on Thursday. Google launched Gemini 3.7 Flash, while OpenAI opened a preview of GPT-5.6 Sol Ultrafast, a service layer that runs its most capable model at up to 750 tokens per second.
The Gemini 3.7 Flash is a generally available model focused on coding and autonomous agents. Google says it accepts up to 1 million input tokens and returns up to 64,000 output tokens, in addition to processing text, images, video, audio, and PDFs, with the ability to use tools and control a computer. The company says the model completed a test coding task in 2 minutes 13 seconds, while the previous Flash version took more than 5 minutes.
The promotional price of Gemini 3.7 Flash is US$ 0.75 per million input tokens and US$ 3.75 per million output tokens until the end of the year. According to Google, that amounts to half the initial rate of Gemini 3.6 Flash. After December 31, the price becomes US$ 1.50 and US$ 7.50.
GPT-5.6 Sol Ultrafast
GPT-5.6 Sol Ultrafast is not a new model but a faster version of GPT-5.6 Sol, running on Cerebras chips. OpenAI said the speed is up to 14 times faster than the standard, with generation of up to 750 tokens per second, about 560 words. Access was opened to a select group of API customers, with expansion as capacity permits.
While Google published its own benchmarks placing Gemini 3.7 Flash ahead of competitors such as Claude Sonnet 5 and GPT-5.6 Terra in 11 of the 18 categories tested, OpenAI did not release direct comparisons. The company cited customer assessments, such as Jane Street engineer John Crepezzi and Podium product lead Courtland Lykins, who highlighted speed gains.
The announcements come amid the race between the companies to make AI fast enough for real-time agents. Gemini 3.7 Flash is already available in more than 160 countries; GPT-5.6 Sol Ultrafast remains invite-only.



