A OpenAI launched GPT 6.1 Sol this Tuesday (29), just one week after the debut of GPT 6 Sol. The new version arrives with advances in programming, computer use and professional tasks and, according to the company, brings its performance close to that of GPT 6 Astra with input and output prices equivalent to one-fifth of those charged for the most advanced model.

The model is available to Plus, Pro, Business, Enterprise and Edu users in ChatGPT Work and Codex, in addition to the API. For now, GPT 6.1 Sol is not available in standard ChatGPT.

In the API, OpenAI charges US$ 2 per million input tokens and US$ 10 per million output tokens. Inputs retrieved from cache cost US$ 0.10 per million tokens, half the price charged for GPT 6 Sol and 95% below uncached input.

The reduction is especially relevant for agents that reuse large volumes of context across different requests, since a larger share of processing can rely on information already stored.

GPT 6.1 Sol approaches Astra in OpenAI tests

In DeepSWE v1.1, focused on complex software engineering tasks in real projects, OpenAI states that GPT 6.1 Sol matched GPT 6 Astra for approximately one-fifth of the cost per task. The result was also 6.4 points above GPT 6 Sol's best performance, even while using a lower level of reasoning and cost.

The company reports similar gains in automation and computer use tasks. In OSWorld 2.0, the new model was 2.1 points below Astra at the maximum reasoning level, but with a cost per task close to one-seventh. In AutomationBench, used to evaluate multi-step corporate workflows, the improvement over GPT 6 Sol was 4.8 points at the same reasoning level.

There was also an improvement in factual accuracy in internal tests with deliberately difficult questions. At a low reasoning level, the share of responses containing at least one factual error fell from 11.4% in GPT 6 Sol to 7.7% in GPT 6.1 Sol, a reduction of approximately 32%. OpenAI itself emphasizes that this test set was built to provoke errors and does not represent everyday use of the model.

Astra still maintains an advantage in some more demanding tasks. In Terminal Bench Science, for example, it achieved the highest score among the evaluated models, and OpenAI continues to recommend its use for more complex scientific work.

The company also states that GPT 6.1 Sol improved in alignment evaluations, with greater transparency about limitations and a lower incidence of unauthorized actions in tasks with agents. The published results, however, were obtained in OpenAI's own research and API environments and may differ from the behavior observed in final products.

In the coming days, the company plans to launch the GPT 6.1 Sol Ultrafast, a version that could generate tokens up to eight times faster than the standard speed in Codex.

More from Radar