On September 30, Google introduced Gemini 4 Argon, its new frontier model, after months in which the company stopped releasing Gemini 3.5 Pro while OpenAI and Anthropic updated their most advanced models. Argon arrives with competitive results in professional work, programming, and security, plus an announced limit of 1 million output tokens. But the most relevant feature of the launch is not in knowing whether Google won more rows in a benchmark table.

The question is different: how far a model can remain working alone on a long task, using tools, querying large volumes of information, correcting its own steps, and reaching a verifiable result?

It is at this point that Argon offers the most interesting signal. Google is trying to turn prolonged reasoning capacity into a commercial feature of the model. If this performance holds up outside tests, competition among leading labs will depend less and less on the quality of a single response and more on the cost, reliability, and duration of entire processes executed by agents.

The most important result is not on Google's scoreboard

The launch materials place Argon ahead of competitors in several tests, but comparisons produced by the developer itself must be treated with caution. Different models may use different tools, levels of effort, and execution structures, and small differences in score do not always represent perceptible advantages in production.

Gemini 4 Argon leads several benchmarks but still trails rivals in some categories.
Comparison of Gemini 4 Argon with OpenAI and Anthropic models on knowledge, programming, science, long-context, and security benchmarks.

Vals AI evaluated Argon in its own suite and placed the model in first position on the Vals Index, with 68.9%, against 67.04% for Claude Sonnet 5.5 and 66.97% for Claude Opus 5.5. The index combines tasks in finance, programming, law, and taxation and tries to bring the test closer to professional work with economic value.

Another particularly useful result comes from Zapier's AutomationBench. In this test, the agent must operate simulated enterprise environments, consult CRM, spreadsheets, emails, and other tools, and correctly modify the state of these systems. It is not enough to produce a convincing answer: the task must actually end with the correct records changed.

Argon at the High level reached 51.29% success. Claude Opus 5.5 reached 42.47% and GPT 6 Astra, 41.4% in their best configurations presented in the same ranking. This still means that almost half of the tasks are not fully completed, but the difference is relevant precisely because the test measures end-to-end execution, not just knowledge or text generation.

This may be Argon's most important indicator: the advance appears precisely in the type of task that companies would like to delegate to agents.

This does not mean universal leadership. Vals's own evaluation places Argon only in fifth place on Terminal Bench 4.0 and near the end of the small CUA Bench sample, which tests computer use. Google also showed results in which competing models remain superior in certain terminal, science, and software engineering tasks.

There are, however, external signs that make the launch more relevant.

Therefore, the reading most supported by the data is not that Google built a better model at everything. It is that Argon seems especially competitive when reasoning, a large amount of context, and multistep execution appear together.

One million tokens changes more than the size of the response

Google raised from 64,000 to 1 million tokens, the announced maximum output limit of Argon. This is different from merely offering a large context window.

Context determines how much material the model can receive. Output determines how much work it can produce within a trajectory. For agents, this difference matters.

A prolonged run may involve reading files, tool calls, forming hypotheses, tests, inspecting results, corrections, and new attempts. Each of these steps consumes tokens. Smaller limits force complex systems to interrupt trajectories, summarize the state, or start new sessions.

Google argues that the new limit allows complex problems to be kept within a single much longer trajectory.

The internal examples released help explain what the company intends to do with this. According to Google, agents with Argon analyzed server telemetry and identified optimizations capable of freeing more than 300 TiB of memory when deployed. The company also says it is using the model in migrating C and C++ codebases to Rust, including more than 800,000 lines of the Zircon kernel, from Fuchsia. In another case, the model worked on 32,000 lines of SIMD code from a video decoder and produced a Rust implementation 2.7 times faster than the previous port.

These numbers are results released by Google itself and still do not amount to independent validation. The strategic point, however, is clear: the product Google is trying to sell is not just a chatbot that answers better. It is computational capacity applied for longer to a problem.

This even changes the economic unit by which frontier models may begin to be compared.

Cost per task begins to matter more than cost per token

Argon will initially launch at US$ 2 per million input tokens and US$ 10 per million output tokens. After the promotional period, the prices become US$ 4 and US$ 20.

The final price will be equal to that of Claude Opus 5.5 per token and twice the price of Claude Sonnet 5.5. GPT 6 Astra, in turn, costs US$ 10 per million input tokens and US$ 50 per million output tokens at the standard price published by OpenAI.

But this comparison becomes progressively less useful for agents.

A cheap model that makes 20 tool calls, repeatedly returns to the same documents, and needs several attempts may end up more expensive than another with a higher price per token but capable of completing the work in fewer steps.

Gemini 4 Argon in Vals AI's independent evaluation.
Gemini 4 Argon metrics in the Vals Index, including accuracy, cost, and latency.

AutomationBench already allows this difference to be seen. Considering the standard price and not the promotion, a task with Argon High cost an average of US$ 1.70 in the test, practically the same as the US$ 1.73 for Astra Max, but correctly completed a larger proportion of tasks. During the promotion, Zapier calculates about US$ 0.85 per task for Argon High.

In Vals, Argon showed an average cost of US$ 15.68 per test in the Vals Index, compared with US$ 21.34 for Sonnet 5.5 and US$ 32.14 for Opus 5.5 in the configurations evaluated. In long agentic tasks, however, costs rise quickly: the same evaluation records US$ 57.82 per test in code migration and US$ 193.78 in the CUA Bench.

This difference helps explain why tokens alone are ceasing to be a sufficient measure. The economically relevant indicator tends to be the cost of arriving at a correct result.

Cybersecurity has become a laboratory for the next generation of agents

There is another unusual element in the launch: most developers will not be able to use Argon immediately.

First access is being given to selected participants in the Fairwind Program, a Google initiative created for governments, critical infrastructure operators, and security partners. The program already had more than 650 participants when it was announced in September.

For these approved defenders and for internal teams, Google says it will make Argon available without the cyber restrictions applied to normal access, allowing more advanced defensive activities. According to the company, the model can search for, validate, and fix vulnerabilities autonomously.

Gemini 4 Argon reaches 68% on CWE Bench v1 in vulnerability-fixing tasks.
Performance comparison of Gemini 4 Argon and rival models on the CWE Bench v1 cybersecurity benchmark.

Google says Wiz has already used the system to find a critical vulnerability in software used by hospitals, a flaw that previous models had not identified. On CWE Bench v1, focused on fixing vulnerabilities, Argon scored 68%.

More important than an isolated case is the distribution architecture that is emerging.

OpenAI structured Daybreak to grant additional security capabilities to verified defenders and says that GPT 6 Astra has already reached its level Critical of cyber capability. Anthropic also uses verification programs to release broader capabilities to approved professionals.

This points to a market model different from the one that marked the first years of large language models.

Not every capability of a frontier model will necessarily be delivered to all users at the same time. Identity, purpose, execution environment, and level of control may begin to define which practical version of intelligence a customer actually receives.

Argon turns this mechanism into a central part of the launch.

The paradox is that Google's most interesting model still cannot be widely tested

This limitation also prevents a stronger conclusion about Argon's performance.

Google has not yet announced a public availability date. The expansion should begin with paying API customers and Google AI Ultra subscribers, after data collection from the first users and the strengthening of the system's protections.

This matters because the announcement comes after a difficult period for the company's advanced model line.

Gemini 3.5 Pro, which Sundar Pichai had indicated for June, will no longer be released. During that interval, OpenAI and Anthropic introduced new models, while Google reorganized parts of DeepMind. Argon is therefore not only a new product but also the company's attempt to return to the dispute over the most advanced models after months without a new flagship available.

In this context, limited access works at the same time as a security mechanism and as a limitation on the available evidence.

There are independent evaluations of the model, such as Vals and Zapier, but there is not yet the diversity of tests, real applications, integrations, and production workloads that emerges when thousands of developers begin using a model.

Even the 1 million token limit needs to go through this process. Vals, for example, evaluated Argon with a maximum of 262,144 output tokens in its configuration. This does not contradict the limit announced by Google, but it shows that the model's most extreme feature has not yet been widely exercised in external evaluations.

Generating hundreds of thousands of tokens also creates problems that a technical specification does not solve on its own: latency, cost, degradation along the trajectory, accumulated errors, and difficulty verifying an enormous amount of work produced autonomously.

An agent that can work for longer is only more useful if it remains correct for longer.

The next competition will be over the reliability of the entire trajectory

Argon's launch reinforces a change that was already appearing in products from OpenAI, Anthropic, and Google itself.

Leading models are being designed less and less to answer a question and more and more to receive an objective.

This difference changes the necessary benchmarks. A model can write excellent code in one response and still fail to modify a repository over four hours. It can understand a financial document and still make an error when transferring the result to a spreadsheet. It can identify a vulnerability and introduce another when trying to fix it.

That is why evaluations such as AutomationBench gain importance. Even the current leader can fully complete only slightly more than half of the evaluated tasks.

Gemini 4 Argon leads AutomationBench with 51.3% success in enterprise workflow automation.
Comparison of the performance of Gemini 4 Argon, GPT 6 Astra, Claude Fable 5.1, and Claude Opus 5.5 on the AutomationBench benchmark.

There is a lot of space between “the model can do” and “a company can hand this to the model without supervision”.

Argon reduces this distance in some categories, according to the first data. It does not demonstrate that it has disappeared.

What really needs to be observed now

The next relevant milestone will not be another chart published by Google.

It will be the opening of Argon to developers and companies at a scale sufficient to allow independent comparisons of duration, cost, and completion rate on real tasks.

It will also be important to observe whether applications can take advantage of very long trajectories without cost growing faster than the value produced; whether AutomationBench performance repeats in real enterprise environments; whether the advantage observed in professional work remains when Sonnet, Opus, and Astra are given equivalent agent structures; and how Google manages access to the most sensitive cyber capabilities.

Finally, there will be an economic test: how many customers will choose Argon for the result obtained per dollar, and not for the number of benchmarks won.

Gemini 4 Argon does not end the dispute between Google, OpenAI, and Anthropic. Its launch shows something possibly more important: the definition of a frontier model is beginning to change.

The next commercial leap may not be a system that answers better in a few seconds.

It may be one capable of receiving a problem in the morning, working on it for hours, using dozens of tools, reviewing its own work, and returning a result that can actually be used.

Argon is one of the clearest demonstrations so far that this is the direction Google is betting on.

Tools mentioned

More from Radar