Chinese models processed about 61.2 trillion tokens on OpenRouter between September 7 and 13, nearly three times the 21.8 trillion attributed to American models in the same survey. It was the twentieth consecutive week in which the aggregate Chinese volume was above the American one on the platform. The difference draws attention because leadership in usage does not necessarily accompany leadership in capability: U.S. models remain among the strongest in intelligence evaluations, while Chinese options have gained ground mainly where price and efficiency weigh as much as absolute performance.

The gap opens a more important question than the position in a ranking: if a model does not need to be the best to execute a large part of production workloads, how much is a benchmark advantage worth in the face of several-fold differences in cost per task?

OpenRouter shows a competition different from that of benchmarks

The Yicai survey shows that OpenRouter processed 127 trillion tokens in the week of September 7 to 13. Chinese models accounted for 61.2 trillion, versus 21.8 trillion for the American ones considered in the comparison. Seven Chinese models appeared among the ten highest-volume ones, against only two a year earlier.

This does not mean that a Chinese model individually surpassed all competitors. The GPT-5.6 Luna, from OpenAI, was the single model with the highest volume in that period, with 18.2 trillion tokens. The Chinese advance appears mainly in the sum of Tencent, Z.AI, DeepSeek, Xiaomi and other providers, which came to occupy several positions simultaneously among the most used models.

It is also important to correctly measure what is being compared. OpenRouter's ranking counts input and output tokens processed by its API, and not the number of calls. Requests kept private are excluded, and OpenRouter itself emphasizes that the data does not represent number of users, requests or spending. It also does not include all traffic made directly on OpenAI, Anthropic and Google APIs or on other clouds and private infrastructures.

The data, therefore, does not demonstrate that Chinese models are more used globally. It shows something more specific: in the multicloud and multimodel market visible through OpenRouter, a growing share of workloads is being directed to Chinese providers.

When the model is already good enough, price changes the decision

The economic explanation appears when performance and cost are analyzed together.

Artificial Analysis found, in August, an extreme contrast in DeepSeek V4-Flash. The model cost on average about US$ 0.03 per task in the company's tests, versus US$ 1.86 for GPT-5.6 Sol and US$ 3.15 for Claude Fable 5. At that moment, the DeepSeek model lagged behind OpenAI and Anthropic's strongest options in intelligence, but its cost per task was more than one hundred times lower than that of the Claude used in the comparison.

The difference does not disappear when models with closer performance are compared. In version 4.3 of the Artificial Analysis Intelligence Index, the GLM-5.3 Flash reached the same score of 42 as GPT-5.6 Terra in a given configuration, but presented an average cost of US$ 0.25 per task, versus US$ 1.40 for the OpenAI model — about 18% of the cost.

This type of relationship changes the math for applications that perform thousands or millions of inferences. A small difference in quality can be economically acceptable when the workload involves classification, data transformation, extraction, automations, intermediate agent steps or code tasks that do not require the most capable available model on every call.

In this scenario, the question stops being only “which model gets the highest score?” and starts to include how much it costs to complete the task with sufficient quality.

Cache and agents expand the importance of efficiency

DeepSeek V4.1 Flash shows how this competition is migrating to less visible components of inference.

The model was launched on September 10 with an architecture of 552 billion parameters, but activates only 8 billion of them during input and 16 billion during output. DeepSeek also states it has significantly reduced the memory requirements of its KV cache, a mechanism used to reuse already processed context.

On OpenRouter, V4.1 Flash providers offer cache reads at prices that reach about US$ 0.006 per million tokens. Yicai reports that approximately 90% of the model's volume observed after launch came precisely from cache reads.

This matters especially for agents and long workflows, in which large blocks of context can be reused repeatedly. The greater the number of steps executed by an application, the greater the possibility that seemingly small differences in the cost of a single inference multiply in the complete operation.

The competition, therefore, begins to move away from the nominal price per token and advance toward metrics such as cost per task, cache efficiency, amount of computation required and total cost of executing a workflow.

Capability leadership remains relevant

The change does not eliminate the advantage of frontier models.

In Artificial Analysis's most recent evaluations, Anthropic and OpenAI models remain among the highest-scoring systems, although Chinese competitors already occupy nearby positions and some results change quickly with each new version. OpenRouter itself explicitly separates its usage ranking from capability benchmarks because one metric does not replace the other.

There are tasks in which a few additional performance points can be worth much more than the savings obtained with a cheaper model: complex research, advanced coding agents, scientific problems, decisions with a high cost of error and workflows in which an initial failure propagates through the following steps.

Therefore, the movement observed on OpenRouter does not indicate that companies will abandon the most capable models. It opens space for a different architecture: reserving frontier models for steps that truly require their capability and directing simpler or repetitive tasks to lower-cost models.

This logic favors routing environments, in which applications can switch models according to price, latency and task difficulty.

The pressure shifts from absolute intelligence to production economics

The Chinese strategy is also starting to appear in companies' financial results.

Reuters Breakingviews observed that Chinese labs have been competing aggressively on efficiency, price and open-weight models. Z.AI stated it has reduced its unit inference cost by 80% since the beginning of the year, while the gross margin of its API division reached 25% in the first half. At the same time, DeepSeek, Z.AI, Tencent and Alibaba accounted for 45% of the market observed by OpenRouter in the week ended September 14.

The economic model still has limitations. Low prices compress margins, and Chinese labs continue spending heavily on research and infrastructure. The same Reuters analysis shows that Z.AI and MiniMax still invest high amounts in R&D relative to their revenues. Gaining volume, therefore, does not by itself guarantee a more profitable business.

But the growth in usage creates a different competitive pressure for OpenAI, Anthropic and Google. They do not only need to defend leadership in benchmarks; they need to demonstrate that the additional capability gain justifies its cost in workloads in which cheaper alternatives already reach the necessary level.

The next signal will be cost per task, not just the next benchmark

OpenRouter remains only a window into the market, and a window particularly favorable to comparison among multiple providers. Confirmation of a structural change will depend on signals beyond this platform.

Among the most relevant will be the adoption of Chinese models in enterprise workflows, the evolution of American API prices, the growth of systems that automatically route tasks among models and, mainly, the ability of Chinese companies to transform large inference volume into sustainable margins.

Benchmarks will remain important for measuring the technical frontier. But recent data show that leading the frontier and capturing the largest volume of work are no longer necessarily the same competition.

More from Radar