The Chinese Zhipu AI revealed that the anonymous Ox Alpha model, which gained traction among developers in recent days, was actually the new GLM-5.3-Flash. said on Wednesday (26) that all test traffic was processed on AI accelerators produced in China, while the model reached the top of the OpenRouter and OpenCode platforms before the official launch.
Zhipu's shares closed Thursday (27) up more than 12%, to HK$ 1,160, in Hong Kong. During the period in which it operated anonymously, the model processed about 62 trillion tokens on OpenRouter and OpenCode, according to data released by the company and reported by South China Morning Post.
The launch draws attention not only for the model's performance, but for the infrastructure used to support it. In the technical report, Zhipu says it put the GLM-5.3-Flash into production on a large network of Chinese chips and managed to increase by three times the inference performance compared to the initial configuration used on the same hardware.

Viral model was tested without revealing its origin
Before the official presentation, the GLM-5.3-Flash appeared on OpenCode and OpenRouter under the codename Ox Alpha, without identification of the responsible company.
The strategy allowed the system to be subjected to real developer traffic before its origin was known. According to Zhipu, Ox Alpha quickly became the most popular model of the week on the platforms where it was available.
On OpenRouter, the model had processed more than 11 trillion tokens in the first three days, establishing the largest launch ever recorded by the platform. On Thursday, it accounted for about 10.3 trillion tokens among programming-oriented systems, approximately 31% of the category's weekly volume.
The infrastructure behind this traffic also became a central part of the announcement. Zhipu says in its technical publication that the operation took place on a structure made up of tens of thousands of accelerators developed in China. Information provided by the company to the Chinese press placed the total capacity used at about 100,000 domestic chips.
Zhipu redesigns architecture to reduce cost
The GLM-5.3-Flash is the first natively multimodal model in the GLM-5 family, capable of working with different types of visual and textual information from the pretraining stage.
The system has 320 billion parameters, but activates only 18 billion of them during processing. The company also reduced the number of layers to 45, versus 92 in the GLM-4.5 generation, which used 32 billion active parameters.
Another change is in the mechanism used to handle large context volumes. The architecture combines linear and sparse attention, allowing the model to process nearby information and retrieve relevant parts of a broader context without performing the same amount of computation at each step.
According to tests released by Zhipu, this configuration reduces by three times the attention-related processing and by 4.4 times the size of the so-called KV cache compared to the GLM-5.3. This feature is especially important when the model works with contexts that can reach 1 million tokens.
The company also says it used a set of 30 trillion multimodal tokens in the pretraining of the new system.

Chinese chips support large-scale operation
To run the model on domestic chips, Zhipu developed a specific inference mechanism based on SGLang and adapted different parts of the infrastructure to the memory and bandwidth limitations of the accelerators.
The system also separates tasks such as interpreting multimodal content, initial prompt processing, and response generation into independent groups of machines. This architecture allows different inference stages to be distributed according to demand.
The company says the optimizations brought the cost per token and hardware efficiency close to the levels achieved with conventional GPUs from Nvidia. The result is presented by Zhipu as evidence that Chinese chips can sustain inference of advanced models at production scale.
In the tests published by the company itself, the GLM-5.3-Flash scored 63.4 points on DeepSWE v1.1, versus 46.2 for the GLM-5.2, and 48.8 on AutomationBench, against 26.2 for the previous generation. Zhipu also says the new model approaches Claude Opus 4.8 in different evaluations of programming and AI agents.
The weights of the GLM-5.3-Flash were made publicly available, and the model can already be run using frameworks such as SGLang, vLLM and TokenSpeed. Zhipu also began offering it to subscribers of its programming-focused plan.



