The Chinese company Z.AI launched this Friday (18) the GLM-5.3-FlashX, a new version of its model family aimed at applications that require lower latency. The model is already available via API and reaches up to 200 tokens per second, according to the company's official documentation.

FlashX delivers inference speed five times greater than that of GLM-5.3-Flash, but it arrives at a price 2.5 times higher than that of the original version, according to information released by the company and reported this Friday by Chinese outlets.

The proposal is to offer an alternative for developers who prioritize response speed, especially in AI agents, interactive applications, and services where latency has a direct impact on user experience.

FlashX trades lower cost for more speed

In the developer documentation, Z.AI lists the codes glm-5.3-flash and glm-5.3-flashx for access via the API. FlashX, however, is not yet available in the GLM Coding Plan, the company's subscription aimed at using the models in programming tools.

The GLM-5.3-Flash/FlashX family offers a context window of 1 million tokens and supports outputs of up to 128,000 tokens. The model accepts text, images, video, and files as input, while the output is text.

GLM-5.3-Flash has 320 billion parameters in total, with 18 billion activated, and uses a hybrid architecture that combines sparse and linear attention. According to Z.AI, this design reduces the computational cost of attention and the size of the KV cache compared with GLM-5.3.

Arquitetura do GLM-5.3-Flash e comparação de eficiência em contextos longos.
Diagrama da arquitetura do GLM-5.3-Flash com gráficos de cache KV e custo de atenção.

The launch of FlashX comes amid rising demand for the Flash version. Z.AI says it has expanded its computing capacity to absorb the growth in calls, while Chinese outlets report that the cluster used by the company, made up of about 100,000 accelerators produced in China, quickly reached high utilization after the model's debut.

With the new variant, Z.AI now separates two priorities more clearly: the GLM-5.3-Flash remains focused on cost, while the FlashX offers higher generation speed for applications willing to pay more for lower response time.

More from Radar