Qwen3.8 Max scored 58 points on the Artificial Analysis Intelligence Index, an 11-point improvement over the previous generation and a result equivalent to Claude Opus 4.8. The data was released by Artificial Analysis.

The performance put the Alibaba model above GLM-5.2, which scored 51 points, and one point ahead of Moonshot AI's Kimi K3, with 57.

The difference shows up in test costs. Each task run by Qwen3.8 Max on the index cost US$ 1.14, while Kimi K3 came in at US$ 0.86 per task, about 25% less.

Qwen advances on professional tasks

On GDPval-AA, a benchmark focused on work-related activities, Qwen3.8 Max reached 1,739 Elo points. The previous version had scored 468 points less.

The model also surpassed Kimi K3, which scored 1,685 points. Among the systems cited in the survey, only Claude Opus 5 was higher, with 1,852.

The advance, however, came with greater resource usage. Qwen3.8 Max needed 64 steps per task, versus 14 in the previous generation.

Input token volume also increased roughly 15-fold. The test methodology resends the full conversation history to the model at each new step, which increases consumption as the number of interactions grows.

Cost per task more than doubles

Alibaba cut the prices it charges for tokens on the new model. Input fell from US$ 2.50 to US$ 2 per million tokens, while output went from US$ 7.50 to US$ 6.

The price for cached tokens dropped from US$ 0.50 to US$ 0.25 per million.

Even with the reductions, the larger volume processed raised the average cost on the Intelligence Index. The US$ 1.14 per task figure is more than double the US$ 0.53 recorded by Qwen3.7 Max.

In the same test, Kimi K3 cost US$ 0.86 per task. GLM-5.2 had an even lower cost of US$ 0.57.

Hallucination rate rises to 40%

Qwen3.8 Max also showed declines in two specific Artificial Analysis tests. In AA-LCR, which measures the ability to gather information from long texts, the result fell two points.

The drop reached ten points on AA-Omniscience, an evaluation that checks whether the model correctly answers knowledge questions or recognizes when it lacks sufficient information.

The rate of correct answers remained near 31%. The hallucination rate, meanwhile, went from 23% to 40% compared with the previous version.

The results indicate that the new model answered more frequently even when it did not have enough knowledge to provide a correct answer.

More from Radar