The Nvidia put the Groq 3 LPX into production, an accelerator dedicated to artificial intelligence inference, after the system reached 3,400 tokens per second in an independent test. The company announced the new hardware stage on Monday (24), during Hot Chips 2026.
The Groq 3 LPX integrates the Vera Rubin platform and was developed to accelerate token generation in applications of agentic AI, in which models perform multiple reasoning steps, use tools, and process large volumes of context.
In the benchmark conducted by Artificial Analysis, the system ran the open model Gemma 4 31B with a 100,000-token context window and reached 3,400 tokens per second. According to Nvidia, it was the fastest performance ever recorded for that model.
The company says the result represents a response capacity four times greater than the nearest competing platform. Data analyzed by The Register point to Cerebras as that rival, with about 882 tokens per second in the same scenario.
Comparison depends on the number of chips
The performance difference does not, however, represent a direct comparison between equivalent quantities of hardware. Each Groq 3 LPU has 500 MB of SRAM memory, which requires distributing larger models across multiple accelerators.
The Register calculated that the Gemma 4 31B, running in FP8, needs at least 64 LPUs to store its parameters. Meanwhile, the same model could be accommodated on one or two CS-3 accelerators from Cerebras. The Nvidia did not detail how the chips were distributed during the benchmark.
The architecture can also face greater challenges with models of the type mixture of experts (MoE), which distribute processing among different parameter groups. According to the analysis, a model like DeepSeek V3 could require 1,342 Groq 3 accelerators, equivalent to a little more than five racks.
Nvidia's comparison also considers the current generation of Cerebras used in the tests and not the CS-4, recently presented by the competitor with increases in computing capacity, bandwidth, and interconnection.
Nvidia's strategy is to combine GPUs and LPUs at different stages of inference. GPUs are responsible for the initial processing of prompts, while the Groq 3 LPX focus on the token generation phase, which is more dependent on memory bandwidth.
The Nebius will be the first cloud provider to adopt the Groq 3 LPX, with plans to make it available through the Token Factory platform. The Groq is also among the first companies expected to use the technology.



