Nvidia said its Groq 3 LPX inference chip has entered full production and reached an inference speed of 3,400 tokens per second on the Gemma 4 31B model. The company said that figure is four times faster than Cerebras.
The comparison, though, is not straightforward. According to the report, Nvidia needs at least 64 accelerators to reach that throughput, while Cerebras requires only one to two. That makes raw tokens-per-second figures harder to compare without accounting for system scale.
The report also flagged an open question around architecture scalability. While Nvidia presented the production milestone and throughput result, its ability to scale on large mixture-of-experts, or MoE, models remains to be seen. The report was cited by Techub News, with The Decoder named as the source.
Nvidia said its Groq 3 LPX inference chip has entered full production, posting an inference speed of 3,400 tokens per second on the Gemma 4 31B model.
The company said that pace is four times faster than Cerebras. The report added that the performance comparison is more complicated in practice, because Nvidia needs at least 64 accelerators to reach that speed, while Cerebras needs only one to two.
The report also said the architecture's scalability on large MoE models still needs to be tested.
This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan. Disclaimer:
The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.
Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.