SemiAnalysis says Kimi K3 cuts KV bandwidth, but AI network demand may still rise

SemiAnalysis says Kimi K3 cuts KV bandwidth, but AI network demand may still rise

N
News Editor
2026-07-19 02:20:13
SemiAnalysis said Kimi K3 can sharply reduce KV-cache transfer bandwidth through its KDA-based design, but argued that this does not point to a major contraction in the AI networking market. According to the research firm, roughly three-quarters of Kimi K3’s network layers use KDA, which can cut KV-cache transmission bandwidth by as much as 10x versus a full global-attention model. Even so, the model’s overall infrastructure demands remain large. SemiAnalysis said Kimi K3 has 2.8 trillion parameters and still requires about 1.5TB of HBM bandwidth per forward pass even with MXFP4. To deploy the model profitably while maintaining reasonable interaction speed, operators would still need to connect large numbers of chips through high-bandwidth networking systems such as GB300 NVL72 and rely on WideEP for scaling. The firm added that WideEP distributes 896 expert models across multiple GPUs and performs token dispatch and result merging twice per layer in each forward pass, exceeding 120 operations in a single pass. By comparison, KV-cache transfer between prefilling and decoding happens only once per dialogue round, suggesting the bandwidth saved by KDA may be smaller than the additional scaling demand created by large expert-model architectures.
SemiAnalysisKimi K3AI networkingKV cacheKDAGB300 NVL72WideEP

On July 19, semiconductor and AI research firm SemiAnalysis said that while Kimi K3 uses KDA in roughly three-quarters of its network layers and can reduce KV-cache transfer bandwidth by as much as 10x compared with a full global-attention model, that does not mean the AI network-switch market is set for a sharp contraction.

SemiAnalysis said Kimi K3 has 2.8 trillion parameters. Even with MXFP4, each forward pass still requires about 1.5TB of HBM bandwidth. To deploy the model profitably while keeping interaction speeds at a reasonable level, operators would still need to connect large numbers of chips through high-bandwidth networking platforms such as GB300 NVL72 and depend on WideEP for scale-out services.

According to the firm, WideEP spreads 896 expert models across multiple GPUs and performs token dispatch and output merging twice in every layer during each forward pass, for more than 120 executions in a single pass. KV-cache transfer between prefilling and decoding, by contrast, takes place only once in each dialogue round. On that basis, SemiAnalysis said the bandwidth savings from KDA may be smaller than the network expansion required by large-scale expert models.

SemiAnalysis also said more efficient attention mechanisms could push context length from 1 million tokens to more than 5 million tokens. Under Jevons paradox, it added, efficiency gains may expand AI usage and in turn increase network demand.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
400

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.