BackWideEP

WideEP

SemiAnalysis
2026-07-19 02:20:13

SemiAnalysis says Kimi K3 cuts KV bandwidth, but AI network demand may still rise

SemiAnalysis said Kimi K3 can sharply reduce KV-cache transfer bandwidth through its KDA-based design, but argued that this does not point to a major contraction in the AI networking market. According to the research firm, roughly three-quarters of Kimi K3’s network layers use KDA, which can cut KV-cache transmission bandwidth by as much as 10x versus a full global-attention model. Even so, the model’s overall infrastructure demands remain large. SemiAnalysis said Kimi K3 has 2.8 trillion parameters and still requires about 1.5TB of HBM bandwidth per forward pass even with MXFP4. To deploy the model profitably while maintaining reasonable interaction speed, operators would still need to connect large numbers of chips through high-bandwidth networking systems such as GB300 NVL72 and rely on WideEP for scaling. The firm added that WideEP distributes 896 expert models across multiple GPUs and performs token dispatch and result merging twice per layer in each forward pass, exceeding 120 operations in a single pass. By comparison, KV-cache transfer between prefilling and decoding happens only once per dialogue round, suggesting the bandwidth saved by KDA may be smaller than the additional scaling demand created by large expert-model architectures.

1120
SemiAnalysis says Kimi K3 cuts KV bandwidth, but AI network demand may still rise