‹ BackNewsKV Cache

KV Cache

AI Infrastructure Spending Shifts Toward Optical Interconnects as GPU Share Falls
SemiAnalysis
2026-09-22 04:50:22

SemiAnalysis says memory bandwidth matters more than capacity in AI inference, with scheduling emerging as a core layer

SemiAnalysis has published a report breaking down the underlying architecture of large-model inference services, arguing that the rise of mixture-of-experts, or MoE, models has turned inference into a multi-stage pipeline rather than a single compute task. The report describes that pipeline as consisting of Prefill, Midfill, Decode Attention, and Decode Experts, with each stage placing different demands on compute, memory bandwidth, and networking. Its central conclusion is that, in most inference scenarios, memory bandwidth carries more economic value than raw capacity. According to the report, high-bandwidth memory can improve token generation efficiency, while idle HBM mainly adds cost. SemiAnalysis also projects that by 2027, a single pipeline stage may require roughly 400 GB to 500 GB of local fast memory, though processed KV cache should be moved to CPU DRAM and lower-cost network storage to avoid tying up scarce HBM resources. The report also identifies the scheduling layer as a critical part of inference infrastructure. It says Prefill and Midfill are relatively predictable in runtime, while decode latency is more variable and can create backlog and delay-feedback oscillation. SemiAnalysis further compares integrated and disaggregated architectures, and includes simulator-based projections for Kimi K3 on Nvidia B200, B300, and GB200 systems.

400
SemiAnalysis says memory bandwidth matters more than capacity in AI inference, with scheduling emerging as a core layer
Jev team shares Coding Agent draft centered on per-turn context selection
Wall Street Reprices SanDisk as AI Inference Lifts NAND Into the Infrastructure Trade
Musk Agrees AI Agent Era's Real Bottleneck Is Memory, Not Compute
Marvell unveils AI memory lineup spanning CXL expansion and photonic shared memory
SemiAnalysis says Kimi K3 cuts KV bandwidth, but AI network demand may still rise