SemiAnalysis says memory bandwidth matters more than capacity in AI inference, with scheduling emerging as a core layer

SemiAnalysis says memory bandwidth matters more than capacity in AI inference, with scheduling emerging as a core layer

N
News Editor
2026-09-22 04:50:22
SemiAnalysis has published a report breaking down the underlying architecture of large-model inference services, arguing that the rise of mixture-of-experts, or MoE, models has turned inference into a multi-stage pipeline rather than a single compute task. The report describes that pipeline as consisting of Prefill, Midfill, Decode Attention, and Decode Experts, with each stage placing different demands on compute, memory bandwidth, and networking. Its central conclusion is that, in most inference scenarios, memory bandwidth carries more economic value than raw capacity. According to the report, high-bandwidth memory can improve token generation efficiency, while idle HBM mainly adds cost. SemiAnalysis also projects that by 2027, a single pipeline stage may require roughly 400 GB to 500 GB of local fast memory, though processed KV cache should be moved to CPU DRAM and lower-cost network storage to avoid tying up scarce HBM resources. The report also identifies the scheduling layer as a critical part of inference infrastructure. It says Prefill and Midfill are relatively predictable in runtime, while decode latency is more variable and can create backlog and delay-feedback oscillation. SemiAnalysis further compares integrated and disaggregated architectures, and includes simulator-based projections for Kimi K3 on Nvidia B200, B300, and GB200 systems.

Research firm SemiAnalysis has released a report that breaks down the underlying architecture of large-model inference services. The report says that as mixture-of-experts, or MoE, models become mainstream, AI inference is no longer a single compute task. Instead, it has developed into a more complex pipeline made up of Prefill, Midfill, Decode Attention, and Decode Experts.

Resource demands vary across inference stages

According to the report, those stages place clearly different demands on compute, memory bandwidth, and networking. In most inference scenarios, memory bandwidth has more economic value than capacity. SemiAnalysis says high-bandwidth memory can improve token generation efficiency, while idle HBM only adds cost.

Single pipeline stages may need 400 GB to 500 GB of fast local memory by 2027

SemiAnalysis estimates that by 2027, an individual pipeline stage could require about 400 GB to 500 GB of local fast memory. At the same time, KV cache that has already been processed should be moved promptly to CPU DRAM and lower-cost network storage, so that scarce HBM resources are not occupied unnecessarily.

Scheduling is described as a key infrastructure layer

The report identifies the scheduling layer as a critical part of AI inference infrastructure. Prefill and Midfill have relatively predictable processing times, while decode duration can vary widely, making task backlogs and delay-feedback oscillation more likely. SemiAnalysis says stable buffering, scheduling across different time scales, and tiered KV cache management will directly shape throughput efficiency in inference clusters.

Integrated and disaggregated designs come with different trade-offs

On architecture choices, the report compares integrated and disaggregated approaches. In the integrated design, Prefill and Decode are placed on the same device, which reduces cross-node KV cache transfers but requires hardware that combines high compute density with high memory bandwidth. In the disaggregated design, dedicated nodes are assigned to different stages, allowing hardware optimization for each role, though at the cost of more data movement and higher network overhead.

SemiAnalysis says the trade-off between the two approaches will ultimately depend on whether future accelerators can combine both high compute performance and high memory bandwidth in a single general-purpose design.

Simulator projections cover Kimi K3 on Nvidia systems

The report also uses a simulator to project Kimi K3 performance on Nvidia B200, B300, and GB200 systems. The results show that GB200 has an advantage in some low-latency scenarios, while the peak single-GPU HBM residency in most frontier configurations stays below 80 GB. The report says this adds to the case for prioritizing memory bandwidth and storage orchestration.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
100

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.