‹ BackNewsMoE

MoE

DeepSeek open-sources Ascend infrastructure components with APIs aligned to Nvidia versions
NaiveAI open-sources first model as AI helps drive 151 optimization runs in six days
Vitalik says offline mobile AI apps have improved, but still trail laptop models
SemiAnalysis
2026-09-22 04:50:22

SemiAnalysis says memory bandwidth matters more than capacity in AI inference, with scheduling emerging as a core layer

SemiAnalysis has published a report breaking down the underlying architecture of large-model inference services, arguing that the rise of mixture-of-experts, or MoE, models has turned inference into a multi-stage pipeline rather than a single compute task. The report describes that pipeline as consisting of Prefill, Midfill, Decode Attention, and Decode Experts, with each stage placing different demands on compute, memory bandwidth, and networking. Its central conclusion is that, in most inference scenarios, memory bandwidth carries more economic value than raw capacity. According to the report, high-bandwidth memory can improve token generation efficiency, while idle HBM mainly adds cost. SemiAnalysis also projects that by 2027, a single pipeline stage may require roughly 400 GB to 500 GB of local fast memory, though processed KV cache should be moved to CPU DRAM and lower-cost network storage to avoid tying up scarce HBM resources. The report also identifies the scheduling layer as a critical part of inference infrastructure. It says Prefill and Midfill are relatively predictable in runtime, while decode latency is more variable and can create backlog and delay-feedback oscillation. SemiAnalysis further compares integrated and disaggregated architectures, and includes simulator-based projections for Kimi K3 on Nvidia B200, B300, and GB200 systems.

390
SemiAnalysis says memory bandwidth matters more than capacity in AI inference, with scheduling emerging as a core layer
DeepSeek says it is training a 2 trillion-parameter model and planning an 8 trillion-parameter version
StepFun unveils Step 5 Preview, with full model weights set for Oct. 15 release
Tencent open-sources Hunyuan Hy4 preview with 770B parameters and 1 million-token context
Tencent unveils T1 for long-horizon terminal Agent tasks, with 300-plus tool calls per run