Veteran chip designer and hardware analyst Mr. Bubble said in an interview on the Frictionless Podcast that the core bottleneck in AI compute infrastructure is shifting away from simply increasing chip performance and toward advanced packaging, system interconnect latency and memory hierarchy management.
Mr. Bubble said the cost of advanced process nodes is rising quickly, while gains in transistor density are limited. That, he said, is creating an economic challenge for the traditional Moore’s Law model. Instead of relying only on smaller process nodes, the industry is increasingly using advanced packaging to link multiple chips into a single computing system.
The basic unit of AI compute may move to the rack level
He said the basic unit of AI compute in the future may no longer be a single chip or a single server, but a full rack or even multiple racks. In that setup, PCB design, packaging and interconnect capability would become key limiting factors.
Memory hierarchy is becoming a bigger issue
On memory, Mr. Bubble said that as inference context windows expand, large volumes of historical context should not all occupy expensive high-bandwidth memory, or HBM. He said the industry may increasingly adopt flash offloading, keeping high-frequency data in HBM while moving lower-frequency context to storage tiers with larger capacity and lower cost.
He also said 3D DRAM could become an important direction for next-generation memory technology.
System-level hardware design may gain weight in AI infrastructure
On AI infrastructure investment, Mr. Bubble said system-level hardware design will become more important as prefilling and decoding gradually move toward disaggregated architectures.
He added that compared with directly betting on large model companies, businesses with capabilities in hardware, packaging, memory and other critical infrastructure may play a longer-term role in the AI industry chain.

