AI Infrastructure Spending Shifts Toward Optical Interconnects as GPU Share Falls

AI Infrastructure Spending Shifts Toward Optical Interconnects as GPU Share Falls

N
News Editor
2026-09-29 06:19:19
Two research reports from Jefferies and Bank of America argue that the cost structure of AI infrastructure is changing as rack-scale systems grow larger. Jefferies’ bill-of-materials analysis of NVIDIA AI rack platforms shows GPU costs taking a smaller share of total rack spending, falling from 62.5% in the early GB300 NVL72 generation to 46% in the Rubin-era NVL576 Pod. Over the same period, total procurement cost for a single rack jumped from $4.16 million to $55.25 million, while networking and fiber interconnect rose from 8.6% to 22.4%, or more than $12 million in absolute terms. Bank of America, citing an interview with former Microsoft engineering vice president Fran Cardells, said the shift is tied to inference workloads rather than training. In agentic AI systems, KV cache must hold prior token sequences, enterprise context, guardrail rules, and agent decision logs, and in some cases can reach 10 times the size of model weights. As context windows expand from about 128K tokens toward 1 million tokens and multiple agents run at once, memory demand keeps rising. Cardells said the bottleneck is no longer simply how many GPUs a system has, but how quickly data can move across GPUs, memory, and storage. Both reports point to optical networking, photonics, memory pooling, and related infrastructure as areas the market may still be underestimating.

Research from Jefferies and Bank of America says the economics of AI infrastructure are being reshaped as rack-scale systems get bigger. Compute chips now account for a smaller share of total build cost, memory is holding roughly steady, and optical interconnects are taking a larger slice of spending. Both reports argue that investors still may not be fully pricing in that shift.

Rack economics are changing as systems scale up

Jefferies, in a bill-of-materials analysis of NVIDIA AI rack platforms, found that GPUs made up 62.5% of total rack cost in the early GB300 NVL72 generation. By the Rubin-era NVL576 Pod, that share had dropped to 46%.

The overall price tag, however, moved sharply in the other direction. Total procurement cost for a single rack climbed from $4.16 million to $55.25 million, an increase of more than 13 times. The report said that does not mean GPUs have become cheaper. It means other parts of the system are getting more expensive at a faster rate.

Networking equipment and fiber interconnect show the clearest jump. Their share of rack cost rose from 8.6% in the GB300 era to 22.4% in the NVL576 Pod. In absolute dollars, more than $12 million of rack build cost is now going to optical interconnects and switching networks.

Inference workloads are pushing memory into the center of the stack

Bank of America reached a similar conclusion from a different angle in an interview conducted days earlier with former Microsoft engineering vice president Fran Cardells. He said AI inference infrastructure differs fundamentally from model training. Inference needs near real-time response, very low latency, highly localized memory placement, and GPU utilization that makes economic sense.

Under agentic AI workloads, KV cache has to store prior token sequences, enterprise context, guardrail rules, and agent decision logs. In some scenarios, Cardells said, that cache can grow to 10 times the size of the model weights themselves. As context windows expand from roughly 128K tokens toward 1 million tokens, and as multiple agents operate at the same time, memory demand continues to rise.

The report said HBM4, larger system memory requirements, and the use of 3D NAND as an offload medium for KV cache inside rack configurations have all lifted memory’s share of the bill of materials. That share may later ease back to about 19% as HBM4 production matures, but Cardells said HBM demand is still expected to stay elevated and tight supply is unlikely to clear in the near term.

Photonics and high-speed interconnect were singled out as underappreciated

Once GPU clusters move beyond a single rack into multi-rack pods and then into larger data center fabrics, optical links between racks and external scale-out networks become a major line item. Cardells put it plainly: 「High-speed networking, photonics, memory pooling, on-chip memory innovation, and Ultra Ethernet are the most underestimated AI infrastructure investment themes in the market right now.」

He said inference efficiency depends on keeping GPUs busy. That requires high-bandwidth networks that can move KV cache quickly, offload inactive data to lower-cost memory tiers such as DRAM, NVMe, or flash, and reload compressed information in real time when tasks resume.

In that setup, the main constraint is no longer just the number of GPUs available. It is how fast data can move between GPUs, memory, and storage.

Potential beneficiaries extend beyond the traditional GPU chain

Across the two reports, the list of likely beneficiaries is becoming clearer. Bank of America explicitly named SanDisk (SNDK) as a direct beneficiary of 3D NAND demand. It also said DigitalOcean Holdings (DOCN), Apple (AAPL), DELL, and Hewlett Packard Enterprise (HPE) are positioned for longer-term opportunities if agentic AI inference infrastructure sees broad adoption.

Jefferies’ BOM data included optical transceiver suppliers such as Coherent (COHR) and InnoLight, contract manufacturer Fabrinet (FN), and high-speed network switching chip designers Broadcom (AVGO) and Marvell (MRVL). The firm said order visibility for those names is improving.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
100

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.