Morgan Stanley said in its latest report that AI memory shortages could persist through the cycle, and that Nvidia may not wait for supply expansion to catch up. Instead, the bank said, the company is likely to offer lower-memory versions of its Rubin platform, with HBM4 possibly reduced from 288GB to 192GB and rack-level LPDDR5 cut from 54TB to 28TB in some configurations. The report said demand is not disappearing; it is shifting toward NAND, DRAM and high-speed interconnects.
The report, written by Morgan Stanley semiconductor analyst Joseph Moore, argues that the main constraint in the AI race is moving away from the number of GPUs and toward memory availability. In that framework, the industry response is not to stop building, but to redesign systems so existing hardware can keep supporting AI growth.
Nvidia’s Rubin platform may arrive in lower-memory versions
Morgan Stanley said tight DRAM and NAND supply, along with rising prices, are pushing the AI hardware supply chain to revisit product configurations. For chip vendors, the report said, it may make more sense to offer multiple memory-capacity options than to stick to original specifications and risk missing delivery schedules.
At the rack-memory level, Rubin had originally been planned with about 54TB of LPDDR5, according to the report. Some updated configurations could reduce that to 28TB, with SOCAMM2 module capacity dropping from 192GB to 96GB.
For HBM, Morgan Stanley said a single Rubin GPU had originally been expected to carry 288GB of HBM4 using eight 12-high stacks. The bank now expects Nvidia could add a 192GB HBM4 version by moving to eight-high stacks.
For Rubin Ultra, the report said 512GB is a more appropriate comparison point after a move to a dual compute-chiplet design, and future capacity options could range from about 192GB to 384GB.
Morgan Stanley said these are its own judgments on likely product configurations as of the date of publication, and that not all of the specifications have been formally confirmed by Nvidia.
From a cost perspective, the strategy looks rational in the report’s view. Cutting HBM stack height reduces capacity, while theoretical bandwidth does not necessarily fall at the same time if interface count and pin speed stay unchanged.
Still, capacity and bandwidth solve different problems. For smaller models and shorter contexts, lower capacity may have limited effects. But as models grow and inference tasks become more complex, insufficient memory capacity can force more data movement across storage tiers, raising latency or increasing GPU usage.
Demand is shifting, not disappearing
Morgan Stanley said reducing HBM does not make AI data vanish. It changes where that data sits.
If GPU HBM capacity is not enough, some data may move into main memory. If rack-level LPDDR5 is reduced, part of the demand for KV cache may shift toward NAND flash.
The report described KV cache as the cache that stores historical computation information during large-model inference. As context windows get longer and concurrent requests increase, its capacity requirements also rise.
That leaves Nvidia’s lower-memory approach as a way to ease DRAM supply constraints at least temporarily, while at the same time increasing demand for NAND, interconnect chips and larger GPU clusters.
Morgan Stanley cited historical data from Epoch AI, saying frontier model parameter counts in the large language model era had at one point shown a pattern of doubling roughly every six months. Context windows also expanded quickly. At the same time, more AI applications are moving beyond simple question answering into coding, complex reasoning and long-running agent tasks, all of which add to memory needs.
Because of that, the bank said model size, context length and inference concurrency should keep driving storage demand higher. Lower configurations look more like a transitional response to tight supply than a signal of a long-term fall in AI memory demand.
Prefill and decode are being split apart
Morgan Stanley’s second focus is disaggregated inference, where different parts of the inference workload are separated and assigned to hardware that is better suited to each task.
Traditional large-model inference usually has two stages. The first is prefill, where user prompts, files or historical context are processed with large amounts of parallel computation. That stage tends to rely more on raw compute performance. The second is decode, where the model generates output tokens one by one. That process repeatedly accesses model weights and KV cache, which makes memory bandwidth, capacity and latency more important.
Historically, the same GPU setup often handled both jobs. Morgan Stanley said that may no longer be the most efficient arrangement as inference demand grows. A more effective direction may be to use compute-heavy hardware for prefill, then hand decode to architectures tuned for low latency and high bandwidth.
The report said that could lift hardware utilization and let data centers scale the two kinds of resources separately instead of buying the same GPU-and-HBM setup for every workload.
Cerebras, AMD and AWS were cited as examples
Morgan Stanley said Cerebras is one of the more direct potential beneficiaries of this shift. Cerebras uses a wafer-scale processor architecture that integrates large numbers of compute units and SRAM on one chip. SRAM has lower density than DRAM, but it offers low latency and high bandwidth, which can suit some inference workloads that need frequent data access.
By moving compute closer to data, the report said, Cerebras can reduce the need for external memory access and improve processing speed in certain parts of inference.
Morgan Stanley said Cerebras has already partnered with AMD and Amazon Web Services to explore unified inference systems built from different processors.
In the AMD-Cerebras setup, AMD Helios handles more compute-intensive prefill and large-context processing, while the Cerebras wafer-scale engine handles decode. Morgan Stanley expects that setup to enter production in the fourth quarter of 2026.
AWS is taking a similar route, using Trainium for prefill and Cerebras for decode. The combined offering is expected to enter Amazon Bedrock in the first quarter of 2027.
Based on performance figures disclosed by Cerebras and AMD for specific configurations, the report said the pairing could deliver as much as a fivefold increase in throughput while maintaining Cerebras inference speed. Morgan Stanley said that does not mean every inference workload will see the same gain, but it does show that matching hardware to task type can materially improve system economics.
For Cerebras, the shift could mean more than added chip sales. The company also runs its own inference cloud service, and if the same hardware can generate more tokens, cost per token could fall, improving gross margin or giving the company room to reduce customer pricing.
Nvidia is exploring a similar path
Morgan Stanley said Nvidia is also working in a similar direction. Through the integration of Groq-related technology, the report said, Nvidia is combining HBM-based GPUs with an LPU architecture built around high-speed SRAM.
In the setup described by the bank, Rubin GPUs would handle prefill and part of decode that requires larger cache capacity, while the Groq architecture would take workloads better matched to its low-latency characteristics. NVIDIA Dynamo software would coordinate tasks across the processors.
The report drew a distinction here: Nvidia’s 2025 arrangement with Groq was a technology licensing and talent-acquisition deal, not a full acquisition of the company.
Morgan Stanley said heterogenous compute architectures tailored to different stages of inference could gain more room in the market as inference spending grows faster than training. In its infrastructure spending forecast, the bank expects inference to account for about 56% of AI infrastructure spending by 2030, with training at about 44%.
That changes what may matter in evaluating AI systems. Peak compute alone may not be enough; token generation rate and cost per token may become just as important.
CXL is emerging as another area to watch
Beyond changing how inference work is assigned, Morgan Stanley also focused on using existing memory resources more efficiently. That is where CXL comes in, according to the report.
Compute Express Link is a high-speed interconnect protocol that allows processors to access external memory more flexibly rather than relying only on directly attached DRAM.
In conventional servers, memory resources are usually tied to specific CPUs. Even if one server has a large amount of idle memory, that memory is not easily available to another server. CXL improves overall DRAM utilization in data centers through memory expansion, sharing and pooling. Expansion adds capacity to a single processor, sharing allows multiple processors to access the same memory resource, and pooling organizes distributed memory into a common pool that can be assigned dynamically based on workload.
The report said the technology had been more common in traditional CPU environments, where workloads are better able to tolerate the extra latency that comes with external memory. AI GPUs, by contrast, have long depended on HBM for extremely high data bandwidth, and external DRAM connected over CXL cannot directly replace HBM.
That said, the architecture is becoming more relevant as inference patterns change. Long-context inference requires large amounts of KV cache, but not every piece of that cache has to stay in the fastest HBM all the time. Systems can keep performance-sensitive data in HBM while moving less frequently accessed and less latency-sensitive data to external DRAM.
Morgan Stanley said that can ease pressure on expensive HBM capacity while lifting utilization of available DRAM. In the bank’s view, AI is pushing CXL beyond traditional server memory expansion into a wider market for accelerator memory connectivity and sharing.
The report noted that Astera Labs had previously put the potential market for CXL memory controllers at more than $4 billion. Morgan Stanley now estimates that the addressable market for CXL and related memory-connectivity semiconductors could reach about $6 billion by 2030 as demand rises for AI inference, KV cache offloading and rack-scale memory pooling.
The bank added that this figure is an estimate of potential market size, not realized revenue.
Astera Labs and Marvell were singled out
Morgan Stanley highlighted Astera Labs and Marvell as companies to watch.
For Astera Labs, the opportunity is centered on its Leo memory controller products. The report said the company expects standardized and customized Leo products to begin scaling in 2027 with two large U.S. cloud-service customers, including designs for AI inference KV cache offloading.
Marvell, meanwhile, is targeting memory expansion and rack-scale memory pooling through Structera X and Structera S. The company had previously said CXL-related business could contribute more than $1 billion in revenue around 2028.
Morgan Stanley also said the actual size of the CXL market remains difficult to pin down. Cloud providers could choose non-CXL interconnect technologies, larger-capacity HBM or flash-based cache approaches instead, and adoption could differ materially from current expectations.
For that reason, the investment case around CXL still depends on actual deployment progress at cloud companies, product shipments and revenue tied to those products.
Morgan Stanley kept Overweight on Micron and Sandisk
For the storage supply chain, Nvidia lowering some memory configurations does not look positive at first glance. If HBM capacity per GPU falls and LPDDR5 usage per rack declines, storage vendors may end up selling less memory than originally planned.
Morgan Stanley acknowledged that this could reduce some demand and pricing opportunities for storage companies from a short-term profit-maximization perspective. Even so, the bank said that should not be read as a sign that the storage cycle is ending. Its point is that lower configurations are being driven by supply constraints, not because customers no longer need as much memory.
If AI companies want to buy more DRAM but cannot secure enough supply, lower configurations can help keep system deliveries on track, the report said. If supply expands later, vendors still have an incentive to move configurations higher again, which means some unmet demand today could turn into future purchasing room.
Based on that view, Morgan Stanley maintained Overweight ratings on Micron and Sandisk. The bank said storage tightness is unlikely to end quickly.
It also said the main risk still comes from AI demand itself. If growth in model development, inference applications or compute spending slows sharply, both the compute and storage supply chains could be hit.
The report also described a more complicated case in which AI demand stays strong but data center construction is delayed by land, power or infrastructure constraints. In that scenario, storage products may already have been produced, but delayed server deployment could prevent them from being absorbed as expected, leading to temporary supply-demand imbalances.
Morgan Stanley said these buildout bottlenecks could create added volatility for storage vendors and even cause their short-term performance to diverge from some compute-chip companies. That is why watching DRAM pricing and HBM orders alone is not enough; actual AI infrastructure deployment also matters.
What Morgan Stanley says to watch next
The bank’s central conclusion is that the industry’s real problem is not simply waiting for more memory. It is figuring out how to keep expanding AI compute when memory remains scarce.
The variables Morgan Stanley said to watch include whether Rubin ships with lower real-world configurations, whether disaggregated inference reaches commercial scale, whether CXL products ramp on schedule in 2027, and whether AI data center buildouts keep absorbing new storage supply.
If those technical adjustments can sustain AI system deployment while long-term storage demand continues to grow, storage and interconnect suppliers could benefit together. If AI investment slows or land and power limits keep delaying data center rollouts, the current tightness could be repriced.

