Elon Musk endorsed a post on X arguing that memory, not compute, is the real rate-limiting factor for the AI agent era, replying: 「Few realize this.」
The remark was brief, but it was not new. According to the source material, Musk has made the same point four times this year. It also came just three days after Korean brokerages cut target prices for Samsung Electronics and SK hynix on worries that the memory cycle may be nearing a peak.
Musk has returned to the same argument four times this year
The original post came from XPRIZE founder and futurist Peter H. Diamandis, who said memory, rather than compute, is the true bottleneck in the age of AI agents. His case was that long-term memory, context memory and persistent memory will matter more as agents take on larger workloads. Strong compute alone does not solve the problem if an agent cannot retain what it has already done or retrieve it when needed.
Musk’s reply on X was only one line, but the broader argument has appeared repeatedly in his public comments this year.
The first instance cited in the source was Tesla’s Q1 earnings call on April 22. Musk said that, given the pace of industry growth, logic chips and even more so memory chips were likely to become a bottleneck, and that Tesla would hit limits if it did not make chips itself. The report says that argument became part of the rationale for Terafab, with Tesla committing $3 billion at Giga Texas to pursue Intel 14A, while the total project budget was set at $25 billion to $40 billion.
The second came in late June, when Apple CEO Tim Cook complained that memory price increases were unlike anything seen before. Musk publicly agreed.
The third came in early August during SpaceX’s Q2 earnings call. That was where he added numbers: memory capacity is increasing by only about 20% a year, while demand is rising 200% a year or more. In his framing, memory supply is now the factor constraining AI compute expansion.
The fourth was his response on X to Diamandis.
The source article argues that, for Musk, who has already chosen to move into chip manufacturing, the phrase 「Few realize this」 also serves as an explanation for the $25 billion budget behind that push.
Why AI agents put more pressure on memory than chatbots do
The bottleneck described by Diamandis is tied in the report to a specific technical component: KV cache. Each time a model reads a token, it generates key-value data that represents relationships among tokens. That data needs to stay in memory so the model does not have to recompute everything from scratch.
A chatbot conversation does not consume much in one session by comparison. An AI agent handling coding or research work can run through millions of tokens of context, and it needs to retain the entire path of the task as it goes. The issue, then, is not just how much compute is available. It is also how much the system can remember, how long it can hold that state, and how quickly it can retrieve it.
HBM is costly and limited, so the industry is shifting context down the stack
The source says HBM is both too expensive and too scarce, and that the industry’s answer is not simply to keep adding more HBM. Instead, more context data is being pushed downward in the memory hierarchy.
At CES this year, NVIDIA introduced its Inference Context Memory Storage Platform. Using Dynamo and NIXL software, the platform places KV cache directly on NVMe SSDs, making it part of the memory address space and allowing it to persist across inference tasks. NVIDIA said the setup can deliver up to 5x tokens per second and 5x energy efficiency. Dell, Hewlett Packard Enterprise, IBM, Pure Storage, VAST Data and more than ten other vendors were listed in the source as supporting the approach.
That is why the memory bottleneck is not only a story about HBM demand rising. The report frames it as a reordering of the memory stack itself. Context data is moving from HBM down toward NAND and hard drives, lowering cost and allowing systems to support more total computation.
Sandisk and TrendForce point to new NAND demand from KV cache
The source cites SNDK estimates that KV cache alone will add 75 EB to 100 EB of NAND demand in 2027. That figure could double in 2028, and the report says it is not included in current demand forecasts.
SNDK also estimates that by 2030, KV cache will account for 35% of AI data center NAND workloads.
TrendForce data cited in the report shows that servers already make up more than 40% of NAND bit demand. In 2026, data centers are expected to overtake smartphones and laptops for the first time and become the single largest NAND application.
In that reading, this is also why pure-play NAND names such as Sandisk are being repriced in the current cycle.
Markets are not questioning the importance of memory, but the duration of the cycle
Musk said few people realize the issue. The report, though, argues that recent moves in Korean and wider Asian memory stocks suggest the market is focused on a different question.
On July 29, SK hynix disclosed that it was working with major customers to finalize 2027 HBM supply volumes and pricing. The same day, the KOSPI fell 10.84%, with SK hynix down more than 14% and Samsung Electronics down 13%.
On Aug. 11, Kiwoom Securities cut its target price for Samsung from KRW 390,000 to KRW 350,000 and lowered its target for SK hynix from KRW 2.2 million to KRW 2.1 million. Mirae Asset, Shinhan Investment and Samsung Securities also followed with cuts, citing concerns that the conventional memory cycle may already have passed its high point, although both stocks kept their 「buy」 ratings.
Then on Aug. 14, Sandisk issued long-term guidance for 15% to 20% annual revenue growth from 2028 to 2030, saying long-term pricing agreements with customers would support that outlook. Asian memory shares rebounded after that. Kioxia at one point rose 8.7% in Tokyo, while SK hynix gained as much as 6.5% in Seoul. The Bloomberg Asia semiconductor stock index extended its gains to a fifth straight trading day.
The source article says that a three-week stretch taking the market from a sharp selloff, to target-price cuts, to a rebound shows investors are well aware of how important memory is. The real uncertainty is how long the current run can last.
Supply-demand balance remains the next question
TrendForce estimated in July that NAND supply would still face a 4% to 5% gap in 2026, with supply and demand not returning to balance until the second half of 2027. Over the same period, Chinese manufacturers’ share of bit output is projected to rise to nearly 19%.
The source also says Micron’s new capacity is not expected to come online until as early as mid-2027, with other manufacturers lined up for 2028.
In the source’s framing, Musk’s claim that few people realize the problem and broker target-price cuts are answers to two separate questions: whether memory supply is sufficient for AI expansion, and when the current memory upcycle may start to fade.

