NVMe

Elon Musk
2026-08-14 08:12:12

Musk says memory, not compute, is the real bottleneck for AI agents

Elon Musk backed a claim on X that memory, rather than raw compute, is the real rate-limiting factor in the age of AI agents, replying that 「Few realize this.」 The point is one he has made repeatedly this year. According to the source material, Musk raised the issue at Tesla’s first-quarter earnings call on April 22, echoed complaints about memory price inflation in late June, and then gave a sharper figure at SpaceX’s Q2 earnings call in early August: memory capacity is growing about 20% a year, while demand is rising 200% or more. The report ties that view to a broader shift in AI infrastructure. As agentic systems hold much larger context windows and persistent task histories, KV cache becomes a bigger memory burden. With HBM still expensive and tight in supply, companies are moving more context data down the memory stack into NAND and storage. NVIDIA’s Inference Context Memory Storage Platform, introduced at CES, is cited as one example, using Dynamo and NIXL to place KV cache on NVMe SSDs. The same report also tracks how markets are reading the memory cycle. Korean brokerages cut target prices for Samsung Electronics and SK hynix on concerns that the conventional memory upcycle may have peaked, while Sandisk later issued long-term revenue growth guidance that helped memory shares rebound across Asia.

130
Musk says memory, not compute, is the real bottleneck for AI agents
AI Agents
2026-08-11 00:17:09

AI agents are pushing storage into the runtime loop, reshaping the role of SSDs, HBM and memory tiers

A MarsBit report argues that AI agents are changing storage from a passive persistence layer into part of the execution path itself. As agents continuously observe, reason, call tools, write back results and preserve state, the value of storage is no longer limited to saving data after a task is complete. The report says SSDs are beginning to take on functions tied to model weights, KV cache spillover, indexing, encryption, compression, lifecycle control and long-term memory, pointing to a broader shift toward programmable, functional SSDs. The piece lays out how this transition could play out on both devices and in the cloud. On the edge, SSDs may become the long-lived state layer for personal agents, holding local models, adapters, vector indexes, personal memory and tool traces. In cloud deployments, storage nodes could move closer to the inference path, handling shared prefixes, KV data, adapters, vector search and governance. The report cites Mooncake and NVIDIA CMX as examples of systems where storage is already participating in token production rather than merely holding cold data. It also argues that the rise of agent systems does not diminish HBM. Instead, HBM, HBF, DRAM/CXL and SSDs are likely to be re-tiered by speed, mutability, capacity, cost and governance needs. Existing AI SSD efforts from Phison, Longsys, Maxio and partners are presented as early industrial samples of this shift, where the focus is moving from faster disks for AI workloads to a reallocation of responsibilities across runtime, memory hierarchy, controllers and flash.

690
AI agents are pushing storage into the runtime loop, reshaping the role of SSDs, HBM and memory tiers
AI Storage
2026-08-08 14:13:52

IOSG says AI storage boom is being priced for speed, while decentralized storage keeps its case around trusted cold data

IOSG argues that the current storage rally is being driven by artificial intelligence, but not in the way traditional IT buyers used to think about storage. In its view, the market is no longer rewarding raw capacity first. It is rewarding the ability to keep GPUs fed, move checkpoints quickly, support retrieval-augmented generation with very low latency, and raise overall compute utilization across tightly coupled infrastructure stacks. That shift, the article says, is why components such as HBM, DRAM, CXL, enterprise SSDs, SSD controllers, NVMe pathways, and performance storage software have become central to the AI investment narrative. The piece draws a sharp distinction between AI storage and decentralized storage. AI storage is framed as an efficiency system built for hot data and commercial output. Decentralized storage, by contrast, is described as a trust system for cold data, focused on permanence, censorship resistance, auditability, and public memory. IOSG uses Filecoin and Arweave as the main examples, outlining how the two networks diverge in architecture and product direction, while also listing persistent problems across the sector, including weak enterprise service layers, retrieval limits, supply-demand incentive mismatches, privacy and compliance tensions, and token economics that can amplify market cycles rather than solve product-market fit.

610
IOSG says AI storage boom is being priced for speed, while decentralized storage keeps its case around trusted cold data
Kimi K3
2026-08-08 06:24:42

2.78T-Parameter Kimi K3 Runs on 8GB RAM via Open-Source C99 CPU Engine

An open-source project called kimi-k3-in-c aims to run Kimi K3, a model with 2.78 trillion parameters, on devices with just 8GB of memory. The 176KB codebase is written in pure C99 and performs inference on the CPU only, dropping GPU, CUDA, PyTorch, and BLAS entirely. The approach exploits Kimi K3's MoE architecture: only 16 of the 896 experts per layer are activated, so the developer avoids loading the full ~1.56TB of weights and instead streams most expert weights from NVMe storage on demand. Dense trunk layers are streamed layer by layer as well. Trade-offs remain: generating one token takes about 32.7 seconds, and close to 1.7TB of fast storage is needed. The developer describes the project as an experimental exploration of LLM inference infrastructure rather than a production-ready solution, but the combination of hard-drive streaming and sparse MoE activation points toward new ways to run ultra-large models at low cost.

870
2.78T-Parameter Kimi K3 Runs on 8GB RAM via Open-Source C99 CPU Engine
Marvell
2026-08-06 08:33:11

Marvell unveils Bravera SC6 SSD controller for PCIe Gen6 AI data center storage

Marvell said on Aug. 4 that it has launched the Bravera SC6 SSD Controller, model MV-SF1410, aimed at the PCIe Gen6 NVMe SSD market for AI data centers, hyperscale cloud operators, and enterprise database deployments. The company said the chip is built for a shift in AI system design toward tiered memory architectures, where NVMe SSDs act as a high-capacity memory layer rather than serving only as conventional storage. According to Marvell, that setup can provide tens to hundreds of terabytes of capacity at a lower cost than GPU HBM and DRAM, helping keep GPUs utilized while reducing infrastructure costs. Bravera SC6 supports PCIe Gen6, remains backward compatible with PCIe Gen1 through Gen5, and complies with the NVMe 2.2 specification. It also features 16 NAND channels, transfer speeds of up to 3600 MT/s, support for ONFI and Toggle NAND, and compatibility with SLC, MLC, TLC, and QLC flash. Marvell said samples are expected to become available in the fourth quarter of 2026.

700
Marvell unveils Bravera SC6 SSD controller for PCIe Gen6 AI data center storage
SanDisk
2026-07-20 00:16:10

SanDisk and SK hynix push High Bandwidth Flash as a new memory layer for AI inference

SanDisk and SK hynix are promoting a new architecture called High Bandwidth Flash, or HBF, that applies HBM-style advanced packaging and die stacking to NAND Flash. The idea is not to replace High Bandwidth Memory, but to insert a new layer into AI server memory hierarchies for large, read-heavy data sets used during inference. According to public material cited in IEEE Spectrum, first-generation HBF is targeting up to 16 stacked NAND chips, as much as 512GB per stack, and read bandwidth of up to 1.6TB/s. SanDisk’s roadmap then points to 2TB/s and 3.2TB/s in later generations. SK hynix executive Hoshik Kim said the approach could deliver bandwidth well beyond standard NVMe storage and help relieve the severe capacity pressure faced by HBM. The concept is aimed at inference rather than training. In training, models require constant reads and writes, which is a poor fit for Flash. In inference, model weights are largely fixed, making static, read-intensive data such as billion-parameter weights and precomputed KV cache more suitable for HBF. The companies have already launched standardization work under the Open Compute Project, though IEEE Spectrum reported that shipping is still at least a year away and broad deployment may take several more years.

950
SanDisk and SK hynix push High Bandwidth Flash as a new memory layer for AI inference
Japan AI
2026-07-16 05:52:01

Penguin Solutions eyes Japan launch for CXL AI memory server as alternatives to Nvidia reshape HBM demand

Japan’s AI industry is moving to reduce its reliance on Nvidia GPU servers, according to a Nikkei Asia report, and one of the clearest examples is Penguin Solutions’ plan to introduce its MemoryAI KV cache server in Japan in the fourth quarter of 2026. The product is designed to expand memory capacity for large language model inference by offloading KV cache from GPU-based high-bandwidth memory to external DDR5 memory through CXL, or Compute Express Link. Nikkei Asia said the system can be configured with up to 11 TB and may cut per-GB memory expansion costs to between one-third and one-seventh of a GPU expansion approach, while running 10 times faster than NVMe-based alternatives. The report also pointed to a second track in the same “de-Nvidia” shift: lower-power inference systems built around NPUs. Toyotsu Device has signed an MOU with South Korean AI startup Rebellions and is now conducting a proof of concept in Japan with local AI company Tomorrow Net using servers equipped with Rebellions NPUs. Nikkei Asia argued that these changes could affect demand for high-bandwidth memory, a market where Counterpoint data puts SK hynix at 58%, Samsung at 21%, and Micron at 21% by revenue. As AI workloads move from training toward inference, the report said, memory-optimized and lower-power architectures may emerge as a longer-term challenge to GPU-centered infrastructure.

700
Penguin Solutions eyes Japan launch for CXL AI memory server as alternatives to Nvidia reshape HBM demand