Penguin Solutions eyes Japan launch for CXL AI memory server as alternatives to Nvidia reshape HBM demand

Penguin Solutions eyes Japan launch for CXL AI memory server as alternatives to Nvidia reshape HBM demand

N
News Editor
2026-07-16 05:52:01
Japan’s AI industry is moving to reduce its reliance on Nvidia GPU servers, according to a Nikkei Asia report, and one of the clearest examples is Penguin Solutions’ plan to introduce its MemoryAI KV cache server in Japan in the fourth quarter of 2026. The product is designed to expand memory capacity for large language model inference by offloading KV cache from GPU-based high-bandwidth memory to external DDR5 memory through CXL, or Compute Express Link. Nikkei Asia said the system can be configured with up to 11 TB and may cut per-GB memory expansion costs to between one-third and one-seventh of a GPU expansion approach, while running 10 times faster than NVMe-based alternatives. The report also pointed to a second track in the same “de-Nvidia” shift: lower-power inference systems built around NPUs. Toyotsu Device has signed an MOU with South Korean AI startup Rebellions and is now conducting a proof of concept in Japan with local AI company Tomorrow Net using servers equipped with Rebellions NPUs. Nikkei Asia argued that these changes could affect demand for high-bandwidth memory, a market where Counterpoint data puts SK hynix at 58%, Samsung at 21%, and Micron at 21% by revenue. As AI workloads move from training toward inference, the report said, memory-optimized and lower-power architectures may emerge as a longer-term challenge to GPU-centered infrastructure.
Japan AIPenguin SolutionsNvidiaHBMCXLNPURebellionsMemory market

Japan’s AI industry is accelerating efforts to reduce its dependence on Nvidia GPU servers, according to Nikkei Asia. One of the most prominent moves comes from U.S. AI infrastructure company Penguin Solutions (PENG), which plans to launch its MemoryAI KV cache server in Japan in the fourth quarter of 2026.

Penguin Solutions eyes Japan launch for CXL AI memory server as alternatives to Nvidia reshape HBM demand 2

The system is aimed at a key constraint in large language model, or LLM, inference: memory capacity and bandwidth. Nikkei Asia said the offering could give Japanese AI operators a way to scale inference workloads without expanding Nvidia GPU purchases at the same pace.

Penguin targets the KV cache bottleneck with CXL memory expansion

The report said roughly 70% of AI inference performance bottlenecks come from memory bandwidth and capacity, while only 30% are tied to compute clusters. When the conversational context used by an LLM — the KV cache — exceeds the capacity limit of on-board GPU HBM, operators generally have three choices: recompute discarded tokens, which adds latency; offload cache to NVMe SSDs, which also adds latency; or buy more GPUs, which raises deployment costs.

Penguin Solutions’ MemoryAI KV cache server is built around that problem. It uses CXL, short for Compute Express Link, an open memory interconnect standard, to create a new memory tier outside the GPU cluster. That setup offloads LLM KV cache from GPU HBM to external DDR5 memory modules. Nikkei Asia said the server can be configured with as much as 11 TB of capacity and runs 10 times faster than NVMe-based approaches.

The report described it as the industry’s first mass-production-ready CXL KV cache server. By moving KV cache into external memory, per-GB costs could fall to between one-third and one-seventh of a GPU expansion strategy, based on Nikkei Asia’s estimate. That would give Japanese AI companies an alternative route to scaling LLM inference without broadening Nvidia GPU procurement.

NPU-based inference systems are also being tested in Japan

A second route is developing alongside memory optimization. Nikkei Asia said Toyotsu Device signed a memorandum of understanding with South Korean AI startup Rebellions in October 2025 to jointly promote NPU products and develop the Japanese market. The report added that Toyotsu Device is now conducting a proof of concept with local AI company Tomorrow Net using servers powered by Rebellions NPUs.

The value proposition for NPUs lies in architecture tuned for matrix multiplication and transformer workloads during AI inference. On the same inference tasks, they can often deliver competitive throughput at much lower power consumption than GPUs, cutting electricity use and cooling costs in data centers.

Pressure on HBM demand could hit major memory suppliers

Viewed through the AI memory market, Nikkei Asia said the broader shift away from Nvidia-centered infrastructure could weigh on memory makers as well. The impact of CXL-based systems is straightforward: if customers can expand KV cache capacity without buying more GPUs, they do not need to add as much HBM, because each GPU purchase is tied to HBM demand.

Counterpoint data cited in the report put global HBM market share by revenue at 58% for SK hynix, 21% for Samsung, and 21% for Micron. If alternative architectures reduce overall demand for high-cost HBM, those three suppliers would be the most exposed.

Penguin Solutions eyes Japan launch for CXL AI memory server as alternatives to Nvidia reshape HBM demand 4

Nikkei Asia sees inference as a long-term variable for GPU demand

Nikkei Asia argued that the spread of NPUs and the rise of other non-Nvidia approaches, including Google’s in-house TPUs, could become a major variable for GPU supply and demand over the medium to long term.

The reasoning is tied to a shift in AI workloads from training to inference. Training is intermittent and compute-intensive, where large-scale parallel GPU performance stands out. Inference is continuous and memory-intensive, which makes lighter, lower-power NPU designs and CXL-based memory optimization more attractive on economic grounds.

In that context, Penguin Solutions’ CXL memory platform and Rebellions’ NPU servers represent two different technical paths into the same market transition: one focused on memory optimization, the other on inference silicon. Nikkei Asia said that shift could also open an important strategic entry point into Japan for South Korean AI semiconductor companies.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
300

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.