Japan’s AI industry is accelerating efforts to reduce its dependence on Nvidia GPU servers, according to Nikkei Asia. One of the most prominent moves comes from U.S. AI infrastructure company Penguin Solutions (PENG), which plans to launch its MemoryAI KV cache server in Japan in the fourth quarter of 2026.

The system is aimed at a key constraint in large language model, or LLM, inference: memory capacity and bandwidth. Nikkei Asia said the offering could give Japanese AI operators a way to scale inference workloads without expanding Nvidia GPU purchases at the same pace.
Penguin targets the KV cache bottleneck with CXL memory expansion
The report said roughly 70% of AI inference performance bottlenecks come from memory bandwidth and capacity, while only 30% are tied to compute clusters. When the conversational context used by an LLM — the KV cache — exceeds the capacity limit of on-board GPU HBM, operators generally have three choices: recompute discarded tokens, which adds latency; offload cache to NVMe SSDs, which also adds latency; or buy more GPUs, which raises deployment costs.
Penguin Solutions’ MemoryAI KV cache server is built around that problem. It uses CXL, short for Compute Express Link, an open memory interconnect standard, to create a new memory tier outside the GPU cluster. That setup offloads LLM KV cache from GPU HBM to external DDR5 memory modules. Nikkei Asia said the server can be configured with as much as 11 TB of capacity and runs 10 times faster than NVMe-based approaches.
The report described it as the industry’s first mass-production-ready CXL KV cache server. By moving KV cache into external memory, per-GB costs could fall to between one-third and one-seventh of a GPU expansion strategy, based on Nikkei Asia’s estimate. That would give Japanese AI companies an alternative route to scaling LLM inference without broadening Nvidia GPU procurement.
NPU-based inference systems are also being tested in Japan
A second route is developing alongside memory optimization. Nikkei Asia said Toyotsu Device signed a memorandum of understanding with South Korean AI startup Rebellions in October 2025 to jointly promote NPU products and develop the Japanese market. The report added that Toyotsu Device is now conducting a proof of concept with local AI company Tomorrow Net using servers powered by Rebellions NPUs.
The value proposition for NPUs lies in architecture tuned for matrix multiplication and transformer workloads during AI inference. On the same inference tasks, they can often deliver competitive throughput at much lower power consumption than GPUs, cutting electricity use and cooling costs in data centers.
Pressure on HBM demand could hit major memory suppliers
Viewed through the AI memory market, Nikkei Asia said the broader shift away from Nvidia-centered infrastructure could weigh on memory makers as well. The impact of CXL-based systems is straightforward: if customers can expand KV cache capacity without buying more GPUs, they do not need to add as much HBM, because each GPU purchase is tied to HBM demand.
Counterpoint data cited in the report put global HBM market share by revenue at 58% for SK hynix, 21% for Samsung, and 21% for Micron. If alternative architectures reduce overall demand for high-cost HBM, those three suppliers would be the most exposed.

Nikkei Asia sees inference as a long-term variable for GPU demand
Nikkei Asia argued that the spread of NPUs and the rise of other non-Nvidia approaches, including Google’s in-house TPUs, could become a major variable for GPU supply and demand over the medium to long term.
The reasoning is tied to a shift in AI workloads from training to inference. Training is intermittent and compute-intensive, where large-scale parallel GPU performance stands out. Inference is continuous and memory-intensive, which makes lighter, lower-power NPU designs and CXL-based memory optimization more attractive on economic grounds.
In that context, Penguin Solutions’ CXL memory platform and Rebellions’ NPU servers represent two different technical paths into the same market transition: one focused on memory optimization, the other on inference silicon. Nikkei Asia said that shift could also open an important strategic entry point into Japan for South Korean AI semiconductor companies.

