Google Research announced TurboQuant on March 25, describing it as a training-free compression algorithm that can quantize large language model KV cache down to 3 bits without losing model accuracy, while reducing memory usage by at least 6x. The market reaction was immediate. Micron fell as much as 6.1% intraday before closing at $382.09, its lowest close in three weeks.
Other memory-related names also moved lower. Sandisk dropped 3.5%, Seagate lost 2.59%, and Western Digital fell 1.63%. In Asia, Samsung Electronics opened down 3.6%, while SK Hynix declined 4.5%. Traders appeared to focus on one simple implication: if AI models need less memory, the pricing power created by recent component shortages could weaken.
Google targets KV cache, a core memory bottleneck in LLM inference
KV cache, short for Key-Value Cache, stores previously computed attention data so a model does not have to repeat the same calculations for every token it generates. As context windows get larger, that cache consumes more memory and has become a major constraint in inference workloads.
According to Google, TurboQuant uses a two-stage process. The first step applies PolarQuant to rotate data vectors for higher-quality compression. The second step uses the Quantized Johnson-Lindenstrauss algorithm to remove residual error. Google said conventional vector quantization methods usually introduce roughly 1 to 2 bits of overhead per value in memory, and TurboQuant is designed to eliminate that burden.
H100 benchmark shows 8x speedup for 4-bit version
In benchmark tests run on Nvidia H100 GPUs, Google said the 4-bit version of TurboQuant delivered an 8x performance improvement in attention metric calculations compared with unquantized 32-bit keys, while cutting KV cache memory usage by at least 6x. Google also said the method requires no training or fine-tuning and adds very little runtime overhead, making it suitable for production inference systems and large-scale vector search deployments.
The company said the related paper will be formally presented in April at ICLR 2026.
Analysts split on whether lower memory intensity means lower demand
Not everyone accepted the idea that TurboQuant points to a collapse in long-term memory demand. Some analysts pointed to the Jevons paradox, arguing that when a technology reduces the cost of using a resource, total demand can rise instead of fall. Under that view, lower inference costs could broaden AI adoption and increase aggregate demand for memory over time.
Analysts at Lynx Equity Strategies wrote that the method described by Google is likely to have almost no impact on memory and flash demand over the next 3 to 5 years, because supply remains highly constrained. The firm kept its $700 price target on Micron.

