Tencent Hunyuan Releases Extreme Quantization of Hy4 Preview: 1.5TB Shrunk to 214GB, Heterogeneous Inference 6x Faster
Tencent's Hunyuan team has released an extreme quantization version of the Hy4 preview model, compressing the original 1.5TB weights to about 214GB in GGUF format. Using a layer-wise quantization strategy, the average bit-width is 2.38 bpw, with performance loss of only 0.2-1.6 points on benchmarks. The team also tested heterogeneous joint inference with prima.cpp, achieving 1.02 token/s on a setup combining an RTX 4090 laptop and a quad-RTX A4000 server, 6x faster than laptop-only offloading.







