Tencent Hunyuan Releases Extreme Quantization of Hy4 Preview: 1.5TB Shrunk to 214GB, Heterogeneous Inference 6x Faster

Tencent Hunyuan Releases Extreme Quantization of Hy4 Preview: 1.5TB Shrunk to 214GB, Heterogeneous Inference 6x Faster

N
News Editor
2026-09-01 04:22:52
Tencent's Hunyuan team has released an extreme quantization version of the Hy4 preview model, compressing the original 1.5TB weights to about 214GB in GGUF format. Using a layer-wise quantization strategy, the average bit-width is 2.38 bpw, with performance loss of only 0.2-1.6 points on benchmarks. The team also tested heterogeneous joint inference with prima.cpp, achieving 1.02 token/s on a setup combining an RTX 4090 laptop and a quad-RTX A4000 server, 6x faster than laptop-only offloading.

Tencent’s Hunyuan AI team has put out an aggressively quantized version of its recently open-sourced Hy4 preview model. The original weights came in at about 1.5TB. Huge. The new GGUF quantized build is only around 214GB, which sharply lowers the barrier to running this 770B MoE model locally.

Layer-Wise Quantization Strategy

Tencent did not squash the whole model down to 1.25-bit across the board. Instead, it assigned different quantization precision to each layer depending on how sensitive that layer was. The least sensitive layers were pushed down to as little as 1.31-bit, while the sensitive ones kept 2-bit precision or more. The finished file averaged 2.38 bpw. In four benchmarks Tencent shared, the hit versus the BF16 original was just 0.2 to 1.6 points. Pretty small. That suggests very little was lost.

Heterogeneous Device Joint Inference

Once the quantization work was done, Tencent teamed up with prima.cpp to try heterogeneous joint inference. The test setup used an RTX 4090 laptop plus a server carrying four RTX A4000 cards, for a combined 80GB of VRAM and 64GB of system RAM. The model ran at 1.02 token/s. That is about 6 times faster than offloading on the laptop alone. So yes, devices with different configurations can split the inference load, which pushes down the cost of local deployment even more.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
1500

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.