Tencent’s Hunyuan AI team has put out an aggressively quantized version of its recently open-sourced Hy4 preview model. The original weights came in at about 1.5TB. Huge. The new GGUF quantized build is only around 214GB, which sharply lowers the barrier to running this 770B MoE model locally.
Layer-Wise Quantization Strategy
Tencent did not squash the whole model down to 1.25-bit across the board. Instead, it assigned different quantization precision to each layer depending on how sensitive that layer was. The least sensitive layers were pushed down to as little as 1.31-bit, while the sensitive ones kept 2-bit precision or more. The finished file averaged 2.38 bpw. In four benchmarks Tencent shared, the hit versus the BF16 original was just 0.2 to 1.6 points. Pretty small. That suggests very little was lost.
Heterogeneous Device Joint Inference
Once the quantization work was done, Tencent teamed up with prima.cpp to try heterogeneous joint inference. The test setup used an RTX 4090 laptop plus a server carrying four RTX A4000 cards, for a combined 80GB of VRAM and 64GB of system RAM. The model ran at 1.02 token/s. That is about 6 times faster than offloading on the laptop alone. So yes, devices with different configurations can split the inference load, which pushes down the cost of local deployment even more.

