GGUF

Tencent Hunyu
2026-09-01 04:22:52

Tencent Hunyuan Releases Extreme Quantization of Hy4 Preview: 1.5TB Shrunk to 214GB, Heterogeneous Inference 6x Faster

Tencent's Hunyuan team has released an extreme quantization version of the Hy4 preview model, compressing the original 1.5TB weights to about 214GB in GGUF format. Using a layer-wise quantization strategy, the average bit-width is 2.38 bpw, with performance loss of only 0.2-1.6 points on benchmarks. The team also tested heterogeneous joint inference with prima.cpp, achieving 1.02 token/s on a setup combining an RTX 4090 laptop and a quad-RTX A4000 server, 6x faster than laptop-only offloading.

150
Tencent Hunyuan Releases Extreme Quantization of Hy4 Preview: 1.5TB Shrunk to 214GB, Heterogeneous Inference 6x Faster
Kimi K3
2026-07-19 04:56:59

Kimi K3 local deployment starts at 64 GPUs, with power demand around 45kW

Moonshot AI this week unveiled Kimi K3, a 2.8 trillion-parameter model described in the source as the largest open-weight model to date. The company said full weights are expected to go live before July 27 under a Modified-MIT license, allowing self-hosting. But the hardware bar is far above consumer setups. In Moonshot AI’s deployment guidance, running K3 on your own requires at least 64 accelerators. Even before inference begins, storing the model’s native 4-bit MXFP4 weights takes roughly 1.4TB to 1.5TB. Based on raw weight size alone, that works out to about 19 NVIDIA H100 80GB GPUs, 11 H200s, or 8 B200s, and that does not include KV cache for long-context operation. The article also cites H100 pricing at roughly $25,000 to $33,000 per PCIe card, with SXM versions above $35,000 to $40,000, while an 8-GPU H100 server costs more than $300,000. At Moonshot AI’s stated 64-GPU threshold, hardware procurement alone would exceed $2.4 million. Power use is also heavy: 64 H100 GPUs at about 700W each add up to roughly 45kW, pushing deployment into data-center territory rather than home or hobbyist environments.

2900
Kimi K3 local deployment starts at 64 GPUs, with power demand around 45kW
2026-07-05 15:42:11

Tether Bets on Decentralized Local AI With QVAC and Psychohistory Vision

Tether is expanding from stablecoins into AI infrastructure through QVAC, a local-first and peer-to-peer AI stack. Its MedPsy medical models aim to prove that smaller edge-deployable models can outperform larger rivals, but independent replication remains the key test.

270
Tether Bets on Decentralized Local AI With QVAC and Psychohistory Vision
Tether
2026-06-29 12:30:11

Tether Launches QVAC: Building Decentralized Local AI with Asimov's Psychohistory, Reserves Shift from Dollars to Intelligence

Tether's QVAC project draws on Isaac Asimov's concept of psychohistory from the Foundation series to build a decentralized, local-first AI infrastructure. The article details how Tether leverages its stablecoin reserves (Q1 2026 net profit $1.04B, reserve buffer $8.23B) to fund AI development, the architectural distinction of QVAC: edge-native, peer-to-peer inference, cross-platform SDK. The first model, MedPsy, outperforms Google’s MedGemma-27B on medical benchmarks despite being 7x smaller (1.7B/4B parameters), demonstrating the potential of domain-specific edge-scale models. It also explores the trade-off between convenience and control in local vs. cloud AI, and the critical need for independent replication.

400
Tether Launches QVAC: Building Decentralized Local AI with Asimov's Psychohistory, Reserves Shift from Dollars to Intelligence