Tencent has shared blind-test results for its Hy4 preview model, based on real engineering tasks from inside the company. In a test involving 163 internal experts and 203 tasks, Hy4 scored an average of 2.99 out of 4, edging out GLM-5.3 (2.92) and Kimi K3 (2.94). Against GLM-5.3, Hy4 won 46.8% of comparisons, tied 12.8% and lost 40.4%; against Kimi K3, it won 51.2%, tied 7.9% and lost 40.9%. The tasks were drawn from Tencent's actual engineering workload rather than a standard fixed benchmark. However, Hy4 does not lead across all public benchmarks. The company's results show it trading wins and losses with GLM-5.3 and Kimi K3, and it still trails GLM-5.3 on code and cybersecurity tests such as DeepSWE and CyberGym. Pricing is Hy4's clearest edge. At 6 yuan per million tokens for input and 18 yuan for output, it is 25% and around 36% cheaper than GLM-5.3 on those two metrics, and 70% and 82% cheaper than Kimi K3. Cache-hit inference costs just 0.3 yuan per million tokens, 85% less than the 2 yuan charged by both GLM-5.3 and Kimi K3.
Blind test on internal engineering tasks
Tencent has published blind-test results for Hy4 preview, based on engineering tasks pulled from its own operations. The company said 163 internal experts evaluated models across 203 tasks. Hy4 averaged 2.99 out of 4, slightly ahead of GLM-5.3's 2.92 and Kimi K3's 2.94.
In head-to-head comparisons, Hy4 beat GLM-5.3 in 46.8% of cases, tied 12.8% and lost 40.4%. Against Kimi K3, it won 51.2% of the time, tied 7.9% and lost 40.9%. The set uses real internal engineering scenarios, not a fixed public benchmark.
Not a clean sweep on public benchmarks
Yet Hy4 does not lead across every public benchmark. The same evaluation shows it trading wins and losses with GLM-5.3 and Kimi K3. On code- and security-focused tests such as DeepSWE and CyberGym, it still trails GLM-5.3.
Price is the clearest edge
Hy4 pricing starts at 6 yuan per million input tokens and 18 yuan per million output tokens. That is 25% cheaper than GLM-5.3 on input and roughly 36% cheaper on output. Against Kimi K3, the same figures are 70% and 82% cheaper.
Cache-hit inference is priced at 0.3 yuan per million tokens, 85% below the 2 yuan charged by both GLM-5.3 and Kimi K3.
This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan. Disclaimer:
The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.
Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.