Zhipu AI Weighs In-House Inference Chip as GLM Demand Surges and Compute Tightens

Zhipu AI Weighs In-House Inference Chip as GLM Demand Surges and Compute Tightens

N
News Editor
2026-07-07 13:10:26
Chinese large-model company Zhipu AI is reportedly evaluating the development of a custom inference chip and has already sounded out domestic chip design firms about potential cooperation. The move comes as the company faces both supply constraints and rapidly rising model demand. Zhipu has been placed on the U.S. blacklist, limiting its access to Nvidia’s advanced chips, while usage of its open-source GLM-5.2 model has accelerated sharply. According to the report, daily token consumption for GLM-5.2 on developer platform Vercel jumped 27-fold within a single week, worsening compute shortages. Although Zhipu has already deployed domestic computing capacity, including Huawei hardware, and completed substantial software adaptation work, it is said to be exploring a chip strategy similar to Google’s TPU and OpenAI’s in-house efforts. The project remains at an early discussion stage. If it advances, manufacturing would likely be handled by Chinese foundries, but the timeline from design to tape-out is expected to take more than two years, meaning any relief to near-term capacity pressure would be limited.
Zhipu AIAI chipsGLM-5.2inference chipsdomestic computeVercelchip design

Zhipu AI explores an in-house chip route

According to monitoring by Dongcha Beating, Chinese frontier model company Zhipu AI is considering the development of a custom inference chip and has already approached domestic chip design companies to discuss possible cooperation. The effort is still in an early-stage contact phase, with no indication that it has moved into full execution or production planning.

The reported push toward an internal chip strategy is being shaped by two pressures at once. First, Zhipu AI has been placed on the U.S. blacklist, which prevents it from purchasing Nvidia’s advanced chips. Second, demand for its open-source GLM-5.2 model has risen quickly, making existing compute constraints more visible.

GLM-5.2 usage spikes on Vercel

The report said daily token usage for GLM-5.2 on the developer platform Vercel surged 27 times in one week. That increase has intensified the company’s compute shortage. While Zhipu has already deployed domestic computing resources, including Huawei hardware, and completed extensive software adaptation, the company is still evaluating a deeper hardware move to reduce long-term dependence on constrained supply chains.

Another stated goal is to lower long-term cloud inference costs. In that sense, the strategy is described as following a path similar to Google’s TPU program and OpenAI’s own chip efforts. Even so, any such project would be a long-cycle undertaking rather than a near-term fix.

Manufacturing timeline likely to stretch beyond two years

If the project moves forward smoothly, manufacturing would reportedly be handled by Chinese foundries. However, the timeline from chip design to actual tape-out is expected to take more than two years. That means the initiative, if realized, would be aimed more at medium- to long-term supply resilience and cost control than at solving immediate infrastructure bottlenecks.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
200

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.