Zhipu has moved AI usage quotas onto an e-commerce shelf. On Sept. 2, the company officially opened a flagship store on Tmall and listed its GLM Coding Plan there. The individual plans are priced at RMB 118 per month for Lite, RMB 538 for Pro, and RMB 1,078 for Max, each with different point allowances. The product is built on GLM-5.3 and supports more than 20 mainstream agents, including ZCode, Claude Code, and Codex.

A few months ago, AI coding platforms were mostly competing on price. Now the comparison is shifting to something else: how many tokens a user is actually allowed to consume.
Older plans are being phased out as usage caps become central
According to the report, some users received a notice from Zhipu saying their older plans would be removed. The previous GLM Coding Plan with an unlimited weekly quota will stop auto-renewing, and existing users will receive two months of the new plan as compensation.
The more important point is not simply that Zhipu is selling token-based AI plans on Tmall. It is that AI companies are paying far closer attention to how much each customer uses, even while tokens themselves are getting cheaper.
Similar moves have appeared across the sector in recent months. Kimi suspended subscriptions for new consumer users because of compute pressure. On July 19, Moonshot AI said requests after the launch of Kimi K3 had far exceeded expectations and were nearing the limits of its current cluster capacity, so it would prioritize compute resources for existing paying users.
Alibaba Cloud also adjusted its Coding Plan. Its Lite plan stopped accepting new purchases on March 20, and renewals and upgrades ended on April 13. Tencent Cloud changed pricing for some Hunyuan models in March. For HY2.0 Instruct, the input price rose from RMB 0.0008 per 1,000 tokens to RMB 0.004505, an increase of more than 460%.
Cheaper tokens do not mean lower AI usage costs
For the past two decades, one of the defining lessons of internet products has been that digitization tends to push marginal costs lower over time. Mobile data is a straightforward example. Once base stations, fiber networks, and other infrastructure are in place, the marginal cost of transmitting another gigabyte is low, and prices typically fall as scale grows.

That logic has naturally been applied to AI. If chips keep improving, models become more efficient, and inference cost per token falls, AI should become cheaper too.
The report says the first part of that argument is correct. Zhipu disclosed that, as model architecture and inference infrastructure improved, its unit inference cost per token fell 80% from the beginning of the year. The company also said it has reached low-cost, large-scale inference using domestically produced chips at a 100,000-chip scale.
The problem is that token consumption is rising faster than token costs are falling. A token is not the same as ordinary data traffic. Behind each token sits the compute required for model inference: GPUs or AI accelerator chips, high-bandwidth memory, networking, storage, and the electricity and cooling inside data centers. In that sense, a token is closer to a metered and tradable unit of AI compute.
When a model is only chatting with a user, one request may consume just hundreds or thousands of tokens. Once the model is asked to complete a more complex task, the consumption profile changes completely.
Zhipu’s half-year results show a sharp turn toward MaaS and API revenue
Zhipu’s latest half-year report puts that transition directly into its financials. In the first half of 2026, the company posted revenue of RMB 954 million, up 399.7% year over year. Revenue from its MaaS open platform and API business reached RMB 825 million, up 2,735.7%, and accounted for 86.5% of total revenue.
In the same period a year earlier, that share was only 15.2%. At the same time, revenue from localized deployments dropped from 84.8% of the total to 13.5%. In the span of a year, the company’s business model has almost completely shifted gears.
Previously, customers were buying the model itself. It would be deployed on the customer’s own servers and then customized, integrated, and delivered, with each project generating one round of revenue. Now more customers are calling cloud-based models directly. Each call is billed, and higher usage means higher spending. That is a standard MaaS business.

The report argues that this is not just a story of cutting prices to gain scale. Instead, several trends are appearing at the same time: token call volume is up more than 40 times from the beginning of the year; the average API selling price is up about 101%; unit inference cost per token is down 80%; and gross margin for the MaaS business has risen to 24.6%.
Costs are falling. Revenue is climbing. Usage is expanding. Average API pricing is also moving higher. That combination suggests AI commercialization may be moving beyond the earliest phase of straightforward price wars.
From selling models to selling calls, subscriptions, and outcomes
In the early stage, when model capabilities were still limited, platforms often relied on low prices or even free access to attract users. Once models became capable of handling code generation, tool use, and sustained execution, what customers were really paying for was no longer just the model. They were paying for the model’s ability to get work done.
Zhipu’s management summarized that path as: selling models, then selling calls, then selling subscriptions, and finally selling end-to-end task outcomes. Those four steps map out four stages of AI monetization.
The shift is easier to see in product use. Under the earlier chatbot model, a user asked a question and got an answer. One interaction might use only hundreds or thousands of tokens. In the agent model, a user might say, “Build a website for me,” and the system may need to understand requirements, search for information, write code, invoke tools, run tests, find bugs, revise the code, and test again. One task can consume hundreds of thousands, millions, or even more tokens.
In a future multi-agent setup, one agent may handle planning, another search, another coding, another testing, and another review. Token demand would expand again.
Exploding usage is changing how the industry prices AI
The report cites data showing that global weekly token usage surged 250% to 22.7 trillion in the first quarter of 2026. On the OpenAI platform, token calls rose from about 6 billion per minute in October 2025 to 15 billion per minute by the end of March 2026, a 150% jump in less than six months. A Morgan Stanley report said top large language models are going through a “non-linear capability leap,” while AI’s explosive growth is running into systemic supply bottlenecks.

Against that backdrop, the real shift is not that tokens themselves have become more expensive. It is that every person and every company may end up needing far more of them. The report compares this to the Jevons paradox: when a resource becomes more efficient and cheaper on a per-unit basis, total consumption does not necessarily fall and may increase instead.
AI appears to be following a similar pattern. As tokens get cheaper, users are more willing to let AI handle more work. As models get more capable, users are also more willing to hand over more complex work. The result is a new flywheel for token consumption.
Viewed through Zhipu’s half-year report, the company is not merely selling AI quotas. It is turning model capability into a continuously billable cloud service. In 2025, most of Zhipu’s revenue still came from localized deployments. By the first half of 2026, MaaS and API had reached 86.5% of total revenue.
Xiao Lei, secretary to Zhipu’s board, said on the earnings call, “Compared with revenue growth, the shift in revenue mix deserves more attention.” Chairman Liu Debing said on the same call, “A lead on a single benchmark ranking can be caught quickly. What really matters is who can keep delivering higher intelligence at lower cost.”
The report’s conclusion on direction is straightforward: basic tokens may keep getting cheaper, while higher-level intelligence may become more expensive. A lower per-token price does not mean lower overall AI spending. As models grow stronger, the work assigned to them moves from answering a question, to completing a task, to managing an entire workflow. Token consumption grows with that shift.
Zhipu’s half-year report put it this way: “Each time the model completes a leap in capability, what customers purchase moves one step closer to results, and the company’s revenue structure is rewritten once again.” On that reading, the more important question ahead may not be how many tokens RMB 1 can buy, but how much AI intelligence a person or a company consumes each month. Zhipu putting token-based plans on Tmall may be just the first step in bringing this metered AI economy to a mass market.

