Tencent Research Institute says tokens are becoming the base unit of the AI economy

Tencent Research Institute says tokens are becoming the base unit of the AI economy

N
News Editor
2026-09-21 12:27:12
Tencent Research Institute argues in the introduction to its new book Token Economy that a token should be understood not as storage, compute cycles, or data transfer, but as a unit for machine-produced intellectual work. The piece places tokens alongside earlier economic units such as the bushel, the barrel, the kilowatt-hour, and the bit, arguing that major shifts in economic organization often follow the arrival of a new standard of measurement. The article says token-based AI services have already moved beyond theory. It cites OpenAI’s disclosure in March 2026 that its API was processing 15 billion tokens per minute, up from 6 billion per minute in October 2025. It also points to figures released by China’s National Data Administration showing daily token calls in China rising from about 100 billion in early 2024 to more than 140 trillion by March 2026. IDC, according to the article, reported in April 2026 that global AI infrastructure spending reached $318 billion in 2025, up from $153 billion in 2024. The text also argues that token markets strain standard economic frameworks. It says the same number of tokens can create sharply different economic value depending on the task, while prices are falling quickly even as total spending rises. The article describes this as a version of the Jevons paradox and says existing tools for procurement, pricing, ROI analysis, trade, and regulation need to be recalibrated for a market where machine intelligence is bought and sold as a metered service.

Tencent Research Institute argues in the introduction to its new book Token Economy that understanding tokens is essential to understanding the AI economy. Its central claim is that consuming one token does not mean buying storage, compute cycles, or information transfer. It means paying for one unit of intellectual processing performed by a model. Interpreting a sentence, generating code, or making a judgment, tasks the article describes as activities once tied mainly to the human brain, now have a measurable unit and an executable price.

Tencent Research Institute says tokens are becoming the base unit of the AI economy 2

The article starts with a historical argument about measurement. One simple way to identify an industrial revolution, it says, is to look for the birth of a new unit of account. A measurement standard determines what can be traded, priced, and absorbed into an economic system. Once an output that was previously vague becomes standardized, pricing, exchange, competition, and institutions tend to form around it.

It then reviews several earlier shifts in economic history. In agriculture, the key unit was the bushel. Before standardization, grain could only be priced by the sack, and one sack of wheat could differ from another depending on the size of the bag and the honesty of the seller. The bushel did more than simplify trade, the article says. It made forward contracts possible and allowed a buyer in Chicago to purchase Kansas wheat without seeing it first. Railroads lowered the cost of moving goods across regions, but standardized measurement was what made futures trading workable. Without the bushel, the article argues, wheat could not have become a remotely traded standardized commodity.

In the energy era, the relevant unit was the barrel. One standard oil barrel equals 42 gallons. The article says this began as a temporary choice by Pennsylvania oil field operators and was standardized in 1866, after which oil became a commodity that could be priced across borders and move through global markets. Internal combustion engines and the auto industry created demand, but the barrel was one of the pieces of infrastructure that let oil become an asset that could be hedged, shorted, and traded in futures markets.

In the age of electrification, the article points to the kilowatt-hour. It says British engineers first proposed the concept in 1889, and that in 1948 the International Electrotechnical Commission, or IEC, formally recognized the common term for the kilowatt-hour for unofficial use. From there, the kilowatt-hour gradually became the standard unit for electricity trading. Power could move between generating plants and households, but dispatch systems, electricity markets, and industrial energy management only became possible once electricity had a standard unit of measurement.

For the information age, the article turns to the bit. It cites Claude Shannon’s 1948 paper, A Mathematical Theory of Communication, saying the mathematical framework of information entropy turned the bit into a unit of measurement and made information something that could be measured, transmitted, and stored with precision. Before the bit, information volume was a vague intuition. After it, telephone line capacity, hard-drive storage, and internet bandwidth could all be priced using a common standard.

Tokens measure machine output in intellectual tasks

The article says those four units share one trait: each measures something that can be reduced to a physical process. The bushel measures material goods, the barrel measures volume, the kilowatt-hour measures energy, and the bit measures the encoding and transmission of information. Even the bit, which appears more abstract, still measures how much encoded signal there is. It does not say what the content means or whether it is useful.

Tokens, in the article’s framing, are different. They do not measure matter, energy, or information. They measure the output of machines processing intellectual tasks. That is why the article describes the shift as “intelligence as a service.” In the past, buying intelligence meant buying people, labor hours, job roles, or projects. Intelligence was tied to a specific person and could not be measured on its own. Now, the article says, buyers are purchasing tokens, tasks, workflows, results, and even tireless digital employees. In that sense, intelligence has been separated from a single human brain and turned into something that can be called continuously, measured, and billed.

The text is careful to add that a token measures the output of an intelligence service, but is not intelligence itself. Using an industrial analogy, it says the token stands in relation to intelligence services somewhat the way raw materials and energy stand in relation to industrial production. The quality and quantity of raw materials and energy shape industrial output. In a similar way, the quality and quantity of tokens become a new gauge for observing a cognitive production line.

The article extends the comparison by contrasting large models with transportation. Transportation solves the problem of moving something that already exists from one place to another. Large models, it says, solve the problem of turning an input into a judgment, a plan, or a result. Steam engines and electric motors turned physical labor into something that could be scaled through energy. Large models are doing something similar for cognitive labor. In this framework, data, context, prompts, tool instructions, and enterprise knowledge bases are the input materials; electricity, chips, and inference time are the energy; and model architecture, routing, caching, and workflow orchestration are the process technology.

Value depends on use, not on production alone

The article says buying intelligence services by volume differs in one basic way from buying oil by volume. The value of oil, wheat, and electricity is determined mainly on the production side. The quality of a bushel of wheat is largely set at harvest. The grade of a barrel of crude can be measured when it comes out of the ground. Tokens are the opposite, the article argues. Their value is determined entirely by the use case.

It gives a simple example. Using the same model to review a merger contract worth millions of dollars and using it to chat about the weather may consume a similar number of tokens, but the economic value created can be vastly different. That difference can only be known after the call is completed. It cannot be labeled or graded in advance.

The article also says the marginal cost of switching a token from one use to another is close to zero. The same model and the same API can move from casual conversation to contract review with nothing more than a different instruction. That creates a market in which the production side can be measured in a unified way, while the output delivered on the consumption side remains highly non-standardized. In the article’s view, this breaks a common assumption in traditional supply-and-demand analysis: that production characteristics and consumption value are coupled.

An economy that is already operating

The article says the token economy is not hypothetical. It is already real, and its growth rate has outpaced almost every earlier commodity at a comparable stage.

It cites OpenAI’s disclosure in March 2026 that its API was processing 15 billion tokens per minute. That figure covers only the API channel and does not include usage on ChatGPT consumer products. Less than half a year earlier, in October 2025, OpenAI had said at its developer event that the figure was 6 billion tokens per minute. The move from 6 billion to 15 billion over five months is described in the article as 1.5 times growth.

The article adds that Anthropic, Google, Meta, and dozens of Chinese model providers are each operating inference infrastructure at different scales. Globally, by early 2026, inference accounted for about two-thirds of all AI compute, up from one-third in 2023.

It says growth in China has been even more dramatic. According to figures disclosed by China’s National Data Administration at the China Development Forum in March 2026, average daily token calls in China were about 100 billion at the start of 2024, rose to 100 trillion by the end of 2025, and exceeded 140 trillion by March 2026. The article says that amounts to growth of more than 1,000 times in two years.

On global spending, the article cites an April 2026 report from International Data Corporation, or IDC, saying worldwide AI infrastructure spending reached $318 billion in 2025, up from $153 billion in 2024. Its conclusion is that tokens are turning into a new factor of production, not in some future scenario but in the present.

The article also quotes Nvidia Chief Executive Jensen Huang from an interview after the company’s GTC conference in March 2026. Huang said, “If an engineer making $500,000 a year consumes less than $250,000 of tokens in a year, I would be very worried.” He also said, “If a top engineer says, ‘I plan to work only with paper and pen,’ that is as absurd as a chip designer refusing to use EDA tools.” The article says the point is not the exact dollar amount, because token prices keep falling and the number will age quickly. The more important point is the emerging benchmark behind it: token consumption is becoming a new dimension for measuring the productivity of knowledge workers.

Prices are collapsing while spending keeps rising

Behind those figures, the article says, sits a structural fact that runs against intuition. Unit prices are falling sharply, but total spending is still surging.

Using GPT-4-level capability as the benchmark, it says inference pricing in early 2023 was about $60 per million output tokens. By 2025, market pricing for the same capability level had dropped to less than $1. That is a decline close to two orders of magnitude. Yet over the same period, global AI infrastructure spending more than doubled, and revenue at leading vendors rose as well.

The article describes this through the Jevons paradox. In the 19th century, higher steam-engine efficiency should have reduced coal consumption in theory, but total coal use rose instead. The article says the token market is now showing the same pattern, only faster and on a larger scale.

Why token economics strains standard business analysis

The more interesting issue, the article says, is not just scale or growth. It is that token economics keeps pushing against existing analytical frameworks.

It offers a corporate example. Imagine a strategy executive at a mid-sized technology company being asked by the chief executive to build an enterprise AI procurement system from scratch. The executive would quickly find that the supplier evaluation, cost analysis, and procurement negotiation methods taught in MBA programs often run into serious limits in token markets.

On the supply side, there are dozens of model providers around the world, and their cost structures differ sharply. The article says cumulative investment behind OpenAI has reached into the hundreds of billions of dollars, while some open-source teams have used only a few million dollars to fine-tune models that approach frontier performance on certain tasks.

The demand side is even harder to manage. A legal department may consume 50 million tokens this month reviewing contracts and then triple that next month. An engineering team may deploy several agents and see token consumption jump tenfold within a week. In advance, the company cannot know which department will generate how much demand or when. Traditional procurement can usually aggregate a relatively stable annual demand plan. Token demand, the article says, is naturally fragmented, volatile, and still expanding.

Pricing is another problem. The simplest competition models assume that under full competition, homogeneous products, and free entry and exit, prices are constrained by marginal cost and long-run average cost. The token market does not fit those assumptions well. Providers differ in unit inference cost, capital expenditure amortization, model capability, and service quality, while capability pricing is falling quickly. A buyer may not even know what level of capability the same budget will purchase next quarter.

Value assessment may be the hardest part. The same 10 million tokens can help lawyers avoid legal risk worth millions of dollars, while the marketing team may use 10 million tokens to draft weekly reports and save only half an hour of labor. A company can calculate return on investment for a single use case, but the article says it is difficult to create one unified measure for all token spending even though a chief financial officer may still ask for an “enterprise token ROI.”

Several assumptions break at the same time

The article groups these tensions into a broader point: tokens depart from several of the simplifying assumptions used in traditional economic analysis, and those departures are happening at the same time and reinforcing one another.

First, the same quantity of model calls can look similar in measurement terms but produce economic value that differs by orders of magnitude once applied to different tasks. Traditional goods also vary by use case, but many quality differences are labeled, graded, and priced on the production side. With tokens, the article says, the buyer can see the model, the context, and the price before the call, but the final value appears only inside the task itself. That pushes the tension between homogeneous measurement and heterogeneous value to an extreme.

Second, costs are falling by more than 90% a year, while physical infrastructure cannot expand at the same pace. Algorithmic efficiency can improve without any change in hardware, but a semiconductor fab takes three to four years from construction to production. The article describes this as two clocks running at very different speeds.

Third, token consumption decisions are shifting in part from humans manually asking questions to agents and workflows triggering calls continuously under a set goal. Once agents start deciding which model to call and how many tokens to consume, demand reacts to price, latency, failure rates, and task value in a way that differs from familiar consumer markets.

The article also says current trade and regulatory frameworks often struggle to classify a token call. A single token-consuming service can include compute output, data processing, and intelligence generation at once, while existing frameworks usually treat those issues separately.

Each of these issues has its own cause, the article says, but the combined effect is much larger than the sum of the parts. Falling unit prices and rising capability speed up the entry of agents into workflows. Continuous execution by agents amplifies demand. Demand spikes expose physical supply bottlenecks. Those bottlenecks then feed back into pricing, model choice, and application thresholds.

The article calls for recalibrated tools

In its closing argument, the article says economics has previously explained pricing, investment, and competition in the spread of general-purpose technologies such as electricity, semiconductors, and the internet. Each new technology pushes existing frameworks to their limits, and new tools eventually emerge. This time, it says, will be no different, but the adjustment may need to be larger.

Electricity made energy tradable over distance. Semiconductors reshaped the cost of computing. The internet changed information distribution. Those were deep changes, the article says, but they usually did not combine three features at once: highly heterogeneous intellectual output under a single unit of measurement, systems acting as continuous calling agents, and two clocks running at different speeds for algorithmic efficiency and physical infrastructure. Tokens compress all of those changes into one market.

The article adds that this does not mean traditional economics is wrong. Supply and demand balance, marginal analysis, and transaction-cost theory still hold within their proper range. But if analysts do not first check the assumptions those tools rely on, and simply carry them over into token markets, they risk misunderstanding how the token economy works. Once the object of analysis shifts from material goods to machine-produced intellectual output, assumptions that once looked stable no longer hold automatically, and the tools built on them need recalibration.

This article is the introduction to Tencent Research Institute’s new book Token Economy. It was published via the WeChat public account “Tencent Research Institute” and is credited to Tencent Research Institute.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
200

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.