Z.ai founder Jie Tang says bigger parameter counts no longer tell the full story of model strength

Z.ai founder Jie Tang says bigger parameter counts no longer tell the full story of model strength

N
News Editor
2026-08-19 07:28:20
Jie Tang, founder of Z.ai and a professor at Tsinghua University, argues that asking only how many parameters a model has no longer says much about how strong it is. In his review of the evolution of scaling laws—from GPT-3 to Chinchilla and then Mixture of Experts (MoE)—he says model capability depends on more than parameter count. Training data volume, where compute is spent, and how a model is actually used all matter. Tang’s point is that the old training-first view of scaling is less useful once commercial AI systems are deployed and called billions of times a day. Under that setup, inference cost changes the optimization target. A smaller model trained for longer may make more sense than a larger one trained less efficiently. He cited Llama-2-7B and Gemma-2-9B as examples of models trained far beyond the classic Chinchilla ratio. He also said MoE makes headline parameter numbers even less informative, because total parameters and activated parameters describe different things. For reasoning-heavy workloads, Tang argued that effective depth in a single inference pass and post-training may now be more important scaling dimensions. He described GLM-5.3 as a controlled test of that idea, keeping the base model and parameter counts unchanged from GLM-5.2 while expanding long-horizon environments and reinforcement learning over a month.

When a new AI model is released, one of the first questions the market asks is simple: how many parameters does it have? Jie Tang, founder of Z.ai and a professor at Tsinghua University, says that question on its own has lost much of its meaning.

Tang said the path from GPT-3 to Chinchilla and then Mixture of Experts, or MoE, shows that model capability is shaped by more than parameter count. Training data size matters. So does where compute is spent, and who will use the model under what conditions. As AI shifts from a world of train once, test once to one where models may serve billions of inference calls a day, the best answer under scaling laws keeps changing.

In Tang’s view, once a model reaches the point where it is large enough to "contain world knowledge," adding more total parameters may no longer be the best way to improve it. The next scaling targets may be effective depth in a single inference pass and post-training. He said Z.ai’s latest GLM-5.3 serves as a controlled experiment built around that hypothesis.

From GPT-3 to Chinchilla, the industry followed a trillion-parameter detour

Tang began with the history of scaling laws. In 2020, OpenAI researchers led by Kaplan published Scaling Laws for Neural Language Models. The result was widely read as saying that as compute rises, model parameters should grow faster than training data, at roughly a 2.7:1 rate.

That view had a strong effect on the direction of large-model development. GPT-3, DeepMind Gopher and Microsoft MT-NLG all pushed toward larger parameter counts, and "bigger is better" became a dominant industry instinct.

That understanding was revised in 2022 after DeepMind published the Chinchilla paper. Hoffmann and other researchers reran experiments across about 400 models and found that, if the goal is the best result under fixed training compute, model parameters and training data should scale in a more balanced way. The classic rule of thumb became about 20 training tokens for each parameter.

Under that framework, many of the large models from that period may not have been too small in parameter count. They may instead have been made too large while being fed too little data. Tang said the estimation error in the Kaplan scaling law grows with compute scale, meaning the biggest models of that generation may also have been the ones furthest from the best resource allocation point. Looking back, he described the trillion-parameter race as a detour the whole field took together, then reversed together.

Chinchilla optimized training, not the full lifecycle

Tang said Chinchilla mainly addressed one question: how to optimize training compute. Its core assumption was still close to a train-once-then-evaluate setup. Commercial AI models today are different. They may be called billions of times each day.

Once inference cost is added to the optimization target over a model’s full lifecycle, the preferred configuration changes again. Tang said the more reasonable strategy can become a smaller model trained for longer, which means deliberate over-training.

He cited Meta’s Llama-2-7B and Google’s Gemma-2-9B. Their training loads reached about 290 tokens per parameter, or TPP, and 889 TPP, respectively. Both are far above Chinchilla’s classic 20 TPP benchmark. Tang said this is not simply wasted training compute. In a world where a model is heavily deployed and repeatedly queried, more cost can be paid up front in training to get a smaller and cheaper model at inference time.

MoE changes the meaning of parameter counts again

As MoE models spread, Tang said it becomes even harder to judge a model by total parameters alone. He argued that at least two ideas need to be separated.

  • Total parameters largely determine how much a model can store, including knowledge, facts and long-tail information.
  • Activated parameters and effective depth are closer to what the model can actually compute in one inference pass, and whether it can sustain a longer causal reasoning chain.

For that reason, Tang said directly applying the dense-model rule of 20 tokens per parameter to MoE models may itself be flawed.

He also cited research by Roberts and others in 2025 saying that the best tokens-per-parameter ratio is not a fixed constant and depends on the task. Memory-heavy tasks favor more parameters, while reasoning-heavy tasks benefit more from more data. Later MoE research even found that, under fixed TPP conditions, simply increasing total parameters can reduce reasoning ability. By contrast, raising the number of experts that actually participate in computation can improve reasoning performance more reliably.

That is why the question "how many trillion parameters does this model have?" is becoming less useful as a standalone measure of strength.

Finding vulnerabilities is not the same as retrieving data

Tang used software security flaws to show the difference between knowing a lot and reasoning deeply. If a model is trying to discover a new software vulnerability, he said, the hard part is not memorizing more CVE records. The hard part is whether the model can carry a reasoning chain that may run 20 steps long and still preserve the right context and causal links all the way through.

How much world knowledge a model stores and how deeply it can think through a problem may be two different scaling problems, in his view.

Tang said total parameter count may be especially important only before a certain threshold. A model first has to be large enough to contain the world. After that threshold is crossed, new capability may come more from other scaling dimensions, including effective depth in a single forward pass and post-training.

GLM-5.3 as a controlled experiment

Tang said GLM-5.3 can be viewed as Z.ai’s controlled experiment for this scaling hypothesis. GLM-5.3 uses the same base model, the same architecture, and the same total and activated parameter counts as GLM-5.2.

Instead of enlarging the model again, the team spent one month expanding long-horizon environments and reinforcement learning, or RL. In other words, the experiment intentionally held model size fixed and turned only one knob: post-training scaling.

According to Tang, the resulting gain was not marginal but quite significant. That led him to a broader conclusion. Scaling has not stopped working. What the AI industry may need is a new understanding of what, exactly, is being scaled.

He did not say parameter size no longer matters, nor did he declare pre-training scaling finished. On the contrary, he said there is still room to expand base model size, pre-training data and the compute used in each forward pass, and Z.ai expects to return to those directions as well.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
50

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.