WhiteLine, produced by the WuBlockchain team, used its latest episode to focus on a single question after Kimi K3 opened its model weights: if models move closer to a "free business," who still captures the profit pool in AI? Host Minta’s core takeaway was blunt. Model weights may be free, but inference and distribution at scale are not. The companies most likely to bill for the next trillion tokens, the program said, are infrastructure providers that control GPUs, cloud platforms, and inference access points.
Cheaper API pricing does not automatically mean lower task costs
According to the episode, Kimi K3 is priced at $3 per million input tokens and $15 per million output tokens under its official API schedule. That is about 40% to 70% lower than some leading closed-source U.S. models.
But the program argued that enterprises should not compare models by token sticker price alone. What matters is the total cost of completing the same task. Long-chain reasoning, multi-turn conversations, and tool use can all increase token consumption. Failed runs add retry costs, and in some cases human intervention becomes necessary. On that basis, a cheaper API can still produce a higher overall operating cost for the buyer.
Free weights still leave production inference expensive
WhiteLine said Kimi K3 has 2.8 trillion parameters. Even with open weights, production-grade deployment still calls for large amounts of GPU capacity, high-speed networking, and specialized engineering teams. Those costs do not disappear because the model can be accessed more openly.
The program framed the self-hosting decision around utilization. Running inference in-house only saves money if GPU usage stays high enough. If request volume is too low, idle hardware can cost more than buying API access directly. That is why the main winners from free weights may be cloud vendors and inference platforms with steady traffic rather than smaller users with uneven demand. In WhiteLine’s view, model supply may become more fragmented while production inference remains concentrated on large platforms.
Open weights do not mean the end of commercial charging
The episode also broke down the Kimi K3 licensing structure. It said the license is relatively open for ordinary developers. For MaaS platforms and large commercial products that cross revenue or user thresholds, however, Moonshot still keeps rights tied to separate agreements and brand display.
WhiteLine’s reading of that setup is that Moonshot is using open weights to expand the ecosystem first, while preserving room to negotiate with large-scale distributors later. In other words, it is giving up broad licensing fees, not its bargaining power once scale shows up.
Open and closed models may settle into a layered market
The program said open models are taking on more high-frequency and price-sensitive workloads, but token share should not be confused with profit share. Lower pricing, longer reasoning chains, and Agent calls can all inflate token volume without producing income in the same proportion.
That could push enterprises toward hybrid architectures. In the scenario outlined by the episode, standard tasks would go to open models, complex and high-risk work would stay with closed models, and sensitive data would remain on local systems. Another possible workflow is escalation: a lower-cost model handles the first attempt, and a stronger model is called only if the cheaper one fails.
Whether open models drive the next chip cycle depends on token growth
The final part of the discussion turned to semiconductors. WhiteLine argued that lower prices and a broader user base for open models, combined with higher token consumption from Agent-style requests, could translate into more inference demand.
Most developers will not buy GPUs themselves, the program said. They will access models through cloud vendors and inference platforms via APIs. In that sense, selling tokens is effectively retailing GPU compute by usage. If token consumption rises fast enough, cloud providers can lift GPU utilization, shorten hardware payback periods, and keep buying chips.
There is a counterweight. Open competition can keep pushing token prices lower, which would pressure platform margins. So the question of whether open source can trigger chip FOMO comes down to one variable in the episode’s framework: can token usage keep growing fast enough, and for long enough, to outpace the decline in token pricing?

