Jensen Huang Unveils Token Economics as Nvidia Starts Full Vera Rubin Production

Jensen Huang Unveils Token Economics as Nvidia Starts Full Vera Rubin Production

N
News Editor 01
2026-07-22 13:25:13
At GTC Taipei 2026, Jensen Huang said AI data centers are shifting from hardware sales to monetized compute. Nvidia also said Vera Rubin is now in full production, with Groq-backed inference pushing 1GW annual revenue estimates to $300 billion.
NVIDIAJensen HuangAI Data CentersToken EconomicsVera Rubin

Nvidia CEO Jensen Huang used the GTC Taipei 2026 stage on June 1 to frame AI infrastructure in financial terms, saying “Token is an asset, and token has become a revenue unit for profit.” His message was clear: the business model around AI is moving away from selling GPUs as hardware and toward selling compute output that can be priced directly.

Speaking at the Taipei Music Center during COMPUTEX 2026, Huang laid out the revenue case with a single benchmark. A 1GW AI data center, he said, could move from roughly $30 billion in annual revenue under the Blackwell generation to $300 billion after shifting to Vera Rubin and adding Groq-based decoupled inference. In his description, data centers are no longer just training sites. They are becoming factories that produce tokens.

Token pricing becomes the core of the revenue model

Huang broke the model into five pricing tiers to show how inference output maps to commercial value. The free tier covers basic Q&A and customer service. A lightweight tier is priced at about $5 per million tokens for content generation and summarization. The professional tier reaches $30 per million tokens for coding and data analysis. The enterprise tier rises to $80 per million tokens for compliance and financial modeling. The top tier is set at $150 per million tokens, aimed at scientific research, drug discovery, and real-time inference.

His argument was that once every token can be priced and monetized, AI companies will keep building more compute capacity to generate more tokens. Huang said Taiwan’s demand for AI compute is already rising at what he described as a rocket-like pace.

Vera Rubin enters full-scale production with broad Taiwan supply chain support

On the hardware side, Huang announced that the Vera Rubin architecture has entered full production. He said the supply chain behind it is twice the size of the previous Grace Blackwell generation, with more than 150 Taiwan supply chain partners involved worldwide.

Nvidia’s flagship Vera Rubin NVL72 rack combines 72 Rubin GPUs and 36 Vera CPUs, using a 100% liquid-cooled design for very large AI model deployment. Huang also presented the first roadmap view of the next-generation Feynman architecture, which he said is aimed at pushing inference performance and energy efficiency higher.

Groq joins the inference stack as revenue assumptions expand

Huang also highlighted Nvidia’s coordination with Groq. In his framing, GPUs remain suited for large-scale parallel workloads, while Groq’s LPU targets real-time inference jobs where single-request latency matters most. The presentation said the Groq 3 LPX chip is manufactured by Samsung and is expected to begin shipping in the third quarter.

He used three revenue markers to describe the impact of decoupled inference: about $30 billion a year for a 1GW data center in the Blackwell era, $150 billion with Vera Rubin at the same power level, and up to $300 billion when Vera Rubin is paired with Groq. Near the end of the speech, Huang added that more “surprise” products are planned for the second half of the year.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
200

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.