NVIDIA posts first Vera Rubin NVL72 silicon results, showing up to 30x higher agent throughput per megawatt

NVIDIA posts first Vera Rubin NVL72 silicon results, showing up to 30x higher agent throughput per megawatt

N
News Editor
2026-08-25 02:02:15
NVIDIA has released its first silicon test results for the Vera Rubin NVL72 rack, using DeepSeek-V4-Pro in a coding-focused agent workload rather than a conventional fixed-length inference benchmark. According to the company, Vera Rubin NVL72 delivered up to 30x higher throughput per megawatt than GB300 NVL72 under the AgentX benchmark, while the cost to produce one million tokens fell by as much as 35x. The test framework came from SemiAnalysis and was built to replay full coding sessions with growing context, tool use, and sub-agent generation, reflecting how agent systems behave in production. The company also disclosed that Groq 3 LPX has entered full production as a low-latency inference extension for the Vera Rubin platform. In an Artificial Analysis benchmark, Groq 3 LPX ran Gemma 4 31B with a 100,000-token context window at 3,400 tokens per second, and its peak output in coding workloads reached 6,981 tokens per second. NVIDIA said Vera CPU, a processor built for agent workloads with 88 Olympus cores and 1.2 TB/s of LPDDR5X bandwidth, is also moving into deployment, with SpaceXAI named among early adopters.

NVIDIA has published its first silicon test data for the next-generation Vera Rubin NVL72 rack and used DeepSeek-V4-Pro to run a real agentic coding workload. In the company’s numbers, Vera Rubin NVL72 delivered up to 30x more throughput per megawatt than GB300 NVL72 under the AgentX benchmark, while the cost of producing one million tokens fell by as much as 35x.

NVIDIA posts first Vera Rubin NVL72 silicon results, showing up to 30x higher agent throughput per megawatt 2

Instead of relying on a standard fixed-length inference benchmark, NVIDIA used AgentX from SemiAnalysis. The benchmark replays full coding sessions with growing context, tool calls, and sub-agent creation, aiming to measure a complete agent workflow rather than a single inference request.

The report cited recent OpenRouter data saying that a real-world agent AI task consumes 15 times as many tokens as a standard chat interaction. It gave an example in which an agent researches a company for an investment decision. In that process, the main agent and sub-agents keep reasoning, and the accumulated tokens become input for the next step, making long-context handling a central bottleneck.

On that basis, NVIDIA argued that older AI benchmarks built around fixed sequence lengths no longer capture performance in agent settings and that measurement needs to shift toward end-to-end workflows.

Vera Rubin NVL72 versus GB300 NVL72

For the silicon test, NVIDIA used DeepSeek-V4-Pro (1.6T). Under the AgentX workload, Vera Rubin NVL72 posted up to 30x higher throughput per megawatt than GB300 NVL72.

NVIDIA posts first Vera Rubin NVL72 silicon results, showing up to 30x higher agent throughput per megawatt 3

As a reference point, the report said GB300 NVL72 had already delivered 15x higher throughput per megawatt than H200 in the same DeepSeek-V4-Pro test. It also described the broader progression this way: from H200 to GB300, throughput on real agent workloads rose by as much as 30x while cost roughly doubled; from GB300 to Vera Rubin, throughput climbed by as much as another 30x and rack price roughly doubled again.

The article framed that gain in practical terms for AI factories that are constrained by available power. With the same power budget, Vera Rubin can process far more agent work. For megawatt-scale and gigawatt-scale data centers, higher throughput per megawatt translates into materially better compute utilization.

NVIDIA also said its DSX MaxLPS technology can manage power at the GPU, rack, and workload levels, allowing up to 40% more GPUs to be configured within the same megawatt budget and lifting throughput per megawatt even more.

Token cost drops as throughput rises

The increase in throughput per megawatt flows directly into inference economics. NVIDIA said Vera Rubin NVL72 can cut the cost of generating one million tokens by as much as 35x compared with GB300 NVL72.

NVIDIA posts first Vera Rubin NVL72 silicon results, showing up to 30x higher agent throughput per megawatt 4

The source article tied that change to the economics of agent applications, arguing that sharply lower inference costs would make always-on digital workers and consumer-facing super apps more realistic. It did not provide additional commercial figures beyond the cost comparison.

The gain comes from system-level co-design

The report said the performance jump in Vera Rubin NVL72 does not come from process advances in a single chip alone. It pointed instead to what it called extreme co-design across the stack. The techniques listed in the article included disaggregated serving, distributed KV cache, KV-aware routing, and MegaMoE.

On the model side, NVFP4 quantization compresses model weights to 4-bit precision to reduce memory use and raise throughput. On the interconnect side, sixth-generation NVLink combined with large-model MoE delivers a network that the article described as 10x faster than existing Ethernet and with 3x lower latency. That lets MoE-based models such as DeepSeek route among expert subnetworks across 72 GPUs.

The article’s conclusion on this point was straightforward: NVIDIA is no longer selling a graphics card alone. It is selling an AI factory system.

NVIDIA posts first Vera Rubin NVL72 silicon results, showing up to 30x higher agent throughput per megawatt 5

Groq 3 LPX enters full production

Alongside Vera Rubin NVL72, NVIDIA said at Hot Chips 2026 that Groq 3 LPX has entered full production. The report described Groq 3 LPX as a dedicated extension for the Vera Rubin NVL72 data-center platform, built for low-latency inference acceleration.

According to the article, the technology came from NVIDIA’s $20 billion acquisition of startup Groq in December last year. The design separates large-scale context processing from token generation: Rubin GPUs handle the heavy context work, while Groq 3 LPX handles very fast token output.

In a full rack-scale deployment, as many as 256 LP30 accelerators can work alongside GPUs over an ultra-high-bandwidth interconnect to build an enterprise inference engine.

In an Artificial Analysis benchmark, Groq 3 LPX ran Gemma 4 31B with a 100,000-token context window at 3,400 tokens per second. The article said that made it 4x faster than competing products in latency-sensitive workloads.

NVIDIA posts first Vera Rubin NVL72 silicon results, showing up to 30x higher agent throughput per megawatt 6

It also gave a specific comparison for a 5,000-token generation task: the time required fell from 50 seconds to 1.5 seconds, a 34x gap. In coding workloads, Groq 3 LPX reached a peak output speed of 6,981 token/s.

Jensen Huang said in the source article that LPX pushes ultra-fast token generation and marks 「又一次巨大飞跃」 in AI throughput and responsiveness.

The piece also said Nebius has deployed the chip in Nebius Token Factory to improve responsiveness at each step of the agent loop.

Vera CPU moves into deployment

NVIDIA also introduced Vera CPU as a processor designed for agent workloads. It carries 88 in-house Olympus cores and uses LPDDR5X memory with 1.2 TB/s of bandwidth.

NVIDIA posts first Vera Rubin NVL72 silicon results, showing up to 30x higher agent throughput per megawatt 7

At Hot Chips, an NVIDIA Vera CPU executive said, 「Agentic AI is the most complex computing task in history」, according to the article. The report explained that a single agent task can involve hundreds of steps, including tool calls, Python execution, context retrieval, and large-scale data handling, all of which place substantial scheduling and orchestration load on the CPU.

The article said SpaceXAI has announced a formal large-scale deployment of NVIDIA Vera CPU. It also said NVIDIA had quietly provided the first Vera samples in May to A社, OpenAI, and SpaceXAI for early testing. Three months later, SpaceXAI had moved to full deployment.

The report added that SpaceXAI is building a gigawatt-scale compute factory on the Vera Rubin platform to power Grok. Musk said, 「两家强强联手,设计了一款优化版的Vera Rubin NVL72,将在2028年大规模部署」.

It also said SpaceXAI’s first-generation Starmind AI satellite is planned to use the Vera Rubin NVL72 rack-scale system.

From GPU vendor to full agent infrastructure stack

On the same day, the article described both Vera CPU and Groq 3 LPX as entering full-scale production. Its broader point was that NVIDIA’s role is expanding. In the large-model cycle, the company relied on GPUs to dominate training and inference. In the agent cycle, it is extending that footprint across CPUs, inference accelerators, and NVLink.

The source ended with a simple formulation: NVIDIA is aiming to sell not just a GPU, but a full AI factory capable of producing tokens continuously.

References cited in the source article

  • https://x.com/MinLiBuilds/status/2091915873661686204
  • https://developer.nvidia.com/blog/nvidia-vera-rubin-and-blackwell-set-a-new-standard-for-agentic-ai-performance-per-watt/
  • https://developer.nvidia.com/blog/inside-nvidia-groq-3-lpx-the-low-latency-inference-accelerator-for-the-nvidia-vera-rubin-platform/
  • https://developer.nvidia.com/blog/maximizing-ai-factory-performance-per-watt-with-nvidia-dsx-maxlps/

The article was originally published by the WeChat account 新智元 and credited to ASI启示录, then republished by MarsBit.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
70

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.