Nvidia’s Vera Rubin platform may deliver a sharp jump in efficiency, according to a SemiAnalysis report cited by TMT Breakout on Sept. 15. In pre-release testing, Rubin’s token throughput per unit of power was said to reach as much as roughly 7 times that of Blackwell.
The report said that even if Rubin carries a higher hourly rental cost per chip than Blackwell Ultra, its token output tied to total cost over a three-year lease could still be about 62% higher in an inference scenario running at 80 TPS and P90 latency. It added that the advantage could widen under highly interactive workloads.
System delivery timeline faces a new variable
While performance expectations are rising, the delivery schedule for Rubin Ultra systems may be getting more complicated. The Kyber NVL144 rack, originally planned to carry 144 Rubin Ultra GPUs, could slip by more than 12 months from its original 2027 target because of PCB midplane manufacturing difficulties, moving the timeline to 2028.
A replacement option, NVL72x2, has also reportedly been canceled. The report said the main pressure points are high-density racks, NVLink interconnects, and large-scale system integration.
Cloud expansion pace in 2027 may be constrained
Even if the Rubin chips themselves stay on schedule, cloud providers may still face limits in reaching 144-GPU-scale expansion in 2027. According to the report, AI infrastructure spending may lean more heavily on alternatives such as the existing Oberon architecture.
Nvidia responded to the report by saying only that its product roadmap "remains unchanged." The company has not confirmed a Kyber delay or any broader adjustment to the Rubin Ultra timetable.

