NVIDIA is extending the lifecycle of its Blackwell platform through continuous software stack optimization, while simultaneously preparing for the large-scale deployment of its next-generation Rubin architecture.
Against the backdrop of accelerating the rollout of the new Rubin platform, NVIDIA has disclosed that it managed to quadruple the computational throughput per megawatt (TPS/MW) for the GB200 NVL72 when running the DeepSeek R1 0528 model, achieving this in just three months. This achievement stems from 38 major optimizations completed within a four-month period, supported by over 250,000 simulated configuration tests and a cumulative 1.4 million GPU-hours of optimization validation.
Concurrently, the GB300 "Blackwell Ultra" platform continues to set new records in AI training benchmarks. At a 256-card scale, the GB300 NVL72 set a historical high score of 1,648 TFLOPs per GPU on a DeepSeek-V3 671B pre-training task, representing an approximately 3x improvement over the GB200's 606 TFLOPs. This metric has cumulatively improved by 1.5 times over the past six months.
Substantial Energy Efficiency Leap for GB200, Optimizations Apply Broadly
The company stated that a series of optimizations for the GB200 NVL72 platform resulted in its computational throughput per megawatt growing by a factor of four within three months when operating the DeepSeek R1 0528 model.
This enhancement is underpinned by 38 significant optimization iterations completed by NVIDIA over four months. These optimizations were filtered through more than 250,000 simulated configurations and consumed a total of 1.4 million GPU-hours for practical validation. The company specifically noted that over 90% of these optimization results are transferable and applicable to other models within its AI product portfolio, rather than being custom-tailored for a single task.
This implies that existing data center clients can achieve significant energy efficiency gains through software upgrades without needing to replace hardware—a direct economic benefit for hyperscale cloud providers and enterprise customers highly sensitive to computational costs.
GB300 Sets New Training Record, Performance Up 1.5x in Six Months
On the AI training front, the GB300 NVL72 also demonstrates strong performance momentum. At a 256-card scale, during pre-training of the DeepSeek-V3 671B model using the Megatron Core framework, the GB300 NVL72 achieved a throughput of 1,648 TFLOPs per GPU. This is roughly three times higher than the previous GB200's 606 TFLOPs and sets a global record for this task.
Notably, this figure is not a static hardware limit. The performance of the GB300 NVL72 under the Megatron Core framework has grown from 1,088 TFLOPs/GPU in November 2025 to 1,648 TFLOPs/GPU in June 2026, a cumulative improvement of about 1.5 times over six months.
In terms of collaborative optimization with mainstream AI frameworks, NVIDIA's deep cooperation with the PyTorch and JAX communities has also yielded significant gains. On TorchTitan (the PyTorch-native training stack), the training performance of GB300 NVL72 for DeepSeek-V3 671B improved by 6 times compared to an unoptimized baseline configuration, jumping from 199 TFLOPs/GPU to 1,197 TFLOPs/GPU. The improvement under the JAX framework is even more pronounced; as of July 2026, per-GPU throughput reached 4,082 Tokens/s, corresponding to 1,025 TFLOPs/GPU, representing an approximately 10x improvement over the 418 Tokens/s recorded in January 2026.
Scaling Efficiency Nears Theoretical Limit, 800 Gb/s Network is Key
Scaling efficiency in large-scale training scenarios has long been a core metric for measuring the practicality of AI infrastructure. NVIDIA disclosed that across the scaling range from 256 to 1,024 cards, the GB300 NVL72 maintained scaling efficiency close to the theoretical limit across all three major frameworks: 98.5% for Megatron Core, and 97% for both TorchTitan and JAX.
The company attributes this scaling performance to the 800 Gb/s Scale-Out network chip built into the NVL72 rack. High-speed interconnectivity directly determines communication overhead in multi-machine training, thereby affecting overall scaling efficiency. This networking capability is viewed as the core infrastructure component supporting the high-efficiency performance observed across these frameworks.
Rubin Platform Deployment Accelerates, Blackwell Optimization Continues
These optimization advancements are being announced as NVIDIA's next-generation Vera Rubin platform enters global deployment. Reports indicate that the Vera Rubin NVL72 offers approximately a 10x improvement in token throughput compared to Blackwell, with the GB200 NVL72 achieving about 80,000 Tokens/s, while the Vera Rubin NVL72 can reach 800,000 Tokens/s under the same 150MW power consumption.
However, the Blackwell platform is already deployed in countless data centers worldwide. NVIDIA's strategy mirrors that of the previous Hopper generation—continuing to unlock the potential of deployed hardware through software optimization even as the new platform advances. This approach not only extends the return on investment cycle for existing clients' hardware but also strengthens the platform stickiness of NVIDIA within the AI infrastructure ecosystem. For the market, NVIDIA, with its dual-track advancement of Blackwell and Rubin, is progressively building a software moat that competitors will find difficult to surmount in the short term.
Comments