Renowned industry analyst Ming-Chi Kuo has reported that just as the market began accepting the notion that the Rubin CPX project had been removed from NVIDIA's product roadmap, his latest supply chain checks indicate the company has greenlit the project once more. Production is now slated to commence in the first quarter of 2027, marking a significant strategic pivot.
The revived Rubin CPX design boasts enhanced prefill performance compared to its predecessor, incorporating major overhauls to both the GPU specifications and rack architecture. This shift underscores NVIDIA's intensified focus on prefill solutions, a critical component for handling long-context AI workloads.
In terms of GPU specifications, CPX delivers computational performance approaching that of the standard Rubin, while matching the 2,300-watt maximum power envelope per GPU. However, the memory configuration has been adjusted to 168GB of HBM4, a notable deviation from both Rubin's 288GB HBM4 allocation and the original CPX design which called for 128GB of GDDR7.
The rack architecture has similarly undergone substantial changes. The new CPX iteration utilizes a standalone MGX ETL rack setup, abandoning the previous plan to share a chassis with Rubin. Customers can now select configurations ranging from 64 to 256 CPX GPUs, with each rack housing modular blocks of 64 GPUs. These modules consist of eight compute trays, each containing eight GPUs, paired with a single switching tray.
Connectivity within this new architecture reveals a clear division of labor. NVLink is reserved exclusively for scale-up communication among the eight GPUs within each tray, delivering 1–1.5TB/s per GPU bandwidth—a deliberate reduction from Rubin's 3.6TB/s. For scale-out between trays within a module, NVIDIA employs Spectrum-6 Ethernet with all-copper L1 links. Inter-module expansion is handled via OSFP optical connections routed through each module's Spectrum-6 switches.
The operational framework positions CPX as a complementary component to the Vera Rubin NVL72 platform. NVIDIA recommends a 1:1 pairing of CPX with Rubin GPUs, where CPX is responsible for prefill operations and KV cache construction. This pre-computed cache is then transferred to Rubin via Ethernet RDMA for subsequent decode tasks.
Strategically, this product addresses a growing market need: over 50% of AI inference workloads now involve processing input context and building corresponding KV caches. CPX offers a more flexible and cost-effective prefill solution for these scenarios. Each eight-GPU tray carries approximately 1.34TB of HBM4, sufficient for most long-context prefill workloads and their associated cache requirements.
Comments