
Analyst Ming-Chi Kuo stated yesterday that his latest industry survey shows NVIDIA has restarted the artificial intelligence inference prefilling acceleration GPU project "Rubin CPX." However, its design has undergone major changes, and production is expected to begin in 2027Q1.

At the chip level, the new Rubin CPX delivers computing power close to that of the standard Rubin GPU and further improves prefilling performance. Its maximum power consumption per unit is also 2300W, and it is equipped with 168GB of HBM4 memory (note: the old Rubin CPX had 128GB of GDDR7; 168GB may indicate a 7×24GB design).
At the tray and rack levels, 8 new Rubin CPX chips form 1 compute tray, while 8 compute trays form 1 rack module. Spectrum-6 all-copper Ethernet is used for horizontal expansion (scale-out) within the module. The new Rubin CPX uses dedicated MGX ETL racks, with each rack containing 1~4 rack modules. Horizontal expansion between rack modules uses a Spectrum-6 + OSFP optical solution.
At the system level, NVIDIA recommends pairing Rubin CPX with standard Rubin at a 1:1 ratio. Ethernet-based RDMA connectivity is used between Rubin CPX ETL racks and Vera Rubin NVL72 racks.
The analyst commented that more than 50% of current AI inference workloads involve processing inputs and building the corresponding KV Cache. The new Rubin CPX can handle prefilling workloads with greater deployment flexibility and lower costs. A new Rubin CPX compute tray provides 1.34TB of HBM memory, enough to meet the prefilling and KV Cache creation needs of most long-context workloads.
