NVIDIA's new Rubin server alleviates the landing difficulties for cloud vendors
Complete. Here is the key summaryNVIDIA's new generation of Vera Rubin server cabinets, due to their use of mature hardware architecture, are expected to have installation difficulties far lower than previous products, potentially solving the deployment challenges faced by cloud vendors in data centers. Currently, this cabinet has been delivered in small batches, and a higher power consumption upgraded version is expected to be launched in the second half of next year, which may bring new challenges in heat dissipation and networking
NVIDIA CEO Jensen Huang showcases the Vera Rubin server cabinet
Author: Liu Feibi
Recently, various cloud service providers have faced repeated obstacles in building data centers, but they have finally received good news: NVIDIA's new generation of chips is unlikely to cause deployment issues.
Executives from several cloud vendors purchasing NVIDIA AI servers have confirmed to me that after the first round of practical tests, they expect the upcoming Vera Rubin server cabinet to be much easier to install than the Grace Blackwell cabinet scheduled for 2025. Last year, NVIDIA's top clients were overwhelmed by a host of new technologies they had never encountered before. The related deployment challenges had previously been kept confidential until my colleague reported on them.
Cloud vendors often invest billions of dollars in hardware, and for every additional week spent on debugging and troubleshooting servers, the launch time for high-revenue AI computing resources is delayed by another week; deployment delays can also lead to late order deliveries from data center contractors and compressed profits.
Currently, the Vera Rubin cabinets have just been delivered in small batches to clients, and everything is still in the early stages. The real large-scale deployment test will come next year when clients will receive large quantities, with single deliveries reaching hundreds or even thousands of cabinets.
NVIDIA plans to launch a more powerful upgraded cabinet in the second half of next year, but this high-end product is likely to bring new troubles: the overall power consumption will significantly increase, and the number of switches and cables required for networking between servers will also rise considerably.
At this stage, the basic version of the cabinet being tested by clients contains 72 GPUs and 36 CPUs, essentially belonging to NVIDIA's third-generation mature products, rather than a completely new architecture developed from scratch. The hardware configuration is highly similar to that of the Blackwell cabinet. (The high-performance version of the Vera Rubin cabinet, which can support interconnections of 144 or 576 GPUs, has not yet gone into production.)
An executive from the cloud industry stated, "To be honest, we believe that large-scale deployment of Vera Rubin will not be too tricky; the mechanical structure and cooling solutions have already iterated to the third generation, and the maturity is very high."
NVIDIA claims that the Vera Rubin system is well compatible with the software that clients currently use for the Blackwell servers. However, two cloud vendor executives admitted that companies will still need to spend several months debugging the software to support the stable operation of large clusters of Vera Rubin cabinets.
The industry jokingly refers to deployment as "mystical operations."
The procurement cost of the Vera Rubin cabinet is at least twice that of the first-generation Blackwell cabinet, but NVIDIA's data shows that measured by tokens per second per watt, the new hardware achieves a tenfold improvement in efficiency Of course, the installation of the Vera Rubin cabinet is still not easy. An executive from a cloud vendor stated that many components of this cabinet are newly designed; the total number of components in the entire machine reaches 1.3 million, while the first-generation Blackwell cabinet had 1.2 million.
NVIDIA Vera Rubin NVL72 computing tray, containing 4 Rubin GPUs and 2 Vera CPUs.
At the NVIDIA GTC developer conference in March this year, Jensen Huang showcased a new network switch for interconnecting Rubin chips, admitting that the development of such hardware is extremely challenging: "It's as difficult as climbing to the sky to get this hardware right; this task itself is highly challenging."
Jensen Huang did not specify whether this new chip interconnection network would create manufacturing obstacles for customers, but it was similar network technologies that slowed down the large-scale deployment of the Blackwell cabinet last year. An employee responsible for large computing clusters (including the upcoming Rubin servers) at Oracle stated that troubleshooting network failures in such clusters is extremely difficult, relying solely on experience, which can be described as "mystical troubleshooting."
The power consumption of the Rubin cabinet per unit has increased by 75% compared to the Blackwell cabinet, and it generates more heat during operation. Interestingly, the cooling system that comes with the new cabinet has lower power consumption than the cooling solution of Blackwell, but the installation and debugging process is more complex. An analyst who has been tracking NVIDIA for a long time revealed that engineering issues related to cooling caused the mass production of the first batch of Vera Rubin servers to be delayed by about two months. An NVIDIA official spokesperson responded: "The product planning remains consistent with the information disclosed previously, with no changes."
In addition, the successful launch of the Rubin servers does not only depend on NVIDIA's hardware itself. Many customers are building dedicated data centers according to NVIDIA's recommended room layout, wiring schemes, and management software to adapt to the efficient deployment of the Rubin cabinet.
Kevin Cochran, Chief Marketing Officer of GPU cloud service provider Vltr, stated that customers must build an entirely new infrastructure from scratch to accommodate Rubin devices, as the cost of renovating existing outdated data centers is too high and not worth the investment.
Ultra High-Performance Version Hides New Challenges
Another major variable is the upcoming launch of the Vera Rubin Ultra high-end cabinet next year. This model includes several new designs, and the adaptation and debugging will require a significant amount of time. Conventional server trays are inserted horizontally into the cabinet, while the Ultra model's chip tray is inserted vertically, similar to books on a bookshelf, which will create a completely different load on the entire system due to gravity.
The Ultra cabinet can interconnect up to 576 GPUs, with a single unit's power consumption approaching three times that of the standard version; these changes may give rise to new deployment challenges.
Andrew Bell, Senior Vice President of Hardware at NVIDIA, stated earlier this month that the Blackwell architecture almost completely reconstructed the entire GPU system, requiring a complete overhaul of both hardware and software, making it extremely difficult for customers to deploy In comparison, Vela Rubin's changes are more moderate, with "mass production manufacturing difficulty significantly reduced." Bell provided data: the assembly pass rate for Rubin chip trays can reach 95%, while the initial assembly success rate for the first-generation Blackwell system is only 20%.
Not everyone agrees with this optimistic assessment. Another cloud vendor executive expressed doubts, believing that the high-performance version of Rubin will still repeat the mistakes of Blackwell, with deployment difficulties being comparable
