Report: NVIDIA Plans to Reduce Rubin Ultra Memory Configuration to Address High-End HBM Shortage
Complete. Here is the key summaryThis means customers may need to deploy more GPUs in the future to achieve equivalent computing power, potentially driving up data center construction costs for tech giants like Microsoft, Meta, and Google
The AI infrastructure boom is pushing supply chains to their limits.
According to a report by The Information, NVIDIA is evaluating a rather aggressive adjustment—launching a version of its next-generation AI GPU, Rubin Ultra, with lower memory configurations than originally planned, to alleviate production pressures caused by the shortage of high-end high-bandwidth memory (HBM).
This indicates that even NVIDIA, which dominates the GPU market, must seek a balance between product specifications and supply capabilities.
This change not only reflects that HBM supply has become one of the most critical bottlenecks in the AI industry chain but also suggests that AI server costs may rise further, with pressure on data center construction continuing to transmit downstream.
On Thursday, NVIDIA's stock closed down 0.1%, while SK Hynix's stock fell 4.97%.


Tight HBM Supply Leads NVIDIA to Test Multiple "Reduced Memory" Versions
Sources familiar with the matter revealed that over the past few weeks, NVIDIA has tested at least three different versions of the Rubin Ultra GPU, some of which feature memory capacities lower than initially planned.
A key reason NVIDIA is considering a lower-memory version is the potential inability to secure sufficient quantities of high-end HBM chips to support mass production of the original design.
Rubin Ultra is a crucial component of NVIDIA's next-generation AI computing platform, positioned above the soon-to-be-mass-produced Rubin series, and is regarded as a vital hardware platform for training ultra-large-scale AI models in the future. Under the original plan, it was set to feature HBM with higher capacity and bandwidth to further enhance model training and inference performance.
However, against the backdrop of continued tight HBM supply, NVIDIA is re-evaluating product configurations, hoping to ensure the product launches on schedule by adjusting memory specifications rather than waiting for supply chain capacity expansions.
Reduced Memory Means More GPUs Need to Be Deployed
For AI training, HBM not only determines a GPU's data throughput capability but also directly impacts the scale of large models that can be accommodated.
If Rubin Ultra ultimately adopts a lower memory configuration, customers may need to deploy more GPUs to complete computational tasks that could previously be handled by fewer chips when running large AI workloads such as large language models.
Although NVIDIA can partially offset performance losses by increasing GPU computing power and optimizing interconnect bandwidth, overall, system deployment costs and cluster complexity are likely to increase.
For major cloud providers like Microsoft, Meta, Amazon, and Google, which are continuously expanding their AI capital expenditures, this means data center construction costs could rise further in the future.
AI Boom Exposes the Industry's Biggest Bottleneck
HBM has become one of the most scarce core components in the current AI industry chain.
In recent years, the rapid growth in demand for large model training has driven continuous increases in GPU shipments. However, the HBM required for GPUs is supplied by only a few manufacturers, including Samsung Electronics, SK Hynix, and Micron. Due to the complex manufacturing process, high yield requirements, and long expansion cycles for HBM, supply growth has consistently struggled to keep pace with AI demand.
Industry consensus expects HBM to remain in short supply for the next few years, which is one of the important reasons why AI server prices remain high.
NVIDIA's active consideration of reducing the memory configuration for Rubin Ultra also reflects that even with the strongest bargaining power in the industry chain, its product planning remains constrained by HBM supply.
Cost Pressure Spreads Across the Entire Tech Industry
The impact of rising HBM prices is no longer limited to AI servers.
As key components like GPUs and HBM remain in short supply, hardware costs across the entire tech industry are being pushed up. Corporate budgets for purchasing AI infrastructure are increasing, forcing more companies to raise capital expenditures to meet generative AI deployment needs.
Meanwhile, rising upstream chip costs are beginning to transmit to the consumer electronics sector. Some hardware manufacturers, including Apple, have already absorbed supply chain cost pressures by raising prices on end products.
Market participants believe that as long as demand for AI computing power continues to grow rapidly while the release of new HBM capacity remains limited, competition for high-end memory resources will persist, and HBM supply capacity will continue to be a key factor determining the speed of AI industry expansion.
