NVIDIA Considers Aggressive Strategy: Reducing Memory Capacity on Rubin Ultra Chip

Deep News08-07 18:41

Due to a shortage of advanced high-bandwidth memory chips, NVIDIA is exploring an aggressive strategy: equipping its next-generation GPU, the Rubin Ultra, with less memory than originally planned. According to three sources directly involved in testing, over the past few weeks, NVIDIA has tested at least three versions of the Rubin Ultra GPU, with some models featuring memory capacities below the initially disclosed specifications. The company's consideration of a lower-memory version is partly driven by the potential inability to secure enough high-end memory to support the original design. Reduced memory could impact chip performance, though NVIDIA may mitigate this through other technical means. However, AI companies running large AI models on this downgraded chip would likely need to deploy more units.

NVIDIA's revised plan reflects the AI boom's strain on chip production capacity, which struggles to keep pace with explosive data center demand. The accompanying GPU memory chips remain in short supply, with prices continually rising. This cost escalation has triggered a chain reaction across the tech industry: companies are exceeding hardware budgets, and manufacturers like Apple have raised terminal device prices. NVIDIA has declined to comment on the matter. The situation carries a certain irony: the recent surge in NVIDIA GPU sales has generated massive demand for memory, creating supply constraints, and now NVIDIA itself must adjust its memory configuration strategy. Currently, the three Rubin Ultra samples under testing have memory specifications downgraded compared to the Rubin chips already in mass production and delivered to customers.

Two NVIDIA clients indicated that the reduced memory may not significantly weaken market demand for the Rubin Ultra. One client noted that the lower-memory version chip will likely be priced below the original design, offering a cost-effective option for companies to control expenses. Another client stated that while large memory is crucial for running top-tier frontier models, their company is less concerned with individual chip memory parameters and more focused on long-term partnerships across multiple NVIDIA hardware generations. NVIDIA can also offset the memory reduction through server system design. By leveraging higher-speed interconnects and more efficient data storage solutions, clients can split large models and computing tasks across multiple Rubin Ultra chips for collaborative operation.

Planning for a Shift

The testing of downgraded memory specifications contrasts with NVIDIA's earlier optimistic statements. In mid-July, NVIDIA's Senior Vice President of Hardware Engineering, Andrew Bell, stated at a media briefing that the company has a dedicated team forecasting and resolving supply chain issues years in advance. "We are proactively addressing memory supply challenges and will not face capacity constraints in the short term," Bell said. "Of course, there are global price pressures, and cost increases may be a bigger issue; but on the supply side, we are secure."

An Adjustment Window Remains

Three sources, along with one NVIDIA client, confirmed that the final hardware specifications for the Rubin Ultra have not yet been set. The chip is expected to begin delivery at the end of next year, leaving NVIDIA ample time to adjust the design based on memory supply, costs, and client demand. Industry media outlet SemiAnalysis first reported the potential memory specification downgrade to clients last week. The pricing for the Rubin Ultra has not been determined. Data from Epoch AI shows that high-bandwidth memory accounts for over half of the total component cost of high-end AI chips; a memory-reduced version could be significantly cheaper while still being suitable for most AI applications. NVIDIA CEO Jensen Huang first unveiled the Rubin Ultra at the 2025 NVIDIA GTC conference. He introduced that a single GPU would feature 1TB of HBM4E memory, a next-generation high-bandwidth memory for AI chips offering improved data throughput, storage capacity, and energy efficiency compared to the previous HBM4. Presentation slides indicated that this 1TB memory would consist of 16 memory stacks.

Sources revealed that NVIDIA's current test versions have reduced the number of memory stack layers and per-die memory capacity, with some samples even using the older HBM4. The lowest total memory in test prototypes is only 192GB, with another version at 256GB, far below the originally planned 1TB. For context, NVIDIA's current Vera Rubin chip features a maximum of 288GB of HBM4 memory. Industry insiders explain that reducing memory per chip allows for the allocation of limited memory resources to more GPUs, helping NVIDIA stabilize total chip output.

Manufacturing Challenges

Sources say a core reason for NVIDIA's consideration of memory downgrades is the difficulty memory manufacturers face in producing sufficient HBM4E memory to meet the Rubin Ultra's production timeline. HBM4E offers higher performance but requires denser memory cells and faster circuit interconnects, significantly increasing the complexity of integration during chip packaging. NVIDIA has already taken steps to ease memory supply pressure. Late last month, NVIDIA reached a $500 billion cooperation agreement with the parent company of memory giant SK Hynix to jointly develop various specifications of next-generation HBM4 and HBM4E memory technologies. At a July media briefing, NVIDIA's Vice President of Global AI Cloud and Infrastructure, Raj Mirpuri, stated that this collaboration aims to "secure a stable supply of high-bandwidth memory." He added that SK Hynix will invest in expanding high-bandwidth memory production capacity, prioritizing supply to NVIDIA. SK Hynix had already announced in June plans to double its memory chip production capacity over the next five years.

Disclaimer: Investing carries risk. This is not financial advice. The above content should not be regarded as an offer, recommendation, or solicitation on acquiring or disposing of any financial products, any associated discussions, comments, or posts by author or other users should not be considered as such either. It is solely for general information purpose only, which does not consider your own investment objectives, financial situations or needs. TTM assumes no responsibility or warranty for the accuracy and completeness of the information, investors should do their own research and may seek professional advice before investing.

Comments

We need your insight to fill this gap
Leave a comment