NVIDIA Preemptively Reveals Vera Rubin Performance Data Ahead of Rival's Event

Deep News07-22 08:34

NVIDIA has delivered a strategic blow just before its key competitor's annual product launch, intensively disclosing real-world performance figures for its next-generation Vera Rubin platform and formally revealing the full specifications of its in-house Vera CPU chip.

On Tuesday, NVIDIA disclosed that its Vera CPU delivers nearly double the performance of x86 chips on agentic AI tasks, with a sixfold improvement in latency. Concurrently, early production testing by cloud partner CoreWeave shows that the Vera Rubin NVL72 platform achieves a tenfold increase in token throughput per megawatt when running the DeepSeek R1 model compared to the previous-generation GB200 NVL72 system based on the Blackwell architecture.

The timing of this data release is significant. With Advanced Micro Devices' (AMD) "Advancing AI" annual event scheduled for this Thursday in San Francisco, NVIDIA's concentrated release of performance data is widely seen as a deliberate market-moving tactic. For cloud providers and enterprise customers evaluating next-generation AI infrastructure investments, this data will directly influence their purchasing decisions.

Simultaneously, NVIDIA confirmed that the first shipments of the Vera CPU were completed in June to customers including OpenAI, Anthropic, and Tesla Motors' (TSLA) SpaceX division. This marks NVIDIA's official expansion into the CPU market, posing a direct challenge to the traditional strongholds of AMD and Intel.

Vera CPU Debuts

NVIDIA on Tuesday released the complete specifications, benchmark results, and architectural details for its data center CPU, Vera, providing the critical data needed for potential customer evaluation. The company stated that the Vera chip completed its first deliveries in June to clients such as OpenAI, Anthropic, and SpaceX.

Vera represents NVIDIA's first server CPU designed from the core up, utilizing a custom microarchitecture codenamed "Olympus core," rather than adopting an off-the-shelf design from Arm.

Hannah Coutand, NVIDIA's Vera product marketing lead, explained that the chip's design focuses on single-core speed, high memory bandwidth, and low latency, aiming to "enable the agent to return to the GPU as quickly as possible and keep the GPU highly utilized." In terms of hardware specifications, Vera's power consumption ranges from 250 to 450 watts, with each chip supporting up to 1.5TB of low-power memory.

In the newly released benchmarks, NVIDIA demonstrated that Vera achieves a 1.9x performance gain over x86 chips on agentic AI tasks, with latency reduced by 6x. NVIDIA also stated that Vera outperforms AMD's flagship EPYC Turin CPU by nearly 100% in some industry-standard benchmarks. Previously, NVIDIA had claimed that Vera offers up to 50% better overall performance on AI agent tasks compared to x86 chips.

The launch of Vera is the latest step in NVIDIA's vertical integration strategy. Wolfe Research estimates the average selling price for a Vera chip is around $5,000, with projected shipments this year reaching approximately 1.3 million units. Ian Buck, NVIDIA's vice president of hyperscale computing, stated that agentic AI makes the CPU "more indispensable than ever," predicting that the total server CPU market could eventually reach $200 billion.

Vera Rubin Real-World Testing

CoreWeave completed the industry's first deployment and validation of Vera Rubin NVL72 in early June, covering full-stack confirmation of power, cooling, networking, and compute. The data disclosed represents the first publicly available real-world performance results for Vera Rubin NVL72 silicon.

Using the DeepSeek R1 model as a benchmark and targeting the same interactive response, Vera Rubin NVL72 generated ten times the number of tokens per megawatt per second compared to the GB200 NVL72. CoreWeave also noted that optimizations made for Vera Rubin can be backported to GB200 NVL72 systems, improving their throughput per megawatt by more than fourfold within three months. NVIDIA stated these results were validated through testing of over 250,000 independent configurations and more than 1.4 million GPU-hours.

From an architectural perspective, the Vera Rubin NVL72 rack integrates 72 Rubin GPUs with 36 Vera CPUs, connected via a 260 TB/s fully interconnected NVLink 6 fabric with native NVFP4 precision support. CoreWeave stated this efficiency gain means customers can process more inference traffic within the same power budget or run the same workload with less electricity, thereby reducing cost per token.

CoreWeave also disclosed specific application scenarios: a global cybersecurity company anticipates running threat detection inference with 10x performance per watt; an autonomous programming agent company expects to scale agent tasks with significantly lower token costs; and an AI-native search engine plans to expand real-time search services to more users without exceeding response time limits.

The Power of Co-optimization

NVIDIA emphasized that these performance gains do not rely on the chips alone but are the result of a co-design strategy integrating hardware and software. Through dynamic optimization of the complete infrastructure and energy stack, NVIDIA stated it can deploy up to 40% more GPUs within the same power envelope.

In cooling, NVIDIA employs a 45°C closed-loop liquid cooling system, which saves approximately 4 million gallons of water per megawatt annually compared to standard cooling methods. In networking, the sixth-generation NVLink 6 interconnect architecture delivers 2.3x higher large language model simulated decode throughput compared to Ethernet architectures; the Spectrum-X platform achieves 1.6x faster remote direct memory access bandwidth while reducing switch count by 1.7x, improving optical power efficiency fivefold and reliability tenfold.

NVIDIA's latest-generation Spectrum-6 platform has begun shipping to AI factory customers including CoreWeave, Microsoft (MSFT), Nebius B.V., SpaceXAI Corp., and Tesla Motors (TSLA). Currently, NVIDIA is accelerating shipments to customers and partners including Alphabet's (GOOG) Google Cloud, Microsoft (MSFT) Azure, Meta Platforms, Inc. (META), Oracle Cloud Infrastructure, Dell Technologies, OpenAI, and CoreWeave.

NVIDIA as a CPU Challenger

Despite Vera's impressive performance data, NVIDIA still faces significant market share challenges in the CPU arena. Gartner analyst Kevin Knox noted that AMD is currently the primary competitor in the enterprise AI server CPU space. Reports indicate Intel holds about 66.8% of the server CPU market, with AMD at approximately 33%, and AMD continues to gain share while having established deep partnerships with hyperscale cloud providers.

Buoyed by expectations of rising CPU demand driven by the emergence of agentic AI, AMD and Intel have seen their stock prices surge 128% and 149% year-to-date, respectively, far outpacing NVIDIA's (NVDA) roughly 8% gain over the same period, making them among the best-performing chip stocks of the year.

For now, Oracle is the only major cloud provider listed among NVIDIA's primary cloud service partner announcements for Vera, with others notably absent. Hannah Coutand acknowledged that Vera is still in an "early adopter phase," though OpenAI plans to begin large-scale deployment of Vera chips as early as this quarter.

Karl Freund, founder of Cambrian AI Research, summarized NVIDIA's strategic intent: "Their goal is to wean customers off of Intel or AMD CPUs, and they are determined to capture that revenue. They are choosing to focus on building a unique CPU that no one else in the market offers today."

Disclaimer: Investing carries risk. This is not financial advice. The above content should not be regarded as an offer, recommendation, or solicitation on acquiring or disposing of any financial products, any associated discussions, comments, or posts by author or other users should not be considered as such either. It is solely for general information purpose only, which does not consider your own investment objectives, financial situations or needs. TTM assumes no responsibility or warranty for the accuracy and completeness of the information, investors should do their own research and may seek professional advice before investing.

Comments

We need your insight to fill this gap
Leave a comment