NVIDIA VP Ian Buck on the Evolution of Compute Power for the Agentic AI Era: Insights from the Vera Rubin Platform

Deep News09-20 14:11

NVIDIA Vice President Ian Buck recently shared his perspective on how infrastructure must evolve to meet the demands of the Agentic AI era, marking a fundamental shift from the human-in-the-loop interactions that defined earlier AI workloads. Speaking about the Vera Rubin platform, Buck detailed a full-stack architecture designed not just for raw performance, but for maximizing the revenue-generating capacity of AI factories. His remarks highlight a clear departure from traditional IT procurement logic, where the focus now is on transformative performance gains rather than unit price reductions.

The most effective way to lower the cost per token is not by reducing GPU prices or using cheaper network components, but by delivering dramatically higher compute performance with each new generation. By improving node or rack performance by 10 to 30 times, this enhancement directly impacts the numerator of the cost formula, effectively reducing the cost of every token produced across the entire data center. This approach makes performance-driven cost dilution far more efficient than simply sourcing lower-priced hardware.

Traditional IT procurement habits have long focused on reducing hardware unit costs. However, in the age of AI factories, the exponential 10x to 30x gains in compute performance are now the primary lever for influencing total cost of ownership, surpassing the impact of minor adjustments to silicon prices. Consequently, using performance improvements to spread the unit cost of computing is significantly more effective than relying on cheap hardware purchases alone.

The shift to Agentic AI, which eliminates human intervention, triggers a dramatic surge in compute consumption. By removing humans from the interactive loop, the rate at which computational resources are consumed is entirely dependent on how fast the agent can self-iterate and think. Previously, service designers could rely on the natural pause created by human reading and typing speeds to schedule resources. Now, agents can ask themselves questions and pivot directions at remarkable speeds, causing a 100-fold increase in computational demands as measured by the AgentX benchmark. This transition from Chat mode to Agentic AI represents a paradigm shift in compute requirements, as the traditional cloud scheduling assumptions, which were based on limited human reaction times, are rendered obsolete.

The criteria for evaluating AI chips has also shifted. For each generation of compute products, the consideration is not merely absolute token performance; the only key metric is total token throughput per megawatt. Improving this metric directly enhances an AI factory's ability to generate revenue, as the industry standard for evaluation moves away from single-chip FLOPs or single-card benchmark scores, focusing instead on the efficiency of converting data center electricity into tangible output. The core value of compute is therefore defined by its ability to turn power into revenue, not just peak performance.

Software is also playing a critical role in unlocking hidden capacity within data centers. By dynamically managing power consumption at the rack and firmware level, ensuring it never exceeds thresholds, it becomes possible to deploy up to 40% more racks within the same reserved power envelope. This approach can directly boost throughput to 1.2 billion tokens per second. This breaks the static deployment model where data centers reserve power based on maximum hardware peak consumption. Under constrained power conditions, software-defined power techniques like MaxLPS allow for a 40% increase in GPU deployment density without the need for new data center construction or additional power quotas.

Looking back over the past 25 years, the role of computing has expanded incredibly. From early magnetic core storage to the present, the ability to transform data, extract information, and apply computation to it has been the essence of value creation. NVIDIA's platform now serves millions of developers globally, supporting a diverse range of workloads from computer vision and physical science to natural language and robotics. This extensive ecosystem, built over 25 years, currently hosts over 8,000 major CUDA applications and has grown to support more than 300,000 models on Hugging Face.

The design of Vera Rubin is a response to these evolving workloads. While continuing to support massive-scale training, its primary focus is on next-generation inference and agentic platforms. It is built on a rich ecosystem, with traceable development by over 10 million developers and a supply chain supported by more than 300 major companies. The platform is designed to be universally replaceable, capable of running all workloads and models, whether open-source, closed-source, or frontier models. This supports pre-training, post-training, inference, and agentic AI, all operating simultaneously within agentic workloads. Additionally, the durability of hardware is notable, with GPUs from a decade ago still generating value, demonstrating strong long-term demand.

The nature of workloads has changed dramatically. Early benchmarks were based on Chat, which featured around 1K input tokens and an average of three rounds of conversation, a human-in-the-loop pattern. Today, Agentic AI presents challenges that are 100 times greater numerically. The average input size has reached 142,000 tokens, with dynamic requirements supporting everything from 1K to 200K or longer inputs. Each of these input lengths requires new optimization kernels, tuning, and mathematical adjustments, greatly increasing software optimization complexity. KV Cache sizes have exploded, as models support up to one million tokens, and the context continues to accumulate as agents interact without human pauses. Finally, the number of interaction rounds has skyrocketed, as agents do not need time to read or type, consuming compute at a speed determined solely by their self-iteration rate. This complexity extends to scheduling, with LLMs processing context and observations, another layer handling reasoning, and yet another deciding which sub-agents or tools to call next. All these tools must be managed to prevent agents from hallucinating, wasting time, or generating erroneous thoughts, all while managing a massively expanded context window.

Vera Rubin is a full-stack solution designed for these challenges. The Vera Rubin NVL72 is built on the foundation of Grace Blackwell but extends its capabilities, incorporating accelerated packages based on Groq LPU technology for high-value workloads. The agentic CPU is crucial for fast, efficient, and complete tool calling. All KV Cache must be managed in smart storage, with all context data managed, governed, and integrated into a software-defined network. This is why NVIDIA introduced BlueField for storage, cooperating with storage partners to provide SmartNICs and compute platforms. At the scale-out network level, these models are too large for a single GPU, leading to widespread decoupled inference. The Dynamo software allows for the separation of Prefill and Decode stages, managing all KV Cache above them for optimal token rates and throughput.

The results from the SemiAnalysis agentic benchmark on Vera Rubin show a 30x improvement in AI factory throughput, and in some cases up to 60x, through full-stack co-design across the ecosystem. The key metric is not absolute token performance, but total token throughput per megawatt, as data centers are bound by megawatt power limits. For inference, the goal is to achieve thinking speeds of over 100 tokens per second, moving towards 200-250 tokens per second to enable deeper thinking across a 65-turn interaction without excessive user wait times. By delivering stronger performance each generation, the cost per million tokens drops dramatically, as performance gains directly lower the cost per token across the data center. This is far more effective than simply lowering the price of the GPU chip or using cheaper networking.

For applications requiring extremely fast interactive performance, further enhancements are available. By leveraging Groq technology, performance can be unlocked to achieve thousands of tokens per second per user. For models like Gemma 4 or Qwen 3.8 with 100K long contexts, speeds of up to 3400 and 2529 tokens per user per second can be achieved, respectively. These speeds are critical for agentic applications where longer contexts require more computation to generate each new token.

Tool calling is a critical function, and there is a global shortage of CPUs. NVIDIA has designed the dedicated Vera CPU rack to address this. Vera uses NVIDIA's custom-designed Olympus core, which is built for high speed, running every core at full capacity while keeping memory latency low. The large monolithic compute die ensures high bandwidth between cores without bottlenecks, achieving 40% lower memory latency than other options. This represents a departure from pre-cloud CPUs, which prioritized high single-thread performance with fewer cores, or cloud-era CPUs, which focused on cost per core. Vera is designed to maintain high throughput with 188 cores while also sustaining high single-thread performance for low token latency.

Independent validation from Signal65, using Stanford's Terminal Bench, shows that Vera delivers 1.6x performance improvements. Agents are 60% faster at completing tool calls, and core agent tasks like compiling code, running Git, and code verification are up to 2x faster. Early adopters like Perplexity Space have seen a 1.5x improvement in sandbox startup, execution, and shutdown speeds, while ClickHouse ranked first in its public benchmark suite on Vera. Vera is now officially released, with early adopters including Oracle Cloud Infrastructure and several other system providers.

NVLink is essential for high-speed AI and inference. NVIDIA is opening up NVLink to everyone through the NVLink Fusion initiative, allowing chip designers to integrate their own XPUs with NVLink Chiplet IP and NVHBM technology. Early partners include The Matrix, which is adopting NVLink Fusion for its next-generation Raptor XPU, and Amazon AWS with Annapurna Labs, which is integrating custom IP to free up 30-40% of compute resources on their custom chips while achieving higher bandwidth and lower latency.

At the data center level, there is a core problem of maximizing GPU density within a fixed megawatt footprint. Workload power characteristics are dynamic, with GPU power fluctuating between 600W and 1900W across different stages, leaving reserved power underutilized. To address this, NVIDIA has developed MaxLPS software, which dynamically manages and limits power at the rack and firmware level. This allows operators to deploy up to 40% more racks within the same power envelope, potentially increasing a gigawatt-scale data center from 4,000 to 5,600 racks and boosting token throughput to 1.2 billion tokens per second. Lambda has validated this software stack, noting significant GPU density improvements: around 25% on B200 and up to 40% on Vera Rubin.

Looking ahead, NVIDIA is committed to a strict one-year cadence for its roadmap. Following Vera Rubin, the next architectures will be Vera Rubin Ultra, followed by the Feynman architecture and Feynman Ultra. The company will continue this annual release cycle to advance the development of AI token factories, inference, and data center-scale agentic infrastructure.

Disclaimer: Investing carries risk. This is not financial advice. The above content should not be regarded as an offer, recommendation, or solicitation on acquiring or disposing of any financial products, any associated discussions, comments, or posts by author or other users should not be considered as such either. It is solely for general information purpose only, which does not consider your own investment objectives, financial situations or needs. TTM assumes no responsibility or warranty for the accuracy and completeness of the information, investors should do their own research and may seek professional advice before investing.

Comments

We need your insight to fill this gap
Leave a comment