Morgan Stanley's latest calculations indicate that as the AI computing power supply chain shifts from centralized AI training toward large-scale AI inference, Google's TPU-led AI ASIC computing clusters are evolving from internal cost-reduction and customization tools within Google Cloud into a high-margin AI cloud infrastructure business available for external sales. The commitment of roughly one million TPUs, representing approximately $35 billion in potential mass supply orders, has prompted the Morgan Stanley analyst team to raise its TPU system price assumption from $20 billion per gigawatt to $27 billion per gigawatt, while lifting the gross margin assumption from 20% to 30%.
In a recently published research report, Morgan Stanley projects that Google's TPU-related cloud revenue will reach $84 billion in 2027 and $108 billion in 2028, totaling close to $200 billion combined. This suggests that Google's parent company Alphabet (GOOGL.US) may transition from being a primary bearer of massive AI capital expenditures to a super-platform for leasing and selling AI computing systems, with TPU potentially becoming another scalable revenue growth engine alongside search advertising and public cloud IaaS+SaaS businesses.
For mature models, open-source models, and AI agent workloads with relatively stable architectures, massive call volumes, and sustained high utilization rates, specialized AI chips like Google's TPU—the most typical example of cloud giants' self-developed AI ASIC technology—can be deeply optimized around low-precision matrix computation, memory access, and chip interconnect. This optimization improves unit token costs, per-watt throughput, and total cost of ownership. Meanwhile, Google's TPU computing clusters can also handle large-scale training tasks, so the more accurate trend is a heterogeneous division of labor between GPU and ASIC/TPU.
AI training operator processes and the most complex, rapidly evolving frontier AI workloads still heavily depend on AI GPU clusters—frontier model pre-training, reinforcement learning, and fast-changing new operators continue to rely on GPU programmability, the CUDA ecosystem, and NVLink/NVSwitch cluster capabilities. In contrast, large-scale AI inference workloads centered on mature and open-source AI models, Copilot, and AI agents are increasingly suited to specialized self-developed AI chips.
Specifically, Google has divided its eighth-generation TPU into the training-focused TPU 8t and the inference-focused TPU 8i. The 8i optimizes KV Cache and low-latency inference through larger on-chip SRAM, 288GB HBM, and a dedicated collective communication engine, with official claims of up to 80% better inference price-performance compared to Ironwood.
In the view of institutions like Morgan Stanley and Wedbush Securities that continue to favor AI computing chain investment prospects, the near-unlimited frontier computing demand in the AI inference era and the computing needs surrounding AI agents will enable AI ASIC to grow into a significant component of a second trillion-dollar computing ecosystem without destroying GPU demand. This further reinforces the investment logic that the AI semiconductor supercycle is not a single GPU cycle but rather an exponential increase in silicon content across the entire data center.
The analyst team led by Morgan Stanley analyst Brian Nowak wrote in a recent client report: "Notably, recent media reports indicate that Google, under Alphabet, has committed to supplying approximately one million TPUs at a price of roughly $35 billion—which we estimate equates to about 1.3 gigawatts—translating to approximately $270 billion per gigawatt after our upward revision. Previously, we assumed Google sold racks and systems at $20 billion per gigawatt with a 20% gross margin in these first-party partnerships."
"Furthermore, Google's expanded custom chip development agreement with Marvell Technology—which our semiconductor research team has previously analyzed—is consistent with, and arguably supports, the view that TPU pricing and average selling prices are higher than our prior assumptions. Given this new information, we have significantly raised our TPU first-party sales expectations to $270 billion per gigawatt... and substantially increased the sales gross margin to 30%. Overall, we now expect Google to sell 0.3 gigawatts of TPU systems in the second half of 2026, 3.2 gigawatts in 2027, and 4.2 gigawatts in 2028. Based on these calculations, TPU-related Google Cloud revenue would reach $84 billion and $108 billion in 2027 and 2028, respectively," the Nowak-led analyst team wrote.
The Morgan Stanley team continues to maintain an "Overweight" rating on Google with a reiterated price target of $400. As of Tuesday's early U.S. market session, Alphabet shares were trading near $348.
The token economics of the AI inference era differ fundamentally from AI training workloads. Training involves forward propagation, backward propagation, gradient synchronization, and frequently changing research-oriented operators, prioritizing versatility, software maturity, and cluster scalability. Inference, however, can be broken down into compute-intensive prefill and memory bandwidth, KV cache, and latency-sensitive decoding phases. ASIC/XPU can eliminate unnecessary general-purpose circuits and harden around FP8/FP4 low-precision matrix operations, attention mechanisms, mixture of experts models, KV Cache, and data movement paths, thereby improving per-watt token throughput, per-dollar token counts, and service level agreement determinism.
It's important to emphasize that while AI agents' model invocation patterns are more suited to ASIC than GPU, their planning, tool calling, retrieval augmented generation, and state management still require CPU, memory, network, and storage coordination. The agent era will therefore further strengthen the "heterogeneous computing architecture."
Microsoft (MSFT.US), one of America's tech giants, plans to soon launch its latest self-developed AI accelerator Maia. Combined with Google's new agreement with fabless chip company Marvell Technology covering multiple custom chips including AI accelerators, storage, and memory controllers, this highlights that the AI computing infrastructure buildout led by AI application leaders and hyperscale cloud providers is comprehensively upgrading from "collectively purchasing NVIDIA GPUs" to a heterogeneous computing architecture where NVIDIA GPUs, AMD GPUs, self-developed AI ASICs/XPUs, and CPU/DPUs operate collaboratively at scale.
As AI model architectures gradually stabilize and global token call volumes grow exponentially across industries, high-concurrency AI inference workloads are increasingly suited to reducing unit token costs through custom chips. The new agreement between Google and Marvell Technology comprehensively covers custom chips including AI accelerators, storage, and memory controllers, with potential procurement corresponding to up to approximately $120 billion in revenue for Marvell Technology through fiscal year 2033. This effectively diversifies reliance on Broadcom, the long-time technology partner for Google's TPU and its sole fabless design partner.
Microsoft's Maia 200 leverages TSMC's 3-nanometer process, FP8/FP4 tensor cores, 272MB of on-chip SRAM, and a custom network-on-chip to compress costs for Azure, Copilot, and large model inference. Its core advantage over NVIDIA GPUs is not that it's "faster for any model," but rather that Microsoft can jointly optimize the chip, compiler, model, data center network, and cloud scheduling while avoiding paying the full profit premium to external GPU vendors. A Microsoft research report states that Maia 200 has already achieved a 30% improvement in performance per dollar compared to the latest generation hardware in its fleet, with approximately 40% better performance per watt when its self-developed MAI models run on Maia 200.
This is precisely the most dangerous competitive force of ASIC/XPU in the inference era: when billions of similar token generation tasks can be highly standardized, the flexibility premium of NVIDIA's "universal GPU" may not be worth paying for every inference request.
Comments