Anthropic PBC, the global leader in AI large language models and applications, has recruited Amir Salek, one of the core founding architects behind Alphabet Inc.'s Google TPU custom AI chip program. This strategic hire comes as the world's premier AI laboratory lays critical groundwork for a major push into proprietary AI semiconductor development, also known as custom AI ASIC or XPU technology. The company confirmed on Friday that Salek will join Anthropic's AI computing infrastructure team supporting the Claude AI super platform. Prior to 2022, he led Google's Tensor Processing Unit operations and contributed to the design of the first seven generations of TPU chips. In his new role, Salek will report directly to James Bradbury, Anthropic's head of AI computing infrastructure.
Microsoft (MSFT.US) is preparing to launch its latest in-house AI accelerator, Maia, while Google has inked a new agreement with fabless chip designer Marvell Technology covering custom AI accelerators, storage controllers, and memory controllers. These developments highlight a fundamental shift in AI infrastructure buildout: the industry is moving away from wholesale procurement of NVIDIA GPUs toward a comprehensive heterogeneous computing ecosystem where NVIDIA GPUs, AMD GPUs, custom AI ASICs/XPUs, CPUs, and DPUs operate in large-scale coordination. As AI model architectures stabilize and token call volumes across global industries grow exponentially, high-concurrency AI inference workloads are increasingly suited for custom chips that dramatically reduce per-token costs.
Training operations and the most architecturally complex, rapidly evolving frontier AI workloads still depend heavily on AI GPU clusters. Frontier model pre-training, reinforcement learning, and fast-changing new operators rely on GPU programmability, the CUDA ecosystem, and NVLink/NVSwitch cluster capabilities. However, large-scale inference workloads built around mature and open-source AI models, Copilot, and AI agents are increasingly better served by specialized custom silicon. Morgan Stanley and Wedbush Securities, among other bullish voices on AI computing infrastructure investment, believe that near-limitless frontier computing demands and AI agent-driven capacity requirements will enable AI ASICs to grow into a second trillion-dollar computing ecosystem without destroying GPU demand. Instead, this trend reinforces the investment thesis that the AI semiconductor supercycle is not merely a single GPU cycle, but a broader cycle of rising silicon content across all data centers.
Anthropic Doubles Down on AI Chip Development, Custom Silicon Becomes the Key to Inference Cost Advantage
Anthropic sources AI chips and accelerators through multiple channels including NVIDIA, Google, and Amazon. The company has recently signaled its intent to build an internal custom silicon operation. The San Francisco-based firm has begun recruiting talent and posting relevant positions, with media reports indicating it has initiated talks with Taiwan Semiconductor Manufacturing Company regarding long-term advanced process and advanced packaging foundry partnerships. Anthropic has not yet disclosed the architecture, tape-out timeline, or formal external fabless co-design partners such as Broadcom or Marvell for its latest custom chip initiative. Broadcom previously participated in Anthropic's approximately $35 billion computing infrastructure expansion project, providing custom AI ASICs, networking solutions, and financing support, but has not yet been confirmed as a long-term partner for Anthropic's proprietary chip business.
Like other AI application super-majors, Anthropic is racing to secure sufficient data center infrastructure to support its ambitious goals. Custom AI chip clusters could help the company address global AI chip supply shortages more comprehensively, optimize chip design and performance according to its specific needs, and potentially slash AI inference costs across its data center operations. OpenAI, the company behind ChatGPT and Anthropic's fiercest competitor, is pursuing a similar path. OpenAI has unveiled a custom AI chip architecture called "Jalapeno," co-designed with Broadcom, with plans to begin deployment later this year.
Anthropic is rapidly signing large-scale agreements for AI chips and data center capacity, establishing multi-year partnerships with established suppliers and emerging cloud providers. Recent media reports indicate the company has reached a deal with UK chip startup Fractile for an initial order of approximately $250 million, with plans to expand the contract in the future. Anthropic has also signed large-scale data center capacity agreements with Riot Platforms Inc. and Volta Infra Holdings Ltd. Additionally, the frontier AI developer behind the Claude platform is conducting valuation and IPO fundraising assessments, with management preparing to publicly file necessary documents for a potential mega-IPO as early as the end of this month. Sources familiar with the matter suggest Anthropic management expects its US IPO to match or exceed the record-breaking scale of SpaceX's June listing. This latest development in the AI supply chain underscores the extraordinary demand from institutional and retail investors seeking to profit from the unprecedented artificial intelligence investment boom.
Anthropic's more ambitious IPO target relative to SpaceX reflects how the hottest leaders in AI are reshaping the technology investment landscape. The five-year-old company raised $65 billion at a $965 billion valuation in May, surpassing competitor OpenAI's $852 billion valuation achieved during its $122 billion funding round in March. Anthropic's financial trajectory provides rare verifiability for Wall Street's extremely optimistic earnings path and record IPO expectations. The company's annualized revenue run rate surged from approximately $9 billion at the end of 2025 to $47 billion by May 2026, and exceeded $65 billion by the end of July—a roughly 7.2-fold increase or approximately 622% growth within seven months. Preliminary Q2 revenue exceeded $11.5 billion, approximately 14.6 times the $787 million recorded in the same period of 2025 and about 2.4 times Q1's $4.73 billion, marking the company's first-ever positive adjusted operating profit.
Since leaving Google, Salek has served as a senior managing director at private equity firm Cerberus Capital Management, which was co-founded by Deputy Secretary of Defense Stephen Feinberg. Notably, he previously worked at NVIDIA, the world's leading AI processor manufacturer.
Google Partners with Marvell, Anthropic Poaches TPU Veteran: AI Giants Battle for Silicon Supremacy in the Inference Era
The token economics of AI inference differ fundamentally from AI training workloads. Training involves forward propagation, backward propagation, gradient synchronization, and rapidly changing research operators, prioritizing versatility, software maturity, and cluster scalability. Inference, by contrast, can be broken down into compute-intensive prefill phases and memory bandwidth, KV cache, and latency-sensitive decoding phases. ASIC/XPU designs eliminate unnecessary general-purpose circuits and harden specific capabilities for FP8/FP4 low-precision matrix operations, attention mechanisms, mixture-of-experts models, KV cache, and data movement paths. This enables order-of-magnitude improvements in tokens per watt, tokens per dollar, and service level agreement determinism.
It is important to note that while AI agents' model calling patterns are better suited to ASICs than GPUs, their planning, tool invocation, retrieval-augmented generation, and state management still require CPU, memory, network, and storage coordination. The agent era therefore further reinforces the importance of heterogeneous computing. Specifically, Google has split its eighth-generation TPU into the training-focused TPU 8t and inference-focused TPU 8i. The 8i optimizes KV cache and low-latency inference through larger on-chip SRAM, 288GB of HBM, and dedicated collective communication engines, with Google claiming up to 80% better inference price-performance than Ironwood. Google's new agreement with Marvell covers custom chips including AI accelerators, storage, and memory controllers, potentially corresponding to up to $120 billion in revenue for Marvell through fiscal 2033. This effectively reduces Google's long-standing dependence on Broadcom as its sole fabless design partner for TPU technology.
Microsoft's Maia 200 leverages TSMC's 3-nanometer process, FP8/FP4 tensor cores, 272MB of on-chip SRAM, and a custom network-on-chip to compress Azure, Copilot, and large model inference costs. Its core advantage over NVIDIA GPUs is not that it is "faster for every model," but rather Microsoft's ability to jointly optimize chips, compilers, models, data center networks, and cloud scheduling while avoiding paying external GPU vendors' full profit margins. Chris Caso, strategist at Wolfe Research, noted that the Philadelphia Semiconductor Index (SOXX) roughly doubled over three months before declining approximately 25% from its peak, with recent weakness more reflective of post-rally expectation resets. Caso projects AI chip demand will continue to exceed supply through at least 2028, arguing that concerns about hyperscaler capital expenditure slowdowns have not materialized and that competition around AI agents leaves hyperscalers with "no choice but to invest."
Comments