AI applications continue to expand, but the hyperscale data center investments underpinning this boom are confronting a new demand-side challenge.
According to a Jefferies report, as small language models improve in performance and decline in cost, enterprise AI deployment may shift from centralized cloud computing toward localized solutions, thereby exerting pressure on cloud computing demand and massive data center capital expenditures.
A key piece of evidence for this assessment comes from recent observations of actual AI server usage. In September 2026, Swerve Research Technologies constructed a "congestion index" by tracking wait times in Anthropic's server request queues, finding that relevant server utilization had declined 26% from its peak in January to February 2026. This metric does not directly correspond to company revenue, but if utilization continues to decline, it may indicate that some computing capacity is beginning to sit idle.
Meanwhile, AI infrastructure investment continues to expand rapidly. The market currently expects that combined capital expenditures for four major U.S. tech giants — Meta Platforms, Inc. (NASDAQ: META), Alphabet (NASDAQ: GOOGL), Amazon.com (NASDAQ: AMZN), and Microsoft (NASDAQ: MSFT) — could reach $990 billion in 2027. The report argues that if AI computing demand growth ultimately fails to match such a massive investment scale, some data centers could face insufficient returns or even become "stranded assets."
Server Utilization Declines, AI Computing Supply Faces Scrutiny
After Meta Platforms, Inc. (NASDAQ: META) launched its personal AI product Muse on September 8, the company's stock price rose 21% within weeks, becoming a bright spot in recent AI consumer applications. But the report argues that the success of a single application is not sufficient to prove that the entire AI infrastructure investment cycle can be sustained.
More notably, AI companies themselves have also released some signals that diverge from optimistic market expectations. Anthropic's CEO, OpenAI's CEO Sam Altman, and Elon Musk have all recently publicly discussed the necessity of slowing AI development; Anthropic's originally planned IPO has been postponed to November.
At the same time, Anthropic had previously signed a $45 billion, three-year computing contract with SpaceX, and reached multi-billion-dollar computing partnerships with Amazon.com (NASDAQ: AMZN) and Google Cloud. If AI revenue growth slows while these large-scale computing contracts still need to be fulfilled, corporate cost pressures could rise further.
According to reports, Meta Platforms, Inc. (NASDAQ: META) was also in negotiations with Anthropic in July to lease up to $10 billion worth of data center computing capacity to it over two years. The report argues that major tech companies beginning to actively seek external customers to absorb computing capacity also reflects that, as infrastructure supply expands rapidly, how to improve asset utilization is becoming a new problem.
Small Models Catching Up in Performance, Local Deployment May Divert Cloud Demand
Compared with short-term changes in server utilization, the report pays more attention to changes taking place in AI model architecture.
Research titled "Intelligence per Watt," jointly published by Stanford University and Together AI, shows that small language models are rapidly catching up to large language models in performance, while energy consumption and computing costs are 50% to 85% lower. If this trend continues, both the scale of computing power and the cost structure required for enterprise AI deployment could change accordingly.
In the past, enterprises relied more on centralized cloud-based large models to complete AI tasks; as small model capabilities improve, some scenarios may shift toward local deployment. For industries with high data security requirements such as banking, localized deployment already has strong demand; as small models further reduce deployment costs, the applicable scope of this model may continue to expand.
This means that growth in AI applications does not necessarily correspond to proportional growth in cloud computing demand. Enterprises may expand AI applications while simultaneously reducing their dependence on ultra-large model APIs and centralized cloud computing power.
This is also an important reason why the report questions nearly trillion-dollar capital expenditures: if more AI applications in the future are carried by low-cost, small-scale, localized models, then the current investment scale built around hyperscale data centers may become mismatched with actual computing demand.
Under this scenario, what truly needs to be reassessed is not whether AI demand exists, but how much centralized computing power AI demand actually requires, and whether these data centers can generate returns commensurate with their investment scale.
Comments