The enormous power demands of AI are well known, but a more subtle and pressing issue is emerging: drastic fluctuations in data center electricity consumption are damaging critical equipment, threatening facility reliability, and exporting instability to the broader power grid.
According to an August 6 report, core systems in AI computing facilities, including batteries, generators, and cooling units, are failing or wearing out prematurely due to abnormal power surges. At xAI's Colossus supercomputer facility in Memphis, Tennessee, gas turbines have developed cracks. Similar cracking issues have been reported in several smaller data centers in the UK. Some installed voltage-stabilizing batteries are requiring replacement within weeks or months under high-intensity operation.
The market impact of this problem is becoming apparent. A source involved in data center financing revealed that the actual uptime for some facilities has dropped to around 80%, far below the design expectation of 24/7 operation. If the issue remains unresolved, it could directly impact investors in related projects within the next 12 to 24 months. Separately, the North American Electric Reliability Corporation (NERC) issued a rare Level 3 alert earlier this year, demanding that large data centers immediately address risks to grid stability.
The Root of Power Fluctuations: 'Millisecond-Level' Shocks from GPU Clusters
The power consumption pattern of AI data centers fundamentally differs from traditional facilities. Traditional data centers have relatively stable power usage, while AI computing, particularly during model training, drives hundreds of thousands of GPUs to start and stop synchronously, causing power demand to spike or drop dramatically within milliseconds.
Shannon Miller, founder and president of Mainspring Energy Inc., described this phenomenon vividly: a 1-gigawatt data center, the size of Boston, could have half its power usage flashing on and off every few seconds. Meanwhile, some planned AI campuses in Texas and the U.S. Midwest exceed 5 gigawatts, with average consumption approaching that of New York City.
Drew Baglino, a former Tesla executive and founder of Heron Power Electronics Co., noted that the instantaneous power draw of an AI data center can sometimes surge to over 50% of its designed capacity. "A 1-gigawatt facility could momentarily consume 1.5 gigawatts." His company is developing power fluctuation management equipment for Nvidia's next-generation servers, expected in 2027.
Amber Villegas-Williamson, Principal Advisor at the Uptime Institute in the UK, compared the impact to repeatedly stomping on the gas pedal while driving: "It's like constantly revving the engine beyond its limit, which wears it out much faster than steady cruising."
Equipment Damage is a Reality: From Cracks to Arc Flashes
According to interviews with over 30 power professionals in the U.S. and Europe, the physical stress on equipment from AI data centers is already visible. The report states that multiple sources say crankshafts in small natural gas internal combustion engines used in data centers have fractured. At xAI's Colossus facility, gas turbines developed cracks, leading operators to install batteries to smooth power fluctuations and reduce turbine load. Andrew Cunningham, CEO of GeoPura Ltd., confirmed that similar turbine cracking has occurred in several smaller data centers in the UK.
Jennifer Scanlon, CEO of UL Solutions Inc., pointed out that equipment cracking or wear could trigger arc flashes—where electricity jumps between conductors—potentially damaging AI chips. The Uptime Institute and several industry insiders say batteries installed for voltage stabilization sometimes need replacement within weeks or months. While stabilizing equipment like batteries, capacitors, transformers, and flywheels exists, their deployment in new data centers is clearly insufficient amid the rush to build AI computing power.
Notably, this problem is global, occurring from the Middle East and Africa to Europe and the United States.
Declining Reliability Hits Revenue: Losses Can Reach Hundreds of Thousands of Dollars Per Minute
The direct consequence of equipment damage is facility downtime, which carries an extremely high cost. Jason Hoffman, Chief Strategy Officer at Switch, stated that the financial cost of premature equipment failure "isn't mainly the cost of replacing a pump or a circuit breaker; it's the lost revenue from expensive computing assets being offline and unable to generate income." Estimates suggest that revenue losses from downtime range from thousands to hundreds of thousands of dollars per minute, depending on the facility type and workload.
A specific case cited in the report involves a 2.67-gigawatt AI campus in West Texas, jointly developed by Joulent Inc. and Chevron. To meet Microsoft's requirement for 99.999% reliability, the project had to incorporate extra engineering time, delaying its power-on date from 2027 to 2028. Chris James, CEO of Joulent, said this adjustment was made specifically to address the engineering challenges posed by power fluctuations.
The report states that a source involved in data center financing revealed that while the industry typically assumes facilities will operate non-stop after launch, the actual uptime for some facilities is already approaching 80%. If the problem cannot be solved, it will have a material impact on project investors within the next 12 to 24 months.
Analysts note that the hundreds of billions of dollars in capital expenditure from hyperscale cloud providers have already made investors and lenders nervous. The issue of accelerated asset depreciation from equipment damage, combined with prior market skepticism about the depreciation rate of GPU racks, deepens concerns about whether AI data centers can deliver on their profitability promises.
Even a few minutes of downtime can directly erode the revenue of data center developers. If actual uptime remains below projections for an extended period, the financial models for these projects will face fundamental revaluation. As the boom in AI infrastructure investment continues, this structural risk is becoming a new variable that investors must confront.
Grid Stability at Risk: Regulators Issue Rare Alerts
Notably, the power fluctuations from AI data centers threaten not only the facilities themselves but also export instability to the broader power grid. Sreemant Roy, a power quality expert and global product manager at Schneider Electric, warned that such highly dynamic loads "can lead to grid instability, and if not corrected, could cause blackouts or power outages." He specifically noted that data centers could trigger sub-synchronous oscillations, damaging connected equipment elsewhere on the grid, "a concern that has deeply worried power companies worldwide."
On the regulatory front, NERC has issued multiple warnings over the past two years, listing data centers as one of the biggest risks to grid stability. In a September report, NERC assessed over 33 gigawatts of operating data centers in the U.S. and found that about three-quarters of the load models were "insufficient to reflect the dynamic behavior of data centers." Earlier this year, NERC issued a rare Level 3 alert, requiring large data centers to submit response plans by August 3.
In response to these challenges, various parts of the industry chain are actively seeking solutions. Nvidia stated that since the release of its Blackwell GPU in 2024, it has increased collaboration with power experts. Dion Harris, Senior Director of Hyperscale Infrastructure Solutions at Nvidia, said the company is working towards smoother deployment in data center construction, design, engineering, and power delivery.
At the operational level, some data center users have adopted "side-computing" techniques—running virtual computing tasks not part of the training process—to keep GPUs running steadily and smooth power usage. However, critics point out that this method wastes additional electricity at a time when power demand is already high.
On the policy side, the U.S. Department of Energy last year commissioned the National Renewable Energy Laboratory (NREL) near Denver, Colorado, to establish a test platform specifically for research on safely integrating AI into the grid. Martha Symko-Davies, NREL Laboratory Program Manager, stated that the platform is equipped with GPUs and power generation equipment, allowing developers to test whether their systems can handle the volatility of AI loads. It will be used to test batteries, software, and other stabilizing devices. "We have a chance to get this right now," she said.
Comments