Infrastructure · Investigation
The grid can’t keep up: inside the year AI ran out of power
Utilities on four continents have quietly rewritten their load forecasts. What they found has turned electricity — not chips — into the binding constraint on artificial intelligence.
The first sign that something had broken, engineers at a mid-sized utility in central Ohio said, was not a blackout. It was a spreadsheet. In late 2023 the company’s ten-year load forecast — the document that determines what gets built, where, and at whose expense — projected demand growth of roughly half a percent a year. It was the same number, give or take, that the utility had printed every year for three decades.
By the spring of 2026 that figure had been revised four times. The current version projects growth of just under nine percent annually through 2031. Nothing about the region’s population changed. What changed is that eleven data centre developers filed interconnection requests inside eighteen months, and the smallest of them asked for more electricity than the city of Dayton.
“We spent thirty years planning for a flat curve,” said a transmission planner at the utility, who asked not to be named because they were not authorised to discuss forecasting internals. “Then somebody handed us an exponential and asked why we hadn’t built for it.”
That conversation is now happening in Dublin, in Santiago, in Johannesburg, and in a dozen American states. Across the industry, the constraint on scaling artificial intelligence has quietly migrated. It is no longer the availability of accelerators — supply of those has loosened considerably since the crunch of 2024. It is the far more stubborn problem of getting a gigawatt of firm power to a specific patch of ground, on a schedule set by a training run rather than by a regulator.
The turbine backlog nobody modelled
Ask any developer what the real bottleneck is and, eventually, they will say the same two words: gas turbines. The heavy-duty machines that anchor most new firm generation are built by a very short list of manufacturers, and every one of them is effectively sold out. Order books at the major suppliers now extend into the early 2030s. A turbine ordered today is, in practical terms, a turbine for a facility that will train models nobody has designed yet.
Where the new load is going
Announced data centre capacity by market, 2025–2031 pipeline (GW)
Compiled from utility interconnection filings and operator disclosures reviewed by The Compute Desk. Announced capacity includes projects not yet financed; historical attrition in this category runs 30–40%.
The scarcity has produced a secondary market that would have seemed absurd three years ago. Slots in a manufacturer’s production queue are now traded, informally, between developers. Two people involved in such transactions described payments in the low tens of millions of dollars simply to move up eighteen months — a price that only makes sense if the alternative is a finished building with no way to switch it on.
We used to compete for land, then for chips. Now we compete for a spot in a queue at a factory in Greenville. — A senior development executive at a large cloud operator, speaking on condition of anonymity
Who pays for the wires
The more consequential fight is happening in front of state utility commissions, and it is not technical. When a data centre requires a new substation and forty miles of high-voltage line, someone must pay for that infrastructure. Traditionally, network upgrade costs are socialised — spread across every ratepayer in the territory, on the theory that a stronger grid benefits everyone.
That theory holds up poorly when a single customer accounts for a third of projected load growth. Consumer advocates in at least nine states have filed challenges arguing that ordinary households are subsidising private compute. Utilities, caught between a lucrative customer and a hostile public, have begun proposing a new category of tariff: large-load agreements with minimum take obligations, exit fees, and terms measured in decades.
The industry has largely accepted them. It is a striking reversal. The same companies that spent the 2010s insisting on flexibility and short commitments are now signing fifteen-year contracts for power they may not need, because the alternative is not being able to build at all.
Also in this series
The efficiency question
There is a counter-argument, and it deserves a fair hearing. Every generation of accelerator has delivered meaningfully more computation per watt, and the efficiency curve has not flattened. Inference workloads — which now dominate total energy use, having overtaken training sometime in 2025 — have proven unusually responsive to optimisation. Distillation, quantisation, speculative decoding and aggressive caching have, in several documented cases, cut the energy cost of serving a query by an order of magnitude within a year of deployment.
The trouble is that efficiency gains have so far been reinvested rather than banked. Cheaper inference has meant more inference: longer reasoning traces, agents that call themselves in loops, video generation moving from novelty to default. It is a textbook rebound effect, and it means per-unit efficiency tells you very little about aggregate demand.
“Everyone quotes the joules-per-token number, and it genuinely is improving,” said Dr. Anneke Vermeulen, an energy systems modeller who has published on data centre load flexibility. “But joules per token is not the variable a grid operator cares about. They care about megawatts at 6pm on the hottest Tuesday in August. Those two curves have almost nothing to do with each other.”
Flexibility, finally
The most interesting development of the past year is also the least publicised. Under pressure from regulators, several operators have begun signing agreements that allow grid operators to curtail their load during scarcity events — pausing or throttling training runs for a few hours a year in exchange for dramatically faster interconnection.
Technically this is not hard. Training is among the most interruptible large loads ever connected to a power system; checkpointing already happens continuously, and a four-hour pause costs a schedule, not a dataset. Culturally it has been very hard indeed, because it requires an industry built on the premise of infinite availability to admit that sometimes the answer is wait.
Early results are promising enough that at least two regulators are drafting rules to make curtailable interconnection the default fast track. If that becomes standard, it would represent the first genuine architectural concession the AI industry has made to the physical world it runs on.
Whether it arrives in time is a different question. The forecast that started this reporting — the Ohio spreadsheet with nine percent growth — assumes every announced project gets built. Historically, most do not. But the planner who showed it to us was not comforted by that.
“Thirty percent of these will die,” they said. “I just don’t know which thirty percent, and I have to build the wires either way.”
A note on this page. The Compute Desk is a fictional publication and this article is a design demonstration — the reporting, sources and figures are illustrative, not real. Photographs are genuine stock images and are not of the places described. Read more.