A GPU order placed today can ship in weeks. The transformer that feeds it can take two years. That gap, not chip allocation, is what now decides when your cluster comes online.
For three years the industry told a single story about scarcity: there were not enough accelerators, and whoever held allocation held the market. That story is going stale. Supply of high-end accelerators has loosened at the margins, secondary capacity is being resold, and buyers who once waited months for a quote now get three of them in a week. What has not loosened is everything underneath the chip. Grid interconnection, on-site substations, switchgear, chillers, and liquid-ready white space have all become harder to get than the silicon they exist to serve.
This matters to anyone signing a compute contract, because the constraint has moved somewhere your procurement process probably does not look. You can diligence a provider's GPU inventory in an afternoon. Diligencing whether they actually control the megawatts to run it takes different questions entirely.
The International Energy Agency put global data centre electricity consumption at roughly 415 TWh in 2024, about 1.5 percent of world electricity demand, and projected it to more than double to around 945 TWh by 2030 in its Energy and AI report. In the United States the concentration is sharper still: the Electric Power Research Institute modelled data centers consuming between 4.6 percent and 9.1 percent of national generation by 2030, against roughly 4 percent when the analysis was published.
Aggregate numbers understate the problem, because demand is not spread evenly. It clusters in a handful of markets with fiber density, tax treatment, and existing campuses: Northern Virginia, Dallas, Phoenix, Columbus, and a short list of European and Nordic equivalents. Those are precisely the markets where utilities have stopped quoting fast interconnection dates. The result is a market where the accelerator is available and the place to plug it in is not.
That inversion changes what a provider's inventory page means. A neocloud advertising next-quarter availability is making a claim about two different supply chains, and only one of them is visible to you. It is also part of why the market has repriced operators whose growth assumed unlimited buildout, a dynamic we covered when CoreWeave's revenue miss signalled a turning point for AI infrastructure.
The density curve is the mechanism behind all of this. A conventional enterprise rack in 2022 drew 5 to 10 kW and was cooled with air, in a hall designed around that assumption. An 8-way H100 node draws roughly 10.2 kW on its own, so a single node now consumes the power budget of an entire legacy rack. NVIDIA's GB200 NVL72, a rack-scale system with 72 GPUs in one liquid-cooled enclosure, draws on the order of 120 kW.
Ten to fifteen times the density in one hardware generation is not something a facility absorbs by tuning. It invalidates the floor plan. Halls built for 8 kW racks cannot host 120 kW racks by adding cooling capacity, because the power distribution, the floor loading, the busway ratings, and the heat rejection path were all sized for a different physics problem. Most existing colocation inventory in the world is the wrong shape for the hardware being manufactured right now, which is a very different situation from the familiar choice between GPU and CPU compute for a given workload.
| Deployment | Rack draw | Cooling regime | What breaks first |
|---|---|---|---|
| Enterprise rack, 2022 | 5 to 10 kW | Air, hot aisle containment | Nothing, this is the design point |
| 8-way H100 node | 10.2 kW | Air with rear door heat exchange | Airflow and floor loading |
| Dense H100 or H200 rack | 40 to 60 kW | Direct-to-chip liquid | Busway rating and heat rejection |
| GB200 NVL72 rack-scale | About 120 kW | Liquid only, facility water loop | Utility feed and switchgear capacity |
Air has a hard ceiling. Somewhere between 30 and 50 kW per rack, depending on containment and inlet temperature, moving enough air becomes physically impractical and the fan power required starts eating the efficiency you were trying to protect. Above that line, direct-to-chip liquid cooling stops being an optimization and becomes the only option. The current generation of rack-scale systems is shipped on that assumption.
The operational consequences are underrated by buyers. A liquid-cooled hall introduces a facility water loop, coolant distribution units, quick-disconnect fittings at every node, leak detection, water chemistry management, and a maintenance discipline most enterprise IT teams have never run. It also introduces a new single point of failure: a CDU outage takes down everything downstream of it in minutes, not hours, because there is no thermal mass to coast on.
Efficiency has been flat while all this happened. The Uptime Institute's 2024 global survey put the average reported PUE at about 1.56, essentially unchanged for several years, because the fleet average is dominated by older halls, not by the handful of new hyperscale builds achieving far better numbers. That flat industry average is the number that lands on your invoice when you rent from an operator running legacy floor space.
Behind the facility sits the grid, and the grid has a queue. Lawrence Berkeley National Laboratory's Queued Up series has tracked well over 2,000 GW of generation and storage capacity waiting in United States interconnection queues, with typical waits from request to commercial operation measured in years rather than months. Large load interconnection, which is what a new campus needs, sits in an adjacent process that is no faster.
Equipment lead times compound it. High-voltage transformers, switchgear, and gas turbines have all been quoted at multi-year deliveries since the 2023 demand surge, and the manufacturing base for them expands slowly because the buyers are utilities with long capital cycles. None of this responds to a compute buyer's urgency.
The practical effect for buyers is a widening gap between providers who already hold energized capacity and providers who hold a signed land deal and a utility study. Both can put a 2027 date on a slide. Only one of them controls whether that date holds.
It is worth converting megawatts into the unit you actually buy in. Take a 1,000-GPU H100 cluster built from 8-way nodes at 10.2 kW each. That is 125 nodes and roughly 1,275 kW of IT load. At the industry average PUE of 1.56, facility draw is about 1,989 kW, which works out to almost exactly 2.0 kWh of facility energy per GPU hour.
At 8 cents per kWh, a competitive industrial rate in a low-cost market, that is 16 cents per GPU hour in electricity alone. At 20 cents per kWh, a realistic rate in a constrained metro, it is 40 cents. Against on-demand H100 rental prices that have been trading in the low single dollars per hour, power alone can be a tenth to a fifth of the sell price before anyone has paid for the GPU, the building, the network, or the staff.
Now change one variable. A modern liquid-cooled hall running at a PUE of 1.10 instead of 1.56 drops facility energy to about 1.40 kWh per GPU hour. On a 1,000-GPU cluster at 14 cents per kWh, that difference is roughly 736,000 dollars a year, for identical GPUs doing identical work. Efficiency and location are not facility trivia. They are a line item large enough to swing a build versus rent decision on their own.
| Electricity rate | PUE 1.56 | PUE 1.10 | Annual delta, 1,000 GPUs |
|---|---|---|---|
| 8 cents per kWh | $0.160 | $0.112 | $420,000 |
| 14 cents per kWh | $0.280 | $0.196 | $736,000 |
| 20 cents per kWh | $0.400 | $0.280 | $1,051,000 |
This is also the clearest argument for why cheap capacity in a far-away market is not automatically cheap. Move a cluster to chase a power price and your data has to follow it, and egress on a multi-petabyte training corpus can erase a year of electricity savings in a single migration. Storage that does not tax you for leaving, from providers like Backblaze B2 or Wasabi, is what keeps a power-driven relocation economically reversible.
The vetting question that has served this audience well for GPUs, do you own the hardware, needs a facility-layer companion: do you control the power. Ask providers the following, and treat vague answers as answers.
Installed capacity describes equipment. Contracted load describes what the utility has agreed to deliver. A provider quoting a 50 MW campus with 12 MW energized is selling you a roadmap. Get the energized number, the contracted number, and the date the delta lands.
Design PUE is a specification achieved under ideal conditions. Annualized measured PUE includes summer, includes partial load, and is the number that determines your pass-through. If power is billed as a pass-through, the difference between 1.15 and 1.45 is money moving directly from you to the utility.
All-in dollars per GPU hour, dollars per kW-month plus metered energy, and pass-through at cost plus a margin behave completely differently when tariffs move. Only one of the three leaves the price risk with the provider. Know which one you signed, and model your spend both ways in a cloud cost calculator before committing to a term.
For any liquid-cooled deployment: what is the supply water temperature, is there CDU redundancy, what is the leak detection and response procedure, and who is on site at 3am. These questions separate operators from resellers faster than any specification sheet.
Where capacity is scarce, prices are high and the market is moving, so optionality is worth paying for. Where power is genuinely abundant and the operator controls it, a longer commitment buys a rate you will not beat later. Match contract length to the local power situation, not to a blanket procurement policy, and price the alternatives the way you would when weighing reserved, on-demand and spot instances.
The scarce asset in AI infrastructure is no longer the accelerator. It is an energized, liquid-ready megawatt with a date on it.
Compare AWS, Google Cloud, Azure, and alternatives like Backblaze B2 Discover how much you could save in seconds