Only 28 percent of enterprises have deployed AI in production at scale, across multiple functions with measurable impact. The other 72 percent are still in pilot, proof-of-concept, or limited deployment. The GPU shortage everyone is living through right now is being caused by less than a third of enterprise demand actually showing up yet.
Only 28 percent of enterprises have deployed AI in production at scale, meaning across multiple business functions with measurable impact, according to McKinsey's most recent Global Survey on AI. The other 72 percent are still somewhere in pilot, proof-of-concept, or limited deployment. Separately, Forrester and Anaconda's 2026 research found that 88 percent of enterprise AI agent pilots never make it to production at all.
Sit with that for a second. The GPU shortage everyone has been living through, the 36-to-52-week lead times, the reserved capacity pools locked up months in advance, the on-demand pricing running two to three times higher than reserved rates, is being caused by less than a third of enterprise AI demand actually showing up in production. Consumer AI exploded. Enterprise AI, the workloads that will actually consume the most compute at the largest scale, is still mostly stuck in the pilot stage. When that changes, and it will, the compute math changes with it.
Enterprise AI adoption headlines look impressive on their own. Roughly 78 percent of Global 2000 companies report at least one AI workload in production as of early 2026, up from 41 percent two years earlier. But "at least one AI workload" and "AI deployed at scale across the business" are very different claims, and the gap between them is where most of the real compute demand is still sitting untapped.
McKinsey's numbers make the distinction explicit: 28 percent of enterprises have reached production deployment at scale. The rest are experimenting. Forrester and Anaconda's pilot data tells the same story from a different angle: 88 percent of agent pilots specifically never graduate to production, with evaluation gaps, governance friction, and model reliability cited as the top blockers, not lack of ambition or lack of budget.
This is not a story about AI failing to deliver value. It's a story about most enterprises still being in the experimentation phase of a technology that gets dramatically more compute-hungry the moment it moves into production at scale.
Two structural constraints explain most of the gap between enthusiasm and production deployment, and neither one is close to being solved.
AI and machine learning talent demand currently outpaces supply by roughly 3.2 to 1 across key technical roles. Senior AI engineers, the people who can actually take a promising pilot and turn it into a production system with proper evaluation harnesses, reliability guarantees, and governance, take 90 to 120 days to hire on average, against roughly 25 days for a generic technical role. Sunny Smith, founder and CTO of Massed Compute, made a similar point on the DataStorage.com Podcast: only tens of thousands of people worldwide can build AI infrastructure end to end, and that number is not growing nearly as fast as enterprise ambition is.
A pilot needs to demonstrate a concept works. Production needs governance, monitoring, evaluation, cost controls, and reliability guarantees that most pilots were never built with in mind. Rebuilding a promising prototype to meet those bars is often more work than building the original pilot, which is exactly why so many pilots stall rather than graduate.
Russ Artzt, co-founder of CA Technologies, described a similar historical pattern on the podcast: every major infrastructure shift, mainframe to SaaS to cloud to AI, has followed the same curve, with early enthusiasm running well ahead of the operational maturity needed to deploy at real scale. AI is not moving through that curve unusually slowly. It is moving through it at a completely normal pace, and normal pace still means the large majority of the eventual demand has not arrived yet.
The compute math on the other side of that 28 percent is not a modest step up. Two dynamics compound each other.
A pilot might run a few hundred queries a day against a small model. A production deployment across a real business function runs continuously, at real user volume, often with multiple models chained together. Agentic workloads in particular consume orders of magnitude more compute than simple chat interfaces, since an agent may make dozens of model calls to complete a single task a human would have done in one step. This is closely tied to how AI agents are breaking the per-seat SaaS business model entirely.
McKinsey's infrastructure forecasts project AI will account for roughly 70 percent of data center demand by 2030, up from about 33 percent in 2025, part of a broader 156 gigawatts of AI data center capacity requiring an estimated $5.2 trillion in capital expenditure. Those numbers are not built assuming today's 28 percent production rate stays flat. They are built assuming it doesn't.
Put the two together and a 10x multiple on current enterprise GPU demand is not a hype-cycle exaggeration, it's a reasonable extrapolation from where adoption sits today. If enterprise production deployment merely triples from 28 percent to something closer to universal adoption, and each of those production deployments consumes several times more compute than the pilots that preceded them because agentic workloads are compute-hungry in ways chat interfaces never were, the combined effect compounds well past a simple doubling.
The uncomfortable part of this thesis is that GPU supply is already tight before the enterprise wave has meaningfully arrived.
Data center GPU lead times currently run 36 to 52 weeks. High-bandwidth memory, the component every modern AI accelerator depends on, is supply-constrained because memory manufacturers have redirected capacity away from standard DRAM toward HBM production, and TSMC's CoWoS advanced packaging capacity, required to bond HBM onto GPU substrates, is fully allocated through at least mid-2027. Microsoft, Google, Meta, and Amazon have already placed multi-billion-dollar forward orders for Blackwell-generation GPUs that consume most of NVIDIA's available allocation through the end of 2026 and into 2027, crowding out exactly the mid-market and enterprise buyers who are about to need capacity as their own pilots mature.
NVIDIA controls an estimated 78 to 82 percent of the AI training chip market, concentrating this supply crunch behind a single dominant vendor with limited ability to expand output faster than its own supply chain allows. None of this improves meaningfully before enterprise production adoption catches up to enterprise experimentation. It gets worse.
If your organization has AI pilots that are on a credible path to production in the next 12 to 18 months, the capacity planning conversation needs to start now, not when the pilot is approved for scale-up. Reserved capacity secured today is priced against today's constrained but not-yet-catastrophic supply picture.
With major hyperscalers absorbing the bulk of new GPU allocation through forward orders, neoclouds and specialized providers increasingly represent the more realistic path to capacity for mid-market and enterprise buyers who aren't first in line at AWS, Azure, or Google Cloud. Vetting those providers on ownership and reliability matters more, not less, as competition for their capacity increases too.
A pilot's compute bill is not a reliable predictor of a production deployment's compute bill, especially for agentic workloads. Budget planning that assumes a linear scale-up from pilot spend routinely underestimates production costs by a wide margin, since agent-based systems can consume many multiples of the compute a simple pilot used.
If the binding constraint on your AI production timeline is engineering talent rather than compute availability, throwing more GPU capacity at the problem won't accelerate anything. Address the talent gap directly, through upskilling existing staff, targeted senior hiring, or infrastructure partners who bring more of that expertise in-house, before assuming compute is the bottleneck.
The GPU shortage you're living through right now is what less than a third of enterprise AI demand looks like. Plan accordingly.
Compare AWS, Google Cloud, Azure, and alternatives like Backblaze B2 Discover how much you could save in seconds