In April 2026, Meta expanded its compute contract with CoreWeave to 21 billion dollars, and Nebius signed commitments with Meta reported at up to 27 billion dollars. Six months earlier, either number would have looked like confirmation that GPU demand had no ceiling.
Read today, against Meta's own move to start selling surplus GPU capacity through Meta Compute, those contracts read differently: as the last big deals signed under the old scarcity rules, just as the rules are changing.
This report pulls together where the GPU cloud market actually stands in August 2026: what is loosening, what is tightening, who owns what they rent out, and what buyers should do differently in the next contract cycle.
For three years, the defining fact of AI infrastructure was that demand for GPUs exceeded supply at almost any price. That is no longer uniformly true. Meta Compute's decision to resell surplus capacity, alongside reports of SpaceX doing the same with its own GPU reserves, signals that at least some of the largest buyers over-provisioned against a demand curve that has not yet fully materialized.
This does not mean GPU demand is falling. It means the negotiating leverage that hyperscalers and top-tier neoclouds held for three years is starting to shift toward buyers, at least for certain chip generations and contract lengths. Enterprises locking in 12 month or shorter commitments are now finding room to negotiate that did not exist in 2024 or 2025.
The practical shift: multi-provider strategies and shorter contracts are no longer just a hedge against vendor risk. They are now a legitimate way to capture pricing improvement as supply loosens, without betting the whole workload on one provider's roadmap.
NVIDIA's Blackwell generation (B200, B300, GB200 NVL72) is where the announcements and the marketing budget concentrate. But H100 and H200 capacity, NVIDIA's Hopper generation, still accounts for the majority of production inference workloads running today, and pricing on that older capacity has been falling faster than Blackwell pricing as newer supply comes online.
This creates a genuinely two speed market. Teams training frontier scale models or running the largest inference fleets are chasing Blackwell allocation, often through providers with direct NVIDIA relationships. Teams running steady state production inference, fine tuning, or mid size training runs are frequently better served staying on Hopper, where pricing is now materially more favorable and availability is no longer the constraint it was through most of 2024 and 2025.
The mistake to avoid is treating chip generation as a status symbol rather than a cost input. Match the chip to the workload's actual requirements, then shop that generation across providers rather than defaulting to whichever chip is loudest in the press.
Not every company calling itself a neocloud owns the hardware it rents out. A meaningful share of the market is brokers reselling capacity they lease from someone else, a point Massed Compute founder and CTO Sunny Smith made directly on the DataStorage.com Podcast: the first question to ask any provider is simply whether they own the GPU.
Ownership is not a branding detail. It determines who actually handles a support ticket when a node fails at 2 a.m., how much margin is baked into the price you pay, and how much transparency you get into utilization and hardware health. When support runs through an intermediary, tickets get relayed rather than resolved, and resolution time stretches accordingly.
It also explains why GPU ownership functions as financial engineering as much as infrastructure strategy. A single B300 server runs in the neighborhood of 750,000 dollars, a figure Smith cited on the podcast to illustrate why most market participants choose to rent rather than own: the capital intensity is the barrier, not the technical difficulty. Vetting a provider on ownership before signing anything longer than a quarter is now table stakes procurement practice, not a nice to have.
While the industry has spent three years fixated on GPU scarcity, a quieter shortage has been building underneath it. NVMe storage prices have roughly tripled, and scarcity is expected to persist into late 2027, according to Massed Compute's account on the DataStorage.com Podcast. Storage adjacent to the GPU, with sufficient east west bandwidth, has gone from a casual add on to a planning constraint in its own right.
This matters because GPU and storage decisions cannot be made separately anymore. Moving a petabyte of training data between data centers because a GPU contract moved providers is prohibitively expensive and slow, which is exactly why storage has become the anchor of the AI infrastructure stack. Zero egress storage architecture, the kind offered by providers like Backblaze and Wasabi, has become a genuine infrastructure decision rather than a cost saving footnote. A GPU switch that looks cheap on paper can be erased entirely by the egress bill and the NVMe premium required to re stage the data at the new location.
Checkpoint storage is a concrete example. A single large model checkpoint can run into the range of a terabyte or more depending on model size and precision, a figure that should be treated as a modeled estimate rather than a sourced constant, since it varies significantly by architecture. At current NVMe pricing, the cost of storing rolling checkpoints across a long training run is no longer a rounding error in the infrastructure budget.
Here is the tension sitting underneath every one of the trends above: consumer and developer facing AI usage has exploded, but enterprise production deployment of AI, the kind that runs core business processes rather than pilots, has barely started. The blockers are not primarily technical. They are talent scarcity, since only a limited pool of engineers can build AI infrastructure end to end, and production readiness cycles that move slower than the hype cycle around them.
That gap is why the market can simultaneously show softening signals, like Meta Compute's resale move and CoreWeave's own revenue miss earlier in 2026, while also carrying the possibility that GPU demand could grow many multiples higher once enterprise deployment actually accelerates. Both things are true at once. The current loosening is real, but it is happening in a market that has not yet absorbed its largest future source of demand.
Buyers reading current softness as a permanent trend are making the same mistake as buyers who assumed 2024's scarcity would never end. Neither extreme has held up.
There is also a talent constraint hiding inside the demand constraint. Even companies with budget approved for large AI deployments are finding that the pool of engineers who can build production grade AI infrastructure end to end, not just call a model API, is small and expensive. That bottleneck slows the pace at which announced GPU commitments turn into utilized GPU capacity, which is part of why utilization concerns keep surfacing even as contract values keep climbing.
Three concrete adjustments follow from the state of the market right now. First, default to shorter commitments on Hopper generation capacity where pricing is softening, and reserve longer commitments for Blackwell allocation you genuinely cannot get any other way. Second, treat GPU ownership as a mandatory vetting question for any new provider relationship, not an assumption. Third, model storage and egress costs alongside GPU costs in the same spreadsheet, not a separate one, before signing a multi provider or multi region strategy.
| Dimension | 2023 to 2025 Playbook | Mid to Late 2026 Playbook |
|---|---|---|
| Contract length | 12 to 36 months, locked early to guarantee allocation | 3 to 12 months on Hopper capacity, longer only where Blackwell requires it |
| Provider vetting | Availability and price were the only filters | Do you own the GPU is the first question, before price |
| Storage planning | Treated as a separate line item from compute | Modeled in the same contract decision, given NVMe scarcity |
| Multi provider strategy | Optional hedge against outages | Active lever for capturing loosening prices and avoiding egress lock in |
The market is not settling into a new equilibrium. It is mid transition, from a shortage driven seller's market to something closer to balanced, with a large wave of enterprise demand still offshore. Contracts signed in the next two quarters should be built to flex, not to lock in today's snapshot as permanent.
The GPU shortage is easing. The storage shortage underneath it is not, and it is the one nobody budgeted for.
Compare AWS, Google Cloud, Azure, and alternatives like Backblaze B2 Discover how much you could save in seconds