Spot GPU Pricing in 2026: The Discounts Are Shrinking and Sometimes Inverting

Picture of DataStorage Editorial Team

DataStorage Editorial Team

Cloud Cost & Pricing Transparency 11 min read  ·  August 2026
At a 40 percent headline discount, on a training run that checkpoints hourly, anything more frequent than one interruption every 86 minutes means you are paying more per unit of finished work than you would have paid on demand.

A 40 percent spot discount on an H100 sounds like a straightforward win. Work out how often the job actually gets interrupted and the picture changes. Most teams never run that calculation. They see the discount, book the savings in the forecast, and never reconcile it against the compute they actually consumed.

Spot GPU capacity in 2026 is not the bargain it was in 2021. The discounts have compressed, the pricing has stopped behaving like a market, and on a growing number of workloads the effective cost has crossed above the on demand rate. This is what that looks like in numbers, and what to do about it.

90%
the spot saving still advertised for the general EC2 fleet
AWS EC2 Spot 2026
$12.29
AWS p5.48xlarge on demand cost per H100 GPU hour
AWS list pricing 2026
86 min
how often you can be interrupted before a 40 percent discount stops paying
DataStorage.com model
5%
discount floor below which spot never breaks even at all
DataStorage.com model
GPU Marketplace
Check What an H100 Hour Actually Costs Today
The GPU Price Explorer tracks live rental pricing across 30+ providers, updated daily. Compare on demand and interruptible rates for H100, H200, B200, L40S and MI300X before you commit to a spot strategy.
Explore GPU Pricing  →

The 90 Percent Promise and Where It Went

Spot was one of the cleanest bargains in cloud. Providers had idle capacity, buyers had interruptible work, and the market cleared at a steep discount. AWS still advertises savings of up to 90 percent off on demand rates for EC2 Spot, and for general purpose CPU fleets in unfashionable regions that figure is often close to real.

GPUs never behaved that way, and in 2026 they behave less like it than ever. A spot pool exists because a provider is holding capacity nobody is paying full price for. Since 2023 that condition has almost never held for H100, H200 and B200 class hardware. Accelerators are the constrained resource in every hyperscaler fleet, and constrained resources do not get dumped into a discount bin.

What buyers see instead is a discount that has drifted away from the advertised ceiling into a band that commonly sits between 20 and 50 percent, moves without warning, and arrives with interruption rates that make the headline number close to meaningless. The interesting question is no longer whether spot is cheaper on the price sheet. It is whether the discount is still large enough to pay for the interruptions it brings with it.


Hyperscalers Routed GPU Scarcity Away From the Spot Pool

Three structural changes explain most of the compression, and none of them are going to reverse in the next twelve months.

Reserved capacity products absorbed the demand spot used to serve

EC2 Capacity Blocks for ML let buyers reserve GPU capacity for a defined window at a defined price. Before that product existed, a team needing two weeks of accelerator time with no annual commitment had one realistic option: bid on spot and engineer around eviction. Now they have a first class product that is not spot. That is a better outcome for the buyer and a better revenue outcome for the provider, and it removes precisely the population whose bidding used to make the spot market liquid. Google Cloud and Azure have shipped comparable reservation mechanics. Spot is no longer the escape hatch for short term GPU access, and it is not priced as though it were.

Spot prices stopped being an auction

The word spot still implies a live clearing price. It has not meant that for years. Google Cloud documents that Spot VM prices change at most once every 30 days. AWS smoothed its spot pricing long ago and caps it at the on demand rate, so the price moves in gradual administered steps rather than bidding wars. Azure Spot runs on price caps with eviction notices measured in seconds.

This matters more than it sounds. An auction transmits scarcity in both directions: when capacity frees up, the price falls fast. An administered discount does not. When accelerators are tight, the provider has no commercial reason to widen the spread, and the discount stays narrow long after the underlying supply picture has changed. Buyers who assume spot pricing tracks real time availability are reading a number that was set weeks ago.

The slack that does exist is in the wrong generation

There is genuine idle capacity in the market, but it is concentrated in older accelerators. V100 and A100 class hardware has a long rental tail and still runs near capacity years after launch, which means the discount available on the silicon you want is not the discount available on the silicon there is a surplus of. If your workload can run on an A100, spot economics still look reasonable. If your model needs the memory bandwidth of an H200 or a B200, you are shopping in the tightest part of the market and the spot pool reflects that.


The Discount You See Is Not the Discount You Get

Every interruption on a long running job levies three separate taxes. You lose the work completed since the last checkpoint. You pay GPU time to write checkpoints frequently enough to limit that loss. And you wait for capacity to come back, which costs wall clock time even when it does not cost compute.

That gives a simple model for what a spot hour is really worth. If C is the checkpoint interval in hours, t is the GPU time each checkpoint write costs, and M is the mean time between interruptions, then the share of paid GPU time that produces no finished work is roughly (C divided by 2, divided by M) plus (t divided by C). The first term is the work you redo, assuming an interruption lands on average halfway through a checkpoint interval. The second is the write overhead itself.

Effective cost per useful GPU hour is therefore the spot price divided by one minus that wasted share. Take realistic values for a multi node training run: a one hour checkpoint interval, three minutes of GPU time per checkpoint write, and restart latency excluded entirely, which is a deliberately generous assumption in favour of spot. Set the effective cost equal to the on demand rate and solve, and the break-even mean time between interruptions is 0.5 divided by (the discount minus 0.05).

Headline discount Spot price per GPU hour Break-even interval Realised discount
90 percent $1.23 35 min 87%
70 percent $3.69 46 min 62%
50 percent $6.15 67 min 36%
40 percent $7.37 86 min 23%
30 percent $8.60 120 min 11%
20 percent $9.83 200 min minus 2%
Spot price derived from the AWS p5.48xlarge on demand rate of $12.29 per GPU hour. Realised discount assumes one interruption every three hours. Model assumes a one hour checkpoint interval, three minutes of GPU time per checkpoint write, lost work of half a checkpoint interval per interruption, and restart latency excluded.

The result is not linear, and that is the whole story. A 90 percent discount buys enough headroom to absorb an interruption every 35 minutes. A 40 percent discount collapses that tolerance to 86 minutes. A 20 percent discount means anything more frequent than one interruption every three hours and twenty minutes puts you underwater. Below roughly a 5 percent discount, checkpoint write overhead alone consumes the entire saving and spot can never break even no matter how stable the capacity is.

Run the same model the other way, holding interruptions at one every three hours, and the realised discount is consistently far below the advertised one.

Realised discount at one interruption every three hours
90% headline
87%
70% headline
62%
50% headline
36%
40% headline
23%
30% headline
11%
20% headline
-2%
A 50 percent headline discount returns about 36 percent. A 20 percent headline returns minus 2 percent, which is to say it costs more than on demand while carrying all of the operational risk of spot.
$
Free Tool
See What You Are Actually Paying Across Providers
Use the Cloud Cost Calculator to compare real storage, egress and compute pricing across AWS, Azure, GCP, Backblaze and Wasabi, side by side, in seconds.
Try the Free Calculator  →

Three Ways Spot Pricing Inverts

Inversion is not a rhetorical flourish. It happens in three distinct and independently verifiable ways, and most teams only recognise the third one.

Effective cost inversion

This is the common case and the invisible one. The invoice shows a discount. The unit economics do not, because a meaningful share of the GPU hours purchased were spent redoing work or writing checkpoints. Nothing in a standard billing export separates useful GPU hours from repeated ones, so the loss never surfaces in a cost report. It surfaces as a training run that took nine days instead of six.

Bid based marketplaces clearing above list

On genuine bid based GPU marketplaces, the interruptible price is set by whoever is willing to pay most for the same physical host. During demand spikes the clearing bid can rise to meet, and occasionally exceed, the operator's own fixed on demand rate for equivalent hardware. This is a literal price inversion, visible on the screen, and it happens because interruptible buyers competing for scarce hosts are not bidding against a list price, they are bidding against each other. Anyone running automated bidding without a hard ceiling is exposed to it.

Cross provider inversion

The largest inversion has nothing to do with interruptions. Hyperscaler spot pricing routinely sits above neocloud on demand pricing for the same accelerator. A p5.48xlarge carries a published on demand rate of $98.32 per hour for eight H100s, which is $12.29 per GPU hour. Even a 40 percent spot discount only brings that to $7.37. Published on demand list pricing from specialist GPU providers has been sitting near $2.99 per GPU hour for the same 8x H100 SXM configuration. The discounted hyperscaler price is more than twice the undiscounted specialist price, for identical silicon, with eviction risk attached.

Purchase mode Per GPU hour vs AWS on demand
On demand, AWS p5.48xlarge, 8x H100 $12.29 baseline
Spot at 40 percent off, AWS p5.48xlarge $7.37 40% lower
Spot at 20 percent off, AWS p5.48xlarge $9.83 20% lower
On demand, specialist provider, 8x H100 SXM $2.99 76% lower
List prices as published at time of writing. Verify current figures in the GPU Price Explorer before relying on them for a purchase decision.

That comparison is the one worth running, and almost nobody runs it. Teams benchmark hyperscaler spot against hyperscaler on demand, congratulate themselves on the delta, and never test the assumption that the baseline is the right baseline. The hyperscaler premium survives the spot discount comfortably.

DataStorage.com Podcast
AI Infrastructure Is Changing Everything, with Russ Artzt on GPUs, Neo-Clouds and the Future of Cloud
The co-founder of CA Technologies on why neoclouds exist at all, how GPU economics differ from the CPU era, and why most engineers still do not know the alternative providers are there.
Listen to the Episode
The DataStorage.com Podcast / Episode 5

What to Do With Spot GPU Capacity in 2026

None of this means abandon spot. It means stop treating the headline discount as the decision input. A few rules hold up well in practice.

Calculate the break-even interruption interval before you bid, not after the run. It takes one line of arithmetic and it converts an abstract discount into a concrete reliability requirement you can actually test.

Measure your real mean time between interruptions before you commit a pipeline to spot. Run a throwaway job on the target instance type in the target region for a week and count the evictions. Provider published interruption frequencies are fleet wide averages and tell you very little about the specific accelerator and region you need.

Treat checkpoint size as a cost variable rather than a fixed property of the model. Halving checkpoint write time moves the break-even point materially, and sharded or asynchronous checkpointing is usually cheaper to implement than the compute it saves.

Never put latency sensitive inference on spot. The break-even model assumes work can be redone. Missed inference requests cannot be, and no discount prices that correctly.

Compare against the right baseline. If the effective spot price at your interruption rate lands above a specialist provider's on demand rate, the spot discount is not the decision, the provider is. And when the discount on offer is under 30 percent, a short committed term at a lower rate is almost always the better instrument: you get the price certainty and you drop the eviction engineering entirely.

Spot GPU Decision Rules
  • Calculate the break-even interruption interval before bidding, not after the run.
  • Measure your real interruption rate for a week on the target instance type and region. Fleet wide averages tell you nothing useful.
  • Treat checkpoint size as a cost variable. Halving write time moves the break-even point materially.
  • Never put latency sensitive inference on spot. Missed requests cannot be redone.
  • Benchmark against specialist provider on demand pricing, not only against hyperscaler on demand pricing.
  • When the discount is under 30 percent, a short committed term usually beats spot on both price and operational load.

Key Takeaways

What to Remember
  • Spot GPU discounts have compressed from the advertised ceiling of around 90 percent into a band that commonly sits between 20 and 50 percent, because accelerators are the constrained resource in every hyperscaler fleet and constrained resources do not enter discount pools.
  • The break-even mean time between interruptions is 0.5 divided by (discount minus 0.05) hours, on a one hour checkpoint interval with a three minute write. At 40 percent off that is 86 minutes. At 20 percent off it is 200 minutes.
  • Realised discounts run far below headline discounts. Holding interruptions at one every three hours, a 50 percent headline returns about 36 percent and a 20 percent headline returns minus 2 percent.
  • Effective cost inversion is invisible in billing data, because no cost export distinguishes useful GPU hours from hours spent redoing lost work.
  • Hyperscaler spot pricing frequently sits above specialist provider on demand pricing for identical hardware, which makes cross provider comparison more valuable than any spot optimisation.

FAQ

Is spot GPU pricing still worth using in 2026?
Only for interruptible batch work where the discount clears your measured break-even point. Calculate the break-even interruption interval, measure your real interruption rate for a week, and proceed only if there is margin between them.
What is a realistic spot discount on H100 capacity right now?
Commonly 20 to 50 percent off the on demand rate, not the up to 90 percent advertised for the general fleet. Older A100 and V100 class hardware still sees deeper discounts, because that is where the idle capacity sits.
How do I calculate whether spot is actually cheaper than on demand?
Divide the spot price by one minus the wasted share of GPU time. The wasted share is half the checkpoint interval divided by the mean time between interruptions, plus checkpoint write time divided by the checkpoint interval. If the result exceeds the on demand rate, the discount is not covering the interruptions.
What is the break-even interruption interval for spot GPUs?
It is 0.5 divided by (discount minus 0.05) hours. At 40 percent off that is 86 minutes. Interruptions more frequent than that cost you money.
Can spot prices actually exceed on demand prices?
On AWS the quoted spot price is capped at on demand, so not on the price sheet. It happens three other ways: bid based marketplaces clearing above list, effective cost once interruption overhead is counted, and hyperscaler spot sitting above specialist provider on demand rates.
Why did GPU spot discounts shrink when general cloud spot discounts did not?
A spot pool needs capacity nobody is paying full price for, and accelerators have been the scarcest resource in every large fleet since 2023. Reserved capacity products also absorbed the short term demand that used to keep the GPU spot market liquid.
Spot stopped being a market and became an administered discount. Price it like one: measure the interruptions, do the arithmetic, and compare against every provider rather than the one you already use.
Weekly Newsletter
Stay Ahead in Cloud Infrastructure
Join 1,200+ CTOs, architects, and cloud professionals who get our weekly briefing on storage strategy, GPU compute, and cloud cost intelligence.
Subscribe Free  →

References

Share this article

🔍 Browse by categories

Free Cloud Cost Calculator

Compare AWS, Google Cloud, Azure, and alternatives like Backblaze B2 Discover how much you could save in seconds

🔥 Trending Articles

Newsletter

Stay Ahead in Cloud
& Data Infrastructure

Get early access to new tools, insights, and research shaping the next wave of cloud and storage innovation.