The exact same NVIDIA H100 rents for $1.49 an hour on one platform and $6.98 an hour on another. Identical silicon, a 4.7x price difference, depending entirely on which provider you pick and how you shop.
The exact same NVIDIA H100 rents for $1.49 an hour on one platform and $6.98 an hour on another. Identical silicon, a 4.7x price difference, depending entirely on which provider you pick and how you shop for it. This report pulls together current mid-2026 rental pricing across H100, H200, and B200 GPUs, from hyperscalers, neoclouds, and peer-to-peer marketplaces, so you know what a fair price actually looks like before you sign anything.
The pricing story has a twist worth understanding before the numbers: H100 rental rates fell sharply through 2024 and 2025 as supply caught up with the initial shortage, then reversed and started climbing again in late 2025 and into 2026 as demand accelerated faster than new capacity came online. If your last pricing benchmark is more than six months old, it's probably wrong in the wrong direction.
H100 remains the most widely available and most heavily shopped GPU on the market, which makes its pricing spread the clearest illustration of how much provider choice matters.
| Metric | Rate | Notes |
|---|---|---|
| Market median (GPU-hour) | $2.95 to $3.46 | Cross-provider benchmark |
| Cheapest (peer-to-peer) | $1.49 (Vast.ai) | Spot/preemptible risk |
| Neocloud on-demand | $2.00 to $2.99 | GMI Cloud, Lambda, RunPod |
| Hyperscaler on-demand | $3.90 to $6.98 | AWS, GCP, Azure |
| Multi-GPU node, normalized | $6.16 to $10.00 | CoreWeave, Oracle Cloud |
AWS cut its H100 pricing roughly 44 percent in mid-2025, bringing EC2 P5 instances down to around $3.90 per GPU-hour, still well above neocloud rates but a meaningful improvement from earlier hyperscaler pricing. That downward trend reversed by early 2026: multiple market trackers report H100 rental rates climbing nearly 40 percent since October 2025, from roughly $1.70 to $2.35 an hour on the affected platforms, as available capacity tightened faster than expected. Spot and preemptible pricing remains the cheapest path in, as low as $1.20 an hour, but comes with the risk of your instance being reclaimed mid-job.
H200 offers the same compute architecture as H100 with significantly more high-bandwidth memory, 141GB versus 80GB, which matters most for large model inference and memory-bound workloads rather than raw training throughput.
| Metric | Rate | Notes |
|---|---|---|
| Neocloud on-demand | $2.60 to $4.49 | GMI Cloud, CoreWeave, Lambda, RunPod |
| Typical premium over H100 | 20% to 40% | Same provider, same commitment tier |
Pricing for H200 varies meaningfully by provider ownership model. GMI Cloud, which owns its fleet outright, prices H200 at $2.60 per GPU-hour on demand, no minimum commitment. CoreWeave and Lambda run higher, generally in the $3.89 to $4.49 range, reflecting a combination of InfiniBand networking, managed support tiers, and brand positioning rather than a difference in the underlying chip. The gap between the cheapest and priciest H200 on-demand rate can run over $2,600 a month per GPU at continuous use, a difference that compounds fast on any multi-GPU cluster.
B200, NVIDIA's Blackwell-generation flagship, carries the highest hourly rate of the three, but the calculus changes once you price by task rather than by hour.
| Metric | Rate | Notes |
|---|---|---|
| Broad market range | $4.50 to $7.00 | Q2 2026 |
| Neocloud on-demand | $3.99 to $5.50 | DataCrunch, Lambda, CoreWeave |
| Hyperscaler capacity block | $9.36 | AWS |
| Spot | $5.34 | Spheron |
B200 delivers roughly 2.5 times the training performance of H100 despite typically costing less than double the hourly rate, which means cost-per-training-result frequently favors B200 even at its higher sticker price. B200 availability remains the real constraint rather than price: many providers still offer reservation-only access for meaningful cluster sizes as of mid-2026, and waitlists are common for large multi-GPU configurations even though single-GPU on-demand access has improved.
Four factors explain nearly all of the spread between the cheapest and priciest listing for identical silicon.
Hyperscalers and several neoclouds sell multi-GPU nodes as the base unit, not individual GPUs. CoreWeave's 8-GPU H100 node lists at $49.24 an hour, which normalizes to $6.16 per GPU-hour, a very different number than the headline node price suggests if you don't do the division. Always normalize to a per-GPU rate before comparing across providers.
Providers that own their GPU fleet outright, rather than reselling capacity through intermediaries, can price closer to their actual cost basis. Reseller markup stacks on top of whatever the underlying owner is already charging, which is part of why the neocloud ownership question matters as much for pricing as it does for support quality.
Hyperscaler platforms typically run 10 to 15 percent virtualization overhead compared to bare-metal or lightly virtualized neocloud offerings, meaning an hour of hyperscaler GPU time produces less usable compute than an hour on a provider running closer to the metal, even before comparing sticker prices.
Reserved and committed-use pricing typically runs 20 to 40 percent below on-demand rates in exchange for a one-month to twelve-month commitment. Spot and preemptible pricing goes further, often 60 to 91 percent below on-demand on platforms like Google Cloud, in exchange for the risk of losing the instance with little notice.
For sustained, moderate-to-high utilization workloads, purchasing hardware outright can still beat renting on total cost. Current market analysis puts the breakeven point for owning an H100 at roughly 6 to 14 months of continuous, moderate-utilization use, after which the amortized cost of ownership undercuts ongoing rental. For spiky, experimental, or short-duration workloads, renting remains the clearer choice, since idle owned hardware is a sunk cost that rented capacity simply doesn't carry.
Whether a provider quotes a node price, a cluster price, or a per-GPU price, convert everything to the same unit before drawing conclusions. A cheap-looking node price can hide an expensive per-GPU rate, and vice versa.
Ask directly. Providers reselling capacity through intermediaries carry markup and support latency that owned-fleet providers don't, and that difference tends to show up in both price and reliability over time.
On-demand pricing is the most expensive per hour but carries no lock-in. Reserved pricing saves 20 to 40 percent for workloads with a predictable, sustained duration. Spot pricing saves the most but should only go toward workloads that can tolerate interruption without losing meaningful work.
A higher hourly rate on a faster chip can still be the cheaper option once you account for how much less time the workload takes to finish. This matters most for B200, where the throughput premium over H100 often outweighs the higher sticker price for training workloads.
GPU rental pricing has moved in both directions within the same 18-month window, falling through 2024 and 2025, then climbing again through late 2025 and into 2026. A benchmark from six months ago is not a reliable guide to what you'll pay today.
The sticker price on a GPU listing tells you almost nothing until you know who owns the hardware, what unit they're pricing, and how recently the number was updated.
Compare AWS, Google Cloud, Azure, and alternatives like Backblaze B2 Discover how much you could save in seconds