Rent an H100 inside an AWS p5 instance at on-demand list and you are paying about $12.29 per GPU-hour. Rent the same accelerator from a specialist provider and the number starts with a 2. That is a spread of more than 6x on identical hardware.
An NVIDIA H100 SXM is the same part wherever it is racked. 80 GB of HBM3, roughly 3.35 TB per second of memory bandwidth, 900 GB per second of NVLink. The silicon does not know whose data center it is sitting in.
The price knows. Rent one inside an AWS p5 instance at on-demand list and you are paying about $12.29 per GPU-hour. Rent the same accelerator from a specialist GPU provider and the number starts with a 2, sometimes a 1. That is not a regional quirk or a rounding error. It is a spread of more than 6x on identical hardware.
Most teams find this out the way you find a leak: after the bill. The useful question is not why hyperscalers are expensive, but what exactly the premium buys and whether your workload needs any of it. Some genuinely do. Most of the ones burning six figures a month do not.
The comparison below normalizes everything to cost per GPU-hour at published on-demand rates as of mid-2026. That matters, because hyperscalers do not sell you a GPU. They sell a node with eight of them, so the headline instance price has to be divided by eight before it means anything.
Three tiers fall out of the data, not the two most people assume. Hyperscalers cluster between roughly $11 and $12.50 per GPU-hour. Tier one neoclouds that own their fleet and run a real support organization sit between about $2.99 and $6.16. Below that runs a long tail of marketplaces and resellers under $2.50. Each tier sells a materially different product, and the gap between tier two and tier three is where most buying mistakes happen. One caveat before the arithmetic: these are list prices, and almost no large hyperscaler customer pays list.
| Provider and offering | Per GPU-hour | Vs cheapest |
|---|---|---|
|
Hyperscaler
AWS p5.48xlarge, 8x H100 SXM at $98.32/hr
|
$12.29 | 6.2x |
|
Hyperscaler
Azure ND H100 v5, 8x H100 at about $98/hr
|
$12.24 | 6.1x |
|
Hyperscaler
Google Cloud a3-highgpu-8g at about $88/hr
|
$11.03 | 5.5x |
|
Tier one neocloud
CoreWeave HGX H100, published on-demand
|
$6.16 | 3.1x |
|
Tier one neocloud
Lambda Labs 8x H100 SXM on-demand
|
$2.99 | 1.5x |
|
Tier one neocloud
Vultr H100 on-demand
|
$2.99 | 1.5x |
|
Marketplace long tail
Resold and spot style capacity
|
$1.99 | 1.0x |
The premium is not one thing. It is five things stacked together, and they unbundle differently depending on what you run.
A p5.48xlarge bundles eight H100s with 192 vCPUs, 2 TB of system memory, 30 TB of local NVMe, and 3,200 Gbps of EFA networking. If your inference service needs two GPUs, you still rent eight. At 25 percent utilization, a $12.29 list rate becomes an effective $49 per GPU-hour actually used, and no negotiated discount touches that. Minimum purchase unit is the most underpriced variable in GPU budgeting, and it is invisible on the invoice because the invoice is correct.
The hyperscaler rate includes identity, VPC networking, dozens of regions, audited compliance boundaries, integration with the managed services your pipeline already calls, and capacity that appears without a sales conversation. For a regulated enterprise whose control plane already lives in one account, moving GPU workloads elsewhere means rebuilding identity, networking, logging, and audit evidence in a second place. That project has a price, and it is rarely in the spreadsheet that shows the 6x.
Ingress is free, egress is taxed. Moving a 500 TB training corpus out of S3 at list egress rates costs roughly $45,000 in a single exit. Tiered pricing and committed discounts bring that down, but the shape holds: cheaper compute is only cheap once your data is next to it. Teams that leave the corpus with the old provider and stream it across the internet each epoch pay that toll repeatedly, which is how a migration that modeled as a 70 percent saving lands as a 10 percent one. It belongs on the same list as every other line item your provider is not volunteering.
Committed use discounts, savings plans and reserved capacity, and enterprise agreements routinely take 30 to 60 percent off published GPU rates for one to three year commitments. So the honest comparison for an enterprise buyer is not list against list. It is a negotiated hyperscaler rate against a committed neocloud rate, and that narrows a 6x gap to something closer to 2x or 3x. Still very large, but the negotiated number is the one to take into the meeting. It also comes with its own cost: you have just committed to a specific GPU generation for three years in a market where the next generation ships roughly every 18 months.
Purpose built GPU halls do not cross subsidize a catalog of 200 other services. They site for power rather than latency to every enterprise office, they finance hardware across a five to seven year useful life rather than three, and they carry occupancy risk a hyperscaler does not. An idle GPU is a total loss for a specialist, so they price to fill the hall. That is the actual reason the number is low, and it is neither generosity nor a promotional rate that expires. The same logic explains why accelerator selection and provider selection are increasingly the same decision.
It is also why the long tail deserves suspicion. A meaningful share of GPU cloud brands do not own a single accelerator. They resell capacity from an operator upstream, add margin, and pass your support ticket along the chain. So the first vetting question for anyone quoting a number that looks too good is blunt: do you own the GPU?
Take one 8x H100 node running continuously for 30 days. At AWS on-demand list, $98.32 per hour across 720 hours is $70,790. The same node at a tier one neocloud around $3.00 per GPU-hour is $24 per hour, or $17,280. A long tail marketplace at $2.10 per GPU-hour is $12,096. The gap between the first two numbers is $53,510 per node per month, and $642,000 per node per year.
Now add the data. If the training set is that 500 TB corpus in S3, the one time move costs about $45,000. Against a single 30 day job, that consumes 84 percent of the saving and the migration was close to pointless. Against a node running twelve months, it is 7 percent and the decision is obvious. Duration and data volume, not hourly rate, decide this.
The corollary is worth internalizing: the cheaper the compute, the more storage architecture matters. Zero egress object storage such as Backblaze B2, Wasabi, or Cloudflare R2 exists precisely so the next provider switch is a compute decision rather than a data decision. In a market where GPU availability moves quarterly, that optionality is worth more than the per hour delta.
Five situations make the hyperscaler rate the correct answer, and pretending otherwise costs more than it saves. First, genuinely bursty demand, where you need capacity within minutes and cannot wait on a contract. Second, real data gravity: petabytes already resident with one provider and a job measured in weeks rather than quarters. Third, regulatory and procurement constraints, such as an existing BAA, a FedRAMP boundary, or an approved vendor list that takes nine months to amend. Fourth, tight coupling to managed training and inference services, where re-platforming burns more engineering time than the price gap returns. Fifth, an enterprise discount program with committed spend you will forfeit if you move the workload out.
Outside those five, the premium buys convenience, and convenience at $642,000 per node per year deserves a decision rather than a default.
A low per hour rate is a claim, not a product. Six questions separate operators from resellers, and all six should be answered before a pilot rather than after. Do you own the GPU, or are you reselling capacity from an operator upstream? Ownership determines whether support is a conversation or a relay. What is the interconnect? Eight GPUs in a chassis over PCIe is a different product from InfiniBand across 64 nodes, irrelevant for single node fine tuning and decisive for distributed training. Where does storage sit, and what does a full read cost? Adjacent NVMe with high east-west bandwidth behaves nothing like an object store across a wide area link. What is the reclaim policy, and how much notice do you get before an instance is pulled? What is actually in the contract: minimum term, egress terms, support SLA, and whether the rate is locked for the duration? And finally, what happens at 3am on a Sunday, which is the only support question that has ever mattered.
This takes an afternoon and produces a number you can defend in a budget review. Start by pulling GPU-hours consumed against GPU-hours paid for last month, split by workload. The gap between those two figures is your real utilization, and it usually dwarfs the provider price difference. Then price the same node shape at two tier one neoclouds and one marketplace, so you are working from a current market rate rather than a remembered one. Calculate the egress cost of moving one full dataset once, and divide the monthly saving into it to get your break even in months. If break even lands under three months and the workload carries no regulatory constraint, run a pilot on a single node.
Then do the thing that pays for the exercise: take the quote back to your incumbent account team before the pilot finishes. A credible alternative is the only lever that reliably moves an enterprise discount, and the renegotiated rate is often worth more than the migration you were pricing.
The 6x is real, but it is not a discount waiting to be claimed. It is a bill of materials. Read it, decide which line items your workload actually needs, and pay for those.
Compare AWS, Google Cloud, Azure, and alternatives like Backblaze B2 Discover how much you could save in seconds