nvidia-logo

Nvidia GB200

Grace Blackwell superchip combining a Blackwell GPU with an Arm-based Grace CPU for large-scale AI training and inference systems.

Release

2024

GPU Class

AI Superchip / Data Center Accelerator

Architecture

Grace Blackwell

PRICE SNAPSHOT

Loading price comparison...
Loading live GPU prices...

On-premise Module

~$60k–$100k*

Turnkey System

~$2M–$3M+

Cloud Pricing
(per GPU/hr)

~$3.00–$10.00/hr

chip identity

Nvidia GB200

On-premise Module

GB200

GPU Class

AI Superchip / Data Center Accelerator

Release

2024

Architecture

Grace Blackwell

Target Workload

  • Trillion-parameter LLM training
  • Hyperscale generative AI infrastructure
  • Multi-modal AI models
  • High-throughput inference systems

Large distributed AI clusters

Compatible Platforms

  • GB200 NVL72 AI systems
  • HGX Blackwell infrastructure
  • Hyperscale AI clusters built by cloud providers

Interconnect: NVLink 5 + NVLink Switch

Software Stack: CUDA, cuDNN, NVIDIA AI Enterprise, NVIDIA AI frameworks

Ideal Buyer Profile

Organizations that typically deploy GB200 systems include:

  • Hyperscale cloud providers
  • AI research labs developing foundation models
  • Enterprises building large generative AI platforms
  • Institutions training multi-trillion parameter models

The architecture is particularly suited for organizations requiring rack-scale AI infrastructure with integrated CPU-GPU compute resources.

Availability Notes

The GB200 is typically deployed through complete AI infrastructure systems rather than as a standalone accelerator card. Early deployments are concentrated among hyperscale cloud providers and large research organizations building advanced AI clusters.

Because the platform targets extremely large training systems, it is most commonly integrated into pre-configured rack-scale AI systems.

Recent Developments

  • Announcement of GB200 NVL72 rack-scale AI infrastructure systems.
  • Adoption by hyperscale cloud providers building large generative AI clusters.
  • Expansion of the Grace Blackwell architecture across NVIDIA’s AI platform ecosystem.

overview

The NVIDIA GB200 is a next-generation AI superchip that combines a Blackwell GPU with a Grace CPU in a tightly integrated module designed for large-scale AI infrastructure. This architecture is intended to accelerate generative AI workloads while improving memory bandwidth and system efficiency.

Unlike traditional GPU-only accelerators, the GB200 integrates a high-performance Arm-based CPU with the GPU using high-bandwidth interconnects. This design enables faster data movement between CPU and GPU resources, reducing bottlenecks in large AI training pipelines.

The GB200 is typically deployed as part of large-scale AI systems such as the GB200 NVL72, which connects dozens of GPUs using NVLink switching technology. These systems are designed to support the training and deployment of massive generative AI models across hyperscale infrastructure.

Key specifications

GB200 Superchip Specifications

GB200 Superchip Specifications

Specification Value
Architecture NVIDIA Grace Blackwell
GPU Component Blackwell GPU
CPU Component NVIDIA Grace Arm CPU
Memory 192 GB HBM3e (GPU)
Memory Bandwidth ~8 TB/s
Interconnect NVLink 5
CPU-GPU Link NVLink-C2C
Form Factor Superchip Module
Max TGP ~1000+ W (combined module)
Precision Support FP64, TF32, FP32, FP16, BF16, FP8, FP4
Typical AI Compute ~20 PFLOPS (FP4 Tensor, estimated)
Process Node TSMC 4NP
Transistor Count ~200+ Billion (GPU component)
Multi-GPU Scaling NVLink Switch fabric
System Deployment NVL rack-scale clusters

Performance Summary

  • AI Training: The GB200 is designed to train extremely large generative AI models efficiently across multi-node clusters.
  • Memory Throughput: HBM3e memory provides extremely high bandwidth for large model parameters and training datasets.
  • CPU-GPU Integration: NVLink-C2C provides high-speed communication between the Grace CPU and Blackwell GPU.
  • Cluster Scaling: NVLink Switch technology enables large rack-scale AI systems with dozens of GPUs.

Compared with standalone GPUs, the GB200 provides a more integrated architecture optimized for large AI systems where CPU-GPU coordination and distributed scaling are critical.

primary use case

  • Training trillion-parameter language models
  • Large-scale generative AI infrastructure
  • Multi-modal AI model training
  • Hyperscale AI inference clusters
  • Enterprise AI research platforms

The superchip is designed for environments where AI training systems scale across dozens or hundreds of GPUs.

Alternatives & Upgrade Path

Comparable NVIDIA GPUs:

  • B200: Blackwell GPU accelerator used in AI training clusters.
  • H200: Hopper-generation GPU with large HBM3e memory capacity.

Competing AI Hardware:

These accelerators compete in the market for large-scale AI infrastructure and hyperscale training deployments.

Related Chips & Providers

Related NVIDIA GPUs:

  • B200
  • B100
  • H200

Complementary Silicon:

  • NVIDIA Grace CPU
  • NVLink Switch networking infrastructure

SUMMARY

The NVIDIA GB200 is a next-generation AI superchip that integrates a Blackwell GPU with an Arm-based Grace CPU to support massive generative AI workloads. Designed for hyperscale infrastructure, it combines high-bandwidth memory, advanced tensor compute, and high-speed CPU–GPU interconnects.

Deployed in rack-scale systems such as the GB200 NVL72, the architecture enables efficient training and deployment of large language models and next-generation generative AI systems across large GPU clusters.

 

Newsletter

Stay Ahead in Cloud
& Data Infrastructure

Get early access to new tools, insights, and research shaping the next wave of cloud and storage innovation.