NVIDIAEstimated
GB200 Grace Blackwell
Blackwell + Grace · Blackwell · Superchip (2× B200) · 2024
A Grace CPU tightly coupled to two Blackwell GPUs via NVLink-C2C. Designed for rack-scale NVL72 training and inference.
LLM trainingLLM inferenceMultimodalHPC
Precision fingerprint
6432t3216BF168i8
Memory
384 GB
HBM3e
Bandwidth
16 TB/s
peak
TDP
2700 W
liquid
Max model
~160B
FP16, planning est.
Compute throughput
| FP64 | 80 TFLOPS |
| FP32 | 80 TFLOPS |
| TF32 | 2.20 PFLOPS |
| FP16 | 4.50 PFLOPS |
| BF16 | 4.50 PFLOPS |
| FP8 | 9.00 PFLOPS |
| INT8 | 9.00 PFLOPS |
Platform & software
InterconnectNVLink 5 — 1.8 TB/s/GPU; NVLink-C2C to Grace
PCIePCIe 6.0
Coolingliquid
MIGSupported
PartitioningMIG per GPU
VirtualizationMIG
FrameworksCUDA, TensorRT-LLM, Triton, NeMo, vLLM
AvailabilityAzure, GCP, OCI, bare-metal
Known limitations
- ·Rack-scale, liquid-cooled system — not a drop-in card
- ·Grace CPU couples compute and host memory