NVIDIAEstimated

GB200 Grace Blackwell

Blackwell + Grace · Blackwell · Superchip (2× B200) · 2024

A Grace CPU tightly coupled to two Blackwell GPUs via NVLink-C2C. Designed for rack-scale NVL72 training and inference.

LLM trainingLLM inferenceMultimodalHPC
Precision fingerprint
6432t3216BF168i8
Memory

384 GB

HBM3e

Bandwidth

16 TB/s

peak

TDP

2700 W

liquid

Max model

~160B

FP16, planning est.

Compute throughput

FP6480 TFLOPS
FP3280 TFLOPS
TF322.20 PFLOPS
FP164.50 PFLOPS
BF164.50 PFLOPS
FP89.00 PFLOPS
INT89.00 PFLOPS

Platform & software

InterconnectNVLink 5 — 1.8 TB/s/GPU; NVLink-C2C to Grace
PCIePCIe 6.0
Coolingliquid
MIGSupported
PartitioningMIG per GPU
VirtualizationMIG
FrameworksCUDA, TensorRT-LLM, Triton, NeMo, vLLM
AvailabilityAzure, GCP, OCI, bare-metal
Known limitations
  • ·Rack-scale, liquid-cooled system — not a drop-in card
  • ·Grace CPU couples compute and host memory