Operate

Live Telemetry

Current fleet state from the configured telemetry provider — per-cluster utilization, thermals, power, and ECC, with alerts assessed against the reference bands and deep-linked to their troubleshooting workflows.

Provider sample — a synthetic snapshot derived from the sample fleet + GPU catalog (no cluster required). Set TELEMETRY_PROVIDER=prometheus to scrape a real DCGM/ROCm Prometheus.

sample · synthetic
GPUs

1,048

9 clusters

Avg util

65%

fleet-wide

Draw

379 kW

board power

Critical

1

1 clusters

Warnings

87

3 clusters

ECC / throttle

0 / 3

uncorrectable / capped

Firing alerts
Cluster health
ClusterVendorUtilizationPeak tempPowerFaultsState

prod-train-us

H100 SXM5 · 256 GPU

NVIDIA
82%
92°C
154.5 kWthrottle 3critical

prod-train-eu

H100 SXM5 · 128 GPU

NVIDIA
78%
77°C
74.7 kWwatch

prod-train-us2

Instinct MI300X · 64 GPU

AMD
70%
81°C
37.1 kWwatch

prod-inf-us

L40S · 192 GPU

NVIDIA
64%
71°C
49.3 kWwatch

prod-inf-eu

L40S · 96 GPU

NVIDIA
57%
68°C
23.1 kWhealthy

rag-us

A100 80GB SXM · 64 GPU

NVIDIA
73%
74°C
20.4 kWhealthy

vision-us

L4 · 120 GPU

NVIDIA
55%
66°C
5.8 kWhealthy

dev-us

A30 · 48 GPU

NVIDIA
23%
51°C
3.5 kWhealthy

legacy-us

A100 40GB PCIe · 80 GPU

NVIDIA
37%
58°C
10.8 kWhealthy

Snapshot 2026-07-26 18:59:29 UTC · bands from the observability reference.