low riskScaling & replacement

Cluster scaling (add GPU capacity)

Add GPU nodes to absorb growing demand, validating enablement before admitting workloads.

Change scope

New nodes / node pool

Maintenance impact

None — additive.

Prerequisites
  • ·Capacity forecast (see Executive Center)
  • ·Power/cooling/network headroom (see Compute/Networking)
Pre-checks
  • ·Fabric + power capacity confirmed for the new nodes
Procedure
  1. 1

    Provision the new nodes / node pool (see Cluster Bootstrap).

  2. 2

    Validate GPU enablement and labels/taints before admitting workloads.

Validation
  • New nodes Ready; GPUs allocatable
  • Queue wait times drop
Rollback

Cordon/drain and remove the new nodes if they misbehave.

Communication

Announce added capacity; update quotas if needed.