GPU Platform Hub

Discover

  • Home
  • GPU Catalog
  • Compare
  • Recommend

Design & Deploy

  • Architecture Center
  • Kubernetes
  • Cluster Bootstrap
  • Compute
  • Networking
  • Storage

Operate

  • AI Workloads
  • DevSecOps & MLOps
  • Security
  • Observability
  • Day-Two Ops
  • Troubleshooting

Intelligence

  • AI Operations
  • KnowledgeOps
  • Knowledge Base
  • Labs
  • Executive Center
  • Administration
HomeLabs
Intelligence

Labs

Hands-on exercises for GPU operators — from validating GPU availability to serving an LLM and diagnosing a pending pod. Work through the tasks and track your progress.

Enablement
beginner 15m

Validate GPU availability in Kubernetes

Confirm that GPUs are discovered, enabled, and schedulable on your cluster.

Start lab
beginner 25m

Deploy the NVIDIA GPU Operator

Install the GPU Operator with Helm and validate GPU enablement.

Start lab
intermediate 30m

Configure MIG on an A100/H100

Partition a GPU into MIG instances and schedule onto them.

Start lab
Serving
intermediate 30m

Serve an LLM with vLLM

Deploy vLLM and send an inference request.

Start lab
intermediate 40m

Build a minimal RAG pipeline

Stand up a vector store, ingest documents, and answer a grounded query.

Start lab
Training
advanced 45m

Run a distributed training job

Launch a multi-GPU training job with gang scheduling.

Start lab
Operations
beginner 20m

Configure DCGM GPU monitoring

Scrape GPU metrics with DCGM-exporter and view them.

Start lab
beginner 15m

Diagnose a Pending GPU pod

Find and fix why a GPU pod won't schedule.

Start lab
Platform Assistant
Context

/labs

Try asking
How many GPUs for a 405B training run?NVLink vs Infinity Fabric — when does it matter?Draft a bill of materials for an inference cluster

Answers are computed by the platform's own engines and labelled as confirmed facts vs. assumptions. Agents never execute changes without a confirmed, gated action. Works with no API key; connect a model to upgrade the reasoning.

Open AI Operations Center