GPU Cluster Reliability Checklist for Production AI Infrastructure
A production GPU cluster fails differently than the CPU infrastructure most platform engineers learned on. Silent memory corruption, NVLink degradation, and thermal throttling can all masquerade as "slow training" for weeks before anyone notices. This checklist covers the hardware validation, monitoring, runbook discipline, and operational patterns that separate a reliable GPU platform from an expensive science experiment.
Pre-Deployment: Hardware Validation
Every GPU that enters production should pass a validation suite before it joins the scheduler pool. This is not paranoia—it's the difference between catching a flaky NVLink or marginal memory module during acceptance testing versus discovering it three weeks into a training run.
Burn-In and Memory Testing
Run a memory test that exercises the full framebuffer under thermal load. The NVIDIA diagnostic toolkit includes nvidia-smi -q for basic health and dcgmi diag -r 3 for a longer diagnostic sweep that stresses compute, memory, and interconnect. A level-3 diagnostic typically takes 15–30 minutes per GPU and will catch most marginal hardware before it causes a production incident.
Check for remapped rows and existing ECC errors. On A100 and newer, row remapping is logged via XID 63 (successful remapping recorded) and XID 64 (remapper exhausted). A GPU that has already remapped rows is not necessarily bad, but a high count or a failure to remap is a red flag. Query ECC counters with nvidia-smi --query-gpu=ecc.errors.corrected.aggregate.total,ecc.errors.uncorrected.aggregate.total --format=csv and establish a baseline.
NVLink and Interconnect Validation
NVLink bandwidth degrades silently. A single failed link in an NVSwitch topology can cut intra-node bandwidth by a seventh or more, and the training job will simply run slower without an obvious error. Use nvidia-smi nvlink --status to verify all links are active, then run a bandwidth test across the full node topology. The nccl-tests all_reduce_perf benchmark is the standard: it will report algbw (algorithm bandwidth, the effective rate seen by the application) and busbw (bus bandwidth, which for ring all-reduce is roughly twice algbw for large node counts).
For intra-node validation, expect aggregate NVLink bandwidth per GPU around 600 GB/s on A100 (third-generation NVLink) and around 900 GB/s on H100 (fourth-generation). If you see significantly lower numbers, check nvidia-smi topo -m for topology issues and nvidia-smi nvlink -e for per-link error counters. XID 74 indicates an NVLink error; investigate immediately.
Network Fabric Validation
For multi-node collectives, the network is the long pole. Run nccl-tests across a representative multi-node job (at least four nodes, ideally a full rail or pod). On 200 Gb/s HDR InfiniBand, expect busbw approaching the per-GPU NIC line rate—roughly 23–25 GB/s per GPU for a network-bound all-reduce, not 40–45 GB/s. The busbw ceiling is the physical link rate divided by eight (200 Gb/s ≈ 25 GB/s) minus protocol overhead.
If busbw is significantly lower, check for misconfigured NCCL topology (the NCCL_TOPO_FILE or auto-detected topology), verify that GPUDirect RDMA is enabled (NCCL_NET=IB, NCCL_IB_GID_INDEX set correctly), and confirm that the InfiniBand subnet manager has all links up. Use ibstat and ibdiagnet to validate fabric health.
Monitoring: The Metrics That Matter
DCGM (Data Center GPU Manager) is the standard telemetry source. The dcgm-exporter DaemonSet exposes metrics to Prometheus, and the key is knowing which metrics actually tell you something.
Utilization Metrics and the GPU_UTIL Trap
DCGM_FI_DEV_GPU_UTIL is a coarse "GPU busy" percentage: it reports high if the SM had at least one active kernel during the sampling window, even if that kernel was blocked on memory or used only a fraction of the available parallelism. A GPU can show 95% GPU_UTIL while delivering 20% of its potential throughput because the workload is memory-bound or the batch size is too small.
The profiling (DCP) fields give you the real picture:
DCGM_FI_PROF_SM_ACTIVE: the fraction of time the SMs were actively executing instructions.DCGM_FI_PROF_PIPE_TENSOR_ACTIVE: the fraction of time the Tensor Cores were active (critical for transformer training).DCGM_FI_PROF_DRAM_ACTIVE: memory controller utilization, which tells you if you're memory-bound.DCGM_FI_PROF_SM_OCCUPANCY: average active warps per SM, a proxy for how well the kernel is hiding latency.
Alert on low SM_ACTIVE or PIPE_TENSOR_ACTIVE when GPU_UTIL is high—that's the signature of a starved or misconfigured workload. Also monitor DCGM_FI_DEV_MEM_COPY_UTIL (memory controller utilization) as a simpler memory-bound indicator.
Hardware Health and Error Counters
Track DCGM_FI_DEV_XID_ERRORS, which reports the last XID error code. The critical XIDs for hardware failures are:
| XID Code | Meaning | Action |
|---|---|---|
| 48 | Double-bit ECC error (uncorrectable memory error) | Drain node, RMA GPU |
| 79 | GPU fell off the bus (inaccessible) | Immediate node removal, RMA |
| 63 | ECC page retirement recorded (A100+) | Log and monitor; check row-remap count |
| 64 | Row remapper exhausted | Drain node, RMA GPU |
| 94/95 | Contained/uncontained ECC error (A100+) | Drain node, RMA GPU |
XID 13 (graphics engine exception) and XID 31 (memory page fault) are usually application errors (illegal memory access, out-of-bounds), not hardware faults. XID 43 (GPU stopped processing) is a software-level stop, not a hardware memory fault. Investigate the workload first; if the error is reproducible across different workloads or nodes, suspect hardware.
Monitor ECC counters directly: DCGM_FI_DEV_ECC_SBE_VOL_TOTAL (single-bit errors, correctable) and DCGM_FI_DEV_ECC_DBE_VOL_TOTAL (double-bit errors, uncorrectable). A rising SBE count is normal over time, but a sudden spike or any DBE is a hardware red flag.
Thermal and Power
Set alerts on DCGM_FI_DEV_GPU_TEMP (GPU temperature in Celsius) and DCGM_FI_DEV_POWER_USAGE (power draw in watts). Thermal throttling is silent—clock speeds drop (DCGM_FI_DEV_SM_CLOCK, DCGM_FI_DEV_MEM_CLOCK) and throughput falls, but the job keeps running. If temperature consistently exceeds the thermal design point (typically mid-80s Celsius for most datacenter GPUs under load), check cooling: airflow obstructions, failed fans, or inadequate facility cooling.
An idle GPU does not draw the same power as a busy one—idle draw is a fraction of TDP (tens of watts versus hundreds under load). The A100 SXM is around 400 W at full load, the H100 SXM around 700 W. The dominant cost of an idle GPU is not electricity; it's the amortized capital cost or the reserved hourly price, which you pay regardless of utilization. This is why utilization metrics matter more than wattage for cost control.
Operational Runbooks
Draining Nodes Safely
Use kubectl drain with the Eviction API, not direct deletion. The correct incantation for a GPU node is:
kubectl drain <node> \
--ignore-daemonsets \
--delete-emptydir-data \
--grace-period=300
The --ignore-daemonsets flag is required because DaemonSet pods are not evicted (they are managed by the DaemonSet controller, which will reschedule them when the node is uncordoned). The --delete-emptydir-data flag is required to evict pods using an emptyDir volume, which is common for scratch space in training jobs.
The --grace-period overrides the pod's terminationGracePeriodSeconds. The default --timeout is zero, which means wait forever (no timeout), not "give up immediately." If you need a timeout, set it explicitly (e.g., --timeout=600s).
Never use --disable-eviction unless you understand the consequences: it deletes pods directly, bypassing PodDisruptionBudgets and force-removing running workloads. This is appropriate only for a node that is unresponsive or in a failure state where the Eviction API does not work.
Handling XID Events
When an XID alert fires, the first step is triage: is this a hardware fault (48, 79, 64, 94, 95) or an application issue (13, 31, 43)? For hardware faults, cordon the node immediately (kubectl cordon <node>), then drain it. For application issues, inspect the workload logs and the GPU state with nvidia-smi and dcgmi diag.
If the GPU is still accessible, capture diagnostics before draining: nvidia-smi -q, nvidia-bug-report.sh, and dcgmi diag -r 3. These logs are critical for RMA and root-cause analysis. If the GPU has fallen off the bus (XID 79), the node may be unresponsive; you may need to force-delete pods and reboot the node out-of-band.
Firmware and Driver Updates
GPU firmware and driver updates carry risk. Test updates on a canary node or a non-production cluster first, and validate with the same burn-in suite you use for new hardware. A driver regression can introduce silent correctness issues (wrong results) or performance regressions that are hard to attribute after the fact.
Coordinate driver updates with the CUDA version and the container base images your workloads use. The NVIDIA driver is backward-compatible (a newer driver supports older CUDA toolkits), but forward compatibility is limited. Pin driver versions in your node image or use the NVIDIA GPU Operator to manage driver lifecycle, and roll updates in waves with validation between waves.
Capacity and Scheduling Discipline
Resource Requests and Limits
The GPU extended resource is nvidia.com/gpu, and it is integer-only: you cannot request a fraction of a physical GPU in the standard device plugin. Kubernetes requires that for extended resources, request equals limit (you specify the limit, and the request is set equal). This means a pod requesting one GPU will be scheduled only on a node with at least one free GPU, and it will hold that GPU exclusively unless you configure MIG or time-slicing.
MIG (Multi-Instance GPU) partitions a physical GPU into isolated instances with hardware-level memory and fault isolation. A GPU is divided into seven compute slices and eight memory slices, so the maximum number of MIG instances on one GPU is seven (seven 1g profiles). Profile naming is <compute-slices>g.<memory>gb. For example, on an A100-80GB, the profiles are 1g.10gb, 2g.20gb, 3g.40gb, 4g.40gb, and 7g.80gb. The compute-slice budget is seven, so the maximum count of a profile per GPU is floor(7 / compute-slices): you can fit seven 1g instances, three 2g instances, two 3g instances, or one 4g or 7g instance per GPU.
Time-slicing multiplexes the whole GPU in time across processes. It has no memory isolation and no fault isolation—an OOM or crash in one workload can affect others sharing the GPU, and there is no hard memory cap per client. Time-slicing is useful for inference or low-priority jobs where isolation is less critical, but it is not a substitute for MIG in multi-tenant training environments.
PodDisruptionBudgets for Training Jobs
Set a PodDisruptionBudget (PDB) for multi-GPU training jobs so that kubectl drain does not evict the entire job at once. A typical PDB for a distributed training job allows zero voluntary disruptions (maxUnavailable: 0), forcing the operator to pause the job or wait for a checkpoint before draining nodes. This prevents the job from being split across a maintenance event and losing hours of training progress.
For inference or batch workloads that can tolerate restarts, set a more permissive PDB (e.g., minAvailable: 50%) to allow rolling maintenance. The key is aligning the PDB policy with the workload's tolerance for disruption and the cost of a restart.
Failure Modes and Detection Latency
The insidious failure modes in GPU clusters are the ones that do not trigger an immediate alert. A single degraded NVLink can slow a training job by a few percent, and the slowdown is attributed to "variance" until someone runs a bandwidth test. A GPU with a marginal memory module may pass nvidia-smi but fail under sustained load hours into a job, producing a cryptic CUDA error.
The defense is continuous validation: run lightweight health checks (memory bandwidth, NVLink bandwidth, ECC counters) on a schedule (e.g., daily or before every job) and compare against baseline metrics. Use the DCGM health check feature (dcgmi health -g <group> -c) to automate this. Alert on deviations from baseline, not just hard failures.
For correctness, consider checksum validation or known-answer tests in your training pipeline. Silent data corruption (from a flaky GPU or a bit flip in transit) is rare but not impossible, and the symptom is a model that diverges or produces nonsense results. A periodic validation pass on a known dataset can catch this before you waste weeks of compute.
Conclusion
Reliability in a GPU cluster is not a one-time setup; it is a discipline of validation, monitoring, and operational rigor. The checklist is: validate hardware before it enters production, monitor the metrics that reveal real utilization and hardware health (not just GPU_UTIL), use drain and eviction procedures that respect workload boundaries, and build runbooks for the failure modes you will see—XID events, thermal issues, and silent degradation. The goal is to catch problems early, when they are cheap to fix, rather than discovering them three weeks into a training run.
For a human to find the waste on your actual cluster, book a GPU Cost & Reliability Audit →