GPU FinOps: Reducing AI Infrastructure Cost Without Waste
GPU clusters represent the largest line item in most AI infrastructure budgets, yet the dominant cost driver is not electricity—it is underutilization. A reserved or owned GPU costs the same per hour whether it runs at 10% or 90% utilization, and the capital or rental amortization dwarfs the incremental power cost. Reducing waste means improving utilization through right-sizing, batch optimization, and eliminating idle allocation, all guided by correct measurement of what the GPU is actually doing.
The Dominant Cost: Amortized Capital and Reserved Capacity
The first discipline in GPU FinOps is understanding where the money actually goes. An A100 SXM GPU has a TDP around 400 W and draws tens of watts when idle versus hundreds under load, but the incremental electricity cost of that difference is a small fraction of the total. The overwhelming cost is the amortized capital expenditure if you own the hardware, or the reserved hourly price if you rent from a cloud or colo provider. That cost is fixed per hour regardless of whether the GPU runs a single inference request or sustains 95% utilization on a training job.
This means the primary lever for cost reduction is utilization: ensuring that every GPU-hour you pay for produces useful work. Underutilization—GPUs sitting idle in a reserved pool, or running workloads that use only a fraction of the available compute—is pure waste of the capital or rental cost. Improving utilization does not mean running the GPU hotter or drawing more power; it means filling the capacity you have already paid for.
Measuring Real Utilization with DCGM Profiling Metrics
The standard GPU utilization metric exported by DCGM is DCGM_FI_DEV_GPU_UTIL, which reports the percentage of time the GPU had at least one active kernel. This metric is coarse: it can read 80% or 90% even when the streaming multiprocessors (SMs) are lightly loaded, because any kernel activity—no matter how small—marks that sample period as "busy." For cost optimization you need to know whether the SMs are genuinely saturated or just occasionally tickled.
DCGM profiling fields (the DCP metrics) provide true utilization:
DCGM_FI_PROF_GR_ENGINE_ACTIVE: fraction of time the graphics engine was active.DCGM_FI_PROF_SM_ACTIVE: fraction of time at least one warp was active on the SMs.DCGM_FI_PROF_SM_OCCUPANCY: average fraction of maximum theoretical occupancy achieved across active cycles.DCGM_FI_PROF_PIPE_TENSOR_ACTIVE: fraction of time the Tensor Core pipeline was active (critical for mixed-precision training and transformer inference).DCGM_FI_PROF_DRAM_ACTIVE: fraction of time the DRAM controller was active (memory-bound indicator).
When DCGM_FI_DEV_GPU_UTIL is high but DCGM_FI_PROF_SM_ACTIVE or DCGM_FI_PROF_PIPE_TENSOR_ACTIVE are low, the workload is underutilizing the compute resources—perhaps due to small batch size, CPU-bound preprocessing, or inefficient kernels. Conversely, high DCGM_FI_PROF_DRAM_ACTIVE with lower SM activity suggests memory bandwidth saturation, which is a different optimization problem. These metrics, collected via the dcgm-exporter and scraped into Prometheus, give you the ground truth needed to identify waste.
Right-Sizing Allocation: MIG and Time-Slicing
Many inference and fine-tuning workloads do not require a full GPU. Allocating a whole A100 to a small model wastes the unused capacity. NVIDIA Multi-Instance GPU (MIG) and time-slicing are the two mechanisms for sharing a physical GPU across multiple workloads.
MIG Geometry and Instance Counts
MIG partitions one physical GPU into isolated instances with hardware-level memory, cache, and fault isolation. Internally, a GPU is divided into 7 compute slices (GI slices) and 8 memory slices. The maximum number of MIG instances on one GPU is therefore 7 (seven 1g profiles). Profile naming follows <compute-slices>g.<memory>gb, and the compute-slice budget of 7 determines the maximum count per GPU:
| Profile | Compute Slices | Max Instances per GPU | Example (A100-40GB) |
|---|---|---|---|
| 1g | 1 | 7 | 1g.5gb |
| 2g | 2 | 3 | 2g.10gb |
| 3g | 3 | 2 | 3g.20gb |
| 4g | 4 | 1 | 4g.20gb |
| 7g | 7 | 1 | 7g.40gb |
On A100-40GB, the memory slice is 5 GB, so profiles are 1g.5gb, 2g.10gb, 3g.20gb, 4g.20gb, and 7g.40gb. On A100-80GB and H100-80GB, the memory slice is 10 GB, so the 3-slice profile is 3g.40gb (not 3g.20gb). A common mistake is assuming you can fit four 3g instances on one GPU; you cannot, because 4 × 3 = 12 slices exceeds the budget of 7. The maximum is two 3g instances per GPU.
MIG provides true isolation: an out-of-memory error or crash in one instance does not affect the others, and each instance has a fixed memory cap enforced in hardware. This makes MIG suitable for multi-tenant inference or for isolating batch jobs that should not interfere with each other.
Time-Slicing: No Isolation, Lower Overhead
Time-slicing multiplexes the whole GPU in time across multiple processes or pods. It has no memory isolation and no fault isolation—an OOM or crash in one workload can affect others sharing the GPU, and there is no hard memory cap per client. Time-slicing is configured through the NVIDIA device plugin and allows fractional GPU requests in Kubernetes, but the lack of isolation means it is appropriate only for trusted, cooperative workloads or for oversubscription scenarios where you accept the risk.
The cost trade-off: MIG adds a small amount of scheduling and context-switch overhead compared to a dedicated GPU, but the isolation is often worth it. Time-slicing has lower overhead but no protection. Choose MIG when you need guaranteed memory and fault boundaries; choose time-slicing when you need maximum flexibility and the workloads are trusted.
Batch Size and Continuous Batching
For inference workloads, batch size is the primary knob for GPU utilization. A small batch size leaves the SMs underutilized and increases the per-token cost. Increasing batch size improves throughput (tokens per second) up to the point where memory or compute saturates, but it also increases latency for each request in the batch. The optimal batch size balances throughput, latency, and cost.
Modern inference engines like vLLM use continuous batching: new requests are added to the active batch as soon as earlier requests finish generating tokens, rather than waiting for the entire batch to complete. This keeps the GPU busy without forcing you to choose a single static batch size. vLLM's --max-num-seqs flag controls the maximum number of concurrent sequences per iteration (the batch width), and --max-num-batched-tokens sets the token budget per scheduler step. Tuning these parameters to match your latency SLA and request rate directly impacts GPU utilization and cost per token.
vLLM also exposes --gpu-memory-utilization (default 0.9), which sets the fraction of GPU memory reserved for the KV cache and model weights. Lowering this value leaves headroom for other processes or for oversubscription, but reduces the maximum batch size or context length you can serve. The cost implication: a lower memory utilization means fewer concurrent requests, which may require more GPU instances to meet the same throughput target.
Eliminating Idle Allocation
The simplest and most impactful cost reduction is eliminating GPUs that sit idle in a reserved pool. This requires visibility into allocation versus actual usage. In Kubernetes, the extended resource nvidia.com/gpu is integer-only and requires request == limit, so a pod that requests 1 GPU holds that GPU exclusively even if the workload inside uses only 10% of it. If the pod is long-lived (a notebook, a development environment, a model server with low traffic), the GPU is effectively idle but still costs you the full hourly rate.
Strategies to recover idle capacity:
- Autoscaling and ephemeral jobs: Use Kubernetes autoscaling (HPA, KEDA, or custom controllers) to scale inference deployments down to zero replicas during low-traffic periods, and scale batch-training jobs to run only when there is work in the queue. This requires that your workloads tolerate cold-start latency, but it eliminates the cost of idle standby capacity.
- Shared development clusters with time limits: Enforce time limits on notebook and interactive sessions, and automatically terminate or suspend them after a period of inactivity. Combine this with checkpointing so users can resume work without losing state.
- Bin-packing and fragmentation reduction: Use node affinity, taints, and the Kubernetes scheduler's bin-packing behavior to concentrate workloads onto fewer nodes, leaving entire nodes empty so they can be drained and powered down (or returned to a cloud provider if using spot or on-demand instances). Fragmentation—many nodes each running one small pod—wastes the unused capacity on each node.
Draining Nodes Safely for Maintenance and Cost Optimization
When you need to remove a GPU node from the cluster—either to return unused capacity or to perform maintenance—use kubectl drain to evict running pods safely. The command cordons the node (marks it unschedulable) and then evicts pods via the Eviction API, which respects PodDisruptionBudgets:
kubectl drain <node> --ignore-daemonsets --delete-emptydir-data --grace-period=300
Key flags and their real behavior:
--ignore-daemonsets: required in practice because DaemonSet pods are not evicted by drain (they are managed by the DaemonSet controller).--delete-emptydir-data: required to evict pods that use anemptyDirvolume, acknowledging that the ephemeral data will be lost.--grace-period: overrides the pod'sterminationGracePeriodSeconds, giving the application time to shut down cleanly.--timeout: the default is0, which means wait forever (no timeout), not "give up immediately." If a pod cannot be evicted (e.g. due to a PodDisruptionBudget that cannot be satisfied), drain will block indefinitely unless you set an explicit timeout.--disable-eviction: makes drain delete pods directly instead of using the Eviction API, which bypasses PodDisruptionBudgets and force-removes running pods. This is a last resort for stuck nodes; it does not mean "only delete completed pods."
After draining, the node remains cordoned. If you are returning cloud capacity, terminate the instance. If you are performing maintenance, uncordon the node with kubectl uncordon <node> when ready.
Network and Interconnect Costs
For multi-node training, the interconnect—InfiniBand or high-speed Ethernet—is a significant cost component and a performance bottleneck. Understanding realistic bandwidth helps you size the network correctly and avoid over-provisioning.
InfiniBand generations: HDR is 200 Gb/s, NDR is 400 Gb/s, and EDR is 100 Gb/s per link. Convert line rate to bytes per second by dividing by 8: 200 Gb/s ≈ 25 GB/s theoretical, with achievable bandwidth around 23–24 GB/s after protocol overhead. For network-bound collective operations like all-reduce, the per-GPU NIC bandwidth caps the achievable bus bandwidth reported by nccl-tests. On a cluster with 200 Gb/s HDR, expect busbw for ring all-reduce to approach 23–25 GB/s per GPU, not 40–45 GB/s (which would require a faster network or multiple NICs per GPU).
Intra-node, NVLink provides much higher bandwidth: A100 aggregate NVLink bandwidth per GPU is around 600 GB/s (3rd-gen NVLink), and H100 is around 900 GB/s (4th-gen NVLink). NVSwitch connects all GPUs in a node at full NVLink speed, so intra-node collectives are not network-limited. The cost implication: scaling to multiple nodes requires investing in high-speed interconnect (HDR or NDR InfiniBand), and the return on that investment depends on whether your training workload is communication-bound. If gradient all-reduce time is a small fraction of the iteration time, a cheaper network may suffice.
Monitoring and Continuous Optimization
GPU cost optimization is not a one-time exercise. Workload characteristics change, new models are deployed, and utilization patterns drift. Instrument your cluster with DCGM metrics, track utilization and idle time per GPU and per workload, and set alerts for sustained low utilization (e.g., DCGM_FI_PROF_SM_ACTIVE < 0.3 for more than an hour). Review the data weekly or monthly to identify underutilized capacity, right-size allocations, and adjust autoscaling policies.
The cost structure of GPU infrastructure—dominated by fixed capital or reservation costs—means that small improvements in utilization have large financial impact. A 10-percentage-point increase in average utilization across a 100-GPU cluster can defer the need to purchase or rent additional capacity, saving the cost of those incremental GPUs entirely. The mechanisms described here—correct measurement with profiling metrics, right-sizing with MIG, batch optimization, and eliminating idle allocation—are the levers that deliver that improvement without adding complexity or operational risk.
For a human to find the waste on your actual cluster, book a GPU Cost & Reliability Audit →