GPU FinOps: Reducing AI Infrastructure Cost Without Waste

GPU clusters represent the largest line item in most AI infrastructure budgets, yet the dominant cost driver is not electricity—it is underutilization. A reserved or owned GPU costs the same per hour whether it runs at 10% or 90% utilization, and the capital or rental amortization dwarfs the incremental power cost. Reducing waste means improving utilization through right-sizing, batch optimization, and eliminating idle allocation, all guided by correct measurement of what the GPU is actually doing.

💡 Put a number on it first: the free GPU Idle-Cost Calculator turns fleet size + utilization into an annual waste figure. Two minutes, no signup.

The Dominant Cost: Amortized Capital and Reserved Capacity

The first discipline in GPU FinOps is understanding where the money actually goes. An A100 SXM GPU has a TDP around 400 W and draws tens of watts when idle versus hundreds under load, but the incremental electricity cost of that difference is a small fraction of the total. The overwhelming cost is the amortized capital expenditure if you own the hardware, or the reserved hourly price if you rent from a cloud or colo provider. That cost is fixed per hour regardless of whether the GPU runs a single inference request or sustains 95% utilization on a training job.

This means the primary lever for cost reduction is utilization: ensuring that every GPU-hour you pay for produces useful work. Underutilization—GPUs sitting idle in a reserved pool, or running workloads that use only a fraction of the available compute—is pure waste of the capital or rental cost. Improving utilization does not mean running the GPU hotter or drawing more power; it means filling the capacity you have already paid for.

Measuring Real Utilization with DCGM Profiling Metrics

The standard GPU utilization metric exported by DCGM is DCGM_FI_DEV_GPU_UTIL, which reports the percentage of time the GPU had at least one active kernel. This metric is coarse: it can read 80% or 90% even when the streaming multiprocessors (SMs) are lightly loaded, because any kernel activity—no matter how small—marks that sample period as "busy." For cost optimization you need to know whether the SMs are genuinely saturated or just occasionally tickled.

DCGM profiling fields (the DCP metrics) provide true utilization:

When DCGM_FI_DEV_GPU_UTIL is high but DCGM_FI_PROF_SM_ACTIVE or DCGM_FI_PROF_PIPE_TENSOR_ACTIVE are low, the workload is underutilizing the compute resources—perhaps due to small batch size, CPU-bound preprocessing, or inefficient kernels. Conversely, high DCGM_FI_PROF_DRAM_ACTIVE with lower SM activity suggests memory bandwidth saturation, which is a different optimization problem. These metrics, collected via the dcgm-exporter and scraped into Prometheus, give you the ground truth needed to identify waste.

Right-Sizing Allocation: MIG and Time-Slicing

Many inference and fine-tuning workloads do not require a full GPU. Allocating a whole A100 to a small model wastes the unused capacity. NVIDIA Multi-Instance GPU (MIG) and time-slicing are the two mechanisms for sharing a physical GPU across multiple workloads.

MIG Geometry and Instance Counts

MIG partitions one physical GPU into isolated instances with hardware-level memory, cache, and fault isolation. Internally, a GPU is divided into 7 compute slices (GI slices) and 8 memory slices. The maximum number of MIG instances on one GPU is therefore 7 (seven 1g profiles). Profile naming follows <compute-slices>g.<memory>gb, and the compute-slice budget of 7 determines the maximum count per GPU:

ProfileCompute SlicesMax Instances per GPUExample (A100-40GB)
1g171g.5gb
2g232g.10gb
3g323g.20gb
4g414g.20gb
7g717g.40gb

On A100-40GB, the memory slice is 5 GB, so profiles are 1g.5gb, 2g.10gb, 3g.20gb, 4g.20gb, and 7g.40gb. On A100-80GB and H100-80GB, the memory slice is 10 GB, so the 3-slice profile is 3g.40gb (not 3g.20gb). A common mistake is assuming you can fit four 3g instances on one GPU; you cannot, because 4 × 3 = 12 slices exceeds the budget of 7. The maximum is two 3g instances per GPU.

MIG provides true isolation: an out-of-memory error or crash in one instance does not affect the others, and each instance has a fixed memory cap enforced in hardware. This makes MIG suitable for multi-tenant inference or for isolating batch jobs that should not interfere with each other.

Time-Slicing: No Isolation, Lower Overhead

Time-slicing multiplexes the whole GPU in time across multiple processes or pods. It has no memory isolation and no fault isolation—an OOM or crash in one workload can affect others sharing the GPU, and there is no hard memory cap per client. Time-slicing is configured through the NVIDIA device plugin and allows fractional GPU requests in Kubernetes, but the lack of isolation means it is appropriate only for trusted, cooperative workloads or for oversubscription scenarios where you accept the risk.

The cost trade-off: MIG adds a small amount of scheduling and context-switch overhead compared to a dedicated GPU, but the isolation is often worth it. Time-slicing has lower overhead but no protection. Choose MIG when you need guaranteed memory and fault boundaries; choose time-slicing when you need maximum flexibility and the workloads are trusted.

Batch Size and Continuous Batching

For inference workloads, batch size is the primary knob for GPU utilization. A small batch size leaves the SMs underutilized and increases the per-token cost. Increasing batch size improves throughput (tokens per second) up to the point where memory or compute saturates, but it also increases latency for each request in the batch. The optimal batch size balances throughput, latency, and cost.

Modern inference engines like vLLM use continuous batching: new requests are added to the active batch as soon as earlier requests finish generating tokens, rather than waiting for the entire batch to complete. This keeps the GPU busy without forcing you to choose a single static batch size. vLLM's --max-num-seqs flag controls the maximum number of concurrent sequences per iteration (the batch width), and --max-num-batched-tokens sets the token budget per scheduler step. Tuning these parameters to match your latency SLA and request rate directly impacts GPU utilization and cost per token.

vLLM also exposes --gpu-memory-utilization (default 0.9), which sets the fraction of GPU memory reserved for the KV cache and model weights. Lowering this value leaves headroom for other processes or for oversubscription, but reduces the maximum batch size or context length you can serve. The cost implication: a lower memory utilization means fewer concurrent requests, which may require more GPU instances to meet the same throughput target.

Eliminating Idle Allocation

The simplest and most impactful cost reduction is eliminating GPUs that sit idle in a reserved pool. This requires visibility into allocation versus actual usage. In Kubernetes, the extended resource nvidia.com/gpu is integer-only and requires request == limit, so a pod that requests 1 GPU holds that GPU exclusively even if the workload inside uses only 10% of it. If the pod is long-lived (a notebook, a development environment, a model server with low traffic), the GPU is effectively idle but still costs you the full hourly rate.

Strategies to recover idle capacity:

Draining Nodes Safely for Maintenance and Cost Optimization

When you need to remove a GPU node from the cluster—either to return unused capacity or to perform maintenance—use kubectl drain to evict running pods safely. The command cordons the node (marks it unschedulable) and then evicts pods via the Eviction API, which respects PodDisruptionBudgets:

kubectl drain <node> --ignore-daemonsets --delete-emptydir-data --grace-period=300

Key flags and their real behavior:

After draining, the node remains cordoned. If you are returning cloud capacity, terminate the instance. If you are performing maintenance, uncordon the node with kubectl uncordon <node> when ready.

Network and Interconnect Costs

For multi-node training, the interconnect—InfiniBand or high-speed Ethernet—is a significant cost component and a performance bottleneck. Understanding realistic bandwidth helps you size the network correctly and avoid over-provisioning.

InfiniBand generations: HDR is 200 Gb/s, NDR is 400 Gb/s, and EDR is 100 Gb/s per link. Convert line rate to bytes per second by dividing by 8: 200 Gb/s ≈ 25 GB/s theoretical, with achievable bandwidth around 23–24 GB/s after protocol overhead. For network-bound collective operations like all-reduce, the per-GPU NIC bandwidth caps the achievable bus bandwidth reported by nccl-tests. On a cluster with 200 Gb/s HDR, expect busbw for ring all-reduce to approach 23–25 GB/s per GPU, not 40–45 GB/s (which would require a faster network or multiple NICs per GPU).

Intra-node, NVLink provides much higher bandwidth: A100 aggregate NVLink bandwidth per GPU is around 600 GB/s (3rd-gen NVLink), and H100 is around 900 GB/s (4th-gen NVLink). NVSwitch connects all GPUs in a node at full NVLink speed, so intra-node collectives are not network-limited. The cost implication: scaling to multiple nodes requires investing in high-speed interconnect (HDR or NDR InfiniBand), and the return on that investment depends on whether your training workload is communication-bound. If gradient all-reduce time is a small fraction of the iteration time, a cheaper network may suffice.

Monitoring and Continuous Optimization

GPU cost optimization is not a one-time exercise. Workload characteristics change, new models are deployed, and utilization patterns drift. Instrument your cluster with DCGM metrics, track utilization and idle time per GPU and per workload, and set alerts for sustained low utilization (e.g., DCGM_FI_PROF_SM_ACTIVE < 0.3 for more than an hour). Review the data weekly or monthly to identify underutilized capacity, right-size allocations, and adjust autoscaling policies.

The cost structure of GPU infrastructure—dominated by fixed capital or reservation costs—means that small improvements in utilization have large financial impact. A 10-percentage-point increase in average utilization across a 100-GPU cluster can defer the need to purchase or rent additional capacity, saving the cost of those incremental GPUs entirely. The mechanisms described here—correct measurement with profiling metrics, right-sizing with MIG, batch optimization, and eliminating idle allocation—are the levers that deliver that improvement without adding complexity or operational risk.

Start with the number: the free GPU Idle-Cost Calculator shows what your current utilization is costing per year, in two minutes.

For a human to find the waste on your actual cluster, book a GPU Cost & Reliability Audit →
Get the next field note. Practical, occasional notes on GPU/K8s cost, scheduling and reliability — the traps and the numbers. No spam.