It is often said that time is our scarcest resource. The one commodity we cannot create more of. I bet whoever said that had never tried to get their hands on GPUs.
In 2026, GPUs have become a precious resource, powering so much of the software we use and build. Getting them is one problem. Knowing where the money goes once you have them is another.
Our last few releases have focused on helping you tag your resources, integrate with existing platforms, and even replace your existing tooling completely.
Naturally, a big next step for us was allowing you to track, attribute, and rightsize your GPUs.
Starting today, this is now a reality.
Who is this for?
Based on feedback from users, we built this feature with two types of teams in mind.
Inference providers. If you serve models, GPU spend is a big part of what it costs to do that. You need to connect the hardware to the models running on it, understand the cost of serving tokens, and see whether capacity is keeping up with demand or sitting unused between requests.
Teams with a growing GPU footprint. A few cards become a fleet spread across machines, clusters, and workloads. You need to know who is using them, how much compute and memory they actually use, and where there is room to reclaim capacity before buying more.
So what can you do with this?
See your GPU fleet in one place
The new GPUs page brings accelerators across your instances and Kubernetes workloads into one view. See the hardware model, where each card runs, GPU and VRAM utilization, partitioning, and when it last reported.
Filter by provider, model, utilization, or placement to find the cards you care about. Fleet spend estimates and monthly projections help put that capacity in context, using catalog pricing where it is available.
Follow the hardware to the workload
Open a GPU to see the machine, cluster, and namespace it belongs to. Where workload reports are available, you can follow the card to the pods and containers holding it, including holders of individual MIG partitions.
For a connected model-serving deployment, CostGraph brings that relationship together with serving metrics such as token throughput, latency, and queue depth. When accelerator pricing and output-token throughput are available, cost per million output tokens helps you understand the GPU cost behind serving a model.
A shared GPU needs more care: showing which workloads hold a card does not mean its bill can be split evenly between them. CostGraph keeps that distinction visible instead of inventing a per-model allocation.
The detail view connects hardware identity and placement with utilization evidence. Available attribution depends on the workload and serving telemetry reported.
Find capacity worth reclaiming
A GPU can be full of model weights and still have very little work to do. Looking at memory alone will not tell you that.
CostGraph brings compute utilization, VRAM, and the available power and activity signals together. Usage patterns and observation windows help distinguish a quiet moment from sustained underuse, while the verdict explains the evidence behind a recommendation.
That gives you a starting point for rightsizing: investigate an idle card, review an underused replica, or check whether a workload still needs its current allocation. You make the change, with the workload's needs in mind.
The detail screenshot above shows one example: a card that ran no work over its observation window, with no VRAM claimed. Missing measurements are shown separately; an unreported card is not automatically an idle one.
Catch models holding capacity they are not using
Alongside these other features, CostGraph's model-serving insights help you spot models holding GPU capacity they are not using. A model can keep a GPU reserved long after its last request, and these insights help you decide whether that capacity is still needed.
In one example, a Qwen2.5-1.5B-Instruct deployment held a GPU while serving no requests for effectively the entire 24-hour observation window. Cache utilization stayed at zero. The insight suggested checking whether the deployment was still needed, reclaiming the card, or scaling the replica to zero between uses.
That is a useful conversation to have before adding more hardware. If you deliberately keep the model warm, that context matters; sustained traffic would make this a sizing question instead. The example did not have enough pricing data to quantify the cost impact.
Give it a try
Open the GPUs page to explore your fleet. If you're new to CostGraph, create an account, or book a demo and we'll walk through your setup together.
We built this from conversations with teams trying to make better use of expensive hardware. We'd love to hear what you want to see next.
Frequently asked questions
Does this work with both Kubernetes and virtual machines?
Yes. The inventory includes accelerators on Kubernetes nodes and standalone machines. The detail available depends on the integration and telemetry: discovering a GPU's model does not by itself provide utilization measurements.
Can I see what it costs to serve a model?
For connected model-serving deployments, CostGraph can show accelerator cost per million output tokens when pricing and token-throughput data are available. This measures the GPU component of serving cost, not every expense involved in running the model. Shared cards are not automatically divided into exact per-model bills.
Does CostGraph support MIG partitions?
Yes. Where NVIDIA Multi-Instance GPU partitions are reported, you can inspect partitions and their workload holders separately from the physical card, and identify idle partitions.
Are the GPU cost figures my cloud invoice?
Fleet estimates and projections use available catalog rates. They are separate from billed machine spend. Missing pricing or attribution is shown as unavailable rather than as zero cost.
Will CostGraph resize or stop my GPUs automatically?
No. CostGraph shows utilization evidence and recommendations so you can decide what to change. Check the workload's memory needs, traffic patterns, and any capacity you deliberately keep warm before reclaiming it.
Why might a GPU show no utilization data?
Inventory and utilization come from different signals. A discovered card may not have reported measurements, may have stopped reporting, or may not support a particular metric. CostGraph distinguishes those cases from measured zero utilization.