Cost on Kubernetes is an allocation problem, not a billing problem. The cluster costs what it costs; the question is which namespace, workload or label is responsible for how much of it, and the answer is computed from requests, usage and time, out of the same Prometheus metrics you already collect.
make up obsOrientation
You already built the measuring half in section 1.2. This section adds the money view and the loop that connects them: measure usage, compare with requests, adjust requests, confirm nothing degraded. That loop is the competency; OpenCost is just the instrument.
For each container, for each time window: cost = max(request, usage) × unit_price × hours, summed over CPU, memory, GPU, storage and network, then attributed to the pod's namespace, labels or owner. Everything OpenCost shows is that expression sliced differently. Once you see it, the "why is my namespace expensive when nothing is running" question answers itself: because request won the max().
What drives spend, in the order it matters
- Requests, not usage. Capacity planning and most internal pricing allocate by what you reserved. A deployment requesting 4 CPU and using 200m costs 4 CPU. The gap between the two is the "efficiency" number OpenCost reports, and closing it is right-sizing.
- Idle capacity. Nodes are bought whole; unrequested space is idle cost. OpenCost can either show it separately or spread it across tenants, and knowing that both accounting choices exist is the exam-relevant fact. Showing idle separately makes the platform team accountable for bin-packing; spreading it makes tenants feel the true cost of the cluster they share. Neither is wrong; they answer different questions.
- Object sprawl. LoadBalancers, PVs, snapshots, forgotten namespaces and orphaned PVCs bill whether or not traffic flows. This is why the team-a quota caps
services.loadbalancersat 1, and why quota is a cost tool as much as a fairness tool. - Node sizing and lifecycle. The lever above the workload: instance families, spot/preemptible capacity, consolidation (Karpenter-style), and scale-to-zero for dev environments. You cannot demonstrate these in kind, but "right-size the pods, then right-size the nodes" is the correct order to say out loud.
The same gap, priced. Notional rates, real arithmetic; note how quickly trimming turns risky once usage approaches the request:
Showback reports cost to teams; chargeback actually bills them. Allocation is dividing shared cost by a defensible key (namespace, label, owner). Efficiency is usage ÷ request. Unit economics is cost per business thing (per tenant, per request). A "reduce cost" scenario is usually asking for allocation first: you cannot cut what you cannot attribute.
The loop as FinOps says it, and the levers by layer
FinOps phrases the loop as inform (allocate and show), optimize (right-size, commit, clean up) and operate (make it routine with budgets and alerts). Map it onto layers and you have a complete answer to "how would you reduce this cluster's cost":
| Layer | Waste looks like | Lever | Where the number is |
|---|---|---|---|
| Container | efficiency far below 1 | VPA recommendation: status.recommendation.containerRecommendations[].target (with lowerBound, upperBound, uncappedTarget); resourcePolicy.containerPolicies to bound it or set mode: Off per container | kubectl cost ... --show-efficiency, VPA status |
| Namespace | quota far above usage, LoadBalancers, retained PVs | quota as a budget, services.loadbalancers, per-class storage quota, ttl on Jobs | aggregate=namespace, kubectl get pvc -A |
| Node | high __idle__, nodes below 50% requested | consolidation (Karpenter WhenEmptyOrUnderutilized, CA scale-down), right-sized instance families, fewer DaemonSets | includeIdle=true, idleByNode=true |
| Purchase | on-demand for interruptible work | spot/preemptible pools with a taint and a PDB for stateless workers; commitments for the steady base | spot prices in the price list; cloud billing export |
Showback is the inform phase done for teams; chargeback is the same report wired into budgets. The maturity model (3.1) treats "teams value platform capabilities enough to pay for them, for example via a chargeback system" as a Scalable-level marker, so a scenario that mentions chargeback is asking about organizational maturity as much as about OpenCost.
How OpenCost actually works here
OpenCost is a small controller plus a query API. It reads pod and node inventory from the API server, joins it with usage series from Prometheus (container_cpu_usage_seconds_total, container_memory_working_set_bytes, kube_pod_container_resource_requests, node capacity) and applies a price list: cloud billing rates in a real cluster, default on-prem rates in this one. The prices in the lab are notional; the ratios and gaps are real, and those are what you are learning to read.
kube-state-metrics ─┐
cadvisor / kubelet ─┼─▶ Prometheus ─▶ OpenCost ─▶ allocation API ─▶ kubectl cost / UI / Grafana
node inventory ─┘ ▲
price list (cloud rates or defaults)
Two consequences worth predicting before you run anything. First, OpenCost is only as good as the metrics under it, so a broken scrape shows up as missing cost, not as an error, and diagnosing "why is this namespace zero" is a Prometheus problem, not a cost problem. Second, allocation windows matter: a 1-hour window over a freshly created workload shows nothing useful, so use --window 1d when you want stable numbers.
kubectl cost assumes a Kubecost install by default. --opencost is what points it at this stack: different service name, port and API path, all bundled in the one flag. Forgetting it produces a connection error that looks like OpenCost is broken when it is simply not being asked.
The right-sizing loop, as an exam answer
- Measure:
kubectl top, Prometheus, VPA recommendations (1.2). - Compare: requests vs p95-ish usage, per container, not per namespace average.
- Adjust:
kubectl set resourcesor the manifest in git; memory to roughly peak plus headroom, CPU to a sane request with a generous or absent limit. - Confirm: nothing throttling, nothing OOMKilled, nothing Pending, and the efficiency number moved.
Step 4 is where people lose the mark. Right-sizing that breaks scheduling or introduces throttling is worse than the waste it removed, and it is worth saying so unprompted.
The allocation API and its knobs
Everything the UI, Grafana and kubectl cost show is one HTTP query against the OpenCost pod on port 9003 (the UI is a separate deployment on 9090). Knowing the parameters lets you answer a cost question with curl when the plugin is not installed, and lets you read the plugin's flags as what they are: parameter presets.
GET /allocation
| Parameter | Values | Why it matters |
|---|---|---|
| window | today, yesterday, week, lastweek, month, 30m, 12h, 7d, RFC3339 pair, unix pair | required; a fresh workload needs a window that contains it; a long window smooths the request-vs-usage max() |
| aggregate | cluster, node, namespace, controllerKind, controller, service, pod, container, label:<key>, annotation:<key>; comma lists | this is "allocation by a defensible key": aggregate=label:team is showback by team without touching namespaces |
| step | 1h, 1d, ... | splits the window into a time series; default is one set for the whole window |
| resolution | 1m default, 30m | Prometheus query step; coarser is faster and less accurate for short-lived pods |
| includeIdle | true/false | adds an __idle__ row: allocated-but-unrequested capacity |
| shareIdle | true/false | spreads idle across the other rows proportionally to their cost, per resource; the two accounting choices from the panel above, as a flag |
| idleByNode | true/false | compute idle per node instead of per cluster |
Each row in the response carries cpuCost, ramCost, gpuCost, pvCost, networkCost, loadBalancerCost, sharedCost, totalCost, plus cpuEfficiency, ramEfficiency and totalEfficiency (usage divided by the max(request, usage) that was billed, so above 1.0 means the workload used more than it requested). kubectl cost namespace --opencost is exactly --service-port 9003 --service-name opencost --kubecost-namespace opencost --allocation-path /allocation/compute; the subcommands (namespace, deployment, controller, label, pod, node) are the aggregate values; the default output is a projected monthly rate and --historical switches to the total for the window.
Where the price comes from
- Cloud: on-demand list prices from the provider's pricing API, keyed by the node's
providerID, instance type and region; spot nodes are priced as spot when the node is labeled as such (Karpenter'skarpenter.sh/capacity-type: spot, cloud labels). Actual invoices are reconciled only when a billing export is configured (cloud costs integration). - On-prem or kind: nodes without a cloud
providerIDfall back to thecustomprovider anddefault.json: CPU 0.031611 per core-hour, RAM 0.004237 per GB-hour, GPU 0.95 per hour, storage 0.00005479452 per GB-hour, spot CPU 0.006655, described as "based on GCP us-central1". Override with the custom pricing ConfigMap or the HelmcustomPricingvalues when a task says "our CPU costs X"; the ratios between namespaces are unchanged, only the currency figures move.
The spec's definitions, since scenario text quotes them
- Total cluster cost = cluster asset costs (nodes, PVs, attached disks, load balancers, network) + cluster overhead (management fees).
- Workload cost = max(request, usage) for allocation-billed resources (CPU, GPU, RAM), PVC capacity for storage, bytes for network; computed per container, then aggregated by any key.
- Idle cost = asset cost - workload cost; idle % = idle / allocation costs. Idle can be shown as its own line or distributed: uniformly, proportionally to consumption, or by a custom metric.
- Shared costs: kube-system, monitoring, the platform team's own namespaces, spread the same three ways. This is the knob that turns showback into a fair chargeback.
- Pods in
ImagePullBackOffare not charged even though they hold a request; a Pending pod holds nothing.
A namespace that shows 0 with running pods means OpenCost could not join its Prometheus series: check PROMETHEUS_SERVER_ENDPOINT on the opencost Deployment, that kube-state-metrics and cAdvisor targets are up, and that the window covers the pods' lifetime. In a sharded Prometheus, point OpenCost at the global query endpoint (Thanos Query, Mimir), not at one replica.
Exercises
OpenCost's UI is on the LoadBalancer make urls prints, but the CLI is faster and closer to what the exam asks:
kubectl cost namespace --opencost --show-all-resources
kubectl cost namespace --opencost --historical --window 1doutputcaptured 2026-08-26
$ kubectl cost namespace --opencost --show-all-resources
+-----------------+-------------------------------+--------------------+-----------------+
| CLUSTER | NAMESPACE | MONTHLY RATE (ALL) | COST EFFICIENCY |
+-----------------+-------------------------------+--------------------+-----------------+
| default-cluster | kube-system | 22.934101 | 0.673952 |
| | flux-system | 13.663152 | 0.062010 |
| | kyverno | 10.060340 | 0.111306 |
| | gatekeeper-system | 7.597615 | 0.166510 |
| | vpa | 6.984870 | 0.049210 |
| | kro | 6.205071 | 0.077431 |
| | crossplane-system | 6.078437 | 0.413770 |
| | team-b | 6.060347 | 0.015395 |
| | team-a | 3.981019 | 0.053393 |
| | tracing | 3.032553 | 0.100108 |
| | tekton-pipelines | 2.572672 | 0.469501 |
| | tekton-pipelines-resolvers | 2.572672 | 0.036930 |
| | flux-demo | 1.388055 | 0.101211 |
| | trivy-system | 1.186362 | 1.102488 |
| | opencost | 0.779983 | 0.384823 |
| | monitoring | 0.111300 | 2.162723 |
| | default | 0.018600 | 0.930326 |
| | local-path-storage | 0.000000 | 0.000000 |
| | spire | 0.000000 | 0.000000 |
| | argo-rollouts | 0.000000 | 0.000000 |
| | argocd | 0.000000 | 0.000000 |
| __idle__ | __idle__ | 0.000000 | 0.000000 |
| default-cluster | cnpg-system | 0.000000 | 0.000000 |
| | external-secrets | 0.000000 | 0.000000 |
| | argo | 0.000000 | 0.000000 |
| | opentelemetry-operator-system | 0.000000 | 0.000000 |
+-----------------+-------------------------------+--------------------+-----------------+
| SUMMED | | 95.227148 | |
+-----------------+-------------------------------+--------------------+-----------------+
$ kubectl cost namespace --opencost --historical --window 1d
+-----------------+-------------------------------+------------------+-----------------+
| CLUSTER | NAMESPACE | TOTAL COST (ALL) | COST EFFICIENCY |
+-----------------+-------------------------------+------------------+-----------------+
| default-cluster | kube-system | 0.034410 | 0.673952 |
| | flux-system | 0.020500 | 0.062010 |
| | kyverno | 0.013930 | 0.111306 |
| | gatekeeper-system | 0.010520 | 0.166510 |
| | vpa | 0.010480 | 0.049210 |
| | kro | 0.009310 | 0.077431 |
| | crossplane-system | 0.009120 | 0.413770 |
| | team-b | 0.007690 | 0.015395 |
| | tracing | 0.004560 | 0.099890 |
| | team-a | 0.004130 | 0.053393 |
| | tekton-pipelines | 0.003870 | 0.468298 |
| | tekton-pipelines-resolvers | 0.003870 | 0.036836 |
| | monitoring | 0.003710 | 2.162723 |
| | trivy-system | 0.001780 | 1.102547 |
| | flux-demo | 0.001440 | 0.101211 |
| | opencost | 0.001080 | 0.384823 |
| | default | 0.000620 | 0.930254 |
| | local-path-storage | 0.000000 | 0.000000 |
| | external-secrets | 0.000000 | 0.000000 |
| | argo | 0.000000 | 0.000000 |
| __idle__ | __idle__ | 0.000000 | 0.000000 |
| default-cluster | cnpg-system | 0.000000 | 0.000000 |
| | argocd | 0.000000 | 0.000000 |
| | spire | 0.000000 | 0.000000 |
| | opentelemetry-operator-system | 0.000000 | 0.000000 |
| | argo-rollouts | 0.000000 | 0.000000 |
+-----------------+-------------------------------+------------------+-----------------+
| SUMMED | | 0.141020 | |
+-----------------+-------------------------------+------------------+-----------------+monitoring lands before looking: kube-prometheus-stack is the heaviest thing installed, yet it sets almost no requests, so max(request, usage) prices it near the bottom with an efficiency above 1. The namespaces that dominate are the ones full of idle requests. If the command errors instead, diagnose the path: kubectl-cost talks to the opencost service, which talks to Prometheus; kubectl -n opencost get deploy,svc and the pod logs tell you which hop broke.Two views of the same truth. From cost:
kubectl cost namespace --opencost --show-efficiencyoutputcaptured 2026-08-26
$ kubectl cost namespace --opencost --show-efficiency
+-----------------+-------------------------------+--------------------+-----------------+
| CLUSTER | NAMESPACE | MONTHLY RATE (ALL) | COST EFFICIENCY |
+-----------------+-------------------------------+--------------------+-----------------+
| default-cluster | kube-system | 22.940766 | 0.673763 |
| | flux-system | 13.663152 | 0.062010 |
| | kyverno | 10.060340 | 0.111306 |
| | gatekeeper-system | 7.597615 | 0.166510 |
| | vpa | 6.984870 | 0.049210 |
| | kro | 6.205071 | 0.077431 |
| | crossplane-system | 6.078437 | 0.413770 |
| | team-b | 6.060347 | 0.015396 |
| | team-a | 3.981019 | 0.053393 |
| | tracing | 3.039218 | 0.099890 |
| | tekton-pipelines-resolvers | 2.579337 | 0.036836 |
| | tekton-pipelines | 2.579337 | 0.468298 |
| | flux-demo | 1.388055 | 0.101211 |
| | trivy-system | 1.186362 | 1.102609 |
| | opencost | 0.779983 | 0.384823 |
| | monitoring | 0.111300 | 2.162723 |
| | default | 0.018600 | 0.930178 |
| | argo-rollouts | 0.000000 | 0.000000 |
| __idle__ | __idle__ | 0.000000 | 0.000000 |
| default-cluster | spire | 0.000000 | 0.000000 |
| | opentelemetry-operator-system | 0.000000 | 0.000000 |
| | external-secrets | 0.000000 | 0.000000 |
| | argo | 0.000000 | 0.000000 |
| | local-path-storage | 0.000000 | 0.000000 |
| | argocd | 0.000000 | 0.000000 |
| | cnpg-system | 0.000000 | 0.000000 |
+-----------------+-------------------------------+--------------------+-----------------+
| SUMMED | | 95.253808 | |
+-----------------+-------------------------------+--------------------+-----------------+From raw metrics, CPU requested minus used, per namespace (run in the Prometheus UI from make urls):
sum by (namespace) (kube_pod_container_resource_requests{resource="cpu"})
- sum by (namespace) (rate(container_cpu_usage_seconds_total{container!="",container!="POD"}[10m]))outputcaptured 2026-08-26
$ kubectl -n monitoring port-forward svc/prometheus-operated 9090 & # or open the UI from make urls
$ curl -sG http://localhost:9090/api/v1/query --data-urlencode \
'query=sum by (namespace) (kube_pod_container_resource_requests{resource="cpu"}) - sum by (namespace) (rate(container_cpu_usage_seconds_total{container!="",container!="POD"}[10m]))' \
| jq -r '.data.result[] | [.metric.namespace, .value[1]] | @tsv' | sort -t$'\t' -k2 -rn
flux-system 0.5267671215880919
team-b 0.34988733380995113
kro 0.25387000718459796
kube-system 0.24557372718547377
team-a 0.22489119384540543
crossplane-system 0.185862717786078
gatekeeper-system 0.1480803033971353
vpa 0.146653874778916
tracing 0.09838868995846957
tekton-pipelines-resolvers 0.09818722856439036
tekton-pipelines 0.0866530147060892
flux-demo 0.045067453280133396
trivy-system 0.034891862144704494
opencost 0.017759279317309074
monitoring -0.1752915664261709
kyverno -0.7670737681734979The label filters matter: cadvisor also emits pod-level and node-level aggregate series (empty container) and, on many versions, the pause container as container="POD". Without {container!="",container!="POD"} you count the same CPU two or three times and the gap you compute is wrong.
Deploy the demo app with a deliberate oversize:
kubectl apply -k examples/demo-app/base
kubectl set resources deploy demo --requests=cpu=500m,memory=512Mi --limits=cpu=1,memory=1Gioutputcaptured 2026-08-26
$ kubectl apply -k examples/demo-app/base
service/demo created
deployment.apps/demo created
$ kubectl set resources deploy demo --requests=cpu=500m,memory=512Mi --limits=cpu=1,memory=1Gi
deployment.apps/demo resource requirements updatedWait ten minutes for metrics to accumulate, get the VPA recommendation for it (section 1.2), then apply sane numbers with kubectl set resources again.
kubectl cost namespace --opencost --show-efficiency for default improves between the two states, and the rollout left nothing Pending.Compute what team-a can maximally cost: its quota hard-caps requests at 2 CPU / 4Gi. That number is a budget expressed in Kubernetes objects. Write the one-sentence explanation of why a platform team sets quotas even when nobody fights over capacity.
The UI is a client of one endpoint. Query it yourself, because the exam question is never "open the dashboard", it is "what does this namespace cost and how much of the cluster is nobody's".
kubectl -n opencost port-forward deploy/opencost 9003:9003 & PF1=$!
sleep 5
curl -sG localhost:9003/allocation -d window=1d -d aggregate=namespace -d includeIdle=true | jq '.data[0] | keys'
curl -sG localhost:9003/allocation -d window=1d -d aggregate=namespace -d includeIdle=true | jq '.data[0]["__idle__"].totalCost'
kill $PF1outputcaptured 2026-09-12
$ kubectl -n opencost port-forward deploy/opencost 9003:9003 & PF1=$!
$ sleep 5
Forwarding from 127.0.0.1:9003 -> 9003
Forwarding from [::1]:9003 -> 9003
$ curl -sG localhost:9003/allocation -d window=1d -d aggregate=namespace -d includeIdle=true | jq '.data[0] | keys'
Handling connection for 9003
[
"__idle__",
"argo",
"argo-rollouts",
"argocd",
"cnpg-system",
"crossplane-system",
"external-secrets",
"flux-demo",
"flux-system",
"gatekeeper-system",
"kro",
"kube-system",
"kyverno",
"local-path-storage",
"monitoring",
"opencost",
"opentelemetry-operator-system",
"payments-dev",
"spire",
"team-a",
"team-b",
"tekton-chains",
"tekton-pipelines",
"tekton-pipelines-resolvers",
"tracing",
"trivy-system",
"vpa"
]
$ curl -sG localhost:9003/allocation -d window=1d -d aggregate=namespace -d includeIdle=true | jq '.data[0]["__idle__"].totalCost'
Handling connection for 9003
0
$ kill $PF1__idle__, and on this cluster the __idle__ entry's totalCost is 0. Idle is the cluster capacity no namespace requested, priced from node assets; with the custom price sheet this lab runs there is nothing left over to attribute, so the bucket exists and is empty. How the answer is built is what you are learning, not the number.Namespace is the platform's unit. A label is the product's unit, and the two rarely line up. The same endpoint aggregates by either, and the leftover bucket tells you how much of the cluster is unlabeled.
kubectl -n opencost port-forward deploy/opencost 9003:9003 & PF1=$!
sleep 5
curl -sG localhost:9003/allocation -d window=1d -d aggregate=label:app | jq '.data[0] | keys'
curl -sG localhost:9003/allocation -d window=1d -d aggregate=label:app | jq '.data[0] | map_values(.totalCost)'
kill $PF1outputcaptured 2026-09-12
$ kubectl -n opencost port-forward deploy/opencost 9003:9003 & PF1=$!
$ sleep 5
Forwarding from 127.0.0.1:9003 -> 9003
Forwarding from [::1]:9003 -> 9003
$ curl -sG localhost:9003/allocation -d window=1d -d aggregate=label:app | jq '.data[0] | keys'
Handling connection for 9003
[
"__unallocated__"
]
$ curl -sG localhost:9003/allocation -d window=1d -d aggregate=label:app | jq '.data[0] | map_values(.totalCost)'
Handling connection for 9003
{
"__unallocated__": 0.24568
}
$ kill $PF1__unallocated__ holds everything without the label. If that bucket is the biggest one, your chargeback story is a labeling project, not a cost project.Whether idle is a line item or is spread across tenants is a policy decision with a flag. Run both and watch every tenant's number move.
kubectl -n opencost port-forward deploy/opencost 9003:9003 & PF1=$!
sleep 5
curl -sG localhost:9003/allocation -d window=1d -d aggregate=namespace -d includeIdle=true -d shareIdle=false | jq '.data[0] | map_values(.totalCost)'
curl -sG localhost:9003/allocation -d window=1d -d aggregate=namespace -d includeIdle=true -d shareIdle=true | jq '.data[0] | map_values(.totalCost)'
kill $PF1outputcaptured 2026-09-12
$ kubectl -n opencost port-forward deploy/opencost 9003:9003 & PF1=$!
$ sleep 5
Forwarding from 127.0.0.1:9003 -> 9003
Forwarding from [::1]:9003 -> 9003
$ curl -sG localhost:9003/allocation -d window=1d -d aggregate=namespace -d includeIdle=true -d shareIdle=false | jq '.data[0] | map_values(.totalCost)'
Handling connection for 9003
{
"__idle__": 0,
"argo": 0,
"argo-rollouts": 0,
"argocd": 0,
"cnpg-system": 0,
"crossplane-system": 0.01695,
"external-secrets": 0,
"flux-demo": 0.00002,
"flux-system": 0.0381,
"gatekeeper-system": 0.0212,
"kro": 0.01731,
"kube-system": 0.06396,
"kyverno": 0.02805,
"local-path-storage": 0,
"monitoring": 0.00697,
"opencost": 0.00218,
"opentelemetry-operator-system": 0,
"payments-dev": 0.00741,
"spire": 0,
"team-a": 0.00046,
"team-b": 0.0009,
"tekton-chains": 0,
"tekton-pipelines": 0.00718,
"tekton-pipelines-resolvers": 0.00718,
"tracing": 0.00847,
"trivy-system": 0.00006,
"vpa": 0.01949
}
$ curl -sG localhost:9003/allocation -d window=1d -d aggregate=namespace -d includeIdle=true -d shareIdle=true | jq '.data[0] | map_values(.totalCost)'
Handling connection for 9003
{
"argo": 0,
"argo-rollouts": 0,
"argocd": 0,
"cnpg-system": 0,
"crossplane-system": 0.01695,
"external-secrets": 0,
"flux-demo": 0.00002,
"flux-system": 0.0381,
"gatekeeper-system": 0.0212,
"kro": 0.01731,
"kube-system": 0.06396,
"kyverno": 0.02805,
"local-path-storage": 0,
"monitoring": 0.00697,
"opencost": 0.00218,
"opentelemetry-operator-system": 0,
"payments-dev": 0.00741,
"spire": 0,
"team-a": 0.00046,
"team-b": 0.0009,
"tekton-chains": 0,
"tekton-pipelines": 0.00718,
"tekton-pipelines-resolvers": 0.00718,
"tracing": 0.00848,
"trivy-system": 0.00006,
"vpa": 0.01949
}
$ kill $PF1__idle__ key is present under shareIdle=false and gone under shareIdle=true, and every other total is the same to the cent. Idle is zero on this cluster, so there is nothing to spread and the only movement is a sub-cent rounding difference on one namespace. On a cluster with real idle the other totals would rise to absorb it, and the flag would be the whole argument.On-prem there is no cloud bill, so OpenCost costs from a price sheet and a Prometheus endpoint. Both are configuration, and a wrong one makes every number on the page confidently wrong.
kubectl -n opencost get deploy opencost -o jsonpath='{.spec.template.spec.containers[0].env}' | jq '[.[] | {name, value}]'
kubectl -n opencost exec deploy/opencost -- sh -c 'cat /models/default.json 2>/dev/null || ls /models'
kubectl -n opencost get cm -o nameoutputcaptured 2026-09-13
$ kubectl -n opencost get deploy opencost -o jsonpath='{.spec.template.spec.containers[0].env}' | jq '[.[] | {name, value}]'
[
{
"name": "LOG_LEVEL",
"value": "info"
},
{
"name": "CUSTOM_COST_ENABLED",
"value": "false"
},
{
"name": "INSTALL_NAMESPACE",
"value": "opencost"
},
{
"name": "PROMETHEUS_QUERY_RESOLUTION_SECONDS",
"value": "300"
},
{
"name": "API_PORT",
"value": "9003"
},
{
"name": "PROMETHEUS_SERVER_ENDPOINT",
"value": "http://prometheus-kube-prometheus-prometheus.monitoring.svc.cluster.local:9090"
},
{
"name": "INSECURE_SKIP_VERIFY",
"value": "false"
},
{
"name": "CLUSTER_ID",
"value": "default-cluster"
},
{
"name": "RESOLUTION_1D_RETENTION",
"value": "15"
},
{
"name": "RESOLUTION_1H_RETENTION",
"value": "49"
... 38 more lines
$ kubectl -n opencost exec deploy/opencost -- sh -c 'cat /models/default.json 2>/dev/null || ls /models'
Defaulted container "opencost" out of: opencost, opencost-ui
{
"provider": "custom",
"description": "Default prices based on GCP us-central1",
"CPU": "0.031611",
"spotCPU": "0.006655",
"RAM": "0.004237",
"spotRAM": "0.000892",
"GPU": "0.95",
"storage": "0.00005479452",
"zoneNetworkEgress": "0.01",
"regionNetworkEgress": "0.01",
"internetNetworkEgress": "0.12",
"natGatewayEgress": "0.045",
"natGatewayIngress": "0.045"
}
$ kubectl -n opencost get cm -o name
configmap/kube-root-ca.crt
configmap/opencost-ui-nginx-configVPA says what a workload should request. OpenCost says what the gap is costing. Neither is an argument on its own; together they are the right-sizing case you would take to a team.
# make up installs the VPA components but creates no VerticalPodAutoscaler, so there is nothing to read
kubectl apply -f - <<'EOF'
apiVersion: autoscaling.k8s.io/v1
kind: VerticalPodAutoscaler
metadata: { name: staging-demo, namespace: team-a }
spec:
targetRef:
apiVersion: apps/v1
kind: Deployment
name: staging-demo
updatePolicy: { updateMode: "Off" }
EOF
# the recommender needs a few minutes of history before it will answer
sleep 300
kubectl -n team-a get vpa staging-demo
kubectl -n team-a get vpa staging-demo -o jsonpath='{.status.recommendation.containerRecommendations[0]}' | jq
kubectl -n opencost port-forward deploy/opencost 9003:9003 & PF1=$!
sleep 5
curl -sG localhost:9003/allocation -d window=1d -d aggregate=namespace | jq '.data[0]["team-a"] | {cpuCoreRequestAverage, cpuCoreUsageAverage, ramByteRequestAverage, ramByteUsageAverage, totalCost}'
kill $PF1
kubectl -n team-a delete vpa staging-demooutputcaptured 2026-09-13
$ # make up installs the VPA components but creates no VerticalPodAutoscaler, so there is nothing to read
$ kubectl apply -f - <<'EOF'
apiVersion: autoscaling.k8s.io/v1
kind: VerticalPodAutoscaler
metadata: { name: staging-demo, namespace: team-a }
spec:
targetRef:
apiVersion: apps/v1
kind: Deployment
name: staging-demo
updatePolicy: { updateMode: "Off" }
EOF
verticalpodautoscaler.autoscaling.k8s.io/staging-demo created
$ # the recommender needs a few minutes of history before it will answer
$ sleep 300
$ kubectl -n team-a get vpa staging-demo
NAME MODE CPU MEM PROVIDED AGE
staging-demo Off 15m 100Mi True 5m
$ kubectl -n team-a get vpa staging-demo -o jsonpath='{.status.recommendation.containerRecommendations[0]}' | jq
{
"containerName": "web",
"lowerBound": {
"cpu": "15m",
"memory": "100Mi"
},
"target": {
"cpu": "15m",
"memory": "100Mi"
},
"uncappedTarget": {
"cpu": "15m",
"memory": "100Mi"
},
"upperBound": {
"cpu": "63m",
"memory": "136359041"
}
}
$ kubectl -n opencost port-forward deploy/opencost 9003:9003 & PF1=$!
$ sleep 5
Forwarding from 127.0.0.1:9003 -> 9003
Forwarding from [::1]:9003 -> 9003
$ curl -sG localhost:9003/allocation -d window=1d -d aggregate=namespace | jq '.data[0]["team-a"] | {cpuCoreRequestAverage, cpuCoreUsageAverage, ramByteRequestAverage, ramByteUsageAverage, totalCost}'
Handling connection for 9003
{
"cpuCoreRequestAverage": 0.05253,
"cpuCoreUsageAverage": 0,
"ramByteRequestAverage": 71900157.30178,
"ramByteUsageAverage": 26387597.79358,
"totalCost": 0.04666
}
$ kill $PF1
$ kubectl -n team-a delete vpa staging-demo
verticalpodautoscaler.autoscaling.k8s.io "staging-demo" deleted from team-a namespaceSelf-check
A namespace runs almost no traffic and still tops the cost report. Give the two most likely explanations.
Oversized requests (billed on reservation, not usage), or expensive attached objects: LoadBalancers, provisioned volumes, retained snapshots. Check efficiency first: very low efficiency points at requests; high efficiency with high cost points at the objects.
Why is container!="" in that PromQL query not optional?
cadvisor emits per-container series plus pod-level and node-level rollups with an empty container label, and often the pause container as container="POD". Summing everything counts the same CPU two or three times, so every aggregate over cadvisor series needs {container!="",container!="POD"} or an equivalent.
Should idle cost be spread across tenants or shown separately? Argue both.
Spread: tenants see the true cost of the cluster they share, which discourages hoarding. Separate: the platform team owns bin-packing and node choice, so making idle visible as their number is what drives consolidation. Mature setups show both: tenant cost at request-level, plus a platform-owned idle line.
You cut a deployment's CPU request from 500m to 50m and latency degrades. What happened?
Requests set the cgroup CPU weight, so under contention the container now gets a much smaller share; and if a limit exists near the old request, throttling is likelier. Right-sizing means matching observed usage with headroom for peaks, not matching the average. Confirm with throttling metrics, not with the cost graph.
Give the four-step right-sizing loop in one breath.
Measure usage → compare with requests → adjust requests → confirm no degradation (no throttling, no OOM, no Pending) and re-measure efficiency. The fourth step is the one that makes it engineering rather than cost-cutting.
Show cost per team when teams span several namespaces and pods carry a team label. Query?
/allocation?window=7d&aggregate=label:team (or kubectl cost label --opencost -l team). Aggregation keys include label:<key> and annotation:<key>, so allocation follows whatever label discipline you enforce; a Kyverno policy requiring the label is what makes the report complete.
What does shareIdle=true change in the numbers, and who should want it?
The __idle__ row disappears and its cost is distributed across the other rows in proportion to their own cost, per resource. Tenants who should feel the true cost of the cluster they share want it; a platform team that owns bin-packing wants includeIdle=true without sharing, so idle stays their line.
Where do on-prem prices come from, and what is the CPU default?
Nodes with no cloud providerID use the custom provider and OpenCost's default.json: CPU 0.031611 per core-hour, RAM 0.004237 per GB-hour, GPU 0.95 per hour, described as based on GCP us-central1. Override them via the custom pricing ConfigMap or Helm values. The figures are notional; the ratios and efficiency numbers are what you read.
A workload shows totalEfficiency 1.4. Is that good?
It means usage exceeded requests by 40%, so the workload is under-requested: it is being billed on usage, cheap on paper, but it is stealing headroom from neighbors and is first in line for eviction under memory pressure. Raise the request toward observed usage; efficiency near but below 1 is the target.
Docs to know your way around
- opencost.io: the allocation API and the efficiency definition; the "how cost is calculated" page is short and is exactly the formula above.
- github.com/kubecost/kubectl-cost: flag reference; it has more views (controller, deployment, label) than the two used here.
- finops.org: the FinOps framework vocabulary (inform, optimize, operate) if a scenario question uses the words.
- Offline:
kubectl cost --help, and the same Prometheus queries from section 4.1. - opencost.io/docs/integrations/api: the /allocation parameters (window, aggregate, step, resolution, includeIdle, shareIdle, idleByNode) and the response fields; opencost.io/docs/specification for the cost definitions quoted in scenario text.
- github.com/opencost/opencost, configs/default.json: the on-prem default prices; opencost.io/docs/configuration/on-prem for overriding them.
make down-obs