The obs layer installs kube-prometheus-stack, which is three things people conflate: Prometheus itself, the Prometheus Operator that configures it through CRDs, and a bundle of exporters and dashboards. On the exam you work through the operator's CRDs, so that is the layer to be fluent in.
make up obsOrientation
Two skills, both testable: get a target scraped, and write a query that answers a question. Everything else in domain 4 (alerting, dashboards, DORA metrics, incident diagnosis) is built on those two.
Prometheus scrapes HTTP endpoints on an interval and stores samples in a local TSDB. Each series is a metric name plus a set of labels, and every distinct label combination is a separate series, which is why cardinality is the resource you actually manage. Service discovery (in Kubernetes: the API server) produces the target list; relabeling rewrites and filters it; the scrape either succeeds (up = 1) or does not (up = 0). Nothing is pushed, so a workload that only exists for ten seconds is invisible unless it pushes to a Pushgateway or emits an event elsewhere.
The operator model
What to scrape is declared, not configured. A ServiceMonitor says "scrape the endpoints of Services matching these labels, on this port name, at this path, every N seconds". A PodMonitor does the same without a Service. A ScrapeConfig (newer) covers targets outside Kubernetes. A Probe drives blackbox-exporter checks. The operator watches these CRDs and rewrites Prometheus's configuration live.
ServiceMonitor (selector: app=example, port: web, path: /metrics)
│ matched by Prometheus.spec.serviceMonitorSelector ← the filter people forget
▼
Service app=example, ports: [{name: web, port: 8080}]
│ its EndpointSlice supplies the actual pod IPs
▼
pod:8080/metrics ──scraped every 30s──▶ series {job="example", namespace=…, pod=…}
Prometheus only picks up ServiceMonitors matching its own serviceMonitorSelector, which on a stock kube-prometheus-stack matches the Helm release label (release: prometheus). A perfectly correct monitor without that label is silently ignored: no error, no event, no target. This lab deliberately disables that filter (serviceMonitorSelectorNilUsesHelmValues=false), so its selector is {} and every monitor everywhere is picked up. Either way, read the selector first, then write the monitor.
kubectl -n monitoring get prometheus -o jsonpath='{.items[0].spec.serviceMonitorSelector}'; echo
kubectl -n monitoring get prometheus -o jsonpath='{.items[0].spec.serviceMonitorNamespaceSelector}'; echooutputcaptured 2026-08-26
$ kubectl -n monitoring get prometheus -o jsonpath='{.items[0].spec.serviceMonitorSelector}'; echo
{}
$ kubectl -n monitoring get prometheus -o jsonpath='{.items[0].spec.serviceMonitorNamespaceSelector}'; echo
{}Prints {} on this cluster and the release-label selector on a stock install. That one check is the answer to "my target does not appear" everywhere. The exam's cluster has its own selector, so check it rather than assuming. The namespace selector is the second half of the same trap: a monitor in a namespace Prometheus does not watch is equally invisible.
The other three ways a target goes missing
- Port name, not number.
endpoints[].portis the Service's port name. An unnamed Service port cannot be referenced, and the monitor matches while scraping nothing. - No endpoints. The Service selector matches no ready pods (section 1.1). Same symptom, different layer.
- RBAC or network. Prometheus's ServiceAccount must be able to list endpoints in that namespace, and NetworkPolicy must allow monitoring → workload. In a default-deny tenant namespace, that allow is a thing someone has to write.
Diagnosis order: Status → Targets in the UI (it shows the reason for a failed scrape), then up{job="…"}, then the four causes above.
Relabeling, at exam depth
relabelings act on discovered targets before the scrape (drop targets, rewrite the job or namespace labels); metricRelabelings act on samples after (drop expensive series). The one pattern worth remembering: dropping a high-cardinality metric at ingestion with action: drop on a __name__ regex is how you rescue a Prometheus drowning in one bad exporter.
PromQL, the working subset
| Type | Read it with | Example |
|---|---|---|
| Counter (only climbs, resets to 0) | rate() · increase() | rate(http_requests_total[5m]) |
| Gauge (up and down) | read it directly | kube_pod_status_ready |
| Histogram (buckets) | histogram_quantile() | histogram_quantile(0.95, sum by (le) (rate(x_bucket[5m]))) |
| Summary (pre-computed quantiles) | read the quantile label | x{quantile="0.99"} |
- Selectors:
up{namespace="argocd"}, with=,!=,=~,!~. Regexes are fully anchored. - Ranges:
[5m]turns an instant vector into a range vector: required byrate/increase, and forbidden everywhere else. "Expected type instant vector" almost always means a stray range selector. - Aggregation:
sum by (label) (...)andsum without (pod) (...).bykeeps only what you name;withoutkeeps everything else. Aggregating beforerateis wrong:ratefirst, thensum. - Arithmetic between vectors matches on labels; mismatched label sets produce an empty result, which is why
on()/ignoring()andgroup_leftexist. You used vector arithmetic for the cost gap in section 1.5. - Emptiness:
absent(x)is 1 when x has no series (the way to alert on "the metric stopped existing"), and… or vector(0)is how you make a query return a number instead of nothing. ratevsirate:rateaverages over the window (use it for alerts and dashboards),irateuses the last two samples (spiky, for zooming in).increaseisrate × window, useful for "how many in 24h".
Why the raw counter is useless, and what the window does to the answer:
up metric
1 per healthy target, 0 per failing one, which makes count(up == 0) the cluster's own health check. make validate runs exactly that query and demands zero; the lab patched kubeadm's localhost-bound control-plane metrics to make them reachable (make fix-cp-metrics). Any "is monitoring healthy" question starts with this query.
What is already collected, for free
| Source | Gives you | Typical series |
|---|---|---|
| kube-state-metrics | object state from the API server | kube_pod_info, kube_deployment_status_replicas_unavailable, kube_pod_container_status_restarts_total |
| cAdvisor (kubelet) | container resource usage | container_cpu_usage_seconds_total, container_memory_working_set_bytes |
| node-exporter | host metrics | node_load1, node_filesystem_avail_bytes |
| kubelet / apiserver | control-plane health and latency | apiserver_request_duration_seconds_bucket |
Knowing these exist means you rarely need to instrument anything to answer an exam question about the cluster itself.
The rest of the operator's CRDs
Every operator CRD is picked up through a pair of selectors on the Prometheus object, and every pair has the same two failure modes: the object's labels do not match, or its namespace is not watched. Learn the pairs as a table and you never lose time to them again.
| Kind | Selected by | Scrapes | Field that decides |
|---|---|---|---|
| ServiceMonitor | serviceMonitorSelector / serviceMonitorNamespaceSelector | endpoints of matching Services | endpoints[].port is a port name; targetPort exists for numbers |
| PodMonitor | podMonitorSelector / podMonitorNamespaceSelector | pods directly, no Service needed | podMetricsEndpoints[].port is the container port name; portNumber for numbers |
| Probe | probeSelector / probeNamespaceSelector | blackbox-exporter checks | prober.url (the exporter), targets.staticConfig.static[] or targets.ingress, module |
| ScrapeConfig | scrapeConfigSelector / scrapeConfigNamespaceSelector | anything: staticConfigs, kubernetesSDConfigs, httpSDConfigs, fileSDConfigs, dnsSDConfigs, consul and cloud SD | a raw scrape_config as a CRD; use it for a VM, an external database, or a role: node discovery |
| PrometheusRule | ruleSelector / ruleNamespaceSelector | nothing; evaluated | see 4.2 |
| AlertmanagerConfig | alertmanagerConfigSelector / alertmanagerConfigNamespaceSelector (on the Alertmanager object) | nothing; merged into routing | see 4.2 |
- Namespace selector semantics. On the Prometheus object, a nil namespace selector means "this namespace only",
{}means every namespace, andmatchLabelsnarrows by namespace label (kubernetes.io/metadata.name: team-aselects one by name). On a ServiceMonitor,spec.namespaceSelectoris the second hop:matchNames: [argocd]orany: true; leaving it out means the monitor's own namespace, which is why a monitor inmonitoringpointed at Services inargocdscrapes nothing without it. - Labels you get for free.
jobLabelnames the Service label whose value becomesjob;targetLabelsandpodTargetLabelscopy Service or pod labels onto every series. Without them the only identity on the series isnamespace,pod,service,endpoint, so "group by team" needs the team label copied here first. - Endpoints to EndpointSlice. Since Kubernetes 1.33 the
v1 EndpointsAPI is deprecated. Newer operators supportspec.serviceDiscoveryRole: EndpointSliceon ServiceMonitors (or globally on the Prometheus object) and needdiscovery.k8s.io/endpointslicesin Prometheus's RBAC. A warning in the operator log aboutv1 Endpoints is deprecatedis this migration, not a fault. - Limits that protect you.
sampleLimit,labelLimit,labelNameLengthLimitandlabelValueLengthLimiton a monitor drop a whole scrape that exceeds them (the target shows down with the reason in Status → Targets), which is the blunt tool whenmetricRelabelingsis too slow to write.scrapeIntervalandscrapeTimeoutper endpoint override the global 30s; timeout must be shorter than interval or the config is rejected. - Relabel actions.
replace(default),keep/dropon aregexoversourceLabels,labelmap(copy__meta_kubernetes_pod_label_(.+)to plain labels),labeldrop/labelkeep,hashmod. The__address__,__scheme__,__metrics_path__targets are how a relabel rewrites where to scrape; aScrapeConfigwithkubernetesSDConfigsalmost always needs them because discovery gives you metadata, not a metrics URL.
Before writing any monitor: read the six selector pairs off the Prometheus object with kubectl -n monitoring get prometheus -o yaml, then kubectl get svc <name> -o jsonpath='{.spec.ports[*].name}' for the port name, then kubectl get endpointslices -l kubernetes.io/service-name=<name> for ready addresses. Three reads, no guessing.
PromQL beyond rate, and the metrics platform tools give you
| Function / form | Use it for | Remember |
|---|---|---|
| increase(c[1h]) | "how many in the last hour" | extrapolated; can be non-integer, and it is rate × window |
| irate(c[5m]) | the last two samples only | zooming in, never alerting |
| changes(g[1h]) / resets(c[1h]) | flapping detection; counter restarts | changes(kube_pod_container_status_ready[10m]) > 3 is a crash-loop detector |
| absent(x) / absent_over_time(x[10m]) | alert when a series disappears | returns 1 with the labels you wrote in the selector |
| predict_linear(g[1h], 4*3600) | disk or certificate expiry in N seconds | gauges only; predict_linear(node_filesystem_avail_bytes[6h], 24*3600) < 0 |
| label_replace(v, "dst", "$1", "src", "(.*)") | derive or rename a label for joins | regex anchored; label_join concatenates |
| topk(5, ...) / bottomk | worst offenders | instant queries; in a range query it can return more than k series over time |
| count by (namespace) (kube_pod_info) | counting objects | count counts series, sum adds values; for kube-state-metrics info series they differ |
| histogram_quantile(0.99, sum by (le) (rate(b[5m]))) | classic histograms | keep le; on native histograms drop the by (le) and the _bucket suffix |
| histogram_count(h) / histogram_sum(h) / histogram_fraction(0, 0.3, h) | native histograms | Prometheus 3.8+ stable; scraping still needs scrape_native_histograms: true (operator: scrapeNativeHistograms) |
| x offset 1w / x @ end() | compare with last week; pin to a time | offset goes after the selector, not the function |
| a / on(pod) group_left(team) b | joins with mismatched labels | on names the join keys; group_left copies labels from the many side |
| (sum(...) or vector(0)) | a number instead of empty | what a dashboard "single stat" needs |
Storage, cardinality, scale
- Retention.
--storage.tsdb.retention.time(default 15d when neither is set) and--storage.tsdb.retention.size; the operator exposes them asspec.retentionandspec.retentionSize. Size should stay at roughly 80-85% of the volume because compaction briefly overshoots. - Finding cardinality. Status → TSDB Status lists the top series counts by metric name and label; the query form is
topk(10, count by (__name__) ({__name__=~".+"})). Fix at the source (metricRelabelingswithaction: dropon__name__, orlabeldropon the offending label) rather than by buying memory. - Beyond one server.
remoteWrite(protocol 2.0 in Prometheus 3.x) ships samples to Thanos Receive, Mimir, Cortex or a vendor; Thanos sidecar plus object storage is the other long-term pattern;PrometheusAgentis the operator's scrape-only mode that exists solely to remote-write. Federation (/federate?match[]=...scraped withhonor_labels: true) pulls aggregated series from another Prometheus and is the older, smaller-scale answer. In a question, "long-term retention across clusters" is remote write or Thanos, not federation. - Native histograms in Kubernetes. Kubernetes 1.36 added the
NativeHistogramsfeature gate (alpha, off) for control-plane metrics; 1.37 makes it beta and on. Until then the components expose classic buckets only, so_bucketqueries againstapiserver_request_duration_secondsremain right.
Metric names the platform tools expose
| Tool | Key series | Scrape via |
|---|---|---|
| Argo CD | argocd_app_info{sync_status,health_status}, argocd_app_sync_total{phase}, argocd_app_reconcile, argocd_appset_info | Services *-metrics in argocd, port http-metrics |
| Flux | gotk_reconcile_duration_seconds (per kind/name/namespace), controller_runtime_reconcile_total{controller,result}, and gotk_resource_info{ready,suspended,revision}, which comes from kube-state-metrics with a CustomResourceState config, not from the controllers | PodMonitor on the flux-system controllers, port http-prom |
| Tekton | tekton_pipelines_controller_pipelinerun_total{status}, ..._pipelinerun_duration_seconds (histogram or last value, per config-observability), ..._running_pipelineruns, ..._taskrun_total | tekton-pipelines-controller Service, port 9090 |
| Crossplane | crossplane_managed_resource_ready, ..._synced, ..._first_time_to_readiness_seconds, ..._drift_seconds (providers, port 8080); core needs metrics.enabled=true | PodMonitor on provider pods, or annotations via DeploymentRuntimeConfig |
| Any CRD | kube_customresource_* built from --custom-resource-state-config (groupVersionKind, metrics[].each.gauge.path, labelsFromPath) | kube-state-metrics, plus RBAC to list the kind |
| Kyverno / Gatekeeper | kyverno_policy_results_total{policy_validation_mode,rule_result}, kyverno_admission_review_duration_seconds; gatekeeper_violations{enforcement_action}, gatekeeper_audit_last_run_time | ServiceMonitors on their metrics Services |
The recurring task is "expose tool X's metrics to Prometheus and write a query that shows Y". It is 4.1's selector and port-name drill against an unfamiliar Service, followed by one query from the table above. Discover the Service and its port names first; the metric names you can read from the /metrics page itself once the target is up.
Exercises
Prometheus's UI address comes from make urls.
In the UI under Status → Targets, or via API: count the scrape pools, find which ServiceMonitor each corresponds to (kubectl -n monitoring get servicemonitors), and confirm zero targets down.
curl -s 'http://<prom>/api/v1/query?query=count(up==0)' | jq -r '.data.result[0].value[1] // "0"' prints 0. If not, the down target's job name points at the responsible ServiceMonitor.Deploy an app that exposes metrics and monitor it, first wrong, then right:
kubectl create deploy example --image=quay.io/brancz/prometheus-example-app:v0.5.0 --port=8080
kubectl expose deploy example --port=8080 --name=example
kubectl apply -f - <<'EOF'
apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata: { name: example, namespace: default }
spec:
selector: { matchLabels: { app: example } }
endpoints: [{ port: "8080" }]
EOFoutputcaptured 2026-08-26
$ kubectl create deploy example --image=quay.io/brancz/prometheus-example-app:v0.5.0 --port=8080
deployment.apps/example created
$ kubectl expose deploy example --port=8080 --name=example
service/example exposed
$ kubectl apply -f - <<'EOF'
apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata: { name: example, namespace: default }
spec:
selector: { matchLabels: { app: example } }
endpoints: [{ port: "8080" }]
EOF
servicemonitor.monitoring.coreos.com/example createdWait a minute, search Targets: nothing. The endpoint's port must be the Service's port name, and this Service has none, so the monitor matches the Service and then finds no endpoint. It fails silently. Fix it: patch the Service so the port is named (kubectl patch svc example --type=json -p '[{"op":"add","path":"/spec/ports/0/name","value":"web"}]') and set port: web in the monitor.
up{job="example"} returns 1 in the UI. Then add the release: prometheus label to your monitor anyway and confirm nothing changes here, while being able to say why it would be the difference between working and ignored on a stock install.Generate a little traffic (kubectl run curl --image=curlimages/curl:8.11.1 --restart=Never -- sh -c 'for i in $(seq 100); do curl -s example.default.svc:8080; done'), then write, without copying: the request rate (rate(http_requests_total{job="example"}[5m])), summed by status code, and the p95 duration from http_request_duration_seconds_bucket.
Answer three questions using only metrics already collected: how many pods per namespace (kube_pod_info), which container restarts most (kube_pod_container_status_restarts_total), and each node's CPU pressure (node_load1 vs allocatable).
ServiceMonitor and PodMonitor only reach things Kubernetes knows about. A ScrapeConfig is how the operator scrapes a static target, which is what you need for a switch, an appliance, or anything outside the cluster.
kubectl -n monitoring get prometheus -o jsonpath='{.items[0].spec.scrapeConfigSelector}{"\n"}'
kubectl -n default create deployment fakemetrics --image=ghcr.io/stefanprodan/podinfo:6.7.1
kubectl -n default expose deployment fakemetrics --port=9898
kubectl -n default get svc fakemetrics -o jsonpath='{.spec.clusterIP}{"\n"}'
kubectl apply -f - <<'EOF'
apiVersion: monitoring.coreos.com/v1alpha1
kind: ScrapeConfig
metadata: { name: static-target, namespace: monitoring, labels: { release: prometheus } }
spec:
staticConfigs:
- targets: ["fakemetrics.default.svc.cluster.local:9898"]
labels: { source: static }
metricsPath: /metrics
scrapeInterval: 30s
EOF
sleep 60
kubectl -n monitoring port-forward svc/prometheus-kube-prometheus-prometheus 9090:9090 & PF1=$!
sleep 5
curl -s localhost:9090/api/v1/targets | jq '.data.activeTargets[] | select(.labels.source=="static") | {scrapeUrl, health, lastError}'
kill $PF1
kubectl delete scrapeconfig static-target -n monitoring
kubectl -n default delete deployment fakemetrics
kubectl -n default delete svc fakemetricsoutputcaptured 2026-09-12
$ kubectl -n monitoring get prometheus -o jsonpath='{.items[0].spec.scrapeConfigSelector}{"\n"}'
{"matchLabels":{"release":"prometheus"}}
$ kubectl -n default create deployment fakemetrics --image=ghcr.io/stefanprodan/podinfo:6.7.1
deployment.apps/fakemetrics created
$ kubectl -n default expose deployment fakemetrics --port=9898
service/fakemetrics exposed
$ kubectl -n default get svc fakemetrics -o jsonpath='{.spec.clusterIP}{"\n"}'
10.96.20.183
$ kubectl apply -f - <<'EOF'
apiVersion: monitoring.coreos.com/v1alpha1
kind: ScrapeConfig
metadata: { name: static-target, namespace: monitoring, labels: { release: prometheus } }
spec:
staticConfigs:
- targets: ["fakemetrics.default.svc.cluster.local:9898"]
labels: { source: static }
metricsPath: /metrics
scrapeInterval: 30s
EOF
scrapeconfig.monitoring.coreos.com/static-target created
$ sleep 60
$ kubectl -n monitoring port-forward svc/prometheus-kube-prometheus-prometheus 9090:9090 & PF1=$!
$ sleep 5
Forwarding from 127.0.0.1:9090 -> 9090
Forwarding from [::1]:9090 -> 9090
$ curl -s localhost:9090/api/v1/targets | jq '.data.activeTargets[] | select(.labels.source=="static") | {scrapeUrl, health, lastError}'
Handling connection for 9090
{
"scrapeUrl": "http://fakemetrics.default.svc.cluster.local:9898/metrics",
"health": "up",
"lastError": ""
}
$ kill $PF1
$ kubectl delete scrapeconfig static-target -n monitoring
scrapeconfig.monitoring.coreos.com "static-target" deleted from monitoring namespace
$ kubectl -n default delete deployment fakemetrics
deployment.apps "fakemetrics" deleted from default namespace
$ kubectl -n default delete svc fakemetrics
service "fakemetrics" deleted from default namespacescrapeConfigSelector first: if it does not match your labels, the object is valid, ignored, and completely silent.A Probe answers "can this be reached", which is a different question from any metric the app exports. It needs blackbox-exporter, which the lab does not install, so this block installs it.
Uninstall it afterwards if you are tight on memory; leave it if you intend to write availability alerts in 4.2.
helm repo add prometheus-community https://prometheus-community.github.io/helm-charts
helm repo update prometheus-community
helm upgrade --install blackbox prometheus-community/prometheus-blackbox-exporter -n monitoring --wait
kubectl apply -f - <<'EOF'
apiVersion: monitoring.coreos.com/v1
kind: Probe
metadata: { name: demo-reachable, namespace: monitoring, labels: { release: prometheus } }
spec:
interval: 30s
module: http_2xx
prober:
url: blackbox-prometheus-blackbox-exporter.monitoring.svc:9115
targets:
staticConfig:
static:
- http://prometheus-kube-prometheus-prometheus.monitoring.svc:9090/-/healthy
- http://staging-demo.team-a.svc.cluster.local
EOF
sleep 120
kubectl -n monitoring port-forward svc/prometheus-kube-prometheus-prometheus 9090:9090 & PF1=$!
sleep 5
curl -sG localhost:9090/api/v1/query --data-urlencode 'query=probe_success' | jq '.data.result[] | {instance: .metric.instance, value: .value[1]}'
# Prometheus reloads its config and scrapes before any of this appears
sleep 120
curl -sG localhost:9090/api/v1/query --data-urlencode 'query=probe_duration_seconds' | jq '.data.result[] | {instance: .metric.instance, value: .value[1]}'
kill $PF1
kubectl -n monitoring delete probe demo-reachable
helm -n monitoring uninstall blackboxoutputcaptured 2026-09-12
$ helm repo add prometheus-community https://prometheus-community.github.io/helm-charts
"prometheus-community" already exists with the same configuration, skipping
$ helm repo update prometheus-community
Hang tight while we grab the latest from your chart repositories...
...Successfully got an update from the "prometheus-community" chart repository
Update Complete. ⎈Happy Helming!⎈
$ helm upgrade --install blackbox prometheus-community/prometheus-blackbox-exporter -n monitoring --wait
Release "blackbox" does not exist. Installing it now.
NAME: blackbox
LAST DEPLOYED: Sun Sep 13 12:44:55 2026
NAMESPACE: monitoring
STATUS: deployed
REVISION: 1
DESCRIPTION: Install complete
TEST SUITE: None
NOTES:
See https://github.com/prometheus/blackbox_exporter/ for how to configure Prometheus and the Blackbox Exporter.
1. Get the application URL by running these commands:
export POD_NAME=$(kubectl get pods --namespace monitoring -l "app.kubernetes.io/name=prometheus-blackbox-exporter,app.kubernetes.io/instance=blackbox" -o jsonpath="{.items[0].metadata.name}")
export CONTAINER_PORT=$(kubectl get pod --namespace monitoring $POD_NAME -o jsonpath="{.spec.containers[0].ports[0].containerPort}")
echo "Visit http://127.0.0.1:8080 to use your application"
kubectl --namespace monitoring port-forward $POD_NAME 8080:$CONTAINER_PORT
$ kubectl apply -f - <<'EOF'
apiVersion: monitoring.coreos.com/v1
kind: Probe
metadata: { name: demo-reachable, namespace: monitoring, labels: { release: prometheus } }
spec:
interval: 30s
module: http_2xx
prober:
url: blackbox-prometheus-blackbox-exporter.monitoring.svc:9115
targets:
staticConfig:
static:
- http://prometheus-kube-prometheus-prometheus.monitoring.svc:9090/-/healthy
- http://staging-demo.team-a.svc.cluster.local
EOF
probe.monitoring.coreos.com/demo-reachable created
$ sleep 120
$ kubectl -n monitoring port-forward svc/prometheus-kube-prometheus-prometheus 9090:9090 & PF1=$!
$ sleep 5
Forwarding from 127.0.0.1:9090 -> 9090
Forwarding from [::1]:9090 -> 9090
$ curl -sG localhost:9090/api/v1/query --data-urlencode 'query=probe_success' | jq '.data.result[] | {instance: .metric.instance, value: .value[1]}'
Handling connection for 9090
{
"instance": "http://prometheus-kube-prometheus-prometheus.monitoring.svc:9090/-/healthy",
"value": "1"
}
{
"instance": "http://staging-demo.team-a.svc.cluster.local",
"value": "0"
}
$ # Prometheus reloads its config and scrapes before any of this appears
$ sleep 120
$ curl -sG localhost:9090/api/v1/query --data-urlencode 'query=probe_duration_seconds' | jq '.data.result[] | {instance: .metric.instance, value: .value[1]}'
Handling connection for 9090
{
"instance": "http://prometheus-kube-prometheus-prometheus.monitoring.svc:9090/-/healthy",
"value": "0.003463362"
}
{
"instance": "http://staging-demo.team-a.svc.cluster.local",
"value": "5.002871734"
}
$ kill $PF1
$ kubectl -n monitoring delete probe demo-reachable
probe.monitoring.coreos.com "demo-reachable" deleted from monitoring namespace
$ helm -n monitoring uninstall blackbox
release "blackbox" uninstalledprobe_success returns two series. The monitoring target is 1. The team-a target is 0 with probe_duration_seconds at the 5 second module timeout, because team-a's default-deny NetworkPolicy has no rule for ingress from monitoring. A blackbox probe measures reachability from where the prober sits, which is why it disagrees with the target's own metrics.A ServiceMonitor only looks in its own namespace unless you tell it otherwise. That single omission is the most common reason a correct-looking monitor produces no targets and no error.
kubectl -n monitoring delete servicemonitor argocd-metrics --ignore-not-found
kubectl apply -f - <<'EOF'
apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata: { name: argocd-metrics, namespace: monitoring, labels: { release: prometheus } }
spec:
selector:
matchLabels: { app.kubernetes.io/name: argocd-metrics }
endpoints:
- port: http-metrics
EOF
sleep 60
kubectl -n monitoring port-forward svc/prometheus-kube-prometheus-prometheus 9090:9090 & PF1=$!
sleep 5
sleep 60
curl -s localhost:9090/api/v1/targets | jq '[.data.activeTargets[] | select(.scrapePool | contains("argocd-metrics"))] | length'
kubectl -n monitoring patch servicemonitor argocd-metrics --type merge -p '{"spec":{"namespaceSelector":{"matchNames":["argocd"]}}}'
sleep 60
sleep 60
curl -s localhost:9090/api/v1/targets | jq '[.data.activeTargets[] | select(.scrapePool | contains("argocd-metrics"))] | length'
kill $PF1
kubectl -n monitoring delete servicemonitor argocd-metricsoutputcaptured 2026-09-13
$ kubectl -n monitoring delete servicemonitor argocd-metrics --ignore-not-found
$ kubectl apply -f - <<'EOF'
apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata: { name: argocd-metrics, namespace: monitoring, labels: { release: prometheus } }
spec:
selector:
matchLabels: { app.kubernetes.io/name: argocd-metrics }
endpoints:
- port: http-metrics
EOF
servicemonitor.monitoring.coreos.com/argocd-metrics created
$ sleep 60
$ kubectl -n monitoring port-forward svc/prometheus-kube-prometheus-prometheus 9090:9090 & PF1=$!
$ sleep 5
Forwarding from 127.0.0.1:9090 -> 9090
Forwarding from [::1]:9090 -> 9090
$ sleep 60
$ curl -s localhost:9090/api/v1/targets | jq '[.data.activeTargets[] | select(.scrapePool | contains("argocd-metrics"))] | length'
Handling connection for 9090
0
$ kubectl -n monitoring patch servicemonitor argocd-metrics --type merge -p '{"spec":{"namespaceSelector":{"matchNames":["argocd"]}}}'
servicemonitor.monitoring.coreos.com/argocd-metrics patched
$ sleep 60
$ sleep 60
$ curl -s localhost:9090/api/v1/targets | jq '[.data.activeTargets[] | select(.scrapePool | contains("argocd-metrics"))] | length'
Handling connection for 9090
1
$ kill $PF1
$ kubectl -n monitoring delete servicemonitor argocd-metrics
servicemonitor.monitoring.coreos.com "argocd-metrics" deleted from monitoring namespacenamespaceSelector to your mental template for every ServiceMonitor you write in a task.A Prometheus that falls over is usually a labels problem, not a volume problem. Two queries and the TSDB status page tell you which metric and which label are responsible, which is the answer a capacity question wants.
kubectl -n monitoring port-forward svc/prometheus-kube-prometheus-prometheus 9090:9090 & PF1=$!
sleep 5
curl -sG localhost:9090/api/v1/query --data-urlencode 'query=topk(10, count by (__name__) ({__name__=~".+"}))' | jq -r '.data.result[] | "\(.metric.__name__) \(.value[1])"'
curl -s localhost:9090/api/v1/status/tsdb | jq '{numSeries: .data.headStats.numSeries, numLabelPairs: .data.headStats.numLabelPairs, topSeriesCountByMetricName: .data.seriesCountByMetricName[:5]}'
curl -s localhost:9090/api/v1/status/tsdb | jq '.data.labelValueCountByLabelName[:5]'
kill $PF1outputcaptured 2026-09-12
$ kubectl -n monitoring port-forward svc/prometheus-kube-prometheus-prometheus 9090:9090 & PF1=$!
$ sleep 5
Forwarding from 127.0.0.1:9090 -> 9090
Forwarding from [::1]:9090 -> 9090
$ curl -sG localhost:9090/api/v1/query --data-urlencode 'query=topk(10, count by (__name__) ({__name__=~".+"}))' | jq -r '.data.result[] | "\(.metric.__name__) \(.value[1])"'
Handling connection for 9090
apiserver_request_duration_seconds_bucket 13550
apiserver_request_body_size_bytes_bucket 13472
apiserver_request_sli_duration_seconds_bucket 8224
etcd_request_duration_seconds_bucket 7560
apiserver_watch_list_duration_seconds_bucket 5760
apiserver_response_sizes_bucket 5000
apiserver_watch_cache_read_wait_seconds_bucket 3850
apiserver_watch_events_sizes_bucket 2628
kubernetes_feature_enabled 2177
apiserver_request_total 1683
$ curl -s localhost:9090/api/v1/status/tsdb | jq '{numSeries: .data.headStats.numSeries, numLabelPairs: .data.headStats.numLabelPairs, topSeriesCountByMetricName: .data.seriesCountByMetricName[:5]}'
Handling connection for 9090
{
"numSeries": 169887,
"numLabelPairs": 11365,
"topSeriesCountByMetricName": [
{
"name": "apiserver_request_duration_seconds_bucket",
"value": 13550
},
{
"name": "apiserver_request_body_size_bytes_bucket",
"value": 13472
},
{
"name": "apiserver_request_sli_duration_seconds_bucket",
"value": 8224
},
{
"name": "etcd_request_duration_seconds_bucket",
"value": 7560
},
{
"name": "apiserver_watch_list_duration_seconds_bucket",
"value": 5760
}
]
}
$ curl -s localhost:9090/api/v1/status/tsdb | jq '.data.labelValueCountByLabelName[:5]'
Handling connection for 9090
[
{
"name": "__name__",
"value": 2476
},
{
"name": "name",
"value": 1391
},
{
"name": "id",
"value": 672
},
{
"name": "resource",
"value": 618
},
{
"name": "le",
"value": 415
}
]
$ kill $PF1Two functions carry more exam weight than their length suggests. One extrapolates a trend into the future, the other alerts on a series that stopped existing, which is the failure a threshold can never catch.
kubectl -n monitoring port-forward svc/prometheus-kube-prometheus-prometheus 9090:9090 & PF1=$!
sleep 5
curl -sG localhost:9090/api/v1/query --data-urlencode 'query=count by (mountpoint,fstype) (node_filesystem_avail_bytes)' | jq -r '.data.result[] | "\(.metric.mountpoint) \(.metric.fstype)"'
curl -sG localhost:9090/api/v1/query --data-urlencode 'query=predict_linear(node_filesystem_avail_bytes[1h], 24*3600)' | jq '[.data.result[] | {instance: .metric.instance, mountpoint: .metric.mountpoint, predicted: .value[1]}] | .[0:3]'
curl -sG localhost:9090/api/v1/query --data-urlencode 'query=absent_over_time(up{job="does-not-exist"}[5m])' | jq '.data.result'
curl -sG localhost:9090/api/v1/query --data-urlencode 'query=absent_over_time(up{job="kubelet"}[5m])' | jq '.data.result'
kill $PF1outputcaptured 2026-09-13
$ kubectl -n monitoring port-forward svc/prometheus-kube-prometheus-prometheus 9090:9090 & PF1=$!
$ sleep 5
Forwarding from 127.0.0.1:9090 -> 9090
Forwarding from [::1]:9090 -> 9090
$ curl -sG localhost:9090/api/v1/query --data-urlencode 'query=count by (mountpoint,fstype) (node_filesystem_avail_bytes)' | jq -r '.data.result[] | "\(.metric.mountpoint) \(.metric.fstype)"'
Handling connection for 9090
/etc/hostname btrfs
/etc/hosts btrfs
/etc/kubernetes/audit btrfs
/etc/resolv.conf btrfs
/usr/lib/modules btrfs
/var btrfs
/run tmpfs
/run/credentials/systemd-journald.service tmpfs
/tmp tmpfs
$ curl -sG localhost:9090/api/v1/query --data-urlencode 'query=predict_linear(node_filesystem_avail_bytes[1h], 24*3600)' | jq '[.data.result[] | {instance: .metric.instance, mountpoint: .metric.mountpoint, predicted: .value[1]}] | .[0:3]'
Handling connection for 9090
[
{
"instance": "172.18.0.4:9100",
"mountpoint": "/etc/hostname",
"predicted": "907166223218.7759"
},
{
"instance": "172.18.0.4:9100",
"mountpoint": "/etc/hosts",
"predicted": "907166223218.7759"
},
{
"instance": "172.18.0.4:9100",
"mountpoint": "/etc/kubernetes/audit",
"predicted": "907166223218.7759"
}
]
$ curl -sG localhost:9090/api/v1/query --data-urlencode 'query=absent_over_time(up{job="does-not-exist"}[5m])' | jq '.data.result'
Handling connection for 9090
[
{
"metric": {
"job": "does-not-exist"
},
"value": [
1789306067.750,
"1"
]
}
]
$ curl -sG localhost:9090/api/v1/query --data-urlencode 'query=absent_over_time(up{job="kubelet"}[5m])' | jq '.data.result'
Handling connection for 9090
[]
$ kill $PF1/ is not among them, because a kind node's root is an overlay mount and node-exporter excludes overlay by filesystem type, so the usual mountpoint="/" selector matches nothing and predict_linear over it returns nothing rather than an error. The second query returns one predicted number per filesystem that does exist, and absent_over_time returns a series for the job that does not exist and nothing for the one that does. That inversion is the whole idea: the alert fires when the query has a result.Retention is two numbers and whichever binds first wins. On a cluster you inherit, read them before you promise anyone a month of history.
kubectl -n monitoring get prometheus -o jsonpath='{.items[0].spec.retention} {.items[0].spec.retentionSize}{"\n"}'
kubectl -n monitoring get prometheus -o jsonpath='{.items[0].spec.storage}' | jq
kubectl -n monitoring get pvc -l app.kubernetes.io/name=prometheus -o custom-columns=NAME:.metadata.name,CAPACITY:.status.capacity.storage,CLASS:.spec.storageClassName
kubectl -n monitoring port-forward svc/prometheus-kube-prometheus-prometheus 9090:9090 & PF1=$!
sleep 5
curl -s localhost:9090/api/v1/status/runtimeinfo | jq '.data | {storageRetention, corruptionCount, timeSeriesCount}'
kill $PF1outputcaptured 2026-09-12
$ kubectl -n monitoring get prometheus -o jsonpath='{.items[0].spec.retention} {.items[0].spec.retentionSize}{"\n"}'
6h
$ kubectl -n monitoring get prometheus -o jsonpath='{.items[0].spec.storage}' | jq
$ kubectl -n monitoring get pvc -l app.kubernetes.io/name=prometheus -o custom-columns=NAME:.metadata.name,CAPACITY:.status.capacity.storage,CLASS:.spec.storageClassName
NAME CAPACITY CLASS
$ kubectl -n monitoring port-forward svc/prometheus-kube-prometheus-prometheus 9090:9090 & PF1=$!
$ sleep 5
Forwarding from 127.0.0.1:9090 -> 9090
Forwarding from [::1]:9090 -> 9090
$ curl -s localhost:9090/api/v1/status/runtimeinfo | jq '.data | {storageRetention, corruptionCount, timeSeriesCount}'
Handling connection for 9090
{
"storageRetention": "6h",
"corruptionCount": 0,
"timeSeriesCount": null
}
$ kill $PF1Flux and Tekton both export metrics that answer delivery questions rather than infrastructure ones. Wiring them is two objects, and the queries afterwards are the raw material for section 4.5.
kubectl apply -f - <<'EOF'
apiVersion: monitoring.coreos.com/v1
kind: PodMonitor
metadata: { name: flux-controllers, namespace: monitoring, labels: { release: prometheus } }
spec:
namespaceSelector: { matchNames: [flux-system] }
selector:
matchExpressions:
- { key: app, operator: Exists }
podMetricsEndpoints:
- port: http-prom
---
apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata: { name: tekton-controller, namespace: monitoring, labels: { release: prometheus } }
spec:
namespaceSelector: { matchNames: [tekton-pipelines] }
selector:
matchLabels: { app: tekton-pipelines-controller }
endpoints:
- port: http-metrics
EOF
sleep 90
kubectl -n monitoring port-forward svc/prometheus-kube-prometheus-prometheus 9090:9090 & PF1=$!
sleep 5
curl -sG localhost:9090/api/v1/query --data-urlencode 'query=gotk_reconcile_duration_seconds_count' | jq '.data.result | length'
curl -sG localhost:9090/api/v1/query --data-urlencode 'query=tekton_pipelines_controller_pipelinerun_total' | jq '.data.result[] | {status: .metric.status, value: .value[1]}'
kill $PF1outputcaptured 2026-09-12
$ kubectl apply -f - <<'EOF'
apiVersion: monitoring.coreos.com/v1
kind: PodMonitor
metadata: { name: flux-controllers, namespace: monitoring, labels: { release: prometheus } }
spec:
namespaceSelector: { matchNames: [flux-system] }
selector:
matchExpressions:
- { key: app, operator: Exists }
podMetricsEndpoints:
- port: http-prom
---
apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata: { name: tekton-controller, namespace: monitoring, labels: { release: prometheus } }
spec:
namespaceSelector: { matchNames: [tekton-pipelines] }
selector:
matchLabels: { app: tekton-pipelines-controller }
endpoints:
- port: http-metrics
EOF
podmonitor.monitoring.coreos.com/flux-controllers created
servicemonitor.monitoring.coreos.com/tekton-controller created
$ sleep 90
$ kubectl -n monitoring port-forward svc/prometheus-kube-prometheus-prometheus 9090:9090 & PF1=$!
$ sleep 5
Forwarding from 127.0.0.1:9090 -> 9090
Forwarding from [::1]:9090 -> 9090
$ curl -sG localhost:9090/api/v1/query --data-urlencode 'query=gotk_reconcile_duration_seconds_count' | jq '.data.result | length'
Handling connection for 9090
8
$ curl -sG localhost:9090/api/v1/query --data-urlencode 'query=tekton_pipelines_controller_pipelinerun_total' | jq '.data.result[] | {status: .metric.status, value: .value[1]}'
Handling connection for 9090
{
"status": "success",
"value": "1"
}
$ kill $PF1Self-check
Your ServiceMonitor is correct and no target appears. List the checks in order.
1) Prometheus's serviceMonitorSelector and serviceMonitorNamespaceSelector: does it even consider your monitor? 2) Does the monitor's selector match the Service's labels? 3) Is endpoints[].port the port name, and does the Service name it? 4) Does the Service have ready endpoints? 5) RBAC and NetworkPolicy between monitoring and the target namespace.
Why is sum(rate(x[5m])) right and rate(sum(x)[5m]) wrong?
Counters reset per series; rate knows how to handle resets within a single series. Summing first merges series (and their resets) into a meaningless line, and the syntax is invalid anyway without a subquery. Rate first, aggregate second, always.
How do you alert on "the metric disappeared entirely"?
absent(up{job="x"} == 1) or absent(metric): an ordinary comparison returns no series when there is nothing to compare, so it can never fire. absent() exists precisely to turn "no data" into a value of 1.
A p95 query returns NaN. Two plausible reasons?
No observations in the window (nothing has been recorded yet, so all buckets are empty), or you aggregated away the le label that histogram_quantile needs. The form that keeps it: histogram_quantile(0.95, sum by (le) (rate(x_bucket[5m]))).
An exporter added a label with one value per request. What happens, and what is the fix?
Cardinality explosion: a new series per unique value, memory and TSDB growth, slow queries. Fix at ingestion with a metricRelabeling that drops the label or the metric, and upstream by removing it from the exporter. High-cardinality labels (user IDs, URLs with IDs, trace IDs) belong in logs and traces, not metric labels.
You need to scrape a database VM outside the cluster with the Prometheus Operator, without editing the Prometheus config by hand. Which CRD, and which selector must match?
A ScrapeConfig with staticConfigs[].targets: ["db.example.internal:9187"] (or an httpSDConfigs/dnsSDConfigs entry), labeled to match the Prometheus object's scrapeConfigSelector, in a namespace its scrapeConfigNamespaceSelector covers. If that selector is nil, the ScrapeConfig must live in Prometheus's own namespace.
Predict that a volume fills within a day, and alert if any series that should exist has vanished. Two expressions?
predict_linear(node_filesystem_avail_bytes{mountpoint="/data"}[6h], 24*3600) < 0 for the volume (a gauge extrapolated linearly). absent_over_time(up{job="payments"}[10m]) for the disappearance: it yields 1 only when no sample existed in the window, whereas a plain comparison returns nothing when there is nothing to compare.
A ServiceMonitor in monitoring selects the right labels and port for Services in argocd, yet no target appears and Prometheus's own selectors are {}. What is missing?
spec.namespaceSelector on the ServiceMonitor itself: matchNames: [argocd] or any: true. Omitted, it means "Services in my own namespace", so the monitor matches nothing in argocd. Prometheus's serviceMonitorNamespaceSelector decides which monitors are read; the monitor's own namespaceSelector decides which Services it may reach.
Docs to know your way around
- prometheus.io: querying basics and the function reference (
rate,increase,histogram_quantile,absent); metric types. - prometheus-operator.dev: the ServiceMonitor troubleshooting page, which is the release-label story in official form; the API reference for ServiceMonitor/PodMonitor/ScrapeConfig.
- Offline:
kubectl explain servicemonitor.spec.endpoints, the Prometheus UI's Status → Targets and Status → Configuration pages, and its built-in expression autocompletion. - prometheus-operator.dev/docs/developer/scrapeconfig: the ScrapeConfig page with the selector and namespace rules; the API reference's
ServiceMonitorsection listsjobLabel,targetLabels, the limits andserviceDiscoveryRole. - prometheus.io/docs/prometheus/latest/querying/functions:
predict_linear,absent_over_time,label_replace, and the native-histogram functions; /storage for the retention flags. - argo-cd.readthedocs.io/en/stable/operator-manual/metrics, fluxcd.io/flux/monitoring/metrics, tekton.dev/docs/pipelines/metrics, docs.crossplane.io/latest/guides/metrics: the metric-name tables, one page each.