The obs layer installs kube-prometheus-stack, which is three things people conflate: Prometheus itself, the Prometheus Operator that configures it through CRDs, and a bundle of exporters and dashboards. On the exam you work through the operator's CRDs, so that is the layer to be fluent in.

needsmake up obs

Orientation

competency 4.1 · monitoring solutions

Two skills, both testable: get a target scraped, and write a query that answers a question. Everything else in domain 4 (alerting, dashboards, DORA metrics, incident diagnosis) is built on those two.

The pull model in one paragraph

Prometheus scrapes HTTP endpoints on an interval and stores samples in a local TSDB. Each series is a metric name plus a set of labels, and every distinct label combination is a separate series, which is why cardinality is the resource you actually manage. Service discovery (in Kubernetes: the API server) produces the target list; relabeling rewrites and filters it; the scrape either succeeds (up = 1) or does not (up = 0). Nothing is pushed, so a workload that only exists for ten seconds is invisible unless it pushes to a Pushgateway or emits an event elsewhere.

The operator model

ServiceMonitor · PodMonitor · selectors

What to scrape is declared, not configured. A ServiceMonitor says "scrape the endpoints of Services matching these labels, on this port name, at this path, every N seconds". A PodMonitor does the same without a Service. A ScrapeConfig (newer) covers targets outside Kubernetes. A Probe drives blackbox-exporter checks. The operator watches these CRDs and rewrites Prometheus's configuration live.

ServiceMonitor (selector: app=example, port: web, path: /metrics)
      │  matched by Prometheus.spec.serviceMonitorSelector   ← the filter people forget
      ▼
Service app=example, ports: [{name: web, port: 8080}]
      │  its EndpointSlice supplies the actual pod IPs
      ▼
pod:8080/metrics  ──scraped every 30s──▶  series {job="example", namespace=…, pod=…}
The silent selector trap

Prometheus only picks up ServiceMonitors matching its own serviceMonitorSelector, which on a stock kube-prometheus-stack matches the Helm release label (release: prometheus). A perfectly correct monitor without that label is silently ignored: no error, no event, no target. This lab deliberately disables that filter (serviceMonitorSelectorNilUsesHelmValues=false), so its selector is {} and every monitor everywhere is picked up. Either way, read the selector first, then write the monitor.

kubectl -n monitoring get prometheus -o jsonpath='{.items[0].spec.serviceMonitorSelector}'; echo
kubectl -n monitoring get prometheus -o jsonpath='{.items[0].spec.serviceMonitorNamespaceSelector}'; echo
outputcaptured 2026-08-26
$ kubectl -n monitoring get prometheus -o jsonpath='{.items[0].spec.serviceMonitorSelector}'; echo
{}
$ kubectl -n monitoring get prometheus -o jsonpath='{.items[0].spec.serviceMonitorNamespaceSelector}'; echo
{}

Prints {} on this cluster and the release-label selector on a stock install. That one check is the answer to "my target does not appear" everywhere. The exam's cluster has its own selector, so check it rather than assuming. The namespace selector is the second half of the same trap: a monitor in a namespace Prometheus does not watch is equally invisible.

The other three ways a target goes missing

  1. Port name, not number. endpoints[].port is the Service's port name. An unnamed Service port cannot be referenced, and the monitor matches while scraping nothing.
  2. No endpoints. The Service selector matches no ready pods (section 1.1). Same symptom, different layer.
  3. RBAC or network. Prometheus's ServiceAccount must be able to list endpoints in that namespace, and NetworkPolicy must allow monitoring → workload. In a default-deny tenant namespace, that allow is a thing someone has to write.

Diagnosis order: Status → Targets in the UI (it shows the reason for a failed scrape), then up{job="…"}, then the four causes above.

Relabeling, at exam depth

relabelings act on discovered targets before the scrape (drop targets, rewrite the job or namespace labels); metricRelabelings act on samples after (drop expensive series). The one pattern worth remembering: dropping a high-cardinality metric at ingestion with action: drop on a __name__ regex is how you rescue a Prometheus drowning in one bad exporter.

PromQL, the working subset

90% of exam queries live here
TypeRead it withExample
Counter (only climbs, resets to 0)rate() · increase()rate(http_requests_total[5m])
Gauge (up and down)read it directlykube_pod_status_ready
Histogram (buckets)histogram_quantile()histogram_quantile(0.95, sum by (le) (rate(x_bucket[5m])))
Summary (pre-computed quantiles)read the quantile labelx{quantile="0.99"}
  • Selectors: up{namespace="argocd"}, with =, !=, =~, !~. Regexes are fully anchored.
  • Ranges: [5m] turns an instant vector into a range vector: required by rate/increase, and forbidden everywhere else. "Expected type instant vector" almost always means a stray range selector.
  • Aggregation: sum by (label) (...) and sum without (pod) (...). by keeps only what you name; without keeps everything else. Aggregating before rate is wrong: rate first, then sum.
  • Arithmetic between vectors matches on labels; mismatched label sets produce an empty result, which is why on()/ignoring() and group_left exist. You used vector arithmetic for the cost gap in section 1.5.
  • Emptiness: absent(x) is 1 when x has no series (the way to alert on "the metric stopped existing"), and … or vector(0) is how you make a query return a number instead of nothing.
  • rate vs irate: rate averages over the window (use it for alerts and dashboards), irate uses the last two samples (spiky, for zooming in). increase is rate × window, useful for "how many in 24h".

Why the raw counter is useless, and what the window does to the answer:

The up metric

1 per healthy target, 0 per failing one, which makes count(up == 0) the cluster's own health check. make validate runs exactly that query and demands zero; the lab patched kubeadm's localhost-bound control-plane metrics to make them reachable (make fix-cp-metrics). Any "is monitoring healthy" question starts with this query.

What is already collected, for free

SourceGives youTypical series
kube-state-metricsobject state from the API serverkube_pod_info, kube_deployment_status_replicas_unavailable, kube_pod_container_status_restarts_total
cAdvisor (kubelet)container resource usagecontainer_cpu_usage_seconds_total, container_memory_working_set_bytes
node-exporterhost metricsnode_load1, node_filesystem_avail_bytes
kubelet / apiservercontrol-plane health and latencyapiserver_request_duration_seconds_bucket

Knowing these exist means you rarely need to instrument anything to answer an exam question about the cluster itself.

The rest of the operator's CRDs

ScrapeConfig, Probe, PodMonitor, the selector family

Every operator CRD is picked up through a pair of selectors on the Prometheus object, and every pair has the same two failure modes: the object's labels do not match, or its namespace is not watched. Learn the pairs as a table and you never lose time to them again.

KindSelected byScrapesField that decides
ServiceMonitorserviceMonitorSelector / serviceMonitorNamespaceSelectorendpoints of matching Servicesendpoints[].port is a port name; targetPort exists for numbers
PodMonitorpodMonitorSelector / podMonitorNamespaceSelectorpods directly, no Service neededpodMetricsEndpoints[].port is the container port name; portNumber for numbers
ProbeprobeSelector / probeNamespaceSelectorblackbox-exporter checksprober.url (the exporter), targets.staticConfig.static[] or targets.ingress, module
ScrapeConfigscrapeConfigSelector / scrapeConfigNamespaceSelectoranything: staticConfigs, kubernetesSDConfigs, httpSDConfigs, fileSDConfigs, dnsSDConfigs, consul and cloud SDa raw scrape_config as a CRD; use it for a VM, an external database, or a role: node discovery
PrometheusRuleruleSelector / ruleNamespaceSelectornothing; evaluatedsee 4.2
AlertmanagerConfigalertmanagerConfigSelector / alertmanagerConfigNamespaceSelector (on the Alertmanager object)nothing; merged into routingsee 4.2
  • Namespace selector semantics. On the Prometheus object, a nil namespace selector means "this namespace only", {} means every namespace, and matchLabels narrows by namespace label (kubernetes.io/metadata.name: team-a selects one by name). On a ServiceMonitor, spec.namespaceSelector is the second hop: matchNames: [argocd] or any: true; leaving it out means the monitor's own namespace, which is why a monitor in monitoring pointed at Services in argocd scrapes nothing without it.
  • Labels you get for free. jobLabel names the Service label whose value becomes job; targetLabels and podTargetLabels copy Service or pod labels onto every series. Without them the only identity on the series is namespace, pod, service, endpoint, so "group by team" needs the team label copied here first.
  • Endpoints to EndpointSlice. Since Kubernetes 1.33 the v1 Endpoints API is deprecated. Newer operators support spec.serviceDiscoveryRole: EndpointSlice on ServiceMonitors (or globally on the Prometheus object) and need discovery.k8s.io/endpointslices in Prometheus's RBAC. A warning in the operator log about v1 Endpoints is deprecated is this migration, not a fault.
  • Limits that protect you. sampleLimit, labelLimit, labelNameLengthLimit and labelValueLengthLimit on a monitor drop a whole scrape that exceeds them (the target shows down with the reason in Status → Targets), which is the blunt tool when metricRelabelings is too slow to write. scrapeInterval and scrapeTimeout per endpoint override the global 30s; timeout must be shorter than interval or the config is rejected.
  • Relabel actions. replace (default), keep/drop on a regex over sourceLabels, labelmap (copy __meta_kubernetes_pod_label_(.+) to plain labels), labeldrop/labelkeep, hashmod. The __address__, __scheme__, __metrics_path__ targets are how a relabel rewrites where to scrape; a ScrapeConfig with kubernetesSDConfigs almost always needs them because discovery gives you metadata, not a metrics URL.
Reflex

Before writing any monitor: read the six selector pairs off the Prometheus object with kubectl -n monitoring get prometheus -o yaml, then kubectl get svc <name> -o jsonpath='{.spec.ports[*].name}' for the port name, then kubectl get endpointslices -l kubernetes.io/service-name=<name> for ready addresses. Three reads, no guessing.

PromQL beyond rate, and the metrics platform tools give you

functions a task uses, storage knobs, tool metric names
Function / formUse it forRemember
increase(c[1h])"how many in the last hour"extrapolated; can be non-integer, and it is rate × window
irate(c[5m])the last two samples onlyzooming in, never alerting
changes(g[1h]) / resets(c[1h])flapping detection; counter restartschanges(kube_pod_container_status_ready[10m]) > 3 is a crash-loop detector
absent(x) / absent_over_time(x[10m])alert when a series disappearsreturns 1 with the labels you wrote in the selector
predict_linear(g[1h], 4*3600)disk or certificate expiry in N secondsgauges only; predict_linear(node_filesystem_avail_bytes[6h], 24*3600) < 0
label_replace(v, "dst", "$1", "src", "(.*)")derive or rename a label for joinsregex anchored; label_join concatenates
topk(5, ...) / bottomkworst offendersinstant queries; in a range query it can return more than k series over time
count by (namespace) (kube_pod_info)counting objectscount counts series, sum adds values; for kube-state-metrics info series they differ
histogram_quantile(0.99, sum by (le) (rate(b[5m])))classic histogramskeep le; on native histograms drop the by (le) and the _bucket suffix
histogram_count(h) / histogram_sum(h) / histogram_fraction(0, 0.3, h)native histogramsPrometheus 3.8+ stable; scraping still needs scrape_native_histograms: true (operator: scrapeNativeHistograms)
x offset 1w / x @ end()compare with last week; pin to a timeoffset goes after the selector, not the function
a / on(pod) group_left(team) bjoins with mismatched labelson names the join keys; group_left copies labels from the many side
(sum(...) or vector(0))a number instead of emptywhat a dashboard "single stat" needs

Storage, cardinality, scale

  • Retention. --storage.tsdb.retention.time (default 15d when neither is set) and --storage.tsdb.retention.size; the operator exposes them as spec.retention and spec.retentionSize. Size should stay at roughly 80-85% of the volume because compaction briefly overshoots.
  • Finding cardinality. Status → TSDB Status lists the top series counts by metric name and label; the query form is topk(10, count by (__name__) ({__name__=~".+"})). Fix at the source (metricRelabelings with action: drop on __name__, or labeldrop on the offending label) rather than by buying memory.
  • Beyond one server. remoteWrite (protocol 2.0 in Prometheus 3.x) ships samples to Thanos Receive, Mimir, Cortex or a vendor; Thanos sidecar plus object storage is the other long-term pattern; PrometheusAgent is the operator's scrape-only mode that exists solely to remote-write. Federation (/federate?match[]=... scraped with honor_labels: true) pulls aggregated series from another Prometheus and is the older, smaller-scale answer. In a question, "long-term retention across clusters" is remote write or Thanos, not federation.
  • Native histograms in Kubernetes. Kubernetes 1.36 added the NativeHistograms feature gate (alpha, off) for control-plane metrics; 1.37 makes it beta and on. Until then the components expose classic buckets only, so _bucket queries against apiserver_request_duration_seconds remain right.

Metric names the platform tools expose

ToolKey seriesScrape via
Argo CDargocd_app_info{sync_status,health_status}, argocd_app_sync_total{phase}, argocd_app_reconcile, argocd_appset_infoServices *-metrics in argocd, port http-metrics
Fluxgotk_reconcile_duration_seconds (per kind/name/namespace), controller_runtime_reconcile_total{controller,result}, and gotk_resource_info{ready,suspended,revision}, which comes from kube-state-metrics with a CustomResourceState config, not from the controllersPodMonitor on the flux-system controllers, port http-prom
Tektontekton_pipelines_controller_pipelinerun_total{status}, ..._pipelinerun_duration_seconds (histogram or last value, per config-observability), ..._running_pipelineruns, ..._taskrun_totaltekton-pipelines-controller Service, port 9090
Crossplanecrossplane_managed_resource_ready, ..._synced, ..._first_time_to_readiness_seconds, ..._drift_seconds (providers, port 8080); core needs metrics.enabled=truePodMonitor on provider pods, or annotations via DeploymentRuntimeConfig
Any CRDkube_customresource_* built from --custom-resource-state-config (groupVersionKind, metrics[].each.gauge.path, labelsFromPath)kube-state-metrics, plus RBAC to list the kind
Kyverno / Gatekeeperkyverno_policy_results_total{policy_validation_mode,rule_result}, kyverno_admission_review_duration_seconds; gatekeeper_violations{enforcement_action}, gatekeeper_audit_last_run_timeServiceMonitors on their metrics Services
How this gets tested

The recurring task is "expose tool X's metrics to Prometheus and write a query that shows Y". It is 4.1's selector and port-name drill against an unfamiliar Service, followed by one query from the table above. Discover the Service and its port names first; the metric names you can read from the /metrics page itself once the target is up.

Exercises

tick the dot when its check passes

Prometheus's UI address comes from make urls.

In the UI under Status → Targets, or via API: count the scrape pools, find which ServiceMonitor each corresponds to (kubectl -n monitoring get servicemonitors), and confirm zero targets down.

verify: curl -s 'http://<prom>/api/v1/query?query=count(up==0)' | jq -r '.data.result[0].value[1] // "0"' prints 0. If not, the down target's job name points at the responsible ServiceMonitor.

Deploy an app that exposes metrics and monitor it, first wrong, then right:

kubectl create deploy example --image=quay.io/brancz/prometheus-example-app:v0.5.0 --port=8080
kubectl expose deploy example --port=8080 --name=example
kubectl apply -f - <<'EOF'
apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata: { name: example, namespace: default }
spec:
  selector: { matchLabels: { app: example } }
  endpoints: [{ port: "8080" }]
EOF
outputcaptured 2026-08-26
$ kubectl create deploy example --image=quay.io/brancz/prometheus-example-app:v0.5.0 --port=8080
deployment.apps/example created
$ kubectl expose deploy example --port=8080 --name=example
service/example exposed
$ kubectl apply -f - <<'EOF'
apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata: { name: example, namespace: default }
spec:
  selector: { matchLabels: { app: example } }
  endpoints: [{ port: "8080" }]
EOF
servicemonitor.monitoring.coreos.com/example created

Wait a minute, search Targets: nothing. The endpoint's port must be the Service's port name, and this Service has none, so the monitor matches the Service and then finds no endpoint. It fails silently. Fix it: patch the Service so the port is named (kubectl patch svc example --type=json -p '[{"op":"add","path":"/spec/ports/0/name","value":"web"}]') and set port: web in the monitor.

verify: up{job="example"} returns 1 in the UI. Then add the release: prometheus label to your monitor anyway and confirm nothing changes here, while being able to say why it would be the difference between working and ignored on a stock install.

Generate a little traffic (kubectl run curl --image=curlimages/curl:8.11.1 --restart=Never -- sh -c 'for i in $(seq 100); do curl -s example.default.svc:8080; done'), then write, without copying: the request rate (rate(http_requests_total{job="example"}[5m])), summed by status code, and the p95 duration from http_request_duration_seconds_bucket.

verify: rates are non-zero and the quantile returns a number, not NaN. If it is NaN, reason about why (too few buckets observed yet); that reasoning is itself testable.

Answer three questions using only metrics already collected: how many pods per namespace (kube_pod_info), which container restarts most (kube_pod_container_status_restarts_total), and each node's CPU pressure (node_load1 vs allocatable).

verify: each answer is one query in the UI. kube-state-metrics and node-exporter are pre-installed sources the exam expects you to know exist.

ServiceMonitor and PodMonitor only reach things Kubernetes knows about. A ScrapeConfig is how the operator scrapes a static target, which is what you need for a switch, an appliance, or anything outside the cluster.

kubectl -n monitoring get prometheus -o jsonpath='{.items[0].spec.scrapeConfigSelector}{"\n"}'
kubectl -n default create deployment fakemetrics --image=ghcr.io/stefanprodan/podinfo:6.7.1
kubectl -n default expose deployment fakemetrics --port=9898
kubectl -n default get svc fakemetrics -o jsonpath='{.spec.clusterIP}{"\n"}'
kubectl apply -f - <<'EOF'
apiVersion: monitoring.coreos.com/v1alpha1
kind: ScrapeConfig
metadata: { name: static-target, namespace: monitoring, labels: { release: prometheus } }
spec:
  staticConfigs:
    - targets: ["fakemetrics.default.svc.cluster.local:9898"]
      labels: { source: static }
  metricsPath: /metrics
  scrapeInterval: 30s
EOF
sleep 60
kubectl -n monitoring port-forward svc/prometheus-kube-prometheus-prometheus 9090:9090 & PF1=$!
sleep 5
curl -s localhost:9090/api/v1/targets | jq '.data.activeTargets[] | select(.labels.source=="static") | {scrapeUrl, health, lastError}'
kill $PF1
kubectl delete scrapeconfig static-target -n monitoring
kubectl -n default delete deployment fakemetrics
kubectl -n default delete svc fakemetrics
outputcaptured 2026-09-12
$ kubectl -n monitoring get prometheus -o jsonpath='{.items[0].spec.scrapeConfigSelector}{"\n"}'
{"matchLabels":{"release":"prometheus"}}
$ kubectl -n default create deployment fakemetrics --image=ghcr.io/stefanprodan/podinfo:6.7.1
deployment.apps/fakemetrics created
$ kubectl -n default expose deployment fakemetrics --port=9898
service/fakemetrics exposed
$ kubectl -n default get svc fakemetrics -o jsonpath='{.spec.clusterIP}{"\n"}'
10.96.20.183
$ kubectl apply -f - <<'EOF'
apiVersion: monitoring.coreos.com/v1alpha1
kind: ScrapeConfig
metadata: { name: static-target, namespace: monitoring, labels: { release: prometheus } }
spec:
  staticConfigs:
    - targets: ["fakemetrics.default.svc.cluster.local:9898"]
      labels: { source: static }
  metricsPath: /metrics
  scrapeInterval: 30s
EOF
scrapeconfig.monitoring.coreos.com/static-target created
$ sleep 60
$ kubectl -n monitoring port-forward svc/prometheus-kube-prometheus-prometheus 9090:9090 & PF1=$!
$ sleep 5
Forwarding from 127.0.0.1:9090 -> 9090
Forwarding from [::1]:9090 -> 9090
$ curl -s localhost:9090/api/v1/targets | jq '.data.activeTargets[] | select(.labels.source=="static") | {scrapeUrl, health, lastError}'
Handling connection for 9090
{
  "scrapeUrl": "http://fakemetrics.default.svc.cluster.local:9898/metrics",
  "health": "up",
  "lastError": ""
}
$ kill $PF1
$ kubectl delete scrapeconfig static-target -n monitoring
scrapeconfig.monitoring.coreos.com "static-target" deleted from monitoring namespace
$ kubectl -n default delete deployment fakemetrics
deployment.apps "fakemetrics" deleted from default namespace
$ kubectl -n default delete svc fakemetrics
service "fakemetrics" deleted from default namespace
verify: the static target appears in the target list with health up and the label you attached. Check scrapeConfigSelector first: if it does not match your labels, the object is valid, ignored, and completely silent.

A Probe answers "can this be reached", which is a different question from any metric the app exports. It needs blackbox-exporter, which the lab does not install, so this block installs it.

Uninstall it afterwards if you are tight on memory; leave it if you intend to write availability alerts in 4.2.

helm repo add prometheus-community https://prometheus-community.github.io/helm-charts
helm repo update prometheus-community
helm upgrade --install blackbox prometheus-community/prometheus-blackbox-exporter -n monitoring --wait
kubectl apply -f - <<'EOF'
apiVersion: monitoring.coreos.com/v1
kind: Probe
metadata: { name: demo-reachable, namespace: monitoring, labels: { release: prometheus } }
spec:
  interval: 30s
  module: http_2xx
  prober:
    url: blackbox-prometheus-blackbox-exporter.monitoring.svc:9115
  targets:
    staticConfig:
      static:
        - http://prometheus-kube-prometheus-prometheus.monitoring.svc:9090/-/healthy
        - http://staging-demo.team-a.svc.cluster.local
EOF
sleep 120
kubectl -n monitoring port-forward svc/prometheus-kube-prometheus-prometheus 9090:9090 & PF1=$!
sleep 5
curl -sG localhost:9090/api/v1/query --data-urlencode 'query=probe_success' | jq '.data.result[] | {instance: .metric.instance, value: .value[1]}'
# Prometheus reloads its config and scrapes before any of this appears
sleep 120
curl -sG localhost:9090/api/v1/query --data-urlencode 'query=probe_duration_seconds' | jq '.data.result[] | {instance: .metric.instance, value: .value[1]}'
kill $PF1
kubectl -n monitoring delete probe demo-reachable
helm -n monitoring uninstall blackbox
outputcaptured 2026-09-12
$ helm repo add prometheus-community https://prometheus-community.github.io/helm-charts
"prometheus-community" already exists with the same configuration, skipping
$ helm repo update prometheus-community
Hang tight while we grab the latest from your chart repositories...
...Successfully got an update from the "prometheus-community" chart repository
Update Complete. ⎈Happy Helming!⎈
$ helm upgrade --install blackbox prometheus-community/prometheus-blackbox-exporter -n monitoring --wait
Release "blackbox" does not exist. Installing it now.
NAME: blackbox
LAST DEPLOYED: Sun Sep 13 12:44:55 2026
NAMESPACE: monitoring
STATUS: deployed
REVISION: 1
DESCRIPTION: Install complete
TEST SUITE: None
NOTES:
See https://github.com/prometheus/blackbox_exporter/ for how to configure Prometheus and the Blackbox Exporter.

1. Get the application URL by running these commands:
  export POD_NAME=$(kubectl get pods --namespace monitoring -l "app.kubernetes.io/name=prometheus-blackbox-exporter,app.kubernetes.io/instance=blackbox" -o jsonpath="{.items[0].metadata.name}")
  export CONTAINER_PORT=$(kubectl get pod --namespace monitoring $POD_NAME -o jsonpath="{.spec.containers[0].ports[0].containerPort}")
  echo "Visit http://127.0.0.1:8080 to use your application"
  kubectl --namespace monitoring port-forward $POD_NAME 8080:$CONTAINER_PORT
$ kubectl apply -f - <<'EOF'
apiVersion: monitoring.coreos.com/v1
kind: Probe
metadata: { name: demo-reachable, namespace: monitoring, labels: { release: prometheus } }
spec:
  interval: 30s
  module: http_2xx
  prober:
    url: blackbox-prometheus-blackbox-exporter.monitoring.svc:9115
  targets:
    staticConfig:
      static:
        - http://prometheus-kube-prometheus-prometheus.monitoring.svc:9090/-/healthy
        - http://staging-demo.team-a.svc.cluster.local
EOF
probe.monitoring.coreos.com/demo-reachable created
$ sleep 120
$ kubectl -n monitoring port-forward svc/prometheus-kube-prometheus-prometheus 9090:9090 & PF1=$!
$ sleep 5
Forwarding from 127.0.0.1:9090 -> 9090
Forwarding from [::1]:9090 -> 9090
$ curl -sG localhost:9090/api/v1/query --data-urlencode 'query=probe_success' | jq '.data.result[] | {instance: .metric.instance, value: .value[1]}'
Handling connection for 9090
{
  "instance": "http://prometheus-kube-prometheus-prometheus.monitoring.svc:9090/-/healthy",
  "value": "1"
}
{
  "instance": "http://staging-demo.team-a.svc.cluster.local",
  "value": "0"
}
$ # Prometheus reloads its config and scrapes before any of this appears
$ sleep 120
$ curl -sG localhost:9090/api/v1/query --data-urlencode 'query=probe_duration_seconds' | jq '.data.result[] | {instance: .metric.instance, value: .value[1]}'
Handling connection for 9090
{
  "instance": "http://prometheus-kube-prometheus-prometheus.monitoring.svc:9090/-/healthy",
  "value": "0.003463362"
}
{
  "instance": "http://staging-demo.team-a.svc.cluster.local",
  "value": "5.002871734"
}
$ kill $PF1
$ kubectl -n monitoring delete probe demo-reachable
probe.monitoring.coreos.com "demo-reachable" deleted from monitoring namespace
$ helm -n monitoring uninstall blackbox
release "blackbox" uninstalled
verify: probe_success returns two series. The monitoring target is 1. The team-a target is 0 with probe_duration_seconds at the 5 second module timeout, because team-a's default-deny NetworkPolicy has no rule for ingress from monitoring. A blackbox probe measures reachability from where the prober sits, which is why it disagrees with the target's own metrics.

A ServiceMonitor only looks in its own namespace unless you tell it otherwise. That single omission is the most common reason a correct-looking monitor produces no targets and no error.

kubectl -n monitoring delete servicemonitor argocd-metrics --ignore-not-found
kubectl apply -f - <<'EOF'
apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata: { name: argocd-metrics, namespace: monitoring, labels: { release: prometheus } }
spec:
  selector:
    matchLabels: { app.kubernetes.io/name: argocd-metrics }
  endpoints:
    - port: http-metrics
EOF
sleep 60
kubectl -n monitoring port-forward svc/prometheus-kube-prometheus-prometheus 9090:9090 & PF1=$!
sleep 5
sleep 60
curl -s localhost:9090/api/v1/targets | jq '[.data.activeTargets[] | select(.scrapePool | contains("argocd-metrics"))] | length'
kubectl -n monitoring patch servicemonitor argocd-metrics --type merge -p '{"spec":{"namespaceSelector":{"matchNames":["argocd"]}}}'
sleep 60
sleep 60
curl -s localhost:9090/api/v1/targets | jq '[.data.activeTargets[] | select(.scrapePool | contains("argocd-metrics"))] | length'
kill $PF1
kubectl -n monitoring delete servicemonitor argocd-metrics
outputcaptured 2026-09-13
$ kubectl -n monitoring delete servicemonitor argocd-metrics --ignore-not-found
$ kubectl apply -f - <<'EOF'
apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata: { name: argocd-metrics, namespace: monitoring, labels: { release: prometheus } }
spec:
  selector:
    matchLabels: { app.kubernetes.io/name: argocd-metrics }
  endpoints:
    - port: http-metrics
EOF
servicemonitor.monitoring.coreos.com/argocd-metrics created
$ sleep 60
$ kubectl -n monitoring port-forward svc/prometheus-kube-prometheus-prometheus 9090:9090 & PF1=$!
$ sleep 5
Forwarding from 127.0.0.1:9090 -> 9090
Forwarding from [::1]:9090 -> 9090
$ sleep 60
$ curl -s localhost:9090/api/v1/targets | jq '[.data.activeTargets[] | select(.scrapePool | contains("argocd-metrics"))] | length'
Handling connection for 9090
0
$ kubectl -n monitoring patch servicemonitor argocd-metrics --type merge -p '{"spec":{"namespaceSelector":{"matchNames":["argocd"]}}}'
servicemonitor.monitoring.coreos.com/argocd-metrics patched
$ sleep 60
$ sleep 60
$ curl -s localhost:9090/api/v1/targets | jq '[.data.activeTargets[] | select(.scrapePool | contains("argocd-metrics"))] | length'
Handling connection for 9090
1
$ kill $PF1
$ kubectl -n monitoring delete servicemonitor argocd-metrics
servicemonitor.monitoring.coreos.com "argocd-metrics" deleted from monitoring namespace
verify: zero targets before the patch and a non-zero count after it, with nothing else changed. Add namespaceSelector to your mental template for every ServiceMonitor you write in a task.

A Prometheus that falls over is usually a labels problem, not a volume problem. Two queries and the TSDB status page tell you which metric and which label are responsible, which is the answer a capacity question wants.

kubectl -n monitoring port-forward svc/prometheus-kube-prometheus-prometheus 9090:9090 & PF1=$!
sleep 5
curl -sG localhost:9090/api/v1/query --data-urlencode 'query=topk(10, count by (__name__) ({__name__=~".+"}))' | jq -r '.data.result[] | "\(.metric.__name__) \(.value[1])"'
curl -s localhost:9090/api/v1/status/tsdb | jq '{numSeries: .data.headStats.numSeries, numLabelPairs: .data.headStats.numLabelPairs, topSeriesCountByMetricName: .data.seriesCountByMetricName[:5]}'
curl -s localhost:9090/api/v1/status/tsdb | jq '.data.labelValueCountByLabelName[:5]'
kill $PF1
outputcaptured 2026-09-12
$ kubectl -n monitoring port-forward svc/prometheus-kube-prometheus-prometheus 9090:9090 & PF1=$!
$ sleep 5
Forwarding from 127.0.0.1:9090 -> 9090
Forwarding from [::1]:9090 -> 9090
$ curl -sG localhost:9090/api/v1/query --data-urlencode 'query=topk(10, count by (__name__) ({__name__=~".+"}))' | jq -r '.data.result[] | "\(.metric.__name__) \(.value[1])"'
Handling connection for 9090
apiserver_request_duration_seconds_bucket 13550
apiserver_request_body_size_bytes_bucket 13472
apiserver_request_sli_duration_seconds_bucket 8224
etcd_request_duration_seconds_bucket 7560
apiserver_watch_list_duration_seconds_bucket 5760
apiserver_response_sizes_bucket 5000
apiserver_watch_cache_read_wait_seconds_bucket 3850
apiserver_watch_events_sizes_bucket 2628
kubernetes_feature_enabled 2177
apiserver_request_total 1683
$ curl -s localhost:9090/api/v1/status/tsdb | jq '{numSeries: .data.headStats.numSeries, numLabelPairs: .data.headStats.numLabelPairs, topSeriesCountByMetricName: .data.seriesCountByMetricName[:5]}'
Handling connection for 9090
{
  "numSeries": 169887,
  "numLabelPairs": 11365,
  "topSeriesCountByMetricName": [
    {
      "name": "apiserver_request_duration_seconds_bucket",
      "value": 13550
    },
    {
      "name": "apiserver_request_body_size_bytes_bucket",
      "value": 13472
    },
    {
      "name": "apiserver_request_sli_duration_seconds_bucket",
      "value": 8224
    },
    {
      "name": "etcd_request_duration_seconds_bucket",
      "value": 7560
    },
    {
      "name": "apiserver_watch_list_duration_seconds_bucket",
      "value": 5760
    }
  ]
}
$ curl -s localhost:9090/api/v1/status/tsdb | jq '.data.labelValueCountByLabelName[:5]'
Handling connection for 9090
[
  {
    "name": "__name__",
    "value": 2476
  },
  {
    "name": "name",
    "value": 1391
  },
  {
    "name": "id",
    "value": 672
  },
  {
    "name": "resource",
    "value": 618
  },
  {
    "name": "le",
    "value": 415
  }
]
$ kill $PF1
verify: you can name the metric with the most series and the label with the most distinct values on this cluster.

Two functions carry more exam weight than their length suggests. One extrapolates a trend into the future, the other alerts on a series that stopped existing, which is the failure a threshold can never catch.

kubectl -n monitoring port-forward svc/prometheus-kube-prometheus-prometheus 9090:9090 & PF1=$!
sleep 5
curl -sG localhost:9090/api/v1/query --data-urlencode 'query=count by (mountpoint,fstype) (node_filesystem_avail_bytes)' | jq -r '.data.result[] | "\(.metric.mountpoint) \(.metric.fstype)"'
curl -sG localhost:9090/api/v1/query --data-urlencode 'query=predict_linear(node_filesystem_avail_bytes[1h], 24*3600)' | jq '[.data.result[] | {instance: .metric.instance, mountpoint: .metric.mountpoint, predicted: .value[1]}] | .[0:3]'
curl -sG localhost:9090/api/v1/query --data-urlencode 'query=absent_over_time(up{job="does-not-exist"}[5m])' | jq '.data.result'
curl -sG localhost:9090/api/v1/query --data-urlencode 'query=absent_over_time(up{job="kubelet"}[5m])' | jq '.data.result'
kill $PF1
outputcaptured 2026-09-13
$ kubectl -n monitoring port-forward svc/prometheus-kube-prometheus-prometheus 9090:9090 & PF1=$!
$ sleep 5
Forwarding from 127.0.0.1:9090 -> 9090
Forwarding from [::1]:9090 -> 9090
$ curl -sG localhost:9090/api/v1/query --data-urlencode 'query=count by (mountpoint,fstype) (node_filesystem_avail_bytes)' | jq -r '.data.result[] | "\(.metric.mountpoint) \(.metric.fstype)"'
Handling connection for 9090
/etc/hostname btrfs
/etc/hosts btrfs
/etc/kubernetes/audit btrfs
/etc/resolv.conf btrfs
/usr/lib/modules btrfs
/var btrfs
/run tmpfs
/run/credentials/systemd-journald.service tmpfs
/tmp tmpfs
$ curl -sG localhost:9090/api/v1/query --data-urlencode 'query=predict_linear(node_filesystem_avail_bytes[1h], 24*3600)' | jq '[.data.result[] | {instance: .metric.instance, mountpoint: .metric.mountpoint, predicted: .value[1]}] | .[0:3]'
Handling connection for 9090
[
  {
    "instance": "172.18.0.4:9100",
    "mountpoint": "/etc/hostname",
    "predicted": "907166223218.7759"
  },
  {
    "instance": "172.18.0.4:9100",
    "mountpoint": "/etc/hosts",
    "predicted": "907166223218.7759"
  },
  {
    "instance": "172.18.0.4:9100",
    "mountpoint": "/etc/kubernetes/audit",
    "predicted": "907166223218.7759"
  }
]
$ curl -sG localhost:9090/api/v1/query --data-urlencode 'query=absent_over_time(up{job="does-not-exist"}[5m])' | jq '.data.result'
Handling connection for 9090
[
  {
    "metric": {
      "job": "does-not-exist"
    },
    "value": [
      1789306067.750,
      "1"
    ]
  }
]
$ curl -sG localhost:9090/api/v1/query --data-urlencode 'query=absent_over_time(up{job="kubelet"}[5m])' | jq '.data.result'
Handling connection for 9090
[]
$ kill $PF1
verify: the first query lists the mountpoints node-exporter actually reports on these nodes. / is not among them, because a kind node's root is an overlay mount and node-exporter excludes overlay by filesystem type, so the usual mountpoint="/" selector matches nothing and predict_linear over it returns nothing rather than an error. The second query returns one predicted number per filesystem that does exist, and absent_over_time returns a series for the job that does not exist and nothing for the one that does. That inversion is the whole idea: the alert fires when the query has a result.

Retention is two numbers and whichever binds first wins. On a cluster you inherit, read them before you promise anyone a month of history.

kubectl -n monitoring get prometheus -o jsonpath='{.items[0].spec.retention} {.items[0].spec.retentionSize}{"\n"}'
kubectl -n monitoring get prometheus -o jsonpath='{.items[0].spec.storage}' | jq
kubectl -n monitoring get pvc -l app.kubernetes.io/name=prometheus -o custom-columns=NAME:.metadata.name,CAPACITY:.status.capacity.storage,CLASS:.spec.storageClassName
kubectl -n monitoring port-forward svc/prometheus-kube-prometheus-prometheus 9090:9090 & PF1=$!
sleep 5
curl -s localhost:9090/api/v1/status/runtimeinfo | jq '.data | {storageRetention, corruptionCount, timeSeriesCount}'
kill $PF1
outputcaptured 2026-09-12
$ kubectl -n monitoring get prometheus -o jsonpath='{.items[0].spec.retention} {.items[0].spec.retentionSize}{"\n"}'
6h 
$ kubectl -n monitoring get prometheus -o jsonpath='{.items[0].spec.storage}' | jq
$ kubectl -n monitoring get pvc -l app.kubernetes.io/name=prometheus -o custom-columns=NAME:.metadata.name,CAPACITY:.status.capacity.storage,CLASS:.spec.storageClassName
NAME   CAPACITY   CLASS
$ kubectl -n monitoring port-forward svc/prometheus-kube-prometheus-prometheus 9090:9090 & PF1=$!
$ sleep 5
Forwarding from 127.0.0.1:9090 -> 9090
Forwarding from [::1]:9090 -> 9090
$ curl -s localhost:9090/api/v1/status/runtimeinfo | jq '.data | {storageRetention, corruptionCount, timeSeriesCount}'
Handling connection for 9090
{
  "storageRetention": "6h",
  "corruptionCount": 0,
  "timeSeriesCount": null
}
$ kill $PF1
verify: you can say how far back a query on this cluster can reach and what would stop it first, time or disk. If the storage block is empty, the answer is "until the pod restarts", and that is worth saying out loud.

Flux and Tekton both export metrics that answer delivery questions rather than infrastructure ones. Wiring them is two objects, and the queries afterwards are the raw material for section 4.5.

kubectl apply -f - <<'EOF'
apiVersion: monitoring.coreos.com/v1
kind: PodMonitor
metadata: { name: flux-controllers, namespace: monitoring, labels: { release: prometheus } }
spec:
  namespaceSelector: { matchNames: [flux-system] }
  selector:
    matchExpressions:
      - { key: app, operator: Exists }
  podMetricsEndpoints:
    - port: http-prom
---
apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata: { name: tekton-controller, namespace: monitoring, labels: { release: prometheus } }
spec:
  namespaceSelector: { matchNames: [tekton-pipelines] }
  selector:
    matchLabels: { app: tekton-pipelines-controller }
  endpoints:
    - port: http-metrics
EOF
sleep 90
kubectl -n monitoring port-forward svc/prometheus-kube-prometheus-prometheus 9090:9090 & PF1=$!
sleep 5
curl -sG localhost:9090/api/v1/query --data-urlencode 'query=gotk_reconcile_duration_seconds_count' | jq '.data.result | length'
curl -sG localhost:9090/api/v1/query --data-urlencode 'query=tekton_pipelines_controller_pipelinerun_total' | jq '.data.result[] | {status: .metric.status, value: .value[1]}'
kill $PF1
outputcaptured 2026-09-12
$ kubectl apply -f - <<'EOF'
apiVersion: monitoring.coreos.com/v1
kind: PodMonitor
metadata: { name: flux-controllers, namespace: monitoring, labels: { release: prometheus } }
spec:
  namespaceSelector: { matchNames: [flux-system] }
  selector:
    matchExpressions:
      - { key: app, operator: Exists }
  podMetricsEndpoints:
    - port: http-prom
---
apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata: { name: tekton-controller, namespace: monitoring, labels: { release: prometheus } }
spec:
  namespaceSelector: { matchNames: [tekton-pipelines] }
  selector:
    matchLabels: { app: tekton-pipelines-controller }
  endpoints:
    - port: http-metrics
EOF
podmonitor.monitoring.coreos.com/flux-controllers created
servicemonitor.monitoring.coreos.com/tekton-controller created
$ sleep 90
$ kubectl -n monitoring port-forward svc/prometheus-kube-prometheus-prometheus 9090:9090 & PF1=$!
$ sleep 5
Forwarding from 127.0.0.1:9090 -> 9090
Forwarding from [::1]:9090 -> 9090
$ curl -sG localhost:9090/api/v1/query --data-urlencode 'query=gotk_reconcile_duration_seconds_count' | jq '.data.result | length'
Handling connection for 9090
8
$ curl -sG localhost:9090/api/v1/query --data-urlencode 'query=tekton_pipelines_controller_pipelinerun_total' | jq '.data.result[] | {status: .metric.status, value: .value[1]}'
Handling connection for 9090
{
  "status": "success",
  "value": "1"
}
$ kill $PF1
verify: both queries return series. If either returns nothing, the port name in the monitor does not match the container's; read the pod spec rather than the chart's documentation, because the chart may have renamed it.

Self-check

answer before opening
Your ServiceMonitor is correct and no target appears. List the checks in order.

1) Prometheus's serviceMonitorSelector and serviceMonitorNamespaceSelector: does it even consider your monitor? 2) Does the monitor's selector match the Service's labels? 3) Is endpoints[].port the port name, and does the Service name it? 4) Does the Service have ready endpoints? 5) RBAC and NetworkPolicy between monitoring and the target namespace.

Why is sum(rate(x[5m])) right and rate(sum(x)[5m]) wrong?

Counters reset per series; rate knows how to handle resets within a single series. Summing first merges series (and their resets) into a meaningless line, and the syntax is invalid anyway without a subquery. Rate first, aggregate second, always.

How do you alert on "the metric disappeared entirely"?

absent(up{job="x"} == 1) or absent(metric): an ordinary comparison returns no series when there is nothing to compare, so it can never fire. absent() exists precisely to turn "no data" into a value of 1.

A p95 query returns NaN. Two plausible reasons?

No observations in the window (nothing has been recorded yet, so all buckets are empty), or you aggregated away the le label that histogram_quantile needs. The form that keeps it: histogram_quantile(0.95, sum by (le) (rate(x_bucket[5m]))).

An exporter added a label with one value per request. What happens, and what is the fix?

Cardinality explosion: a new series per unique value, memory and TSDB growth, slow queries. Fix at ingestion with a metricRelabeling that drops the label or the metric, and upstream by removing it from the exporter. High-cardinality labels (user IDs, URLs with IDs, trace IDs) belong in logs and traces, not metric labels.

You need to scrape a database VM outside the cluster with the Prometheus Operator, without editing the Prometheus config by hand. Which CRD, and which selector must match?

A ScrapeConfig with staticConfigs[].targets: ["db.example.internal:9187"] (or an httpSDConfigs/dnsSDConfigs entry), labeled to match the Prometheus object's scrapeConfigSelector, in a namespace its scrapeConfigNamespaceSelector covers. If that selector is nil, the ScrapeConfig must live in Prometheus's own namespace.

Predict that a volume fills within a day, and alert if any series that should exist has vanished. Two expressions?

predict_linear(node_filesystem_avail_bytes{mountpoint="/data"}[6h], 24*3600) < 0 for the volume (a gauge extrapolated linearly). absent_over_time(up{job="payments"}[10m]) for the disappearance: it yields 1 only when no sample existed in the window, whereas a plain comparison returns nothing when there is nothing to compare.

A ServiceMonitor in monitoring selects the right labels and port for Services in argocd, yet no target appears and Prometheus's own selectors are {}. What is missing?

spec.namespaceSelector on the ServiceMonitor itself: matchNames: [argocd] or any: true. Omitted, it means "Services in my own namespace", so the monitor matches nothing in argocd. Prometheus's serviceMonitorNamespaceSelector decides which monitors are read; the monitor's own namespaceSelector decides which Services it may reach.

Docs to know your way around

study time, not exam time
  • prometheus.io: querying basics and the function reference (rate, increase, histogram_quantile, absent); metric types.
  • prometheus-operator.dev: the ServiceMonitor troubleshooting page, which is the release-label story in official form; the API reference for ServiceMonitor/PodMonitor/ScrapeConfig.
  • Offline: kubectl explain servicemonitor.spec.endpoints, the Prometheus UI's Status → Targets and Status → Configuration pages, and its built-in expression autocompletion.
  • prometheus-operator.dev/docs/developer/scrapeconfig: the ScrapeConfig page with the selector and namespace rules; the API reference's ServiceMonitor section lists jobLabel, targetLabels, the limits and serviceDiscoveryRole.
  • prometheus.io/docs/prometheus/latest/querying/functions: predict_linear, absent_over_time, label_replace, and the native-histogram functions; /storage for the retention flags.
  • argo-cd.readthedocs.io/en/stable/operator-manual/metrics, fluxcd.io/flux/monitoring/metrics, tekton.dev/docs/pipelines/metrics, docs.crossplane.io/latest/guides/metrics: the metric-name tables, one page each.