This competency sounds like management material, but it is testable: "write a query showing deployment frequency" is a concrete task. The trick is knowing which existing metrics stand in for which indicator.

needsmake core obs

Orientation

competency 4.2 · measuring and improving platform efficiency

"Measure platform efficiency" in practice means plumbing the delivery tool's metrics into the monitoring stack and querying them. The wiring is the exercise; the PromQL is five lines.

Two families of measure, do not mix them

Delivery performance (the DORA four) says how well software reaches production. Platform effectiveness (fulfillment latency, adoption, time to first contribution, satisfaction) says how well the platform serves its users. A platform can ship fast and still be miserable to use, and vice versa, which is why the white paper lists both.

The indicators worth naming

and where the raw data lives
IndicatorDefinitionData lives in
Deployment frequencyhow often you release to productionthe CD tool (sync counters), or git tags
Lead time for changescommit → running in productiongit + CD: commit timestamp vs sync timestamp
Change failure rateshare of releases causing degradationrollout outcomes, failed syncs, incident links
Time to restore serviceincident start → recoveredalerting/incident data (firing → resolved)
Fulfillment latencyrequest → capability deliveredyour own CR timestamps (creation → Ready)
Adoption / time to first contributionwho uses the platform, how fast a newcomer shipsoutside the cluster: catalog data and surveys

In an Argo CD shop, the controller's own metrics are the delivery data

  • argocd_app_sync_total: counter per app with a phase label (Succeeded, Failed, Error). Deployment frequency and change failure rate come straight out of this.
  • argocd_app_info: one series per app carrying current sync_status and health_status as labels. Anything about "how many apps are unhealthy right now" is a count over this.
  • argocd_app_reconcile (histogram): controller loop duration; a platform-health signal rather than a delivery one.

Argo CD exposes them, but this Prometheus does not scrape them until you add a ServiceMonitor, which is the first exercise below.

The other half: does the platform run what it shipped?

Workload restart rates (kube_pod_container_status_restarts_total), pods not at desired replicas (kube_deployment_status_replicas_unavailable), pending pods, and the request/usage gap from section 1.5. Deployment metrics tell you the platform ships fast; these tell you it runs what it shipped. A good answer to "how would you measure this platform" names both halves.

A caution about counting syncs

Deployment frequency counted from syncs is an approximation: a self-healing sync of an unchanged app is not a deployment, and one push that updates ten apps is not ten deployments. Say the approximation you are making and what would make it exact (counting distinct revisions, or emitting deployment events from the pipeline).

DORA as currently defined, and SPACE

the 2024 vocabulary, the retired tiers, what the white paper actually lists

The four keys have been renamed and regrouped since most blog posts were written. Use the current wording; a scenario that says "failed deployment recovery time" or "rework rate" is using it.

GroupMetric (2024 name)Older nameDefinition
ThroughputChange lead timeLead time for changescommit to running in production
ThroughputDeployment frequencysamehow often you deploy to production
ThroughputFailed deployment recovery timeTime to restore service / MTTRtime to recover from a deployment that caused a failure (not from every incident)
StabilityChange failure ratesameshare of deployments that cause a failure needing remediation
StabilityDeployment rework rate (added 2024)noneshare of deployments that are unplanned, to fix a user-facing bug

Reliability is measured alongside these as operational performance; DORA's own history page says it was never a "fifth key". The 2024 report's performance clusters were: elite deploys on demand, lead time under a day, recovery under an hour, change failure rate around 5%; high deploys between daily and weekly, lead time a day to a week, recovery under a day, failure rate around 20%; medium weekly to monthly, a week to a month, under a day, around 10%; low monthly to twice a year, one to six months, a week to a month, around 40%. The 2025 report retired that four-tier clustering in favor of team archetypes that combine delivery performance with burnout and friction, so quote tiers as "the 2024 clusters" rather than as a standard.

Computing the keys from what the cluster already knows

  • Deployment frequency. increase(argocd_app_sync_total{phase="Succeeded"}[7d]) per app, with the caveat from the panel above; for Flux, count changes of the revision label on gotk_resource_info{customresource_kind="Kustomization"}; for a pipeline-driven deploy, increase(tekton_pipelines_controller_pipelinerun_total{status="success"}[7d]) on the deploy pipeline.
  • Change lead time. Needs two timestamps: the commit (from git, git log --format=%cI <sha>) and the deploy event (Argo CD's argocd app history row for that revision, or the Flux Kustomization's status.lastAppliedRevision and its condition lastTransitionTime). No single metric holds both; the honest answer is a small job that joins them and emits a gauge or a pushed event.
  • Change failure rate. failed syncs (phase=~"Failed|Error") plus rollbacks (an argocd app rollback is a sync to a previous revision; a Rollout that aborted has rollout_phase{phase="Degraded"} in Argo Rollouts' metrics; Flagger increments flagger_canary_status failures) divided by deployments.
  • Failed deployment recovery time. from the deployment-caused alert firing (ALERTS{alertstate="firing"} for the service) to resolved, or from the failed sync to the next successful one; Alertmanager's notification log or your incident tool holds the human-visible version (MTTA is acknowledge, MTTR is restore, MTTD is detect).
  • Rework rate. deployments labeled as hotfixes; you need a convention (a commit trailer, a PR label, a reason=hotfix annotation on the Application) before you can count it.

SPACE, because the white paper cites it

SPACE is the developer-productivity framework the white paper points at for "user satisfaction and productivity": Satisfaction and well-being, Performance (outcomes, not output), Activity (counts of things done), Communication and collaboration, Efficiency and flow (interruptions, wait time). Its rule is to pick metrics from at least three dimensions and to include at least one perceptual measure (a survey). The lines-of-code question from the self-check is an Activity-only measure, which is exactly what SPACE forbids.

Measuring the platform itself

the white paper's list, maturity-model levels, SLIs for a platform API

The CNCF Platforms white paper names three metric groups. Learn the exact list; a question that asks "which measures would you add for the platform" wants these words.

GroupMeasures namedWhere the data is
User satisfaction and productivityactive users and retention (capabilities provisioned, growth/churn); NPS or other satisfaction surveys; SPACE-style productivity metricsthe portal's usage data, the catalog, surveys
Organizational efficiencylatency from request to fulfillment of a capability; latency to build and deploy a brand-new service into production; time for a new user to submit their first code changeCR timestamps (creation to Ready), template runs to first successful sync, onboarding records
Product and feature deliverythe DORA metrics: deployment frequency, lead time, time to restore, change failure rateCD and alerting systems, as above

The Platform Engineering Maturity Model's Measurement aspect gives the progression: Ad hoc (someone counts something) → Consistent collection (the same measures every period) → Insights (measures drive decisions) → Quantitative and qualitative (surveys and usage together). "We have a dashboard" is level 2; "we cut the request-to-Ready p95 from 20 minutes to 2 after seeing it" is level 3.

Platform KPIs you can compute today

  • Adoption. count of XRs or RGD instances per team (count by (customresource_kind) (kube_customresource_info) once kube-state-metrics has a CustomResourceStateMetrics config for your kinds), and the share of namespaces created through the API rather than by hand (a label the composition stamps, missing on hand-made ones).
  • Golden path usage. Applications generated by the ApplicationSet (they carry the generator's labels) versus Applications created directly; Backstage template runs per week from its scaffolder task list.
  • Request-to-fulfillment latency. the lab's creation-to-Ready delta, exported as a gauge per instance through CustomResourceStateMetrics (a Gauge with path: [status, conditions], labelsFromPath for type, valueFrom: [lastTransitionTime] gives you timestamps to subtract) or the provider's own crossplane_managed_resource_first_time_to_readiness_seconds histogram for the Crossplane half.
  • Ticket volume. the count that should fall: requests to the platform team that a self-service path could have served. Outside the cluster, but it is the number executives ask for.
  • Platform reliability. SLIs for the platform's own control plane: API server availability (apiserver_request_total 5xx ratio), Argo CD sync success ratio, webhook admission latency (apiserver_admission_webhook_admission_duration_seconds), Crossplane and kro reconcile error rates (controller_runtime_reconcile_total{result="error"}). Each gets an SLO and an error budget; for 99.9% monthly the budget is roughly 43 minutes, and the burn-rate rules from 4.2 apply unchanged.
How this gets tested

Two kinds of task. The concrete one: "write a query for deployment frequency / failed syncs / unhealthy apps" against metrics that exist once you scrape the tool. The judgment one: "which indicators would show whether the platform improved developer experience" wants the white paper's list (fulfillment latency, time to first change, adoption, satisfaction) alongside DORA, plus one sentence on why output counts such as lines of code do not qualify.

Exercises

tick the dot when its check passes

Discover what Argo CD exposes, then monitor it:

kubectl -n argocd get svc | grep metrics
outputcaptured 2026-08-26
$ kubectl -n argocd get svc | grep metrics
argocd-application-controller-metrics   ClusterIP      10.96.8.45      <none>        8082/TCP                     75m
argocd-repo-server-metrics              ClusterIP      10.96.30.75     <none>        8084/TCP                     75m
argocd-server-metrics                   ClusterIP      10.96.170.16    <none>        8083/TCP                     75m

Three metrics services (application controller, server, repo server). The lab's gitops layer enables them. If the grep comes back empty on an older install, the fix is helm -n argocd upgrade argocd argo/argo-cd --reuse-values --set controller.metrics.enabled=true --set server.metrics.enabled=true --set repoServer.metrics.enabled=true, itself a fair exam task. Write one ServiceMonitor per service you care about; the controller's is the one with sync metrics:

apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata:
  name: argocd-controller
  namespace: monitoring
  labels: { release: prometheus }
spec:
  namespaceSelector: { matchNames: [argocd] }
  selector: { matchLabels: { app.kubernetes.io/name: argocd-metrics } }
  endpoints: [{ port: http-metrics }]

That endpoint port is http-metrics because that is what the Service names it, not because anyone would guess it. kubectl -n argocd get svc argocd-application-controller-metrics -o jsonpath='{.spec.ports[*].name}' is the thirty-second check that saves an hour of "why is there no target".

verify: argocd_app_info returns rows in the Prometheus UI within a couple of minutes. Everything from 4.1 about selectors and port names applies; this exercise is deliberately a rerun of those skills against an unfamiliar target.

With sync activity from your 2.2/2.6 work in the counters (trigger a couple of argocd app sync demo-staging runs if the range is empty):

sum(increase(argocd_app_sync_total{phase="Succeeded"}[24h]))

sum(increase(argocd_app_sync_total{phase=~"Failed|Error"}[24h]))
  / sum(increase(argocd_app_sync_total[24h]))

count(argocd_app_info{health_status!="Healthy"}) or vector(0)
outputcaptured 2026-08-26
$ kubectl -n monitoring port-forward svc/prometheus-operated 9090 &   # or open the UI from make urls
$ curl -sG http://localhost:9090/api/v1/query --data-urlencode \
  'query=sum(increase(argocd_app_sync_total{phase="Succeeded"}[24h]))' | jq -r '.data.result[0].value[1]'
3.0247315789473683
$ curl -sG http://localhost:9090/api/v1/query --data-urlencode \
  'query=sum(increase(argocd_app_sync_total{phase=~"Failed|Error"}[24h])) / sum(increase(argocd_app_sync_total[24h]))' \
  | jq -r '.data.result[0].value[1] // "no failed syncs in the window"'
no failed syncs in the window
$ curl -sG http://localhost:9090/api/v1/query --data-urlencode \
  'query=count(argocd_app_info{health_status!="Healthy"}) or vector(0)' | jq -r '.data.result[0].value[1]'
0

Deployment frequency, change failure rate, and currently-degraded apps (a time-to-restore ingredient). Then push one bad image tag (the 2.6 drill), sync, revert, and watch the failure rate query move.

verify: numbers that match reality; cross-check the first against argocd app history demo-staging.

Measure the white paper's "request to fulfillment" for the lab's own self-service path: apply a fresh AppEnvironment XR (section 3.5) and time from apply to Ready condition using its lastTransitionTime:

kubectl get appenvironment <name> -o jsonpath='{.status.conditions[?(@.type=="Ready")].lastTransitionTime}'
outputcaptured 2026-08-26
$ kubectl get appenvironment team-c-dev -o jsonpath='{.status.conditions[?(@.type=="Ready")].lastTransitionTime}{"\n"}'
2026-08-27T01:53:06Z
$ kubectl get appenvironment team-c-dev -o jsonpath='{.metadata.creationTimestamp}{"\n"}'
2026-08-27T01:52:58Z

against the object's metadata.creationTimestamp.

verify: a number in seconds, and an opinion about whether it is good. (Sub-minute for namespace-and-quota is; if it took five, something reconciled slowly and section 3.5's trace command finds what.)

The bridge to 4.2: a PrometheusRule that fires when any app stays non-Healthy for 10 minutes (count(argocd_app_info{health_status!="Healthy"}) > 0, for: 10m). Verify with the bad-image trick, then clean up.

verify: the alert reaches Alertmanager. This is what a platform team's pager rule looks like.

kube-state-metrics can expose your CRDs the same way it exposes Deployments, from a config file. That turns "is every tenant environment ready" into a query instead of a script.

cat > /tmp/crs.yaml <<'EOF'
spec:
  resources:
    - groupVersionKind:
        group: platform.lab.local
        version: v1alpha1
        kind: AppEnvironment
      metricNamePrefix: kube_customresource_appenvironment
      labelsFromPath:
        name: [metadata, name]
        namespace: [metadata, namespace]
      metrics:
        - name: condition
          help: "AppEnvironment status conditions"
          each:
            type: StateSet
            stateSet:
              labelName: status
              # without the condition type every condition collapses onto one label set
              labelsFromPath:
                type: [type]
              path: [status, conditions]
              valueFrom: [status]
              list: ["True", "False", "Unknown"]
EOF
kubectl -n monitoring create configmap ksm-customresource --from-file=config.yaml=/tmp/crs.yaml --dry-run=client -o yaml | kubectl -n monitoring apply -f -
# kube-state-metrics needs to read the CRD as well as the custom resource itself
kubectl create clusterrole ksm-appenv --verb=get,list,watch --resource=appenvironments.platform.lab.local,customresourcedefinitions.apiextensions.k8s.io --dry-run=client -o yaml | kubectl apply -f -
kubectl create clusterrolebinding ksm-appenv --clusterrole=ksm-appenv --serviceaccount=monitoring:prometheus-kube-state-metrics --dry-run=client -o yaml | kubectl apply -f -
# the deployment ships with no volumes at all, so these paths are created, not appended to
kubectl -n monitoring patch deploy prometheus-kube-state-metrics --type json -p '[{"op":"add","path":"/spec/template/spec/volumes","value":[{"name":"crs","configMap":{"name":"ksm-customresource"}}]},{"op":"add","path":"/spec/template/spec/containers/0/volumeMounts","value":[{"name":"crs","mountPath":"/etc/ksm"}]},{"op":"add","path":"/spec/template/spec/containers/0/args/-","value":"--custom-resource-state-config-file=/etc/ksm/config.yaml"}]'
kubectl -n monitoring rollout status deploy prometheus-kube-state-metrics --timeout=180s
sleep 60
kubectl -n monitoring port-forward svc/prometheus-kube-prometheus-prometheus 9090:9090 & PF1=$!
sleep 5
curl -sG localhost:9090/api/v1/query --data-urlencode 'query=kube_customresource_appenvironment_condition' | jq -r '.data.result[] | "\(.metric.name) \(.metric.type)=\(.metric.status) \(.value[1])"'
kill $PF1
kubectl -n monitoring patch deploy prometheus-kube-state-metrics --type json -p '[{"op":"remove","path":"/spec/template/spec/volumes"},{"op":"remove","path":"/spec/template/spec/containers/0/volumeMounts"},{"op":"remove","path":"/spec/template/spec/containers/0/args/2"}]'
kubectl delete clusterrolebinding ksm-appenv
kubectl delete clusterrole ksm-appenv
kubectl -n monitoring delete cm ksm-customresource
outputcaptured 2026-09-12
$ cat > /tmp/crs.yaml <<'EOF'
spec:
  resources:
    - groupVersionKind:
        group: platform.lab.local
        version: v1alpha1
        kind: AppEnvironment
      metricNamePrefix: kube_customresource_appenvironment
      labelsFromPath:
        name: [metadata, name]
        namespace: [metadata, namespace]
      metrics:
        - name: condition
          help: "AppEnvironment status conditions"
          each:
            type: StateSet
            stateSet:
              labelName: status
              # without the condition type every condition collapses onto one label set
              labelsFromPath:
                type: [type]
              path: [status, conditions]
              valueFrom: [status]
              list: ["True", "False", "Unknown"]
EOF
$ kubectl -n monitoring create configmap ksm-customresource --from-file=config.yaml=/tmp/crs.yaml --dry-run=client -o yaml | kubectl -n monitoring apply -f -
configmap/ksm-customresource created
$ # kube-state-metrics needs to read the CRD as well as the custom resource itself
$ kubectl create clusterrole ksm-appenv --verb=get,list,watch --resource=appenvironments.platform.lab.local,customresourcedefinitions.apiextensions.k8s.io --dry-run=client -o yaml | kubectl apply -f -
clusterrole.rbac.authorization.k8s.io/ksm-appenv created
$ kubectl create clusterrolebinding ksm-appenv --clusterrole=ksm-appenv --serviceaccount=monitoring:prometheus-kube-state-metrics --dry-run=client -o yaml | kubectl apply -f -
clusterrolebinding.rbac.authorization.k8s.io/ksm-appenv created
$ # the deployment ships with no volumes at all, so these paths are created, not appended to
$ kubectl -n monitoring patch deploy prometheus-kube-state-metrics --type json -p '[{"op":"add","path":"/spec/template/spec/volumes","value":[{"name":"crs","configMap":{"name":"ksm-customresource"}}]},{"op":"add","path":"/spec/template/spec/containers/0/volumeMounts","value":[{"name":"crs","mountPath":"/etc/ksm"}]},{"op":"add","path":"/spec/template/spec/containers/0/args/-","value":"--custom-resource-state-config-file=/etc/ksm/config.yaml"}]'
deployment.apps/prometheus-kube-state-metrics patched
$ kubectl -n monitoring rollout status deploy prometheus-kube-state-metrics --timeout=180s
Waiting for deployment spec update to be observed...
Waiting for deployment spec update to be observed...
Waiting for deployment "prometheus-kube-state-metrics" rollout to finish: 0 out of 1 new replicas have been updated...
Waiting for deployment "prometheus-kube-state-metrics" rollout to finish: 1 old replicas are pending termination...
Waiting for deployment "prometheus-kube-state-metrics" rollout to finish: 1 old replicas are pending termination...
deployment "prometheus-kube-state-metrics" successfully rolled out
$ sleep 60
$ kubectl -n monitoring port-forward svc/prometheus-kube-prometheus-prometheus 9090:9090 & PF1=$!
$ sleep 5
Forwarding from 127.0.0.1:9090 -> 9090
Forwarding from [::1]:9090 -> 9090
$ curl -sG localhost:9090/api/v1/query --data-urlencode 'query=kube_customresource_appenvironment_condition' | jq -r '.data.result[] | "\(.metric.name) \(.metric.type)=\(.metric.status) \(.value[1])"'
Handling connection for 9090
team-c-dev Ready=False 0
team-c-dev Responsive=False 0
team-c-dev Synced=False 0
team-c-dev Ready=True 1
team-c-dev Responsive=True 1
team-c-dev Synced=True 1
team-c-dev Ready=Unknown 0
team-c-dev Responsive=Unknown 0
team-c-dev Synced=Unknown 0
$ kill $PF1
$ kubectl -n monitoring patch deploy prometheus-kube-state-metrics --type json -p '[{"op":"remove","path":"/spec/template/spec/volumes"},{"op":"remove","path":"/spec/template/spec/containers/0/volumeMounts"},{"op":"remove","path":"/spec/template/spec/containers/0/args/2"}]'
deployment.apps/prometheus-kube-state-metrics patched
$ kubectl delete clusterrolebinding ksm-appenv
clusterrolebinding.rbac.authorization.k8s.io "ksm-appenv" deleted
$ kubectl delete clusterrole ksm-appenv
clusterrole.rbac.authorization.k8s.io "ksm-appenv" deleted
$ kubectl -n monitoring delete cm ksm-customresource
configmap "ksm-customresource" deleted from monitoring namespace
verify: kube_customresource_appenvironment_condition returns one series per condition per state, so the AppEnvironment team-c-dev reads Ready=True 1, Synced=True 1 and Responsive=True 1 with the False and Unknown states at 0. Drop the type label from the config and every condition lands on the same label set, which Prometheus rejects as a duplicate sample.

Lead time for change is a subtraction: when was the commit authored, when did it reach the cluster. Do it once by hand for a real revision so the number means something before you automate it.

# the LoadBalancer IP only exists when cloud-provider-kind runs; the NodePort is always there
ARGO=$(kubectl get node cnpe-control-plane -o jsonpath='{.status.addresses[?(@.type=="InternalIP")].address}'):$(kubectl -n argocd get svc argocd-server -o jsonpath='{.spec.ports[?(@.port==80)].nodePort}')
argocd login "$ARGO" --username admin --password "$(kubectl -n argocd get secret argocd-initial-admin-secret -o jsonpath='{.data.password}' | base64 -d)" --plaintext --grpc-web
SHA=$(kubectl -n argocd get app demo-staging -o jsonpath='{.status.sync.revision}')
echo "revision: $SHA"
git clone "http://lab:${GITEA_PASS}@gitea.lab:3000/lab/platform.git" /tmp/platform-leadtime 2>/dev/null || true
git -C /tmp/platform-leadtime fetch --all
git -C /tmp/platform-leadtime log -1 --format=%cI "$SHA"
argocd app history demo-staging
kubectl -n argocd get app demo-staging -o jsonpath='{range .status.history[*]}{.revision} {.deployedAt}{"\n"}{end}'
outputcaptured 2026-09-12
$ # the LoadBalancer IP only exists when cloud-provider-kind runs; the NodePort is always there
$ ARGO=$(kubectl get node cnpe-control-plane -o jsonpath='{.status.addresses[?(@.type=="InternalIP")].address}'):$(kubectl -n argocd get svc argocd-server -o jsonpath='{.spec.ports[?(@.port==80)].nodePort}')
$ argocd login "$ARGO" --username admin --password "$(kubectl -n argocd get secret argocd-initial-admin-secret -o jsonpath='{.data.password}' | base64 -d)" --plaintext --grpc-web
'admin:login' logged in successfully
Context '172.18.0.4:32015' updated
$ SHA=$(kubectl -n argocd get app demo-staging -o jsonpath='{.status.sync.revision}')
$ echo "revision: $SHA"
revision: 83a41e21322e15eff2307bc1e9d5c89e33226d0d
$ git clone "http://lab:${GITEA_PASS}@gitea.lab:3000/lab/platform.git" /tmp/platform-leadtime 2>/dev/null || true
$ git -C /tmp/platform-leadtime fetch --all
From http://gitea.lab:3000/lab/platform
   75be58c..83a41e2  main       -> origin/main
$ git -C /tmp/platform-leadtime log -1 --format=%cI "$SHA"
2026-09-13T00:52:25-04:00
$ argocd app history demo-staging
SOURCE  http://gitea.lab:3000/lab/platform.git
ID      DATE                           REVISION
1       2026-09-13 00:52:26 -0400 EDT  main (83a41e2)
2       2026-09-13 00:54:00 -0400 EDT  main (83a41e2)
3       2026-09-13 00:54:57 -0400 EDT  main (83a41e2)
4       2026-09-13 01:00:12 -0400 EDT  main (83a41e2)
5       2026-09-13 01:00:17 -0400 EDT  main (83a41e2)
6       2026-09-13 01:00:22 -0400 EDT  main (83a41e2)
7       2026-09-13 01:00:57 -0400 EDT  main (83a41e2)
8       2026-09-13 01:01:06 -0400 EDT  main (83a41e2)
9       2026-09-13 01:01:14 -0400 EDT  main (83a41e2)
10      2026-09-13 01:01:44 -0400 EDT  main (83a41e2)
$ kubectl -n argocd get app demo-staging -o jsonpath='{range .status.history[*]}{.revision} {.deployedAt}{"\n"}{end}'
83a41e21322e15eff2307bc1e9d5c89e33226d0d 2026-09-13T04:52:26Z
83a41e21322e15eff2307bc1e9d5c89e33226d0d 2026-09-13T04:54:00Z
83a41e21322e15eff2307bc1e9d5c89e33226d0d 2026-09-13T04:54:57Z
83a41e21322e15eff2307bc1e9d5c89e33226d0d 2026-09-13T05:00:12Z
83a41e21322e15eff2307bc1e9d5c89e33226d0d 2026-09-13T05:00:17Z
83a41e21322e15eff2307bc1e9d5c89e33226d0d 2026-09-13T05:00:22Z
83a41e21322e15eff2307bc1e9d5c89e33226d0d 2026-09-13T05:00:57Z
83a41e21322e15eff2307bc1e9d5c89e33226d0d 2026-09-13T05:01:06Z
83a41e21322e15eff2307bc1e9d5c89e33226d0d 2026-09-13T05:01:14Z
83a41e21322e15eff2307bc1e9d5c89e33226d0d 2026-09-13T05:01:44Z
verify: you can subtract the commit timestamp from the deployed-at timestamp for the same sha and say the number out loud. Then say which part of that interval your platform is responsible for and which part is the developer waiting for a review.

A suspended Kustomization is a promise the platform has stopped keeping, and nothing alerts on it by default. The gauge Flux exports makes it one query, and that query belongs on the platform dashboard.

flux suspend kustomization demo-flux
sleep 60
kubectl -n monitoring port-forward svc/prometheus-kube-prometheus-prometheus 9090:9090 & PF1=$!
sleep 5
# gotk_resource_info comes from kube-state-metrics with a CustomResourceState config this lab does not ship
curl -sG localhost:9090/api/v1/query --data-urlencode 'query=group by (__name__) ({__name__=~"gotk_.+"})' | jq -r '.data.result[].metric.__name__'
curl -sG localhost:9090/api/v1/query --data-urlencode 'query=count(gotk_suspend_status == 1)' | jq -r '.data.result[0].value[1] // "no gotk_suspend_status series"'
curl -sG localhost:9090/api/v1/query --data-urlencode 'query=gotk_suspend_status == 1' | jq '.data.result[] | .metric'
# gotk_resource_info comes from kube-state-metrics with a CustomResourceState config this lab does not ship
curl -sG localhost:9090/api/v1/query --data-urlencode 'query=count(gotk_suspend_status == 1)' | jq -r '.data.result[0].value[1] // "no gotk_suspend_status series"'
curl -sG localhost:9090/api/v1/query --data-urlencode 'query=gotk_suspend_status == 1' | jq '.data.result[] | .metric'
kill $PF1
flux resume kustomization demo-flux
outputcaptured 2026-09-12
$ flux suspend kustomization demo-flux
► suspending kustomization demo-flux in flux-system namespace
✔ kustomization suspended
$ sleep 60
$ kubectl -n monitoring port-forward svc/prometheus-kube-prometheus-prometheus 9090:9090 & PF1=$!
$ sleep 5
Forwarding from 127.0.0.1:9090 -> 9090
Forwarding from [::1]:9090 -> 9090
$ # gotk_resource_info comes from kube-state-metrics with a CustomResourceState config this lab does not ship
$ curl -sG localhost:9090/api/v1/query --data-urlencode 'query=group by (__name__) ({__name__=~"gotk_.+"})' | jq -r '.data.result[].metric.__name__'
Handling connection for 9090
gotk_reconcile_duration_seconds_bucket
gotk_reconcile_duration_seconds_sum
gotk_reconcile_duration_seconds_count
gotk_token_cache_evictions_total
gotk_token_cached_items
gotk_event_http_request_duration_seconds_bucket
gotk_event_http_request_duration_seconds_sum
gotk_event_http_request_duration_seconds_count
gotk_event_http_requests_inflight
gotk_event_http_response_size_bytes_bucket
gotk_event_http_response_size_bytes_sum
gotk_event_http_response_size_bytes_count
$ curl -sG localhost:9090/api/v1/query --data-urlencode 'query=count(gotk_suspend_status == 1)' | jq -r '.data.result[0].value[1] // "no gotk_suspend_status series"'
Handling connection for 9090
no gotk_suspend_status series
$ curl -sG localhost:9090/api/v1/query --data-urlencode 'query=gotk_suspend_status == 1' | jq '.data.result[] | .metric'
Handling connection for 9090
$ # gotk_resource_info comes from kube-state-metrics with a CustomResourceState config this lab does not ship
$ curl -sG localhost:9090/api/v1/query --data-urlencode 'query=count(gotk_suspend_status == 1)' | jq -r '.data.result[0].value[1] // "no gotk_suspend_status series"'
Handling connection for 9090
no gotk_suspend_status series
$ curl -sG localhost:9090/api/v1/query --data-urlencode 'query=gotk_suspend_status == 1' | jq '.data.result[] | .metric'
Handling connection for 9090
$ kill $PF1
$ flux resume kustomization demo-flux
► resuming kustomization demo-flux in flux-system namespace
✔ kustomization resumed
◎ waiting for Kustomization reconciliation
✔ Kustomization demo-flux reconciliation completed
✔ applied revision main@sha1:b793ef22cee9cf7bc5eb3d41e236dda5c43670fb
verify: the first query lists every gotk_ metric this cluster actually has, and gotk_resource_info is not one of them: that metric comes from kube-state-metrics with a CustomResourceState config, which this lab does not ship, so a query against it is empty whatever Flux is doing. If gotk_suspend_status is in the list, the count is 1 while demo-flux is suspended and the last query names it. If it is not, the controllers export no suspension gauge on this Flux version and counting suspended objects means adding the kube-state-metrics config first.

Deployment frequency starts as a counter on the CI controller. increase over a day is the honest form of the number, and knowing which label carries the outcome is the whole exercise.

kubectl -n monitoring port-forward svc/prometheus-kube-prometheus-prometheus 9090:9090 & PF1=$!
sleep 5
curl -sG localhost:9090/api/v1/query --data-urlencode 'query=tekton_pipelines_controller_pipelinerun_total' | jq '.data.result[] | {status: .metric.status, value: .value[1]}'
curl -sG localhost:9090/api/v1/query --data-urlencode 'query=increase(tekton_pipelines_controller_pipelinerun_total{status="success"}[24h])' | jq '.data.result[] | .value[1]'
curl -sG localhost:9090/api/v1/query --data-urlencode 'query=sum by (status) (increase(tekton_pipelines_controller_pipelinerun_total[24h]))' | jq '.data.result[] | {status: .metric.status, value: .value[1]}'
kill $PF1
outputcaptured 2026-09-12
$ kubectl -n monitoring port-forward svc/prometheus-kube-prometheus-prometheus 9090:9090 & PF1=$!
$ sleep 5
Forwarding from 127.0.0.1:9090 -> 9090
Forwarding from [::1]:9090 -> 9090
$ curl -sG localhost:9090/api/v1/query --data-urlencode 'query=tekton_pipelines_controller_pipelinerun_total' | jq '.data.result[] | {status: .metric.status, value: .value[1]}'
Handling connection for 9090
{
  "status": "success",
  "value": "1"
}
$ curl -sG localhost:9090/api/v1/query --data-urlencode 'query=increase(tekton_pipelines_controller_pipelinerun_total{status="success"}[24h])' | jq '.data.result[] | .value[1]'
Handling connection for 9090
"0"
$ curl -sG localhost:9090/api/v1/query --data-urlencode 'query=sum by (status) (increase(tekton_pipelines_controller_pipelinerun_total[24h]))' | jq '.data.result[] | {status: .metric.status, value: .value[1]}'
Handling connection for 9090
{
  "status": "success",
  "value": "0"
}
$ kill $PF1
verify: the counter has a label separating successes from failures and the increase over a day is a plausible number. Say why increase and not the raw counter, and what a counter reset would do to each.

Self-check

answer before opening
Name the DORA four and the system that holds each one's raw data in this lab.

Deployment frequency: Argo CD sync counters. Lead time: git commit timestamps joined with sync timestamps. Change failure rate: sync phases plus rollout outcomes. Time to restore: alerting data (firing to resolved), or incident records. None of them live in one place, which is itself the point.

Why increase() rather than the raw counter for deployment frequency?

The raw counter is cumulative since the controller started and resets on restart. increase(...[24h]) gives the count within the window and handles resets. For a daily figure that is the right unit; rate would give you a per-second number nobody wants to read.

Your new ServiceMonitor for Argo CD produces no target. First two checks?

The endpoint's port against the Service's actual port name (http-metrics, not metrics, not 8082), and the label selector against the Service's real labels. Then, as always, Prometheus's own serviceMonitorSelector and namespace selector.

How would you measure "request to fulfillment" for a self-service API?

Difference between the CR's metadata.creationTimestamp and the lastTransitionTime of its Ready condition. Aggregate it by exporting that delta as a metric (a small exporter or a recording rule over kube-state-metrics custom resource state), then track its p50/p95 over time.

Someone proposes measuring developer productivity by lines of code. Reply in one sentence.

Measure outcomes, not output: DORA's four keys plus platform-effectiveness measures (fulfillment latency, adoption, satisfaction) describe whether the system delivers value, whereas line counts reward volume and are trivially gamed.

Name the five DORA metrics as the 2024 report defines them, and say which group each belongs to.

Throughput: change lead time (commit to production), deployment frequency, failed deployment recovery time. Stability: change failure rate and deployment rework rate (unplanned deployments to fix user-facing bugs, added in 2024). Reliability is tracked as operational performance beside them, not as a sixth key, and the 2025 report replaced elite/high/medium/low clusters with team archetypes.

The white paper's "organizational efficiency" group: which three latencies does it name, and how would you get the first one out of this lab?

Request-to-fulfillment latency for a capability, latency to build and deploy a brand-new service to production, and time for a new user to submit a first code change. The first is metadata.creationTimestamp to the Ready condition's lastTransitionTime on the XR or kro instance, exported through kube-state-metrics CustomResourceStateMetrics so you can graph p50 and p95 over time.

Which platform-tool metric would you alert on to learn that the platform, not the workloads, is degrading, and what kind of SLO fits it?

A control-plane SLI: Argo CD sync success ratio from argocd_app_sync_total, reconcile error ratio from controller_runtime_reconcile_total{result="error"} for Crossplane or kro, or admission webhook latency. Give it an availability SLO (say 99.5% successful reconciles over 30 days) and use the multi-window burn-rate rules from 4.2 rather than a fixed threshold, so a slow bleed and a fast outage both page appropriately.

Docs to know your way around

study time, not exam time
  • argo-cd.readthedocs.io: the metrics page (metric names and labels).
  • dora.dev: definitions, one page each; the exam-usable summaries.
  • tag-app-delivery.cncf.io: the white paper's measurement section, for the platform-specific measures.
  • Offline: the Prometheus UI's metric explorer against the argocd_ prefix, which is faster than any documentation.
  • dora.dev/guides/dora-metrics and dora.dev/insights/dora-metrics-history: the current names (failed deployment recovery time, rework rate) and the clarification that reliability is not a fifth key.
  • tag-app-delivery.cncf.io/whitepapers/platforms/#how-to-measure-the-success-of-platforms: the three groups quoted above; /whitepapers/platform-eng-maturity-model for the Measurement levels.
  • github.com/kubernetes/kube-state-metrics/blob/main/docs/metrics/extend/customresourcestate-metrics.md: how to turn any CR's status into a metric.