The exam will not ask you to write a controller in Go. It will ask you to operate, integrate and above all diagnose operators, which means reading their outputs fluently: status conditions, events, owned resources, and logs, in that order.
make up apiOrientation
Every layer of this lab is operators consuming CRDs: Crossplane, Kyverno, Prometheus Operator, Trivy Operator, CloudNativePG, Argo's controllers, Flux's controllers. An unfamiliar operator on the exam is the same four questions every time: what kind does it watch, what does it create, what does its status say, and where do its logs live?
A controller watches a kind and, for each object, runs the same function: observe actual state, compare with desired spec, take one step toward convergence, write status, requeue. It acts on the state it finds, not on the event that woke it, so missed events cost nothing, duplicate events are harmless, and the loop is safe to run at any time. That is why deleting an operator-owned pod is a non-event: the next reconcile recreates it without knowing or caring that you deleted it.
The conventions that make operators readable
1 · Status conditions
Typed, with type, status (True/False/Unknown), reason (machine-speak, CamelCase), message (for you), lastTransitionTime and observedGeneration. Ready is the summary; the others tell you which phase is stuck. A well-behaved operator also emits Events, which carry the same story with timestamps and counts.
2 · observedGeneration versus metadata.generation
If they differ, the controller has not yet processed your latest edit, and whatever status says describes a previous spec. Checking this first avoids diagnosing stale information, and almost nobody does it. metadata.generation increments on spec changes only, provided the CRD has the status subresource. Without it, status writes bump generation too, and observedGeneration tells you nothing. That is one more reason the subresource is not optional for a real API (section 3.2).
3 · Owner references
Operator-created resources carry ownerReferences pointing at their parent. This answers "what keeps recreating this thing", drives cascade deletion (foreground, background, or orphan), and lets you reconstruct an operator's whole object tree without documentation. Note the rules: an owner must be in the same namespace (or be cluster-scoped), and a namespaced object cannot own a cluster-scoped one, the constraint behind Crossplane's design in 3.5.
4 · Finalizers
A deletion sticks in Terminating while a finalizer is present, because the controller is doing teardown work, or is dead and cannot. A stuck namespace or CR almost always means "find whose finalizer, and why its controller is not running". Removing the finalizer by hand is the last resort, and it leaks whatever the teardown was supposed to clean up (external volumes, cloud resources, DNS records). Accept that trade knowingly before you do it.
Cluster/pg (your CR)
├─ ownerRef ─▶ StatefulSet-ish pods pg-1, pg-2 each with a PVC
├─ ownerRef ─▶ Services pg-rw, pg-ro, pg-r
├─ ownerRef ─▶ Secrets pg-app, pg-superuser ← connection details
└─ status.conditions[] Ready / Initialized / ContinuousArchiving …
observedGeneration must equal metadata.generation
kubectl get <kind> <name> -o jsonpath='{.status.conditions}' | jq
kubectl get events --field-selector involvedObject.name=<name> --sort-by=.lastTimestamp
kubectl get all,pvc,secret -l <the operator's label> -n <ns>
kubectl -n <operator-ns> logs deploy/<controller> --tail=100outputcaptured 2026-08-26
$ kubectl get cluster pg -o jsonpath='{.status.conditions}' | jq
[
{
"lastTransitionTime": "2026-08-27T02:02:29Z",
"message": "Cluster has been bootstrapped",
"reason": "BootstrapCompleted",
"status": "True",
"type": "Initialized"
},
{
"lastTransitionTime": "2026-08-27T02:03:45Z",
"message": "A single, unique system ID was found across reporting instances.",
"reason": "Unique",
"status": "True",
"type": "ConsistentSystemID"
},
{
"lastTransitionTime": "2026-08-27T02:04:50Z",
"message": "Cluster is Ready",
"reason": "ClusterIsReady",
"status": "True",
"type": "Ready"
},
{
"lastTransitionTime": "2026-08-27T02:03:44Z",
"message": "Continuous archiving is working",
"reason": "ContinuousArchivingSuccess",
"status": "True",
"type": "ContinuousArchiving"
}
]
$ kubectl get events --field-selector involvedObject.name=pg --sort-by=.lastTimestamp
LAST SEEN TYPE REASON OBJECT MESSAGE
3m59s Normal CreatingPodDisruptionBudget cluster/pg Creating PodDisruptionBudget pg-primary
3m59s Normal CreatingServiceAccount cluster/pg Creating ServiceAccount
3m59s Normal CreatingRole cluster/pg Creating Cluster Role
3m58s Normal CreatingInstance cluster/pg Primary instance (initdb)
2m36s Normal CreatingInstance cluster/pg Creating instance pg-2
$ kubectl get all,pvc,secret -l cnpg.io/cluster=pg -n default
NAME READY STATUS RESTARTS AGE
pod/pg-1 1/1 Running 0 2m49s
pod/pg-2 1/1 Running 0 108s
NAME TYPE CLUSTER-IP EXTERNAL-IP PORT(S) AGE
service/pg-r ClusterIP 10.96.28.97 <none> 5432/TCP 4m
service/pg-ro ClusterIP 10.96.161.206 <none> 5432/TCP 4m
service/pg-rw ClusterIP 10.96.181.56 <none> 5432/TCP 3m59s
NAME STATUS VOLUME CAPACITY ACCESS MODES STORAGECLASS VOLUMEATTRIBUTESCLASS AGE
persistentvolumeclaim/pg-1 Bound pvc-b4511fca-9bf3-4cbf-bc9d-d56906094336 1Gi RWO standard <unset> 3m59s
persistentvolumeclaim/pg-2 Bound pvc-fae7c17c-a4a7-45d5-b239-2662652ed787 1Gi RWO standard <unset> 2m37s
NAME TYPE DATA AGE
secret/pg-app kubernetes.io/basic-auth 11 4m1s
$ kubectl -n cnpg-system logs deploy/cnpg-cloudnative-pg --tail=100
{"level":"info","ts":"2026-08-27T01:52:15.835497336Z","msg":"Starting EventSource","controller":"pooler","controllerGroup":"postgresql.cnpg.io","controllerKind":"Pooler","source":"kind source: *v1.Role"}
{"level":"info","ts":"2026-08-27T01:52:15.835502695Z","msg":"Starting EventSource","controller":"cluster","controllerGroup":"postgresql.cnpg.io","controllerKind":"Cluster","source":"kind source: *v1.DatabaseRole"}
{"level":"info","ts":"2026-08-27T01:52:15.83547942Z","msg":"Starting EventSource","controller":"pooler","controllerGroup":"postgresql.cnpg.io","controllerKind":"Pooler","source":"kind source: *v1.ServiceAccount"}
{"level":"info","ts":"2026-08-27T01:52:15.835506197Z","msg":"Starting EventSource","controller":"plugin","controllerGroup":"","controllerKind":"Service","source":"kind source: *v1.Service"}
{"level":"info","ts":"2026-08-27T01:52:15.835487001Z","msg":"Starting EventSource","controller":"plugin","controllerGroup":"","controllerKind":"Service","source":"kind source: *v1.Secret"}
{"level":"info","ts":"2026-08-27T01:52:15.835507298Z","msg":"Starting EventSource","controller":"pooler","controllerGroup":"postgresql.cnpg.io","controllerKind":"Pooler","source":"kind source: *v1.ImageCatalog"}
{"level":"info","ts":"2026-08-27T01:52:15.935948428Z","msg":"Starting Controller","controller":"backup","controllerGroup":"postgresql.cnpg.io","controllerKind":"Backup"}
{"level":"info","ts":"2026-08-27T01:52:15.935986228Z","msg":"Starting workers","controller":"backup","controllerGroup":"postgresql.cnpg.io","controllerKind":"Backup","worker count":1}
{"level":"info","ts":"2026-08-27T01:52:16.036525115Z","msg":"Starting Controller","controller":"scheduled-backup","controllerGroup":"postgresql.cnpg.io","controllerKind":"ScheduledBackup"}
{"level":"info","ts":"2026-08-27T01:52:16.036555833Z","msg":"Starting workers","controller":"scheduled-backup","controllerGroup":"postgresql.cnpg.io","controllerKind":"ScheduledBackup","worker count":10}
{"level":"info","ts":"2026-08-27T01:52:16.037638856Z","msg":"Starting Controller","controller":"plugin","controllerGroup":"","controllerKind":"Service"}
{"level":"info","ts":"2026-08-27T01:52:16.037687808Z","msg":"Starting workers","controller":"plugin","controllerGroup":"","controllerKind":"Service","worker count":10}
{"level":"info","ts":"2026-08-27T01:52:16.037735706Z","msg":"Starting Controller","controller":"cluster","controllerGroup":"postgresql.cnpg.io","controllerKind":"Cluster"}
{"level":"info","ts":"2026-08-27T01:52:16.037731587Z","msg":"Starting Controller","controller":"pooler","controllerGroup":"postgresql.cnpg.io","controllerKind":"Pooler"}
{"level":"info","ts":"2026-08-27T01:52:16.037745707Z","msg":"Starting workers","controller":"cluster","controllerGroup":"postgresql.cnpg.io","controllerKind":"Cluster","worker count":10}
{"level":"info","ts":"2026-08-27T01:52:16.03775078Z","msg":"Starting workers","controller":"pooler","controllerGroup":"postgresql.cnpg.io","controllerKind":"Pooler","worker count":10}
{"level":"info","ts":"2026-08-27T01:52:16.037728098Z","msg":"Starting Controller","controller":"database-role","controllerGroup":"postgresql.cnpg.io","controllerKind":"DatabaseRole"}
{"level":"info","ts":"2026-08-27T01:52:16.037778053Z","msg":"Starting workers","controller":"database-role","controllerGroup":"postgresql.cnpg.io","controllerKind":"DatabaseRole","worker count":10}
{"level":"info","ts":"2026-08-27T02:02:26.19113155Z","logger":"cluster-resource","msg":"Defaulting for Cluster","version":"v1","name":"pg","namespace":"default"}
{"level":"info","ts":"2026-08-27T02:02:26.22083863Z","logger":"cluster-resource","msg":"Defaulting for Cluster","version":"v1","name":"pg","namespace":"default"}
{"level":"info","ts":"2026-08-27T02:02:26.32053831Z","logger":"cluster-resource","msg":"Defaulting for Cluster","version":"v1","name":"pg","namespace":"default"}
{"level":"info","ts":"2026-08-27T02:02:26.460210545Z","msg":"no orphan PVCs found, skipping the restored cluster reconciliation","controller":"cluster","controllerGroup":"postgresql.cnpg.io","controllerKind":"Cluster","Cluster":{"name":"pg","namespace":"default"},"namespace":"default","name":"pg","reconcileID":"7652038d-f270-4d4c-9a41-d2fbd1ab01f4"}
{"level":"info","ts":"2026-08-27T02:02:26.716626012Z","msg":"creating service","controller":"cluster","controllerGroup":"postgresql.cnpg.io","controllerKind":"Cluster","Cluster":{"name":"pg","namespace":"default"},"namespace":"default","name":"pg","reconcileID":"7652038d-f270-4d4c-9a41-d2fbd1ab01f4","serviceName":"pg-r","updateStrategy":"patch"}
{"level":"info","ts":"2026-08-27T02:02:26.949154863Z","msg":"creating service","controller":"cluster","controllerGroup":"postgresql.cnpg.io","controllerKind":"Cluster","Cluster":{"name":"pg","namespace":"default"},"namespace":"default","name":"pg","reconcileID":"7652038d-f270-4d4c-9a41-d2fbd1ab01f4","serviceName":"pg-ro","updateStrategy":"patch"}
{"level":"info","ts":"2026-08-27T02:02:27.089001205Z","msg":"creating service","controller":"cluster","controllerGroup":"postgresql.cnpg.io","controllerKind":"Cluster","Cluster":{"name":"pg","namespace":"default"},"namespace":"default","name":"pg","reconcileID":"7652038d-f270-4d4c-9a41-d2fbd1ab01f4","serviceName":"pg-rw","updateStrategy":"patch"}
{"level":"info","ts":"2026-08-27T02:02:27.970787374Z","msg":"Created primary lease","controller":"cluster","controllerGroup":"postgresql.cnpg.io","controllerKind":"Cluster","Cluster":{"name":"pg","namespace":"default"},"namespace":"default","name":"pg","reconcileID":"7652038d-f270-4d4c-9a41-d2fbd1ab01f4","leaseName":"pg"}
{"level":"info","ts":"2026-08-27T02:02:28.411769916Z","msg":"Creating new Job","controller":"cluster","controllerGroup":"postgresql.cnpg.io","controllerKind":"Cluster","Cluster":{"name":"pg","namespace":"default"},"namespace":"default","name":"pg","reconcileID":"7652038d-f270-4d4c-9a41-d2fbd1ab01f4","jobName":"pg-1-initdb","primary":true}
{"level":"info","ts":"2026-08-27T02:02:28.578577487Z","logger":"cluster-resource","msg":"Defaulting for Cluster","version":"v1","name":"pg","namespace":"default"}
{"level":"info","ts":"2026-08-27T02:02:28.64742958Z","msg":"skipping pvc because it has owner metadata","controller":"cluster","controllerGroup":"postgresql.cnpg.io","controllerKind":"Cluster","Cluster":{"name":"pg","namespace":"default"},"namespace":"default","name":"pg","reconcileID":"175abb6a-f53f-48c2-9d02-5df1736425de","step":"get_orphan_pvcs","pvcName":"pg-1"}
{"level":"info","ts":"2026-08-27T02:02:28.647440816Z","msg":"no orphan PVCs found, skipping the restored cluster reconciliation","controller":"cluster","controllerGroup":"postgresql.cnpg.io","controllerKind":"Cluster","Cluster":{"name":"pg","namespace":"default"},"namespace":"default","name":"pg","reconcileID":"175abb6a-f53f-48c2-9d02-5df1736425de"}
{"level":"info","ts":"2026-08-27T02:02:29.251311342Z","logger":"cluster-resource","msg":"Defaulting for Cluster","version":"v1","name":"pg","namespace":"default"}
{"level":"info","ts":"2026-08-27T02:02:35.275772086Z","logger":"cluster-resource","msg":"Defaulting for Cluster","version":"v1","name":"pg","namespace":"default"}
{"level":"info","ts":"2026-08-27T02:02:41.089429634Z","logger":"cluster-resource","msg":"Defaulting for Cluster","version":"v1","name":"pg","namespace":"default"}
{"level":"info","ts":"2026-08-27T02:02:43.811342944Z","logger":"cluster-resource","msg":"Defaulting for Cluster","version":"v1","name":"pg","namespace":"default"}
{"level":"info","ts":"2026-08-27T02:02:44.368388244Z","logger":"cluster-resource","msg":"Defaulting for Cluster","version":"v1","name":"pg","namespace":"default"}
{"level":"info","ts":"2026-08-27T02:02:49.675245765Z","logger":"cluster-resource","msg":"Defaulting for Cluster","version":"v1","name":"pg","namespace":"default"}
{"level":"info","ts":"2026-08-27T02:02:55.257182935Z","logger":"cluster-resource","msg":"Defaulting for Cluster","version":"v1","name":"pg","namespace":"default"}
{"level":"info","ts":"2026-08-27T02:03:00.993060187Z","logger":"cluster-resource","msg":"Defaulting for Cluster","version":"v1","name":"pg","namespace":"default"}
{"level":"info","ts":"2026-08-27T02:03:06.998428892Z","logger":"cluster-resource","msg":"Defaulting for Cluster","version":"v1","name":"pg","namespace":"default"}
{"level":"info","ts":"2026-08-27T02:03:12.204012127Z","logger":"cluster-resource","msg":"Defaulting for Cluster","version":"v1","name":"pg","namespace":"default"}
… (rest of the tail omitted)In that order. Conditions tell you the phase, events tell you the attempts, owned objects tell you what exists, logs tell you why the controller gave up. Jumping to logs first is the most common time sink in this domain.
Vocabulary around operators
- Operator capability levels (the OperatorHub model): 1 basic install → 2 seamless upgrades → 3 full lifecycle (backup, failover) → 4 deep insights (metrics, alerts) → 5 auto-pilot (auto-scaling, auto-tuning). A useful yardstick when a scenario asks whether to adopt an operator or run something yourself.
- OLM (Operator Lifecycle Manager): installs and upgrades operators from catalogs. Not installed here; know the name.
- Kubebuilder / Operator SDK / controller-runtime: the scaffolding.
kubebuilderis on your PATH frommake tools, and scaffolding a controller once (kubebuilder init,kubebuilder create api) is genuinely instructive for understanding what operators are made of. It is also beyond what the CNPE tests. - Reconcile loop hazards worth naming: hot loops (a controller that writes status on every pass and re-triggers itself), requeue backoff, leader election (why an operator's Deployment is usually 1 replica or uses a lease), and cache staleness right after a write.
- Operator vs controller vs webhook: a controller reconciles; an operator is a controller plus domain knowledge shipped with its CRDs; a webhook intercepts writes synchronously. Different failure modes: a dead controller means nothing converges, a dead webhook (with
failurePolicy: Fail) means nothing it matches can be written at all.
"Operators for platform automation and integration" means wiring an operator into the rest of the platform: its CRs live in git (GitOps), its metrics are scraped (ServiceMonitor), its secrets flow to consumers (connection Secrets), its resources are policed (policy engines see CRs like anything else), and its API is exposed to developers through a thinner abstraction (Crossplane, kro, or a Backstage template). Being able to list those five integrations is a complete answer.
Inside a controller: controller-runtime in one page
You will not write Go on the exam, but every operator you diagnose is built on the same dozen concepts, and their names appear in logs, flags and Deployment manifests. Knowing them turns "the operator is weird" into "the cache is stale" or "it lost the lease".
The pieces
- Manager: one process hosting shared informer caches, a client, metrics (
:8080), health probes (:8081,/healthzand/readyz), an optional webhook server (:9443) and the controllers. Flags on the Deployment such as--leader-elect,--metrics-bind-address,--health-probe-bind-addressand--watch-namespaceare manager options. - Controller: a work queue of
namespace/namekeys plus a Reconciler with one method,Reconcile(ctx, req) (Result, error). It receives a key, never an event, which is the code-level reason reconciliation is level-based.MaxConcurrentReconciles(default 1) is why an operator handling hundreds of CRs can look slow. - Builder:
For(&MyKind{})is the primary resource (its key is enqueued);Owns(&appsv1.Deployment{})enqueues the owner when an owned object changes, found throughownerReferenceswithcontroller: true;Watches(...)with a handler maps any other object (a Secret the CR references) to the keys to reconcile.Ownsis why deleting an operator's Deployment or Secret triggers an immediate rebuild and whyownerReferencesis your map of what an operator manages. - Predicates filter events before they hit the queue:
GenerationChangedPredicatedrops updates wheremetadata.generationdid not change (status-only writes),LabelChangedPredicate,AnnotationChangedPredicate, and custom funcs. An operator that "ignores my label change" or "reconciles in a hot loop" usually has a predicate to blame or to add. - Result: return
Result{}when done,Result{RequeueAfter: 5*time.Minute}to poll an external system, or anerrorto retry with exponential backoff (the workqueue's rate limiter: 5 ms doubling up to about 16 minutes, with a global 10 per second budget). "Retrying in 1m4s" style log lines are that backoff. - Cache: reads come from informers, so a write followed by an immediate read can return the old object (cache staleness); well-written reconcilers tolerate this by re-reading on the next pass rather than fighting it.
- Leader election: a
Leasein the operator's namespace named byLeaderElectionID;LeaseDuration15 s,RenewDeadline10 s,RetryPeriod2 s are the defaults. Extra replicas run but do nothing until they win.kubectl get lease -n <operator-ns>showsholderIdentity, which answers "which pod is actually reconciling".
The RBAC an operator needs, and what its absence looks like
| Rule | Why | Symptom when missing |
|---|---|---|
| its own CRD: get, list, watch, update, patch; plus <plural>/status and <plural>/finalizers | watch the primary resource, write status, manage finalizers | startup crash "failed to watch" or Forbidden: ... cannot update resource "widgets/status" in logs; CRs never gain status |
| every owned kind: create, update, patch, delete, get, list, watch | Owns() needs to watch it and write it | CR stays Progressing; logs show is forbidden: User "system:serviceaccount:ns:sa" cannot create ... |
| coordination.k8s.io leases: get, create, update | leader election | "error retrieving resource lock" and no replica ever becomes leader; nothing reconciles |
| events: create, patch | kubectl describe output | silent: conditions update but no Events appear |
| the Secrets or ConfigMaps it reads across namespaces | integration inputs | a Forbidden that looks like the Secret is missing |
Kubebuilder generates these from // +kubebuilder:rbac markers into a ClusterRole; Helm and Ansible operators ship the same rules by hand. A namespaced install (--watch-namespace) can use Roles instead, which is the multi-tenant compromise (1.4).
Admission webhooks inside operators
Many operators register a ValidatingWebhookConfiguration and MutatingWebhookConfiguration for their own kinds (defaulting, cross-field checks CEL cannot do, conversion). They need a serving certificate: cert-manager's cert-manager.io/inject-ca-from annotation, or the operator generating a self-signed pair and patching caBundle. When the operator pod is down, failurePolicy: Fail on its webhook makes every write to its kinds fail with "failed calling webhook"; that includes the GitOps controller's own applies, and it is the first thing to check when a whole API group becomes read-only after an operator upgrade.
kubectl -n <op-ns> get lease (who leads), kubectl -n <op-ns> logs deploy/<op> | grep -i -E 'forbidden|error' (RBAC and backoff), kubectl get validatingwebhookconfiguration -o wide (a webhook that could be blocking), then the four-command diagnosis from the panel above. Restarting the operator is what you do after those, not before.
Installing and judging operators
OLM v0 and OLM v1
| Aspect | OLM v0 (classic) | OLM v1 (operator-controller, GA) |
|---|---|---|
| Catalog | CatalogSource (index image) | ClusterCatalog { spec.source.type: Image, image.ref } with priority and availabilityMode; conditions Serving, Progressing |
| Install request | Subscription (package, channel, installPlanApproval Automatic|Manual) + OperatorGroup for the target namespaces | ClusterExtension { spec.namespace, spec.serviceAccount.name, spec.source.sourceType: Catalog, catalog.packageName, channels, version (semver range), upgradeConstraintPolicy } |
| What runs the install | OLM's own service account with broad rights; a ClusterServiceVersion per version; InstallPlan objects to approve | the ServiceAccount you name, with exactly the RBAC the bundle needs (least privilege is on you) |
| Upgrades | follow the channel's upgrade graph | CatalogProvided (default: only successors the catalog declares) or SelfCertified (any version, including downgrades, at your risk); a CRD upgrade safety preflight blocks scope changes, removed stored versions, removed fields, new required fields, changed defaults or tightened bounds |
| Status | CSV phase Succeeded / Failed | conditions Installed and Progressing; status.install.bundle.name/version |
Outside OLM entirely, operators arrive as Helm charts or plain manifests, and then Flux's HelmRelease or an Argo CD Application is the lifecycle manager; the CRD upgrade caveats from 2.3 (upgrade.crds defaults to Skip) apply.
Capability levels with the question that decides each
| Level | Name | Passes when | Ceiling for |
|---|---|---|---|
| 1 | Basic Install | the CR alone provisions and configures the workload; readiness is written to status | |
| 2 | Seamless Upgrades | bumping the CR (or the operator) upgrades the operand without manual steps or downtime | Helm-based operators top out here |
| 3 | Full Lifecycle | backup and restore, failover, scaling members, reconfiguration flows are operator actions | Ansible operators usually reach this |
| 4 | Deep Insights | the operator exposes metrics, alerts and custom Events for operand and itself | |
| 5 | Auto Pilot | auto-scaling, auto-healing, auto-tuning, anomaly detection from the operand's own signals | Go operators only, in practice |
The three SDK flavors differ in what the reconciler is: a Helm operator renders a chart from the CR's spec as values (no custom logic, level 1 to 2); an Ansible operator runs a playbook or role per reconcile (level 3 is reachable, slow loops); a Go operator is controller-runtime code (anything). A scenario asking "should we adopt this operator or run the database ourselves" wants the level named and matched to the need: a level 1 operator for a stateful service saves nothing on day 2.
Status conditions, the convention in full
metav1.Condition:type(CamelCase, ordomain/CamelCasefor third parties),statusTrue/False/Unknown,reason(CamelCase, required, machine-readable),message,lastTransitionTime(changes only when status flips),observedGeneration. The list islistType: mapkeyed bytype.- Type names describe an observed state (adjective or past participle:
Ready,Available,Succeeded,Degraded), never a transition (Deploying). Long transitions are still states (Resizing,Reconciling) toggled True/False. - Polarity:
Ready=Trueis normal-true;Stalled=TrueorDegraded=Trueis normal-false. You cannot summarize a resource's conditions without knowing each type's polarity, which is whyReady(long-running) orSucceeded(run-to-completion) as a summary condition is the convention. - Absent condition equals
Unknown; a controller should write its conditions on the first pass even as Unknown so consumers know it is alive. kstatus (used by Flux, Argo CD's guidance, kubectl wait) readsReady,ReconcilingandStalledplusobservedGenerationto compute Current, InProgress or Failed.
Reading a stuck operator, in order
metadata.generationvsstatus.observedGeneration: unequal means the controller has not processed your change; if it stays unequal, the controller is not running, not leading, or filtered your event.- Conditions: the
reasonof the False or Stalled one; kstatus-styleStalled=Truemeans it will not retry until spec changes. - Events on the CR:
Warningevents carry the failed API call. - Owned objects via
ownerReferences: which child is missing or unhealthy; its own events (quota, admission, image pull). - Operator pod: running, leader (
kubectl get lease), logs withForbidden(RBAC),connection refused(a dependency), or a backoff loop. Then the webhook configurations it registered. - Finalizers: a CR in Terminating with the operator's finalizer and no operator is the leak trade-off from the panel above; a CR with a finalizer from an operator that was uninstalled needs the operator back, or a deliberate finalizer removal accepting whatever external resource it leaves behind.
"The database CR has been Progressing for twenty minutes" is graded on whether you find the actual cause: an RBAC Forbidden in the operator log, a NetworkPolicy blocking the init Job, a quota refusing the StatefulSet's pods, a webhook with no endpoints. Each leaves a different artefact (log line, Hubble drop, ReplicaSet event, webhook failure text). Naming the artefact and fixing that object is the task; deleting and recreating the CR is a level-based no-op that changes nothing.
Exercises
In default (deliberately; the tenant-namespace variant comes last):
kubectl apply -f - <<'EOF'
apiVersion: postgresql.cnpg.io/v1
kind: Cluster
metadata: { name: pg, namespace: default }
spec:
instances: 2
storage: { size: 1Gi }
EOF
kubectl get cluster pg -woutputcaptured 2026-08-26
$ kubectl apply -f - <<'EOF'
apiVersion: postgresql.cnpg.io/v1
kind: Cluster
metadata: { name: pg, namespace: default }
spec:
instances: 2
storage: { size: 1Gi }
EOF
cluster.postgresql.cnpg.io/pg created
$ kubectl get cluster pg -w
NAME AGE INSTANCES READY STATUS PRIMARY
pg 0s
pg 2s
pg 2s
pg 2s
pg 2s Setting up primary
pg 3s 1 Setting up primary
pg 71s 1 Setting up primary
pg 72s 1 Setting up primary
pg 72s 1 Waiting for the instances to become active
pg 77s 1 Waiting for the instances to become active
pg 78s 1 Waiting for the instances to become active pg-1
pg 78s 1 Waiting for the instances to become active pg-1
pg 79s 1 Waiting for the instances to become active pg-1
pg 84s 1 1 Waiting for the instances to become active pg-1
pg 84s 1 1 Creating a new replica pg-1
pg 85s 2 1 Creating a new replica pg-1
pg 2m12s 2 1 Creating a new replica pg-1
pg 2m13s 2 1 Creating a new replica pg-1
pg 2m13s 2 1 Waiting for the instances to become active pg-1
pg 2m19s 2 1 Waiting for the instances to become active pg-1
pg 2m21s 2 1 Waiting for the instances to become active pg-1
pg 2m23s 2 1 Waiting for the instances to become active pg-1
pg 2m24s 2 2 Waiting for the instances to become active pg-1
pg 2m24s 2 2 Cluster in healthy state pg-1
pg 2m25s 2 2 Cluster in healthy state pg-1
^CWhile it converges, run the diagnosis commands: kubectl get events --field-selector involvedObject.name=pg --sort-by=.lastTimestamp, kubectl get cluster pg -o jsonpath='{.status.conditions}' | jq, and kubectl get pods,pvc,svc -l cnpg.io/cluster=pg.
kubectl delete pod pg-1 and time what happens. Then scale through the spec, kubectl patch cluster pg --type=merge -p '{"spec":{"instances":3}}', and check observedGeneration catches up to generation before you trust the new status.
status.observedGeneration == metadata.generation.The lab ships a deliberately broken manifest: examples/crossplane/pg-cluster.yaml targets team-a, where the default-deny NetworkPolicy silently blocks the instance's access to the API server. (team-a comes from make sec; if you skipped that layer, kubectl apply -f examples/multitenancy/team-a.yaml stages the same trap.) Apply it, and diagnose from evidence only: conditions say Initialized True but Ready False, phase "Setting up primary" forever, and kubectl -n team-a logs job/pg-1-initdb ends in dial tcp 10.96.0.1:443: i/o timeout. The file's comments explain why the obvious ipBlock fix fails (ClusterIP is DNAT'd before policy evaluation) and give the Cilium toEntities: [kube-apiserver] answer; do not read them until you have formed your own theory.
Pick two operators you have installed and, for each, name the watched kind, find one owned resource via ownerReferences, and locate the controller deployment.
kubectl api-resources --api-group=<group> and a one-line answer each. The point is transferable fluency: an unfamiliar operator on the exam is the same four questions.Two replicas of an operator do not reconcile in parallel; one holds a Lease and the other waits. The Lease is a plain object you can read, and it is the fastest way to answer "which pod is actually doing the work".
kubectl -n cnpg-system get lease
kubectl -n cnpg-system get lease -o jsonpath='{range .items[*]}{.metadata.name} {.spec.holderIdentity}{"\n"}{end}'
kubectl -n cnpg-system scale deploy cnpg-cloudnative-pg --replicas=2
kubectl -n cnpg-system rollout status deploy cnpg-cloudnative-pg --timeout=120s
kubectl -n cnpg-system get pods -l app.kubernetes.io/name=cloudnative-pg
kubectl -n cnpg-system get lease -o jsonpath='{range .items[*]}{.metadata.name} {.spec.holderIdentity}{"\n"}{end}'
kubectl -n cnpg-system scale deploy cnpg-cloudnative-pg --replicas=1outputcaptured 2026-09-12
$ kubectl -n cnpg-system get lease
NAME HOLDER AGE
db9c8771.cnpg.io cnpg-cloudnative-pg-7d855d69bb-9c82x_1bf08cef-9c4b-46ce-8407-5408ac771a1b 7h8m
$ kubectl -n cnpg-system get lease -o jsonpath='{range .items[*]}{.metadata.name} {.spec.holderIdentity}{"\n"}{end}'
db9c8771.cnpg.io cnpg-cloudnative-pg-7d855d69bb-9c82x_1bf08cef-9c4b-46ce-8407-5408ac771a1b
$ kubectl -n cnpg-system scale deploy cnpg-cloudnative-pg --replicas=2
deployment.apps/cnpg-cloudnative-pg scaled
$ kubectl -n cnpg-system rollout status deploy cnpg-cloudnative-pg --timeout=120s
Waiting for deployment "cnpg-cloudnative-pg" rollout to finish: 1 out of 2 new replicas have been updated...
Waiting for deployment "cnpg-cloudnative-pg" rollout to finish: 1 of 2 updated replicas are available...
deployment "cnpg-cloudnative-pg" successfully rolled out
$ kubectl -n cnpg-system get pods -l app.kubernetes.io/name=cloudnative-pg
NAME READY STATUS RESTARTS AGE
cnpg-cloudnative-pg-7d855d69bb-4t22l 1/1 Running 0 26s
cnpg-cloudnative-pg-7d855d69bb-9c82x 1/1 Running 0 5h52m
$ kubectl -n cnpg-system get lease -o jsonpath='{range .items[*]}{.metadata.name} {.spec.holderIdentity}{"\n"}{end}'
db9c8771.cnpg.io cnpg-cloudnative-pg-7d855d69bb-9c82x_1bf08cef-9c4b-46ce-8407-5408ac771a1b
$ kubectl -n cnpg-system scale deploy cnpg-cloudnative-pg --replicas=1
deployment.apps/cnpg-cloudnative-pg scaledAn operator is a controller with a ServiceAccount, and everything it can do is in a ClusterRole you can read. The same pattern repeats across every operator: the CRD, its status, its finalizers, and the built-in kinds it composes.
kubectl -n cnpg-system get sa
kubectl get clusterrole -o name | grep -i cnpg
kubectl get clusterrole cnpg-cloudnative-pg -o yaml | grep -B3 -A6 'postgresql.cnpg.io'
kubectl get clusterrole cnpg-cloudnative-pg -o jsonpath='{range .rules[*]}{.apiGroups} {.resources} {.verbs}{"\n"}{end}' | head -20outputcaptured 2026-09-12
$ kubectl -n cnpg-system get sa
NAME AGE
cnpg-cloudnative-pg 81m
default 81m
$ kubectl get clusterrole -o name | grep -i cnpg
clusterrole.rbac.authorization.k8s.io/cnpg-cloudnative-pg
clusterrole.rbac.authorization.k8s.io/cnpg-cloudnative-pg-edit
clusterrole.rbac.authorization.k8s.io/cnpg-cloudnative-pg-view
$ kubectl get clusterrole cnpg-cloudnative-pg -o yaml | grep -B3 -A6 'postgresql.cnpg.io'
- get
- patch
- apiGroups:
- postgresql.cnpg.io
resources:
- clusterimagecatalogs
verbs:
- get
- list
- watch
--
- update
- watch
- apiGroups:
- postgresql.cnpg.io
resources:
- backups
- clusters
- databaseroles
- databases
- poolers
--
- update
- watch
- apiGroups:
- postgresql.cnpg.io
resources:
- failoverquorums
verbs:
- create
- delete
- get
- list
- watch
- apiGroups:
- postgresql.cnpg.io
resources:
- backups/status
- databases/status
- publications/status
... 29 more lines
$ kubectl get clusterrole cnpg-cloudnative-pg -o jsonpath='{range .rules[*]}{.apiGroups} {.resources} {.verbs}{"\n"}{end}' | head -20
[""] ["nodes"] ["get","list","watch"]
["admissionregistration.k8s.io"] ["mutatingwebhookconfigurations","validatingwebhookconfigurations"] ["get","patch"]
["postgresql.cnpg.io"] ["clusterimagecatalogs"] ["get","list","watch"]
[""] ["configmaps","secrets","services"] ["create","delete","get","list","patch","update","watch"]
[""] ["configmaps/status","secrets/status"] ["get","patch","update"]
[""] ["events"] ["create","patch"]
[""] ["persistentvolumeclaims","pods","pods/exec"] ["create","delete","get","list","patch","watch"]
[""] ["pods/status"] ["get"]
[""] ["serviceaccounts"] ["create","get","list","patch","update","watch"]
["apps"] ["deployments"] ["create","delete","get","list","patch","update","watch"]
["batch"] ["jobs"] ["create","delete","get","list","patch","watch"]
["coordination.k8s.io"] ["leases"] ["create","get","list","update","watch"]
["discovery.k8s.io"] ["endpointslices"] ["get","list","watch"]
["monitoring.coreos.com"] ["podmonitors"] ["create","delete","get","list","patch","watch"]
["policy"] ["poddisruptionbudgets"] ["create","delete","get","list","patch","update","watch"]
["postgresql.cnpg.io"] ["backups","clusters","databaseroles","databases","poolers","publications","scheduledbackups","subscriptions"] ["create","delete","get","list","patch","update","watch"]
["postgresql.cnpg.io"] ["failoverquorums"] ["create","delete","get","list","watch"]
["postgresql.cnpg.io"] ["backups/status","databases/status","publications/status","scheduledbackups/status","subscriptions/status"] ["get","patch","update"]
["postgresql.cnpg.io"] ["imagecatalogs"] ["get","list","watch"]
["postgresql.cnpg.io"] ["clusters/finalizers","databaseroles/finalizers","poolers/finalizers"] ["update"]/status, and the one granting /finalizers, and explain why an operator needs all three. Then say which built-in kinds this operator creates on your behalf.A controller with a missing permission does not crash; it retries and complains. The Forbidden line in the log names the verb, the resource and the ServiceAccount, which is everything you need, and the object it was reconciling stays where it was.
kubectl apply -f examples/crossplane/pg-cluster.yaml
sleep 30
# strip the resourceVersion or the restore comes back Conflict and the role stays broken
kubectl get clusterrole cnpg-cloudnative-pg -o json | jq 'del(.metadata.resourceVersion,.metadata.uid,.metadata.creationTimestamp,.metadata.generation,.metadata.managedFields)' > /tmp/cnpg-role.json
kubectl get clusterrole cnpg-cloudnative-pg -o json | jq 'del(.rules[] | select(.resources | index("pods")))' | kubectl apply -f -
kubectl -n team-a delete pod pg-1 --ignore-not-found
sleep 60
kubectl -n cnpg-system logs deploy/cnpg-cloudnative-pg --tail=200 | grep -i forbidden | head -5
kubectl -n team-a get cluster pg -o jsonpath='{.status.phase}{"\n"}'
kubectl replace -f /tmp/cnpg-role.json
sleep 30
kubectl -n team-a get cluster pg -o jsonpath='{.status.phase}{"\n"}'
kubectl get clusterrole cnpg-cloudnative-pg -o json | jq '[.rules[] | select(.resources | index("pods"))] | length'outputcaptured 2026-09-13
$ kubectl apply -f examples/crossplane/pg-cluster.yaml
cluster.postgresql.cnpg.io/pg unchanged
$ sleep 30
$ # strip the resourceVersion or the restore comes back Conflict and the role stays broken
$ kubectl get clusterrole cnpg-cloudnative-pg -o json | jq 'del(.metadata.resourceVersion,.metadata.uid,.metadata.creationTimestamp,.metadata.generation,.metadata.managedFields)' > /tmp/cnpg-role.json
$ kubectl get clusterrole cnpg-cloudnative-pg -o json | jq 'del(.rules[] | select(.resources | index("pods")))' | kubectl apply -f -
clusterrole.rbac.authorization.k8s.io/cnpg-cloudnative-pg configured
$ kubectl -n team-a delete pod pg-1 --ignore-not-found
$ sleep 60
$ kubectl -n cnpg-system logs deploy/cnpg-cloudnative-pg --tail=200 | grep -i forbidden | head -5
{"level":"error","ts":"2026-09-13T12:40:40.616875436Z","logger":"controller-runtime.cache.UnhandledError","msg":"Failed to watch","reflector":"pkg/mod/k8s.io/client-go@v0.36.2/tools/cache/reflector.go:343","type":"*v1.PersistentVolumeClaim","error":"persistentvolumeclaims is forbidden: User \"system:serviceaccount:cnpg-system:cnpg-cloudnative-pg\" cannot watch resource \"persistentvolumeclaims\" in API group \"\" at the cluster scope","stacktrace":"k8s.io/apimachinery/pkg/util/runtime.logError\n\tpkg/mod/k8s.io/apimachinery@v0.36.2/pkg/util/runtime/runtime.go:252\nk8s.io/apimachinery/pkg/util/runtime.handleError\n\tpkg/mod/k8s.io/apimachinery@v0.36.2/pkg/util/runtime/runtime.go:243\nk8s.io/apimachinery/pkg/util/runtime.HandleErrorWithContext\n\tpkg/mod/k8s.io/apimachinery@v0.36.2/pkg/util/runtime/runtime.go:229\nk8s.io/client-go/tools/cache.DefaultWatchErrorHandler\n\tpkg/mod/k8s.io/client-go@v0.36.2/tools/cache/reflector.go:227\nk8s.io/client-go/tools/cache.(*Reflector).RunWithContext.func1\n\tpkg/mod/k8s.io/client-go@v0.36.2/tools/cache/reflector.go:430\nk8s.io/apimachinery/pkg/util/wait.loopConditionUntilContext.func1\n\tpkg/mod/k8s.io/apimachinery@v0.36.2/pkg/util/wait/loop.go:53\nk8s.io/apimachinery/pkg/util/wait.loopConditionUntilContext\n\tpkg/mod/k8s.io/apimachinery@v0.36.2/pkg/util/wait/loop.go:54\nk8s.io/apimachinery/pkg/util/wait.DelayFunc.Until\n\tpkg/mod/k8s.io/apimachinery@v0.36.2/pkg/util/wait/delay.go:39\nk8s.io/client-go/tools/cache.(*Reflector).RunWithContext\n\tpkg/mod/k8s.io/client-go@v0.36.2/tools/cache/reflector.go:428\nk8s.io/client-go/tools/cache.(*controller).RunWithContext.(*Group).StartWithContext.func3\n\tpkg/mod/k8s.io/apimachinery@v0.36.2/pkg/util/wait/wait.go:63\nk8s.io/apimachinery/pkg/util/wait.(*Group).Start.func1\n\tpkg/mod/k8s.io/apimachinery@v0.36.2/pkg/util/wait/wait.go:72"}
{"level":"error","ts":"2026-09-13T12:40:42.032375275Z","logger":"controller-runtime.cache.UnhandledError","msg":"Failed to watch","reflector":"pkg/mod/k8s.io/client-go@v0.36.2/tools/cache/reflector.go:343","type":"*v1.PersistentVolumeClaim","error":"failed to list *v1.PersistentVolumeClaim: persistentvolumeclaims is forbidden: User \"system:serviceaccount:cnpg-system:cnpg-cloudnative-pg\" cannot list resource \"persistentvolumeclaims\" in API group \"\" at the cluster scope","stacktrace":"k8s.io/apimachinery/pkg/util/runtime.logError\n\tpkg/mod/k8s.io/apimachinery@v0.36.2/pkg/util/runtime/runtime.go:252\nk8s.io/apimachinery/pkg/util/runtime.handleError\n\tpkg/mod/k8s.io/apimachinery@v0.36.2/pkg/util/runtime/runtime.go:243\nk8s.io/apimachinery/pkg/util/runtime.HandleErrorWithContext\n\tpkg/mod/k8s.io/apimachinery@v0.36.2/pkg/util/runtime/runtime.go:229\nk8s.io/client-go/tools/cache.DefaultWatchErrorHandler\n\tpkg/mod/k8s.io/client-go@v0.36.2/tools/cache/reflector.go:227\nk8s.io/client-go/tools/cache.(*Reflector).RunWithContext.func1\n\tpkg/mod/k8s.io/client-go@v0.36.2/tools/cache/reflector.go:430\nk8s.io/apimachinery/pkg/util/wait.loopConditionUntilContext.func2\n\tpkg/mod/k8s.io/apimachinery@v0.36.2/pkg/util/wait/loop.go:87\nk8s.io/apimachinery/pkg/util/wait.loopConditionUntilContext\n\tpkg/mod/k8s.io/apimachinery@v0.36.2/pkg/util/wait/loop.go:88\nk8s.io/apimachinery/pkg/util/wait.DelayFunc.Until\n\tpkg/mod/k8s.io/apimachinery@v0.36.2/pkg/util/wait/delay.go:39\nk8s.io/client-go/tools/cache.(*Reflector).RunWithContext\n\tpkg/mod/k8s.io/client-go@v0.36.2/tools/cache/reflector.go:428\nk8s.io/client-go/tools/cache.(*controller).RunWithContext.(*Group).StartWithContext.func3\n\tpkg/mod/k8s.io/apimachinery@v0.36.2/pkg/util/wait/wait.go:63\nk8s.io/apimachinery/pkg/util/wait.(*Group).Start.func1\n\tpkg/mod/k8s.io/apimachinery@v0.36.2/pkg/util/wait/wait.go:72"}
{"level":"error","ts":"2026-09-13T12:40:44.374596211Z","logger":"controller-runtime.cache.UnhandledError","msg":"Failed to watch","reflector":"pkg/mod/k8s.io/client-go@v0.36.2/tools/cache/reflector.go:343","type":"*v1.PersistentVolumeClaim","error":"failed to list *v1.PersistentVolumeClaim: persistentvolumeclaims is forbidden: User \"system:serviceaccount:cnpg-system:cnpg-cloudnative-pg\" cannot list resource \"persistentvolumeclaims\" in API group \"\" at the cluster scope","stacktrace":"k8s.io/apimachinery/pkg/util/runtime.logError\n\tpkg/mod/k8s.io/apimachinery@v0.36.2/pkg/util/runtime/runtime.go:252\nk8s.io/apimachinery/pkg/util/runtime.handleError\n\tpkg/mod/k8s.io/apimachinery@v0.36.2/pkg/util/runtime/runtime.go:243\nk8s.io/apimachinery/pkg/util/runtime.HandleErrorWithContext\n\tpkg/mod/k8s.io/apimachinery@v0.36.2/pkg/util/runtime/runtime.go:229\nk8s.io/client-go/tools/cache.DefaultWatchErrorHandler\n\tpkg/mod/k8s.io/client-go@v0.36.2/tools/cache/reflector.go:227\nk8s.io/client-go/tools/cache.(*Reflector).RunWithContext.func1\n\tpkg/mod/k8s.io/client-go@v0.36.2/tools/cache/reflector.go:430\nk8s.io/apimachinery/pkg/util/wait.loopConditionUntilContext.func2\n\tpkg/mod/k8s.io/apimachinery@v0.36.2/pkg/util/wait/loop.go:87\nk8s.io/apimachinery/pkg/util/wait.loopConditionUntilContext\n\tpkg/mod/k8s.io/apimachinery@v0.36.2/pkg/util/wait/loop.go:88\nk8s.io/apimachinery/pkg/util/wait.DelayFunc.Until\n\tpkg/mod/k8s.io/apimachinery@v0.36.2/pkg/util/wait/delay.go:39\nk8s.io/client-go/tools/cache.(*Reflector).RunWithContext\n\tpkg/mod/k8s.io/client-go@v0.36.2/tools/cache/reflector.go:428\nk8s.io/client-go/tools/cache.(*controller).RunWithContext.(*Group).StartWithContext.func3\n\tpkg/mod/k8s.io/apimachinery@v0.36.2/pkg/util/wait/wait.go:63\nk8s.io/apimachinery/pkg/util/wait.(*Group).Start.func1\n\tpkg/mod/k8s.io/apimachinery@v0.36.2/pkg/util/wait/wait.go:72"}
{"level":"error","ts":"2026-09-13T12:40:48.133726393Z","logger":"controller-runtime.cache.UnhandledError","msg":"Failed to watch","reflector":"pkg/mod/k8s.io/client-go@v0.36.2/tools/cache/reflector.go:343","type":"*v1.PersistentVolumeClaim","error":"failed to list *v1.PersistentVolumeClaim: persistentvolumeclaims is forbidden: User \"system:serviceaccount:cnpg-system:cnpg-cloudnative-pg\" cannot list resource \"persistentvolumeclaims\" in API group \"\" at the cluster scope","stacktrace":"k8s.io/apimachinery/pkg/util/runtime.logError\n\tpkg/mod/k8s.io/apimachinery@v0.36.2/pkg/util/runtime/runtime.go:252\nk8s.io/apimachinery/pkg/util/runtime.handleError\n\tpkg/mod/k8s.io/apimachinery@v0.36.2/pkg/util/runtime/runtime.go:243\nk8s.io/apimachinery/pkg/util/runtime.HandleErrorWithContext\n\tpkg/mod/k8s.io/apimachinery@v0.36.2/pkg/util/runtime/runtime.go:229\nk8s.io/client-go/tools/cache.DefaultWatchErrorHandler\n\tpkg/mod/k8s.io/client-go@v0.36.2/tools/cache/reflector.go:227\nk8s.io/client-go/tools/cache.(*Reflector).RunWithContext.func1\n\tpkg/mod/k8s.io/client-go@v0.36.2/tools/cache/reflector.go:430\nk8s.io/apimachinery/pkg/util/wait.loopConditionUntilContext.func2\n\tpkg/mod/k8s.io/apimachinery@v0.36.2/pkg/util/wait/loop.go:87\nk8s.io/apimachinery/pkg/util/wait.loopConditionUntilContext\n\tpkg/mod/k8s.io/apimachinery@v0.36.2/pkg/util/wait/loop.go:88\nk8s.io/apimachinery/pkg/util/wait.DelayFunc.Until\n\tpkg/mod/k8s.io/apimachinery@v0.36.2/pkg/util/wait/delay.go:39\nk8s.io/client-go/tools/cache.(*Reflector).RunWithContext\n\tpkg/mod/k8s.io/client-go@v0.36.2/tools/cache/reflector.go:428\nk8s.io/client-go/tools/cache.(*controller).RunWithContext.(*Group).StartWithContext.func3\n\tpkg/mod/k8s.io/apimachinery@v0.36.2/pkg/util/wait/wait.go:63\nk8s.io/apimachinery/pkg/util/wait.(*Group).Start.func1\n\tpkg/mod/k8s.io/apimachinery@v0.36.2/pkg/util/wait/wait.go:72"}
$ kubectl -n team-a get cluster pg -o jsonpath='{.status.phase}{"\n"}'
Cluster is unrecoverable and needs manual intervention
$ kubectl replace -f /tmp/cnpg-role.json
clusterrole.rbac.authorization.k8s.io/cnpg-cloudnative-pg replaced
$ sleep 30
$ kubectl -n team-a get cluster pg -o jsonpath='{.status.phase}{"\n"}'
Cluster is unrecoverable and needs manual intervention
$ kubectl get clusterrole cnpg-cloudnative-pg -o json | jq '[.rules[] | select(.resources | index("pods"))] | length'
1system:serviceaccount:cnpg-system:cnpg-cloudnative-pg with the verb and the resource it was denied, and the Cluster's phase does not move while the rule is missing. Nothing about the Cluster object is wrong; that is the whole lesson. pg never reaches a healthy phase in this lab anyway, because team-a is default-deny and initdb cannot reach the API server, so read the phase for movement rather than for a value. The last line must print 1: if the restore came back Conflict, the operator is still missing its permission and every later exercise on this page is running against a broken operator.Most operators ship a validating webhook for their own CRs. Take the operator away and writes to those CRs stop, because the webhook has no endpoints and the failure policy is not forgiving.
kubectl get validatingwebhookconfiguration | grep -i cnpg
kubectl get validatingwebhookconfiguration cnpg-validating-webhook-configuration -o jsonpath='{range .webhooks[*]}{.name} {.failurePolicy}{"\n"}{end}'
# mutation runs before validation, so the refusal below names the mutating one
kubectl get mutatingwebhookconfiguration | grep -i cnpg
kubectl get mutatingwebhookconfiguration cnpg-mutating-webhook-configuration -o jsonpath='{range .webhooks[*]}{.name} {.failurePolicy}{"\n"}{end}'
kubectl -n cnpg-system scale deploy cnpg-cloudnative-pg --replicas=0
kubectl -n cnpg-system rollout status deploy cnpg-cloudnative-pg --timeout=60s || true
kubectl -n team-a patch cluster pg --type merge -p '{"spec":{"instances":3}}'
kubectl -n cnpg-system scale deploy cnpg-cloudnative-pg --replicas=1
kubectl -n cnpg-system rollout status deploy cnpg-cloudnative-pg --timeout=180s
kubectl -n team-a patch cluster pg --type merge -p '{"spec":{"instances":2}}'outputcaptured 2026-09-13
$ kubectl get validatingwebhookconfiguration | grep -i cnpg
cnpg-validating-webhook-configuration 5 17h
$ kubectl get validatingwebhookconfiguration cnpg-validating-webhook-configuration -o jsonpath='{range .webhooks[*]}{.name} {.failurePolicy}{"\n"}{end}'
vbackup.cnpg.io Fail
vcluster.cnpg.io Fail
vscheduledbackup.cnpg.io Fail
vdatabase.cnpg.io Fail
vpooler.cnpg.io Fail
$ # mutation runs before validation, so the refusal below names the mutating one
$ kubectl get mutatingwebhookconfiguration | grep -i cnpg
cnpg-mutating-webhook-configuration 4 17h
$ kubectl get mutatingwebhookconfiguration cnpg-mutating-webhook-configuration -o jsonpath='{range .webhooks[*]}{.name} {.failurePolicy}{"\n"}{end}'
mbackup.cnpg.io Fail
mcluster.cnpg.io Fail
mdatabase.cnpg.io Fail
mscheduledbackup.cnpg.io Fail
$ kubectl -n cnpg-system scale deploy cnpg-cloudnative-pg --replicas=0
deployment.apps/cnpg-cloudnative-pg scaled
$ kubectl -n cnpg-system rollout status deploy cnpg-cloudnative-pg --timeout=60s || true
deployment "cnpg-cloudnative-pg" successfully rolled out
$ kubectl -n team-a patch cluster pg --type merge -p '{"spec":{"instances":3}}'
Error from server (InternalError): Internal error occurred: failed calling webhook "mcluster.cnpg.io": failed to call webhook: Post "https://cnpg-webhook-service.cnpg-system.svc:443/mutate-postgresql-cnpg-io-v1-cluster?timeout=10s": dial tcp 10.96.57.128:443: connect: connection refused
$ kubectl -n cnpg-system scale deploy cnpg-cloudnative-pg --replicas=1
deployment.apps/cnpg-cloudnative-pg scaled
$ kubectl -n cnpg-system rollout status deploy cnpg-cloudnative-pg --timeout=180s
Waiting for deployment "cnpg-cloudnative-pg" rollout to finish: 0 out of 1 new replicas have been updated...
Waiting for deployment "cnpg-cloudnative-pg" rollout to finish: 0 out of 1 new replicas have been updated...
Waiting for deployment "cnpg-cloudnative-pg" rollout to finish: 0 out of 1 new replicas have been updated...
Waiting for deployment "cnpg-cloudnative-pg" rollout to finish: 0 of 1 updated replicas are available...
deployment "cnpg-cloudnative-pg" successfully rolled out
$ kubectl -n team-a patch cluster pg --type merge -p '{"spec":{"instances":2}}'
cluster.postgresql.cnpg.io/pg patched (no change)Internal error occurred: failed calling webhook "mcluster.cnpg.io" ... connection refused while the operator is down, and goes through once it is back. The refusal names the mutating webhook, not the validating one, because mutation runs first; both configurations carry failurePolicy: Fail, which is what turns a scaled-down operator into a write outage on its own custom resources.metav1.Condition gives every controller the same five fields and the same meaning for status. A real operator's condition types are domain-specific but the fields are not, and reading both together is how you learn to skim any operator's status.
kubectl -n team-a get cluster pg -o jsonpath='{.status.conditions[*].type}{"\n"}'
kubectl -n team-a get cluster pg -o jsonpath='{.status.conditions}' | jq '.[] | {type, status, reason, lastTransitionTime}'
kubectl explain cluster.status.conditions
kubectl explain cluster.status.conditions.reasonoutputcaptured 2026-09-12
$ kubectl -n team-a get cluster pg -o jsonpath='{.status.conditions[*].type}{"\n"}'
Initialized ConsistentSystemID Ready
$ kubectl -n team-a get cluster pg -o jsonpath='{.status.conditions}' | jq '.[] | {type, status, reason, lastTransitionTime}'
{
"type": "Initialized",
"status": "True",
"reason": "BootstrapCompleted",
"lastTransitionTime": "2026-09-13T01:54:45Z"
}
{
"type": "ConsistentSystemID",
"status": "False",
"reason": "NotFound",
"lastTransitionTime": "2026-09-13T01:54:42Z"
}
{
"type": "Ready",
"status": "False",
"reason": "ClusterIsNotReady",
"lastTransitionTime": "2026-09-13T01:54:43Z"
}
$ kubectl explain cluster.status.conditions
GROUP: postgresql.cnpg.io
KIND: Cluster
VERSION: v1
FIELD: conditions <[]Object>
DESCRIPTION:
Conditions for cluster object
Condition contains details for one aspect of the current state of this API
Resource.
FIELDS:
lastTransitionTime <string> -required-
lastTransitionTime is the last time the condition transitioned from one
status to another.
This should be when the underlying condition changed. If that is not known,
then using the time when the API field changed is acceptable.
message <string> -required-
message is a human readable message indicating details about the transition.
This may be an empty string.
observedGeneration <integer>
observedGeneration represents the .metadata.generation that the condition
was set based upon.
For instance, if .metadata.generation is currently 12, but the
.status.conditions[x].observedGeneration is 9, the condition is out of date
with respect to the current state of the instance.
reason <string> -required-
reason contains a programmatic identifier indicating the reason for the
condition's last transition.
Producers of specific condition types may define expected values and
meanings for this field,
and whether the values are considered a guaranteed API.
The value should be a CamelCase string.
This field may not be empty.
status <string> -required-
... 5 more lines
$ kubectl explain cluster.status.conditions.reason
GROUP: postgresql.cnpg.io
KIND: Cluster
VERSION: v1
FIELD: reason <string>
DESCRIPTION:
reason contains a programmatic identifier indicating the reason for the
condition's last transition.
Producers of specific condition types may define expected values and
meanings for this field,
and whether the values are considered a guaranteed API.
The value should be a CamelCase string.
This field may not be empty.
kubectl explain describes, and you can say which of them is the machine-readable one and which is for humans. A condition with no reason is a controller bug, not a state.Self-check
You patch a CR and its status still reports the old problem. What do you check before believing it?
status.observedGeneration against metadata.generation. If they differ, the controller has not processed your edit yet and the status describes the previous spec. If they match and the status is still wrong, now it is a real finding.
A namespace has been Terminating for twenty minutes. Method?
Find what is left (kubectl api-resources --verbs=list --namespaced -o name | xargs -n1 kubectl get -n <ns>), find the finalizer on it, and find whether the owning controller is running. The fix is usually reviving the controller; stripping the finalizer is a last resort that leaks whatever it was cleaning up.
Why is deleting an operator-managed pod not a way to "reset" anything?
Because reconciliation is level-based: the controller observes that a pod is missing and creates one, with no memory of the event. If you need a different outcome, change the spec (or the underlying cause); the loop will faithfully restore whatever the spec says forever.
An operator's CRs apply fine and nothing happens at all. Two checks?
Is the controller running (and did it win leader election)? And does its RBAC cover your namespace and your kind? A Forbidden in the controller's logs is the classic silent failure. Only after both: is the CR written in a way the controller ignores (a selector, a class field, a paused annotation)?
Name the five ways an operator integrates with the rest of a platform.
Its CRs are delivered by GitOps; its metrics are scraped and alerted on; its connection Secrets are consumed by apps (or synced by ESO); its CRs are subject to policy at admission; and its API is fronted by a thinner developer-facing abstraction (Crossplane/kro/portal template). That list is the "integration" half of the competency.
An operator Deployment has three replicas and only one does work. How does it decide, and how do you see which?
Leader election on a Lease object in the operator's namespace (LeaderElectionID as the name; 15 s lease, 10 s renew deadline by default). kubectl -n <ns> get lease <id> -o jsonpath='{.spec.holderIdentity}' names the leading pod. Non-leaders run their caches but skip reconciling, so logs on the wrong pod show nothing.
What does Owns() in a controller builder do, and what does it explain about operator behavior you see in the cluster?
It watches the child kind and, when a child changes, enqueues the owner found through its controller ownerReference. That is why deleting an operator-created Deployment or Secret triggers an immediate recreation, why ownerReferences are the map of what an operator manages, and why the operator needs full RBAC on every owned kind.
Install an operator with OLM v1 pinned to versions below 2.0 from channel stable. Which object, and what does upgradeConstraintPolicy: SelfCertified change?
A ClusterExtension with spec.namespace, spec.serviceAccount.name (an SA you created with the bundle's RBAC), and spec.source.catalog: { packageName, channels: [stable], version: "<2.0" }. The default CatalogProvided allows only upgrades the catalog's graph declares; SelfCertified allows any version including downgrades, and the CRD upgrade safety preflight still runs unless disabled.
Why should a condition be named Ready or Degraded rather than Deploying, and what is polarity?
Condition types describe an observed state (adjective or past participle), and a transition is expressed by toggling status, so Reconciling=True rather than a Deploying type. Polarity is whether True is the normal state (Ready) or the abnormal one (Degraded, Stalled); a consumer must know it per type, which is why a summary Ready condition exists.
Docs to know your way around
- kubernetes.io: Operator pattern; Controllers; Owners and Dependents (cascade deletion); Finalizers.
- github.com/kubernetes/community: the API conventions doc's conditions section, for what Reason/Message actually promise.
- cloudnative-pg.io: the Cluster API reference, mostly to practice navigating a big operator's docs quickly.
- Offline:
kubectl explain cluster.status.conditions,kubectl get <res> -o jsonpath='{.metadata.ownerReferences}',kubectl api-resources --api-group=<g>. - book.kubebuilder.io: Watching Resources, Good Practices; pkg.go.dev controller-runtime builder and reconcile: For/Owns/Watches, predicates, Result semantics, leader election options.
- operator-framework.github.io/operator-controller (ClusterExtension, upgrade support, CRD upgrade safety) and sdk.operatorframework.io Operator Capability Levels: the v1 API fields, the two upgrade policies, and the level definitions with their guiding questions.
- github.com/kubernetes/community, api-conventions.md "Typical status properties": the metav1.Condition schema, polarity, and naming rules quoted above.