This is the section that pays back the most study time. Incident tasks are where exam time goes, and make break exists because this skill comes from practice, not reading. The concepts fit on half a page; the rest is reps.
make core obs secOrientation
Under a clock, the difference between five minutes and twenty is not knowledge; it is having a method you follow instead of improvising. Learn the method, then burn it in with drills until it runs without conscious effort.
The method
- Blast radius first. What is broken and what still works.
kubectl get pods -A | grep -v Running,make status, the top-level dashboards. One namespace or all? Data plane or control plane? This bounds everything that follows and stops you debugging a symptom two layers from the cause. - Timeline second. What changed.
kubectl get events -A --sort-by=.lastTimestamp | tail -30, recent syncs in Argo CD,helm list -Afor fresh revisions,kubectl rollout history. Most incidents turn out to be a recent deployment. - Descend one resource at a time. describe → events → logs → previous logs (
--previous, for crash loops) → exec/debug. Resist skipping levels; the fast-looking jump is where wrong theories come from. - Fix the cause, then prove recovery with the same signal that showed the failure. If you found it via a failing curl, the incident ends with that curl succeeding, not with a pod showing Running.
| Symptom | Layer | First command |
|---|---|---|
| Pending | scheduling: resources, taints, PVC, quota, affinity | kubectl describe pod (events at the bottom) |
| ImagePullBackOff / ErrImagePull | registry, tag, pull secret, network | kubectl describe pod |
| CrashLoopBackOff | the app itself | kubectl logs --previous |
| OOMKilled | memory limit (1.2) | …state.terminated.reason / lastState |
| CreateContainerConfigError | missing ConfigMap/Secret key | kubectl describe pod |
| Running but broken | NetworkPolicy, DNS, config, RBAC | Hubble, auth can-i, app logs |
| Terminating forever | finalizer, dead controller (3.3) | kubectl get -o yaml | grep finalizers |
| Forbidden in a controller log | RBAC, not the workload | kubectl auth can-i --as=system:serviceaccount:… |
| Everything failing to admit | a webhook backend is down (5.2) | kubectl get validatingwebhookconfigurations |
"Running but broken" is the hardest row: it is where kubectl get pods stops helping and network, identity and config tooling starts.
Node-level moves
When the fault is the node rather than the workload: kubectl cordon <node> stops new pods landing there, kubectl drain <node> --ignore-daemonsets --delete-emptydir-data evicts the rest (respecting PodDisruptionBudgets, which is why it hangs; see section 1.2), and kubectl uncordon puts it back. Draining is also the safe way to test whether a workload survives losing a node.
Two tools people forget exist
kubectl debug -it <pod> --image=busybox:1.37 --target=<container> # ephemeral container into a distroless pod
kubectl debug node/<node> -it --image=busybox:1.37 # a node shell without SSHoutputcaptured 2026-08-26
$ kubectl -n flux-demo debug -it podinfo-8d8f7547d-tgmmb --image=busybox:1.37 --target=podinfo -- ps # ephemeral container into a distroless pod
Targeting container "podinfo". If you don't see processes from this container it may be because the container runtime doesn't support this feature.
Defaulting debug container name to debugger-jt4fk.
All commands and output from this session will be recorded in container logs, including credentials and sensitive information passed through the command prompt.
If you don't see a command prompt, try pressing enter.
PID USER TIME COMMAND
1 100 0:00 ./podinfo --port=9898 --prefix=/ --cert-path=/data/cert --port-metrics=9797 --grpc-port=9999 --grpc-service-name=podinfo --level=info --random-delay=false --random-error=false
8698 root 0:00 ps
$ kubectl debug node/cnpe-worker -it --image=busybox:1.37 -- chroot /host sh -c 'hostname && uptime' # a node shell without SSH
Creating debugging pod node-debugger-cnpe-worker-ztz6r with container debugger on node cnpe-worker.
cnpe-worker
03:32:46 up 4 days, 7:22, 0 users, load average: 3.45, 3.65, 7.56Both exist in the exam environment. The first is the answer to "the image has no shell"; the second to "I need to look at the node and there is no SSH".
Seven minutes per task is the lab's clock because it is roughly the exam's. If you are four minutes in with no theory that fits the evidence, stop and re-run step 1: widen, do not deepen. And write your diagnosis down before you fix anything; a written theory is falsifiable, a mental one drifts to match whatever you just tried.
The extended signal map
The table in the method panel covers the workload layer. Two more tables finish the map: the exact reasons and codes a pod shows, and the cluster-level incidents whose first symptom is often a workload one.
| What you see | Means | Confirm with |
|---|---|---|
Pending, event 0/3 nodes are available: 3 Insufficient cpu | requests exceed free capacity | kubectl describe node Allocated resources; lower requests or add nodes |
| ... had untolerated taint {...} | taint without toleration | node taints vs pod tolerations |
| ... didn't match Pod's node affinity/selector | nodeSelector or affinity impossible | node labels |
| pod has unbound immediate PersistentVolumeClaims | PVC Pending: no StorageClass, no capacity, wrong access mode | kubectl describe pvc events |
| exceeded quota: ... requested: ..., used: ..., limited: ... | ResourceQuota; the pod is never created, the ReplicaSet event says so | kubectl describe quota -n ns; kubectl describe rs |
ImagePullBackOff with unauthorized / manifest unknown / no such host | pull secret, tag, registry DNS or egress | the exact registry error in describe pod; imagePullSecrets on pod or SA |
| CreateContainerConfigError | a referenced ConfigMap/Secret or key does not exist | event names the missing object |
| CreateContainerError | runtime refused the container spec (bad command, invalid mount) | event text; containerd logs on the node |
ContainerCreating for minutes, event failed to setup network for sandbox / FailedMount / Multi-Attach error | CNI down, or a volume still attached to the old node | CNI DaemonSet pods; kubectl get volumeattachments |
| CrashLoopBackOff | the container keeps exiting; backoff doubles to a 5 minute cap | kubectl logs --previous, then the exit code below |
| OOMKilled | memory limit; reason on lastState.terminated | container_memory_working_set_bytes vs limit |
| Evicted | node pressure (ephemeral storage, memory) | event message names the resource; node conditions |
Terminating forever | finalizer whose controller is gone, or a stuck preStop/grace period | metadata.finalizers; kubectl delete --grace-period=0 --force as a last resort |
Exit codes, read from lastState.terminated.exitCode: 0 ran to completion (a Job, or a container that should not have exited), 1 application error (read the log), 2 shell misuse, 126 command not executable, 127 command not found (wrong command or image), 137 SIGKILL (OOM if the reason says so, otherwise killed after the grace period), 139 segfault, 143 SIGTERM (a graceful stop; if it recurs, the liveness probe is killing a healthy container).
Cluster-level scenarios
| Scenario | Tells | Lever |
|---|---|---|
node NotReady | condition message Kubelet stopped posting node status; pods on it go Unknown, then are evicted after the eviction timeout (5 minutes by default) | kubectl debug node/ and read /host/var/log/kubelet.log and containerd; disk pressure and certificate expiry are the usual causes |
| DNS | everything Running, apps log no such host; nslookup kubernetes.default from a pod fails | CoreDNS pods and its ConfigMap; a NetworkPolicy that forgot UDP/TCP 53 egress to kube-system |
| certificate expiry | x509: certificate has expired or is not yet valid in kubelet or kubectl output; a webhook whose caBundle no longer matches | kubeadm certs check-expiration and kubeadm certs renew; cert-manager for the webhook |
| etcd full | every write fails with etcdserver: mvcc: database space exceeded; reads still work | etcdctl alarm list, compact, defrag, alarm disarm; raise --quota-backend-bytes; find what filled it (events, huge ConfigMaps) |
| admission webhook down | writes time out with failed calling webhook ... context deadline exceeded, reads work | the backend Deployment; as an emergency, delete or scope the webhook configuration (5.2) |
| API server slow | apiserver_request_duration_seconds p99 high; Throttling request took in client logs | a controller in a hot loop (apiserver_request_total by user agent), or a webhook adding latency to every write |
| image pull auth across a namespace | every new pod ImagePullBackOff with 401 Unauthorized | the pull Secret expired or the ServiceAccount lost its imagePullSecrets |
| control-plane tooling | Argo CD Unknown health with ComparisonError; Flux Ready=False with reconciliation failed; Tekton PipelineRun CouldntGetTask; Crossplane Synced=False | each tool's own condition message names the layer; the fix is in that tool's section |
Before anything else on a "cluster is broken" task: kubectl get nodes, kubectl get pods -A | grep -v Running, kubectl get events -A --sort-by=.lastTimestamp | tail -30, and kubectl get validatingwebhookconfigurations,mutatingwebhookconfigurations. Those four outputs place the incident in one row of the tables above nine times out of ten.
The toolbox and the rollback verbs
kubectl debug, the flags
kubectl debug -it pod/x --image=busybox --target=appadds an ephemeral container sharing the process namespace ofapp; ephemeral containers cannot be removed, only the pod deleted.--copy-to=x-debug --share-processes --set-image=app=busyboxcopies the pod with a changed image or command, for a container that crashes before you can attach;--copy-towith--container app -- shoverrides the command.kubectl debug node/n -it --image=busyboxschedules a pod on the node with the host filesystem at/host; it is not privileged unless--profile=sysadmin. The other profiles aregeneral(the default),baselineandrestricted(pass PSS),netadmin(NET_ADMIN and NET_RAW for tcpdump and ip). Delete thenode-debugger-*pod afterwards.kubectl get --raw /readyz?verboseand/livez?verboselist API server health checks by name (etcd, informers, webhooks);kubectl get componentstatusesis gone, do not reach for it.
Rollback, per layer
| Layer | Restore service | Caveat |
|---|---|---|
| Deployment | kubectl rollout undo deploy/x [--to-revision=N]; kubectl rollout history | GitOps will roll it forward again within a sync interval unless you suspend or fix git |
| Argo CD | argocd app rollback x <id> (from argocd app history), or disable auto-sync first (argocd app set x --sync-policy none) | rollback with auto-sync on is refused; revert the commit for the durable fix |
| Flux | flux suspend kustomization x, then kubectl rollout undo or a git revert; flux resume after | suspended objects show suspended="true" in gotk_resource_info; forgetting to resume is a follow-up incident |
| Helm | helm rollback x <revision> | only if Helm still owns the release; not under Argo CD's rendering |
| Progressive delivery | kubectl argo rollouts abort x then undo; Flagger rolls back on its own when analysis fails | an aborted Rollout stays Degraded until you promote or undo |
| Policy | flip validationActions to Audit or enforcementAction to warn; add a PolicyException | deleting a webhook configuration is the last resort and must be re-created after |
| Crossplane / kro | compositionUpdatePolicy: Manual plus the old compositionRevisionRef; crossplane.io/paused: "true" or kro.run/reconcile: suspended to stop the bleeding | pausing stops drift correction too; note it in the postmortem |
| Node | kubectl cordon, drain --ignore-daemonsets --delete-emptydir-data, uncordon | drain honors PodDisruptionBudgets and hangs on them |
Vocabulary and process
MTTD (detect: alert fired minus impact start), MTTA (acknowledge), MTTR (restore), plus the DORA framing of "failed deployment recovery time" from 4.5. A runbook is the document an alert's runbook_url points at: symptom, likely causes in order, the exact commands, the rollback, and who to escalate to. On-call hygiene that scenario questions reward: one severity taxonomy, a rotation with hand-over notes, every page linked to a runbook, an incident channel per incident, and a blameless postmortem within days that produces action items with owners (the five-line format above is the minimum viable version). Remediation restores service (rollback, scale, fail over, pause the policy); resolution removes the cause (fix the code, the composition, the quota). Say which one you did; graders and colleagues both need to know whether the fix is still open.
Incident tasks describe the symptom, not the cause: "the payments service returns 503 since the last deploy", "new pods in namespace X do not start", "nothing can be created in the cluster". Your first four commands should place it in a table row; your last command must reproduce the original symptom and show it gone. Partial credit usually exists for restoring service even if the root cause fix is incomplete, so remediate first, then resolve.
After the fix
Remediation is half the competency; the process around it is the other half, and it compresses to five lines you can write in five minutes:
- Symptom: what a user or a check observed, with a timestamp.
- Blast radius: who and what was affected, and who was not.
- Root cause: the change or condition that produced it, stated as a mechanism, not a culprit.
- Fix: what you did, including anything you did that turned out to be irrelevant.
- Proof of recovery: the command or signal that confirms it, and the follow-up that prevents recurrence (an alert that would have caught it sooner, a policy, a default).
Blameless, mechanism-focused, and short enough that you will actually write it. Vocabulary that may appear in a scenario: MTTD (detect), MTTA (acknowledge), MTTR (restore), error budget (how much unreliability the SLO permits), toil (manual repetitive work that scales with load, the thing a platform is supposed to remove).
On the exam and in life: first restore service (roll back, scale, fail over, restore the policy), then fix the cause properly. Those are two different actions with two different urgencies; say which one you are doing.
Exercises
One session per few days, not all at once:
make break # random fault from the whole library
make break-answer # only after you have committed to a diagnosisoutputcaptured 2026-08-27
$ make break # random fault from the whole library
==> Resetting the drill (heal previous fault, remove old scenario objects)
⏱ Fault injected. Domain: workload. Target: under 7 minutes (exam pace).
Ticket: "team-a: our 'broken' deployment is not healthy and we cannot see why. It worked an hour ago."
Start here:
kubectl -n team-a get pods
kubectl -n team-a describe pod <pod>
kubectl -n team-a get events --sort-by=.lastTimestamp | tail -20
kubectl -n team-a logs <pod> --previous
Reveal the answer when you're done: make break-answer
Auto-repair and see the evidence: make break-fix
$ make break-answer # only after you have committed to a diagnosis
rbac (domain: workload)
the pod was switched to sa/app-sa, which has no RBAC at all. Nothing in pod status shows it; only 'kubectl auth can-i --as=system:serviceaccount:team-a:app-sa' does. Fix: bind a Role with the needed verbs.Then drill the ones that hurt, by name. The classic workload seven: FAULT=image, probe, resources, rbac, quota, netpol, config. The rbac fault is invisible in pod listings and only surfaces under kubectl auth can-i --as=...; the netpol fault leaves everything Running while DNS is dead (the additive-allow-list lesson from 1.4). Beyond those, DOMAIN=gitops, cicd, apis or security scopes the draw to faults in the platform tooling itself: a frozen Argo CD sync, a suspended Flux Kustomization, a canary analysis querying a dead Prometheus, a Tekton pipeline missing its Task, a crashlooping EventListener, a Crossplane provider stripped of RBAC, a paused XR, a Kyverno Deny policy, a PSS flip. Those are the incidents domains 2, 3 and 5 grade.
make break-answer, and the fix confirmed by the failing behavior now succeeding. make break-fix afterwards shows the evidence trail a systematic diagnosis would have followed; compare it with the path you actually took and note where you diverged.Run FAULT=probe make break, but restrict yourself for the first three minutes to Grafana, Prometheus and Loki only: find the symptom in the namespace dashboard (restarts, readiness), the failing pod in kube_pod_container_status_ready, and the evidence in its logs via Explore. Then finish with kubectl.
After any drill, write five lines: symptom, blast radius, root cause, fix, the check that proves recovery.
Course platforms end this module with a scripted "repair a broken stack" lab; this version is unscripted. Pick a layer and delete something the stack depends on but that is easy to overlook (a ServiceMonitor, the Loki datasource ConfigMap, kyverno's webhook… choose while not thinking about diagnosis). Do something else for an hour, then come back and find it via make validate, whose failing checks are your incident ticket.
make validate returns to PASS with FAIL 0. Self-inflicted incidents with a validation suite are the closest thing to a free exam simulator this repo has.etcd's quota is a hard stop: once it trips the cluster goes read-only for writes and stays that way until you compact, defrag and disarm the alarm. Four commands, in an order that has to be right.
This block lowers the quota on the kind control plane so the failure is reachable, then puts it back. Everything happens inside the control-plane container.
ETCD="kubectl -n kube-system exec etcd-cnpe-control-plane -- etcdctl --endpoints=https://127.0.0.1:2379 --cacert=/etc/kubernetes/pki/etcd/ca.crt --cert=/etc/kubernetes/pki/etcd/server.crt --key=/etc/kubernetes/pki/etcd/server.key"
$ETCD endpoint status -w table
docker exec cnpe-control-plane sed -i 's|- etcd$|- etcd\n - --quota-backend-bytes=16777216|' /etc/kubernetes/manifests/etcd.yaml
sleep 90
kubectl create configmap canary --from-literal=a=1
$ETCD alarm list
docker exec cnpe-control-plane sed -i '/--quota-backend-bytes=16777216/d' /etc/kubernetes/manifests/etcd.yaml
sleep 90
REV=$($ETCD endpoint status --write-out=json | jq -r '.[0].Status.header.revision')
$ETCD compact "$REV"
$ETCD defrag
$ETCD alarm disarm
$ETCD alarm list
kubectl create configmap canary --from-literal=a=1 && kubectl delete configmap canaryoutputcaptured 2026-09-12
$ ETCD="kubectl -n kube-system exec etcd-cnpe-control-plane -- etcdctl --endpoints=https://127.0.0.1:2379 --cacert=/etc/kubernetes/pki/etcd/ca.crt --cert=/etc/kubernetes/pki/etcd/server.crt --key=/etc/kubernetes/pki/etcd/server.key"
$ $ETCD endpoint status -w table
+------------------------+-----------------+---------+-----------------+---------+--------+-----------------------+--------+-----------+------------+-----------+------------+--------------------+--------+--------------------------+-------------------+
| ENDPOINT | ID | VERSION | STORAGE VERSION | DB SIZE | IN USE | PERCENTAGE NOT IN USE | QUOTA | IS LEADER | IS LEARNER | RAFT TERM | RAFT INDEX | RAFT APPLIED INDEX | ERRORS | DOWNGRADE TARGET VERSION | DOWNGRADE ENABLED |
+------------------------+-----------------+---------+-----------------+---------+--------+-----------------------+--------+-----------+------------+-----------+------------+--------------------+--------+--------------------------+-------------------+
| https://127.0.0.1:2379 | 26e8d9c7b9b778d | 3.6.8 | 3.6.0 | 75 MB | 72 MB | 5% | 2.1 GB | true | false | 4 | 125621 | 125621 | | | false |
+------------------------+-----------------+---------+-----------------+---------+--------+-----------------------+--------+-----------+------------+-----------+------------+--------------------+--------+--------------------------+-------------------+
$ docker exec cnpe-control-plane sed -i 's|- etcd$|- etcd\n - --quota-backend-bytes=16777216|' /etc/kubernetes/manifests/etcd.yaml
$ sleep 90
$ kubectl create configmap canary --from-literal=a=1
error: failed to create configmap: etcdserver: mvcc: database space exceeded
$ $ETCD alarm list
memberID:175233138742228877 alarm:NOSPACE
$ docker exec cnpe-control-plane sed -i '/--quota-backend-bytes=16777216/d' /etc/kubernetes/manifests/etcd.yaml
$ sleep 90
$ REV=$($ETCD endpoint status --write-out=json | jq -r '.[0].Status.header.revision')
$ $ETCD compact "$REV"
compacted revision 120134
$ $ETCD defrag
Finished defragmenting etcd member[https://127.0.0.1:2379]. took 451.79565ms
$ $ETCD alarm disarm
memberID:175233138742228877 alarm:NOSPACE
$ $ETCD alarm list
$ kubectl create configmap canary --from-literal=a=1 && kubectl delete configmap canary
configmap/canary created
configmap "canary" deleted from default namespaceetcdserver: mvcc: database space exceeded, alarm list shows NOSPACE, and writes only work again after compact, defrag and disarm in that order.Control-plane certificates expire, usually at the least convenient moment, and the command that tells you lives on the node rather than in the API. Run it once so you know where it is.
docker exec cnpe-control-plane kubeadm certs check-expiration
docker exec cnpe-control-plane sh -c 'ls -la /etc/kubernetes/pki/*.crt | head'
kubectl -n kube-system get cm kubeadm-config -o jsonpath='{.data.ClusterConfiguration}' | head -20outputcaptured 2026-09-12
$ docker exec cnpe-control-plane kubeadm certs check-expiration
[check-expiration] Reading configuration from the "kubeadm-config" ConfigMap in namespace "kube-system"...
[check-expiration] Use 'kubeadm init phase upload-config kubeadm --config your-config-file' to re-upload it.
CERTIFICATE EXPIRES RESIDUAL TIME CERTIFICATE AUTHORITY EXTERNALLY MANAGED
admin.conf Sep 12, 2027 16:00 UTC 364d ca no
apiserver Sep 12, 2027 16:00 UTC 364d ca no
apiserver-etcd-client Sep 12, 2027 16:00 UTC 364d etcd-ca no
apiserver-kubelet-client Sep 12, 2027 16:00 UTC 364d ca no
controller-manager.conf Sep 12, 2027 16:00 UTC 364d ca no
etcd-healthcheck-client Sep 12, 2027 16:00 UTC 364d etcd-ca no
etcd-peer Sep 12, 2027 16:00 UTC 364d etcd-ca no
etcd-server Sep 12, 2027 16:00 UTC 364d etcd-ca no
front-proxy-client Sep 12, 2027 16:00 UTC 364d front-proxy-ca no
scheduler.conf Sep 12, 2027 16:00 UTC 364d ca no
super-admin.conf Sep 12, 2027 16:00 UTC 364d ca no
CERTIFICATE AUTHORITY EXPIRES RESIDUAL TIME EXTERNALLY MANAGED
ca Sep 09, 2036 16:00 UTC 9y no
etcd-ca Sep 09, 2036 16:00 UTC 9y no
front-proxy-ca Sep 09, 2036 16:00 UTC 9y no
$ docker exec cnpe-control-plane sh -c 'ls -la /etc/kubernetes/pki/*.crt | head'
-rw-r--r-- 1 root root 1123 Sep 12 16:00 /etc/kubernetes/pki/apiserver-etcd-client.crt
-rw-r--r-- 1 root root 1131 Sep 12 16:00 /etc/kubernetes/pki/apiserver-kubelet-client.crt
-rw-r--r-- 1 root root 1326 Sep 12 16:00 /etc/kubernetes/pki/apiserver.crt
-rw-r--r-- 1 root root 1107 Sep 12 16:00 /etc/kubernetes/pki/ca.crt
-rw-r--r-- 1 root root 1123 Sep 12 16:00 /etc/kubernetes/pki/front-proxy-ca.crt
-rw-r--r-- 1 root root 1119 Sep 12 16:00 /etc/kubernetes/pki/front-proxy-client.crt
$ kubectl -n kube-system get cm kubeadm-config -o jsonpath='{.data.ClusterConfiguration}' | head -20
apiServer:
certSANs:
- localhost
- 127.0.0.1
extraArgs:
- name: runtime-config
value: ""
- name: audit-log-maxsize
value: "100"
- name: audit-log-path
value: /var/log/kubernetes/audit.log
- name: audit-policy-file
value: /etc/kubernetes/audit/policy.yaml
- name: audit-log-maxage
value: "2"
- name: audit-log-maxbackup
value: "2"
extraVolumes:
- hostPath: /etc/kubernetes/audit
mountPath: /etc/kubernetes/auditThe failure text from inside a pod is the thing you are asked to recognize: a resolver that is still configured and still reachable, and a lookup that fails anyway. Produce it deliberately, then put it back.
kubectl -n kube-system scale deploy coredns --replicas=0
until [ -z "$(kubectl -n kube-system get endpointslices -l kubernetes.io/service-name=kube-dns -o jsonpath='{.items[*].endpoints[*].addresses}')" ]; do sleep 5; done; echo 'kube-dns has no endpoints'
kubectl run dnsdead --image=busybox:1.28 --restart=Never -it --rm --command -- nslookup kubernetes.default || true
kubectl -n kube-system get endpointslices -l kubernetes.io/service-name=kube-dns
kubectl -n kube-system scale deploy coredns --replicas=2
kubectl -n kube-system rollout status deploy coredns --timeout=180s
kubectl run dnsalive --image=busybox:1.28 --restart=Never -it --rm --command -- nslookup kubernetes.defaultoutputcaptured 2026-09-12
$ kubectl -n kube-system scale deploy coredns --replicas=0
deployment.apps/coredns scaled
$ until [ -z "$(kubectl -n kube-system get endpointslices -l kubernetes.io/service-name=kube-dns -o jsonpath='{.items[*].endpoints[*].addresses}')" ]; do sleep 5; done; echo 'kube-dns has no endpoints'
kube-dns has no endpoints
$ kubectl run dnsdead --image=busybox:1.28 --restart=Never -it --rm --command -- nslookup kubernetes.default || true
Server: 10.96.0.10
Address 1: 10.96.0.10
nslookup: can't resolve 'kubernetes.default'
Unable to use a TTY - input is not a terminal or the right kind of file
All commands and output from this session will be recorded in container logs, including credentials and sensitive information passed through the command prompt.
If you don't see a command prompt, try pressing enter.
warning: couldn't attach to pod/dnsdead, falling back to streaming logs: unable to upgrade connection: container dnsdead not found in pod dnsdead_default
Server: 10.96.0.10
Address 1: 10.96.0.10
nslookup: can't resolve 'kubernetes.default'
pod "dnsdead" deleted from default namespace
pod default/dnsdead terminated (Error)
$ kubectl -n kube-system get endpointslices -l kubernetes.io/service-name=kube-dns
NAME ADDRESSTYPE PORTS ENDPOINTS AGE
kube-dns-6x2kw IPv4 <unset> <unset> 8h
$ kubectl -n kube-system scale deploy coredns --replicas=2
deployment.apps/coredns scaled
$ kubectl -n kube-system rollout status deploy coredns --timeout=180s
Waiting for deployment "coredns" rollout to finish: 0 out of 2 new replicas have been updated...
Waiting for deployment "coredns" rollout to finish: 0 out of 2 new replicas have been updated...
Waiting for deployment "coredns" rollout to finish: 0 of 2 updated replicas are available...
Waiting for deployment "coredns" rollout to finish: 1 of 2 updated replicas are available...
deployment "coredns" successfully rolled out
$ kubectl run dnsalive --image=busybox:1.28 --restart=Never -it --rm --command -- nslookup kubernetes.default
Server: 10.96.0.10
Address 1: 10.96.0.10 kube-dns.kube-system.svc.cluster.local
Name: kubernetes.default
Address 1: 10.96.0.1 kubernetes.default.svc.cluster.local
Unable to use a TTY - input is not a terminal or the right kind of file
All commands and output from this session will be recorded in container logs, including credentials and sensitive information passed through the command prompt.
If you don't see a command prompt, try pressing enter.
warning: couldn't attach to pod/dnsalive, falling back to streaming logs: unable to upgrade connection: container dnsalive not found in pod dnsalive_default
Server: 10.96.0.10
Address 1: 10.96.0.10 kube-dns.kube-system.svc.cluster.local
Name: kubernetes.default
Address 1: 10.96.0.1 kubernetes.default.svc.cluster.local
pod "dnsalive" deleted from default namespaceThis is the outage that takes a cluster down without touching the API server: an admission webhook that must be called, with nothing to call. Deleting the Deployment is faster than scaling, and the message is the one you must recognize within seconds.
kubectl get validatingwebhookconfiguration -o custom-columns=NAME:.metadata.name,POLICY:.webhooks[*].failurePolicy | grep -i kyverno
# own the webhook rather than scaling a shared component to zero, and scope it to a throwaway namespace
kubectl create ns wh-drill --dry-run=client -o yaml | kubectl apply -f -
kubectl apply -f - <<'EOF'
apiVersion: admissionregistration.k8s.io/v1
kind: ValidatingWebhookConfiguration
metadata: { name: no-backend }
webhooks:
- name: no-backend.example.com
admissionReviewVersions: ["v1"]
sideEffects: None
failurePolicy: Fail
timeoutSeconds: 5
namespaceSelector:
matchLabels: { kubernetes.io/metadata.name: wh-drill }
clientConfig:
service: { name: nope, namespace: wh-drill, path: /validate, port: 443 }
rules:
- { apiGroups: [""], apiVersions: ["v1"], operations: ["CREATE"], resources: ["configmaps"], scope: Namespaced }
EOF
kubectl -n wh-drill create configmap blocked-cm --from-literal=a=1
kubectl get validatingwebhookconfiguration no-backend -o jsonpath='{.webhooks[0].failurePolicy} {.webhooks[0].timeoutSeconds}{"\n"}'
kubectl delete validatingwebhookconfiguration no-backend
kubectl -n wh-drill create configmap blocked-cm --from-literal=a=1
kubectl delete ns wh-drilloutputcaptured 2026-09-13
$ kubectl get validatingwebhookconfiguration -o custom-columns=NAME:.metadata.name,POLICY:.webhooks[*].failurePolicy | grep -i kyverno
kyverno-cel-exception-validating-webhook-cfg Fail
kyverno-cleanup-validating-webhook-cfg Fail
kyverno-exception-validating-webhook-cfg Fail
kyverno-global-context-validating-webhook-cfg Fail
kyverno-policy-validating-webhook-cfg Fail
kyverno-resource-validating-webhook-cfg Fail
kyverno-ttl-validating-webhook-cfg Ignore
$ # own the webhook rather than scaling a shared component to zero, and scope it to a throwaway namespace
$ kubectl create ns wh-drill --dry-run=client -o yaml | kubectl apply -f -
namespace/wh-drill created
$ kubectl apply -f - <<'EOF'
apiVersion: admissionregistration.k8s.io/v1
kind: ValidatingWebhookConfiguration
metadata: { name: no-backend }
webhooks:
- name: no-backend.example.com
admissionReviewVersions: ["v1"]
sideEffects: None
failurePolicy: Fail
timeoutSeconds: 5
namespaceSelector:
matchLabels: { kubernetes.io/metadata.name: wh-drill }
clientConfig:
service: { name: nope, namespace: wh-drill, path: /validate, port: 443 }
rules:
- { apiGroups: [""], apiVersions: ["v1"], operations: ["CREATE"], resources: ["configmaps"], scope: Namespaced }
EOF
validatingwebhookconfiguration.admissionregistration.k8s.io/no-backend created
$ kubectl -n wh-drill create configmap blocked-cm --from-literal=a=1
error: failed to create configmap: Internal error occurred: failed calling webhook "no-backend.example.com": failed to call webhook: Post "https://nope.wh-drill.svc:443/validate?timeout=5s": service "nope" not found
$ kubectl get validatingwebhookconfiguration no-backend -o jsonpath='{.webhooks[0].failurePolicy} {.webhooks[0].timeoutSeconds}{"\n"}'
Fail 5
$ kubectl delete validatingwebhookconfiguration no-backend
validatingwebhookconfiguration.admissionregistration.k8s.io "no-backend" deleted
$ kubectl -n wh-drill create configmap blocked-cm --from-literal=a=1
configmap/blocked-cm created
$ kubectl delete ns wh-drill
namespace "wh-drill" deletedInternal error occurred: failed calling webhook "no-backend.example.com", naming the service that does not exist, and the same create succeeds once the webhook configuration is deleted. The two fields that decide whether this is an outage or a gap are failurePolicy and the match scope, here namespaceSelector plus rules. Under a clock the fix is to delete the webhook configuration and put it back afterwards. Scaling Kyverno to zero does not reproduce this: both creates succeed, so a down policy engine is not by itself a cluster outage on this install.kubectl debug node gives you a pod on the node with a chosen level of access. Knowing which profile grants what saves you from asking for privileged when you needed a network namespace.
kubectl debug node/cnpe-worker -it --image=busybox:1.36 --profile=sysadmin -- chroot /host sh -c 'ls /etc/kubernetes; cat /var/lib/kubelet/config.yaml | head -10'
kubectl debug node/cnpe-worker -it --image=nicolaka/netshoot:latest --profile=netadmin -- ip link
kubectl get pods -o name | grep node-debugger
kubectl delete pod -l app.kubernetes.io/managed-by=kubectl-debug --ignore-not-found
kubectl get pods -o name | grep node-debugger | xargs -r kubectl deleteoutputcaptured 2026-09-12
$ kubectl debug node/cnpe-worker -it --image=busybox:1.36 --profile=sysadmin -- chroot /host sh -c 'ls /etc/kubernetes; cat /var/lib/kubelet/config.yaml | head -10'
Creating debugging pod node-debugger-cnpe-worker-xf265 with container debugger on node cnpe-worker.
Unable to use a TTY - input is not a terminal or the right kind of file
All commands and output from this session will be recorded in container logs, including credentials and sensitive information passed through the command prompt.
If you don't see a command prompt, try pressing enter.
warning: couldn't attach to pod/node-debugger-cnpe-worker-xf265, falling back to streaming logs: unable to upgrade connection: container debugger not found in pod node-debugger-cnpe-worker-xf265_default
kubelet.conf manifests pki
apiVersion: kubelet.config.k8s.io/v1beta1
authentication:
anonymous:
enabled: false
webhook:
cacheTTL: 0s
enabled: true
x509:
clientCAFile: /etc/kubernetes/pki/ca.crt
authorization:
$ kubectl debug node/cnpe-worker -it --image=nicolaka/netshoot:latest --profile=netadmin -- ip link
Creating debugging pod node-debugger-cnpe-worker-shgmz with container debugger on node cnpe-worker.
Unable to use a TTY - input is not a terminal or the right kind of file
All commands and output from this session will be recorded in container logs, including credentials and sensitive information passed through the command prompt.
If you don't see a command prompt, try pressing enter.
warning: couldn't attach to pod/node-debugger-cnpe-worker-shgmz, falling back to streaming logs: unable to upgrade connection: rpc error: code = InvalidArgument desc = tty and stderr cannot both be true
1: lo: <LOOPBACK,UP,LOWER_UP> mtu 65536 qdisc noqueue state UNKNOWN mode DEFAULT group default qlen 1000
link/loopback 00:00:00:00:00:00 brd 00:00:00:00:00:00
2: eth0@if22: <BROADCAST,MULTICAST,UP,LOWER_UP> mtu 1500 qdisc noqueue state UP mode DEFAULT group default
link/ether f6:bf:03:f1:53:11 brd ff:ff:ff:ff:ff:ff link-netnsid 0
3: cilium_net@cilium_host: <BROADCAST,MULTICAST,NOARP,UP,LOWER_UP> mtu 1500 qdisc noqueue state UP mode DEFAULT group default
link/ether fe:c6:be:4f:67:f8 brd ff:ff:ff:ff:ff:ff
4: cilium_host@cilium_net: <BROADCAST,MULTICAST,NOARP,UP,LOWER_UP> mtu 1500 qdisc noqueue state UP mode DEFAULT group default qlen 1000
link/ether ea:a5:4a:b9:92:c8 brd ff:ff:ff:ff:ff:ff
5: cilium_vxlan: <BROADCAST,MULTICAST,UP,LOWER_UP> mtu 1500 qdisc noqueue state UNKNOWN mode DEFAULT group default
link/ether 42:d4:ee:8f:5c:90 brd ff:ff:ff:ff:ff:ff
7: lxc_health@if6: <BROADCAST,MULTICAST,UP,LOWER_UP> mtu 1500 qdisc noqueue state UP mode DEFAULT group default qlen 1000
link/ether 9a:27:33:7a:69:07 brd ff:ff:ff:ff:ff:ff link-netnsid 1
11: lxcd355da451c82@if10: <BROADCAST,MULTICAST,UP,LOWER_UP> mtu 1500 qdisc noqueue state UP mode DEFAULT group default qlen 1000
link/ether 5a:10:28:1c:33:0a brd ff:ff:ff:ff:ff:ff link-netnsid 3
15: lxcc6de2304ea73@if14: <BROADCAST,MULTICAST,UP,LOWER_UP> mtu 1500 qdisc noqueue state UP mode DEFAULT group default qlen 1000
link/ether fa:37:7d:9f:b9:84 brd ff:ff:ff:ff:ff:ff link-netnsid 4
17: lxce8009643b771@if16: <BROADCAST,MULTICAST,UP,LOWER_UP> mtu 1500 qdisc noqueue state UP mode DEFAULT group default qlen 1000
link/ether 02:2e:2b:d8:d7:87 brd ff:ff:ff:ff:ff:ff link-netnsid 5
19: lxc39a31623cd62@if18: <BROADCAST,MULTICAST,UP,LOWER_UP> mtu 1500 qdisc noqueue state UP mode DEFAULT group default qlen 1000
link/ether 52:2a:68:71:5a:c5 brd ff:ff:ff:ff:ff:ff link-netnsid 6
23: lxc3c34b6d00aee@if22: <BROADCAST,MULTICAST,UP,LOWER_UP> mtu 1500 qdisc noqueue state UP mode DEFAULT group default qlen 1000
link/ether 6e:ce:31:d4:84:18 brd ff:ff:ff:ff:ff:ff link-netnsid 2
25: lxc462eb9990672@if24: <BROADCAST,MULTICAST,UP,LOWER_UP> mtu 1500 qdisc noqueue state UP mode DEFAULT group default qlen 1000
link/ether 1e:b9:db:d4:fd:3a brd ff:ff:ff:ff:ff:ff link-netnsid 8
27: lxc9da17d5cb210@if26: <BROADCAST,MULTICAST,UP,LOWER_UP> mtu 1500 qdisc noqueue state UP mode DEFAULT group default qlen 1000
link/ether 56:d5:63:4a:8d:e3 brd ff:ff:ff:ff:ff:ff link-netnsid 9
33: lxc6b7a1dc9d767@if32: <BROADCAST,MULTICAST,UP,LOWER_UP> mtu 1500 qdisc noqueue state UP mode DEFAULT group default qlen 1000
link/ether 12:eb:27:a0:a3:9d brd ff:ff:ff:ff:ff:ff link-netnsid 12
35: lxc9281f9de561b@if34: <BROADCAST,MULTICAST,UP,LOWER_UP> mtu 1500 qdisc noqueue state UP mode DEFAULT group default qlen 1000
link/ether ba:90:29:b4:b9:fe brd ff:ff:ff:ff:ff:ff link-netnsid 13
37: lxc957282906d6f@if36: <BROADCAST,MULTICAST,UP,LOWER_UP> mtu 1500 qdisc noqueue state UP mode DEFAULT group default qlen 1000
link/ether be:61:c4:6d:bc:19 brd ff:ff:ff:ff:ff:ff link-netnsid 14
41: lxc54dcffdf2a27@if40: <BROADCAST,MULTICAST,UP,LOWER_UP> mtu 1500 qdisc noqueue state UP mode DEFAULT group default qlen 1000
link/ether 32:20:ce:9c:c8:c6 brd ff:ff:ff:ff:ff:ff link-netnsid 16
43: lxcee11a1eb7c67@if42: <BROADCAST,MULTICAST,UP,LOWER_UP> mtu 1500 qdisc noqueue state UP mode DEFAULT group default qlen 1000
... 55 more lines
$ kubectl get pods -o name | grep node-debugger
pod/node-debugger-cnpe-worker-shgmz
pod/node-debugger-cnpe-worker-xf265
$ kubectl delete pod -l app.kubernetes.io/managed-by=kubectl-debug --ignore-not-found
pod "node-debugger-cnpe-worker-shgmz" deleted from default namespace
pod "node-debugger-cnpe-worker-xf265" deleted from default namespace
$ kubectl get pods -o name | grep node-debugger | xargs -r kubectl deleteA liveness probe on the wrong port restarts a container that is working perfectly. The evidence is an exit code and an event, and the diagnosis has to separate "the app is broken" from "we are killing it".
kubectl -n default create deployment probed --image=ghcr.io/stefanprodan/podinfo:6.7.1
kubectl -n default patch deployment probed --type merge -p '{"spec":{"template":{"spec":{"containers":[{"name":"podinfo","image":"ghcr.io/stefanprodan/podinfo:6.7.1","livenessProbe":{"httpGet":{"path":"/healthz","port":9999},"initialDelaySeconds":5,"periodSeconds":5,"failureThreshold":2}}]}}}}'
sleep 90
kubectl -n default get pods -l app=probed
POD=$(kubectl -n default get pod -l app=probed -o jsonpath='{.items[0].metadata.name}')
kubectl -n default get pod "$POD" -o jsonpath='{.status.containerStatuses[0].lastState}' | jq
kubectl -n default describe pod "$POD" | sed -n '/Events/,$p'
kubectl -n default delete deployment probedoutputcaptured 2026-09-12
$ kubectl -n default create deployment probed --image=ghcr.io/stefanprodan/podinfo:6.7.1
deployment.apps/probed created
$ kubectl -n default patch deployment probed --type merge -p '{"spec":{"template":{"spec":{"containers":[{"name":"podinfo","image":"ghcr.io/stefanprodan/podinfo:6.7.1","livenessProbe":{"httpGet":{"path":"/healthz","port":9999},"initialDelaySeconds":5,"periodSeconds":5,"failureThreshold":2}}]}}}}'
deployment.apps/probed patched
$ sleep 90
$ kubectl -n default get pods -l app=probed
NAME READY STATUS RESTARTS AGE
probed-5fcd4554f7-8qzz5 1/1 Running 4 (26s ago) 90s
$ POD=$(kubectl -n default get pod -l app=probed -o jsonpath='{.items[0].metadata.name}')
$ kubectl -n default get pod "$POD" -o jsonpath='{.status.containerStatuses[0].lastState}' | jq
{
"terminated": {
"containerID": "containerd://9d61ffcbb66c2a6e74f31775c38f992f2a9b6769a2812941730d973b49fdf5d0",
"exitCode": 0,
"finishedAt": "2026-09-13T03:44:00Z",
"reason": "Completed",
"startedAt": "2026-09-13T03:43:46Z"
}
}
$ kubectl -n default describe pod "$POD" | sed -n '/Events/,$p'
Events:
Type Reason Age From Message
---- ------ ---- ---- -------
Warning PolicyViolation 90s kyverno-admission policy require-resource-requests/ fail: every container must set cpu and memory requests
Normal Scheduled 90s default-scheduler Successfully assigned default/probed-5fcd4554f7-8qzz5 to cnpe-worker
Warning PolicyViolation 59s kyverno-scan policy require-resource-requests/ fail: every container must set cpu and memory requests
Warning Unhealthy 30s (x8 over 80s) kubelet spec.containers{podinfo}: Liveness probe failed: Get "http://10.244.1.218:9999/healthz": dial tcp 10.244.1.218:9999: connect: connection refused
Normal Killing 30s (x4 over 75s) kubelet spec.containers{podinfo}: Container podinfo failed liveness probe, will be restarted
Warning BackOff 26s (x2 over 27s) kubelet spec.containers{podinfo}: Back-off restarting failed container podinfo in pod probed-5fcd4554f7-8qzz5_default(51a2ebe0-4fb8-4789-8b5a-595055e4ddd5)
Normal Pulled 2s (x5 over 88s) kubelet spec.containers{podinfo}: Container image "ghcr.io/stefanprodan/podinfo:6.7.1" already present on machine and can be accessed by the pod
Normal Created 1s (x5 over 88s) kubelet spec.containers{podinfo}: Container created
Normal Started 1s (x5 over 87s) kubelet spec.containers{podinfo}: Container started
$ kubectl -n default delete deployment probed
deployment.apps "probed" deleted from default namespaceLiveness probe failed with connection refused on port 9999 followed by Killing and BackOff, and lastState.terminated records the kill. The exit code is 0 with reason Completed, not 143, because podinfo handles SIGTERM and exits cleanly; 143 is what you see from a process that ignores it. The probe event plus the restart count is the evidence that separates "the app is broken" from "we are killing it", and the exit code alone does not./readyz?verbose lists every readiness check the API server runs and which one is failing. It is the fastest first command on a control plane that feels wrong, and it needs no tooling at all.
kubectl get --raw '/readyz?verbose' | tail -30
kubectl get --raw '/livez?verbose' | tail -10
kubectl get --raw '/healthz/etcd'
kubectl get componentstatuses 2>/dev/null || echo 'componentstatuses is deprecated and may be gone'outputcaptured 2026-09-12
$ kubectl get --raw '/readyz?verbose' | tail -30
[+]poststarthook/priority-and-fairness-filter ok
[+]poststarthook/storage-object-count-tracker-hook ok
[+]poststarthook/start-apiextensions-informers ok
[+]poststarthook/start-apiextensions-controllers ok
[+]poststarthook/crd-informer-synced ok
[+]poststarthook/start-system-namespaces-controller ok
[+]poststarthook/peer-endpoint-reconciler-controller ok
[+]poststarthook/start-cluster-authentication-info-controller ok
[+]poststarthook/start-kube-apiserver-identity-lease-controller ok
[+]poststarthook/start-kube-apiserver-identity-lease-garbage-collector ok
[+]poststarthook/storage-readiness ok
[+]poststarthook/start-legacy-token-tracking-controller ok
[+]poststarthook/start-service-ip-repair-controllers ok
[+]poststarthook/rbac/bootstrap-roles ok
[+]poststarthook/scheduling/bootstrap-system-priority-classes ok
[+]poststarthook/priority-and-fairness-config-producer ok
[+]poststarthook/bootstrap-controller ok
[+]poststarthook/start-kubernetes-service-cidr-controller ok
[+]poststarthook/aggregator-reload-proxy-client-cert ok
[+]poststarthook/start-kube-aggregator-informers ok
[+]poststarthook/apiservice-status-local-available-controller ok
[+]poststarthook/apiservice-status-remote-available-controller ok
[+]poststarthook/apiservice-registration-controller ok
[+]poststarthook/apiservice-discovery-controller ok
[+]poststarthook/kube-apiserver-autoregistration ok
[+]autoregister-completion ok
[+]poststarthook/apiservice-openapi-controller ok
[+]poststarthook/apiservice-openapiv3-controller ok
[+]shutdown ok
readyz check passed
$ kubectl get --raw '/livez?verbose' | tail -10
[+]poststarthook/start-kube-aggregator-informers ok
[+]poststarthook/apiservice-status-local-available-controller ok
[+]poststarthook/apiservice-status-remote-available-controller ok
[+]poststarthook/apiservice-registration-controller ok
[+]poststarthook/apiservice-discovery-controller ok
[+]poststarthook/kube-apiserver-autoregistration ok
[+]autoregister-completion ok
[+]poststarthook/apiservice-openapi-controller ok
[+]poststarthook/apiservice-openapiv3-controller ok
livez check passed
$ kubectl get --raw '/healthz/etcd'
ok
$ kubectl get componentstatuses 2>/dev/null || echo 'componentstatuses is deprecated and may be gone'
NAME STATUS MESSAGE ERROR
scheduler Healthy ok
controller-manager Healthy ok
etcd-0 Healthy ok Self-check
Every pod in a namespace is Running and the app returns nothing. First three moves?
Make something try (a curl or nslookup from inside the namespace) so there is evidence to read; check endpoints/readiness for the Service; check policy verdicts (Hubble DROPPED) and DNS. "Running" says the kubelet is happy, nothing more.
A pod is CrashLoopBackOff and kubectl logs shows nothing useful.
kubectl logs --previous: you are reading the current, not-yet-started container. Then describe for the exit code and reason (137 with OOMKilled points at memory, 1 with app output points at config), then the previous container's lastState.terminated.
Nothing can be created cluster-wide; every apply times out with a webhook error. What happened and what is the lever?
An admission webhook's backend is down and its failurePolicy is Fail, so the API server refuses writes it cannot validate. Restore the backend, or (as a deliberate emergency step) remove/scope the webhook configuration. That is why failurePolicy is a governance decision, not a default to accept blindly (5.2).
You have four minutes left and no theory that fits the evidence. What do you do?
Go back to step 1 and widen: re-check blast radius and timeline rather than digging deeper into the current object. Most stuck diagnoses are the result of over-committing to the first plausible layer. On the exam, flag it and move on; a stuck task costs you two easy ones.
What proves an incident is over?
The same signal that showed the failure, now succeeding: the failing curl returns 200, the alert resolves, the validation check passes. A pod showing Running proves only that a container started.
Every write to the cluster fails with etcdserver: mvcc: database space exceeded but reads work. What happened and what is the order of operations?
etcd hit its backend quota (2 GiB by default) and raised the NOSPACE alarm, which turns the store read-only. Compact old revisions, defragment each member, then etcdctl alarm disarm; raise --quota-backend-bytes if the data is legitimate. Then find the producer (event storms, oversized ConfigMaps or Secrets, a controller writing status in a loop) or it recurs.
A container exits with code 143 every 30 seconds and the app logs look healthy. Diagnosis?
143 is SIGTERM: something is asking it to stop gracefully. On a fixed cadence that is almost always a failing liveness probe (wrong port or path, or a timeout shorter than the app's startup). Read kubectl describe pod for Liveness probe failed events, and fix the probe or add a startup probe. 137 would have pointed at OOM or a forced kill instead.
You roll a Deployment back with kubectl rollout undo and two minutes later it is broken again. Why, and what should you have done?
A GitOps controller re-applied the desired state from git. Either disable or suspend first (argocd app set x --sync-policy none, flux suspend kustomization x) and undo as a stopgap, or go straight to the durable fix: revert the commit and let the controller roll back for you. Say in the incident notes that the suspension must be lifted.
Which kubectl debug invocation gives you tcpdump on a node without full root, and why does the default node debug pod fail at chroot /host?
kubectl debug node/<n> -it --image=<image with tcpdump> --profile=netadmin adds NET_ADMIN and NET_RAW. The default general profile mounts the host filesystem at /host but is not privileged, so operations that need root on the host (chroot, reading some /proc entries) fail; --profile=sysadmin is the privileged one.
Docs to know your way around
- kubernetes.io: Debug Pods / Debug Running Pods (the ephemeral containers page), Troubleshooting Applications, and the Debug Cluster pages.
- sre.google: the postmortem culture chapter, for the vocabulary a scenario question may borrow.
- Your own postmortem notes: by the third drill they are the best incident doc you own.
- kubernetes.io/docs/tasks/debug/debug-application/debug-running-pod (ephemeral containers,
--copy-to,--share-processes) and /debug-cluster/kubectl-node-debug (the/hostmount and profiles). - kubernetes.io/docs/concepts/workloads/pods/pod-lifecycle: container states, the CrashLoopBackOff backoff rules and
restartPolicy; /docs/tasks/debug/debug-cluster for node and control-plane logs. - etcd.io/docs/latest/op-guide/maintenance: compaction, defragmentation and the space quota alarm sequence.
make down-obs