The lab has no storage layer to install because kind ships one: the standard StorageClass backed by the local-path provisioner. That is enough to exercise every concept the exam touches, and its limitations are themselves instructive.
make upmake apiOrientation
Storage tasks on a performance exam are rarely "install a CSI driver". They are: this claim is Pending, say why; this pod lost its data, say why; make this workload survive a reschedule; grow this volume. All four are answered from four fields you can recite.
The graders can only see objects and their status. So the tell for a storage task is almost always a status: PVC Pending, PV Released, pod stuck ContainerCreating with a mount error in events. Learn to map those three states to their causes and you have the competency.
The model: three objects, one relationship
A PersistentVolume is a piece of real storage. A PersistentVolumeClaim is a request for one. A StorageClass is the recipe for making PVs on demand, and its provisioner field names the code that does it. Static provisioning (an admin pre-creates PVs) still exists but dynamic is the assumed default: the PVC references a class, the provisioner makes the PV, the two bind one-to-one and exclusively.
pod ──mounts──▶ PVC ──binds 1:1──▶ PV ──backed by──▶ real disk / host path / cloud volume
│ ▲
└── storageClassName ──▶ StorageClass ──▶ provisioner creates the PV
(binding mode, reclaim policy, expansion, params)
Underneath, CSI is the plugin interface every modern driver implements: a controller component (provision, attach, snapshot) and a node component (mount). You will not be asked to write one, but knowing the split explains error locations: "failed to provision" is a controller-side message, "failed to mount" is node-side, and they point at different logs.
The fields that decide exam tasks
| Field | Values | What it changes |
|---|---|---|
| accessModes | RWO · ROX · RWX · RWOP | A claim binds only to a PV offering what it asks. local-path only does RWO, so an RWX claim here pends forever; recognizing why is the skill. |
| volumeBindingMode | Immediate · WaitForFirstConsumer | WFFC keeps the PVC Pending until a pod uses it, so topology can be considered. The standard class here uses it, so you meet this "problem" immediately, and it is not a problem. |
| persistentVolumeReclaimPolicy | Delete · Retain | Delete throws data away with the claim; Retain keeps the PV in Released, and it will not rebind until someone clears spec.claimRef. Released-but-unusable PVs are a classic troubleshooting scenario. |
| allowVolumeExpansion | true · false | Only classes that set it let you grow a PVC by editing spec.resources.requests.storage. Shrinking is never allowed, anywhere. |
| volumeMode | Filesystem · Block | Block hands the raw device to the container. Databases sometimes want it; nothing else does. |
RWX means "this PV supports many nodes mounting it read-write", and it is a property of the backing storage, not a wish you can express. Nothing in Kubernetes enforces that two pods writing an RWO volume from the same node behave sensibly, and requesting RWX does not turn a local disk into a shared filesystem. RWOP (ReadWriteOncePod) is the strict one: exactly one pod, enforced by the scheduler (the kubelet backstops it at mount), which is how you stop two replicas corrupting a single-writer database.
Lifecycle and the states you will be asked to explain
| Symptom | Usual cause | Check |
|---|---|---|
| PVC Pending, no events | WaitForFirstConsumer, no pod yet | normal: schedule a consumer |
| PVC Pending, provisioner events | unsupported access mode, no capacity, bad class name | kubectl describe pvc |
| Pod ContainerCreating forever | mount/attach failure, node-side | kubectl describe pod → events |
| PV Released, not rebinding | Retain policy leaves claimRef populated | kubectl patch pv … claimRef=null |
| PVC Terminating forever | kubernetes.io/pvc-protection finalizer: a pod still uses it | kubectl get pods -o json | grep claimName |
| Resize stuck | class lacks expansion, or filesystem resize needs a pod restart | status.conditions on the PVC |
Two protection finalizers exist for good reasons and both look like bugs the first time: pvc-protection blocks deleting a claim that a pod mounts, and pv-protection blocks deleting a bound PV. The fix is never to strip the finalizer first; it is to remove the consumer, then let the controller clean up. Stripping finalizers to make an object disappear is the storage equivalent of pulling the disk out.
StatefulSets, where storage semantics become visible
volumeClaimTemplates gives every replica its own PVC, named <template>-<sts>-<ordinal>. Deleting the StatefulSet does not delete those PVCs (unless you set persistentVolumeClaimRetentionPolicy, which newer versions offer), so scaling back up reattaches the old data. That retention is deliberate and it is the entire reason StatefulSets exist rather than "a Deployment with a volume": stable identity, stable storage, ordered rollout.
And the part GitOps cannot do for you: data. "Delete the namespace and the controller rebuilds it" restores manifests, never the contents of a PV. Three things fill that gap. CSI snapshots give point-in-time copies within a cluster. A backup tool (Velero is the common one) snapshots volumes and exports object state for anything that must survive the cluster itself. Database operators bring their own backup story where one exists. Being able to say which of those three you are relying on is the whole DR answer.
Snapshots round out the vocabulary: VolumeSnapshotClass, VolumeSnapshot, VolumeSnapshotContent. They follow the same three-object pattern as class/claim/volume, and they are how a CSI driver exposes point-in-time copies. local-path has no snapshotter, so this lab teaches the nouns and CloudNativePG teaches the backup story instead.
What a platform engineer actually decides
Section 3.1's "APIs as products" idea lands here concretely: StorageClasses are a platform API. You are choosing the menu your tenants order from, and each entry encodes a durability, performance and cost decision they should not have to make.
- Name classes for intent (
fast-ssd,cheap-hdd,shared-rwx), never for the implementation that happens to back them today. - Set exactly one default class (
storageclass.kubernetes.io/is-default-class) and know what it costs, because every PVC without an explicit class silently buys it. If two are marked default, the DefaultStorageClass admission plugin, not the scheduler, picks the most recently created one. That is worse than an error: it is silent, and it changes the next time someone adds a class. - Prefer
WaitForFirstConsumerin any topology-aware environment, or you will provision volumes in zones your pods cannot reach. Deletefor ephemeral tenant workloads,Retainfor anything whose loss ends up in a postmortem.- Cap storage per tenant with quota:
requests.storagefor the total,<class>.storageclass.storage.k8s.io/requests.storageper class. Per-class quota is how you stop everyone ordering the expensive one (section 1.4).
kubectl get sc first, every time: it shows provisioner, reclaim policy, binding mode and expansion in one line, which is four of the five fields above. Then kubectl get pvc -A and look for anything not Bound.
Backup and restore as objects: Velero
When the DR answer is "a backup tool", Velero is the one the exam pool knows, and its model is four CRDs. A BackupStorageLocation names the object store (bucket, prefix, provider plugin); a VolumeSnapshotLocation names where disk snapshots go; a Backup selects what to copy (includedNamespaces, includedResources, labelSelector, ttl, default 30 days) and a Schedule is a Backup template with a cron expression, producing backups named <schedule>-<timestamp>. Volumes are copied either by CSI snapshot (the snapshot is retained only for the life of the backup, then moved or deleted) or by file-system backup, which reads the mounted files from a node agent and works for any volume type at the cost of consistency. Backup hooks (pre and post exec commands, such as flushing a database) are the honest answer to "is this backup consistent". A Restore replays a backup, optionally mapping namespaces (namespaceMapping: {abc: def}), and restored objects carry the label velero.io/restore-name. In a GitOps cluster the split to state out loud is: manifests come back from git, PV contents come back from Velero (or the operator's own backup), and neither replaces the other.
Lifecycle features by version, and copies of data
Every item here is a field on a PVC, StatefulSet or StorageClass that a task can ask you to set, plus the version at which it stopped being optional. The cluster is 1.36, so everything marked GA is simply on.
StatefulSet PVC retention (GA since 1.32)
spec.persistentVolumeClaimRetentionPolicy has two knobs, each Retain (default) or Delete: whenDeleted applies when the StatefulSet is deleted, whenScaled when replicas are reduced. Delete works by putting an ownerReference on the PVC so the garbage collector removes it after the pod is gone; a pod replaced after a node failure keeps its PVC regardless. The common production choice is whenDeleted: Retain, whenScaled: Delete: an accidental delete keeps the data, a deliberate scale-down does not leave orphans.
Expansion and its failure path
- Only with
allowVolumeExpansion: trueon the class and a CSI driver that supports it; you editspec.resources.requests.storageon the PVC, never the PV (editing the PV first makes the controller think the resize already happened). - The block device grows immediately; the filesystem (ext3, ext4, xfs) is resized only when a pod mounts the claim, online if the driver supports it. Until then the PVC shows a
FileSystemResizePendingcondition. "Resize stuck" with no consumer is therefore normal. - Shrinking is never allowed, but a failed grow can be retried smaller:
RecoverVolumeExpansionFailureis GA since 1.34, so you may lower the request to any value abovestatus.capacity. Progress lives instatus.allocatedResourcesandstatus.allocatedResourceStatuses(for exampleControllerResizeInProgress,NodeResizePending).
Snapshots, clones and populators
| Want | Write | Needs |
|---|---|---|
| point-in-time copy | VolumeSnapshot { source: { persistentVolumeClaimName } , volumeSnapshotClassName } | snapshot CRDs + snapshot-controller (installed by the distro), csi-snapshotter sidecar in the driver; a VolumeSnapshotClass with deletionPolicy: Delete or Retain |
| new PVC from a snapshot | dataSource: { kind: VolumeSnapshot, name, apiGroup: snapshot.storage.k8s.io } | same namespace and a class from the same driver; wait for status.readyToUse: true on the snapshot first |
| copy an existing PVC | dataSource: { kind: PersistentVolumeClaim, name } | CSI clone support; same namespace; requested size at least the source size |
| populate from anything else | dataSourceRef: { apiGroup, kind, name } | a volume populator controller for that kind (GA since 1.33); cross-namespace sources are still alpha |
Two protections you will meet as "stuck": a PVC being snapshotted cannot be deleted until the snapshot is readyToUse, and a PV bound to a claim carries kubernetes.io/pv-protection. Both clear themselves; do not strip them.
Ephemeral volumes
- Generic ephemeral volume (GA since 1.23):
volumes[].ephemeral.volumeClaimTemplatein the pod. The controller creates a real PVC named<pod>-<volume>, owned by the pod, so it is deleted with the pod, yet while it exists it can be snapshotted, cloned or expanded like any PVC. UseWaitForFirstConsumerclasses so the scheduler picks the node first. - CSI ephemeral volume:
volumes[].csiinline, no PVC object, for drivers that declare theEphemeralmode (secrets-store style drivers). Data disappears with the pod and never touches a StorageClass. emptyDircounts againstephemeral-storagerequests and limits; withmedium: Memoryit counts against memory instead and can be resized in place on cgroup v2 nodes.
VolumeAttributesClass (GA since 1.34)
A StorageClass decides how a volume is created; a VolumeAttributesClass (driverName plus parameters such as IOPS and throughput) decides how it performs, and pvc.spec.volumeAttributesClassName is mutable, so moving a claim from silver to gold is an edit rather than a migration. Quota can cap it per class (<vac>.volumeattributesclass.storage.k8s.io/...). It is the answer to "change the disk tier without recreating the PVC".
Smaller facts with a version
- Retroactive default class (GA since 1.28): a PVC created with no class while no default exists is patched when a default appears; a PVC with
storageClassName: ""is left alone and binds only to classless PVs. ReadWriteOncePodis GA since 1.29 and CSI-only; the scheduler and kubelet enforce it, unlike the other three modes which are matching hints.HonorPVReclaimPolicyadds anexternal-provisioner.volume.kubernetes.io/finalizerso aDeletePV whose PVC outlives it is still cleaned up in the backend when the PV is deleted first.- StatefulSet
updateStrategy.rollingUpdate.maxUnavailableis beta since 1.35 (default-on in 1.35.0 to 1.35.3, off again from 1.35.4 through 1.36, on by default in 1.37); aRecreateupdate strategy is also feature-gated. Do not assume either in a task; the default is one pod at a time, highest ordinal first.
"Grow the database volume to 20Gi" is graded on the PVC's status.capacity; the edit alone earns nothing. If the class lacks allowVolumeExpansion the edit is rejected outright ("only dynamically provisioned pvc can be resized and the storageclass that provisions the pvc must support resize"); if no pod mounts the claim, capacity stays at the old value with FileSystemResizePending, and the task is not done until a consumer runs.
Exercises
kubectl get storageclass standard -o yaml # read provisioner, bindingMode, reclaimPolicy
cat <<'EOF' | kubectl apply -f -
apiVersion: v1
kind: PersistentVolumeClaim
metadata: { name: scratch, namespace: default }
spec:
accessModes: [ReadWriteOnce]
resources: { requests: { storage: 1Gi } }
storageClassName: standard
EOF
kubectl get pvc scratch # Pending, and that is CORRECToutputcaptured 2026-08-26
$ kubectl get storageclass standard -o yaml # read provisioner, bindingMode, reclaimPolicy
apiVersion: storage.k8s.io/v1
kind: StorageClass
metadata:
annotations:
kubectl.kubernetes.io/last-applied-configuration: |
{"apiVersion":"storage.k8s.io/v1","kind":"StorageClass","metadata":{"annotations":{"storageclass.kubernetes.io/is-default-class":"true"},"name":"standard"},"provisioner":"rancher.io/local-path","reclaimPolicy":"Delete","volumeBindingMode":"WaitForFirstConsumer"}
storageclass.kubernetes.io/is-default-class: "true"
creationTimestamp: "2026-08-27T01:42:26Z"
name: standard
resourceVersion: "317"
uid: c527a04d-959b-4479-90bc-c1e5de4612ad
provisioner: rancher.io/local-path
reclaimPolicy: Delete
volumeBindingMode: WaitForFirstConsumer
$ cat <<'EOF' | kubectl apply -f -
apiVersion: v1
kind: PersistentVolumeClaim
metadata: { name: scratch, namespace: default }
spec:
accessModes: [ReadWriteOnce]
resources: { requests: { storage: 1Gi } }
storageClassName: standard
EOF
persistentvolumeclaim/scratch created
$ kubectl get pvc scratch # Pending, and that is CORRECT
NAME STATUS VOLUME CAPACITY ACCESS MODES STORAGECLASS VOLUMEATTRIBUTESCLASS AGE
scratch Pending standard <unset> 0sNow consume it:
kubectl run writer --image=busybox:1.37 --restart=Never \
--overrides='{"spec":{"containers":[{"name":"writer","image":"busybox:1.37","command":["sh","-c","echo survived > /data/proof && sleep 3600"],"volumeMounts":[{"name":"d","mountPath":"/data"}]}],"volumes":[{"name":"d","persistentVolumeClaim":{"claimName":"scratch"}}]}}'
kubectl get pvc scratch # Bound, seconds after the pod scheduledoutputcaptured 2026-08-26
$ kubectl run writer --image=busybox:1.37 --restart=Never \
--overrides='{"spec":{"containers":[{"name":"writer","image":"busybox:1.37","command":["sh","-c","echo survived > /data/proof && sleep 3600"],"volumeMounts":[{"name":"d","mountPath":"/data"}]}],"volumes":[{"name":"d","persistentVolumeClaim":{"claimName":"scratch"}}]}}'
pod/writer created
$ kubectl get pvc scratch # Bound, seconds after the pod scheduled
NAME STATUS VOLUME CAPACITY ACCESS MODES STORAGECLASS VOLUMEATTRIBUTESCLASS AGE
scratch Bound pvc-5e6b3db5-7d24-4e9f-82bf-8cdad4dec78c 1Gi RWO standard <unset> 11scat /data/proof, and kubectl logs writer must print survived.Create a PVC identical to the above but accessModes: [ReadWriteMany], plus a pod that mounts it (without a consumer, WaitForFirstConsumer keeps the events silent and you learn nothing). Now it pends with a reason: kubectl describe pvc events show the provisioner refusing, because local-path only does RWO.
make api installs the CloudNativePG operator; a Postgres exists once you create one (section 3.3 does, and it is worth jumping ahead for its first exercise). With one running:
kubectl get pvc -A | grep -v Bound # unbound + consumed = a finding (your own scratch claims excepted)
kubectl get pvc -A -o custom-columns='NS:.metadata.namespace,NAME:.metadata.name,SC:.spec.storageClassName,MODE:.spec.accessModes[0],SIZE:.spec.resources.requests.storage'outputcaptured 2026-08-26
$ kubectl get pvc -A | grep -v Bound # unbound + consumed = a finding (your own scratch claims excepted)
NAMESPACE NAME STATUS VOLUME CAPACITY ACCESS MODES STORAGECLASS VOLUMEATTRIBUTESCLASS AGE
$ kubectl get pvc -A -o custom-columns='NS:.metadata.namespace,NAME:.metadata.name,SC:.spec.storageClassName,MODE:.spec.accessModes[0],SIZE:.spec.resources.requests.storage'
NS NAME SC MODE SIZE
default pg-1 standard ReadWriteOnce 1Gi
default pg-2 standard ReadWriteOnce 1Gi
default scratch standard ReadWriteOnce 1Gi
monitoring storage-loki-0 standard ReadWriteOnce 10Gi
spire spire-data-spire-server-0 standard ReadWriteOnce 1Gikubectl delete pod of the instance, and the new pod mounts the same data. That is the operator relying on exactly the PVC semantics above.Create a PV of type hostPath (1Gi, RWO, storageClassName: manual), a PVC requesting it by the same class name, and show they bind with no provisioner involved.
kubectl get pv shows STATUS Bound and CLAIM pointing at your PVC. Then delete the PVC and explain what the PV's new status means given its reclaim policy.A StatefulSet's claims outlive it by default, which is either the safety you wanted or the leak you get billed for. The retention policy splits the two cases apart, and the ownerReference the controller writes on each claim is how the deletion half is carried out.
kubectl apply -f - <<'EOF'
apiVersion: apps/v1
kind: StatefulSet
metadata: { name: keep, namespace: default }
spec:
serviceName: keep
replicas: 3
persistentVolumeClaimRetentionPolicy: { whenScaled: Delete, whenDeleted: Retain }
selector: { matchLabels: { app: keep } }
template:
metadata: { labels: { app: keep } }
spec:
containers:
- name: c
image: busybox:1.36
command: ["sh", "-c", "sleep 86400"]
volumeMounts: [{ name: data, mountPath: /data }]
volumeClaimTemplates:
- metadata: { name: data }
spec:
accessModes: [ReadWriteOnce]
resources: { requests: { storage: 64Mi } }
EOF
kubectl rollout status statefulset/keep --timeout=180s
kubectl get pvc -l app=keep
kubectl get pvc data-keep-2 -o jsonpath='{.metadata.ownerReferences}' | jq # nothing owns it yet
kubectl scale statefulset keep --replicas=1
sleep 15
kubectl get pvc data-keep-2 -o jsonpath='{.metadata.ownerReferences}' | jq # now the condemned pod does
sleep 75
kubectl get pvc -l app=keep
kubectl delete statefulset keep
sleep 15
kubectl get pvc -l app=keep
kubectl delete pvc data-keep-0
kubectl delete pvc -l app=keep --ignore-not-foundoutputcaptured 2026-09-13
$ kubectl apply -f - <<'EOF'
apiVersion: apps/v1
kind: StatefulSet
metadata: { name: keep, namespace: default }
spec:
serviceName: keep
replicas: 3
persistentVolumeClaimRetentionPolicy: { whenScaled: Delete, whenDeleted: Retain }
selector: { matchLabels: { app: keep } }
template:
metadata: { labels: { app: keep } }
spec:
containers:
- name: c
image: busybox:1.36
command: ["sh", "-c", "sleep 86400"]
volumeMounts: [{ name: data, mountPath: /data }]
volumeClaimTemplates:
- metadata: { name: data }
spec:
accessModes: [ReadWriteOnce]
resources: { requests: { storage: 64Mi } }
EOF
statefulset.apps/keep created
$ kubectl rollout status statefulset/keep --timeout=180s
Waiting for statefulset spec update to be observed...
Waiting for 3 pods to be ready...
Waiting for 3 pods to be ready...
Waiting for 2 pods to be ready...
Waiting for 2 pods to be ready...
Waiting for 1 pods to be ready...
Waiting for 1 pods to be ready...
partitioned roll out complete: 3 new pods have been updated...
$ kubectl get pvc -l app=keep
NAME STATUS VOLUME CAPACITY ACCESS MODES STORAGECLASS VOLUMEATTRIBUTESCLASS AGE
data-keep-0 Bound pvc-5d487b38-6213-435f-89fa-e84c958cc0f6 64Mi RWO standard <unset> 24s
data-keep-1 Bound pvc-7e38cb51-e9b5-4d16-9f05-27bc88315451 64Mi RWO standard <unset> 17s
data-keep-2 Bound pvc-cde3b56b-bd6d-4cfb-8233-ad4868f74623 64Mi RWO standard <unset> 9s
$ kubectl get pvc data-keep-2 -o jsonpath='{.metadata.ownerReferences}' | jq # nothing owns it yet
$ kubectl scale statefulset keep --replicas=1
statefulset.apps/keep scaled
$ sleep 15
$ kubectl get pvc data-keep-2 -o jsonpath='{.metadata.ownerReferences}' | jq # now the condemned pod does
[
{
"apiVersion": "v1",
"blockOwnerDeletion": true,
"controller": true,
"kind": "Pod",
"name": "keep-2",
"uid": "a92afd0d-1914-4e04-8ba2-85d6a507a65a"
}
]
$ sleep 75
$ kubectl get pvc -l app=keep
NAME STATUS VOLUME CAPACITY ACCESS MODES STORAGECLASS VOLUMEATTRIBUTESCLASS AGE
data-keep-0 Bound pvc-5d487b38-6213-435f-89fa-e84c958cc0f6 64Mi RWO standard <unset> 114s
$ kubectl delete statefulset keep
statefulset.apps "keep" deleted from default namespace
$ sleep 15
$ kubectl get pvc -l app=keep
NAME STATUS VOLUME CAPACITY ACCESS MODES STORAGECLASS VOLUMEATTRIBUTESCLASS AGE
data-keep-0 Bound pvc-5d487b38-6213-435f-89fa-e84c958cc0f6 64Mi RWO standard <unset> 2m10s
$ kubectl delete pvc data-keep-0
persistentvolumeclaim "data-keep-0" deleted from default namespace
$ kubectl delete pvc -l app=keep --ignore-not-found
No resources foundwhenDeleted: Retain means the set never owns them. Scale down and the two condemned claims each gain an ownerReference to their own pod, which is what deletes them once that pod is gone; the last claim then survives deletion of the set itself. Two halves of one field, opposite outcomes.Expansion is a property of the driver, not of the StorageClass field that advertises it. Setting allowVolumeExpansion on a class whose provisioner cannot expand produces one of the most quoted error strings in storage tasks, so meet it here rather than in a task.
kubectl get storageclass
kubectl patch storageclass standard -p '{"allowVolumeExpansion":true}'
kubectl apply -f - <<'EOF'
apiVersion: v1
kind: PersistentVolumeClaim
metadata: { name: grow, namespace: default }
spec:
accessModes: [ReadWriteOnce]
resources: { requests: { storage: 64Mi } }
EOF
kubectl run grower --image=busybox:1.36 --restart=Never --overrides='{"spec":{"containers":[{"name":"c","image":"busybox:1.36","command":["sh","-c","sleep 86400"],"volumeMounts":[{"name":"d","mountPath":"/d"}]}],"volumes":[{"name":"d","persistentVolumeClaim":{"claimName":"grow"}}]}}'
kubectl wait --for=condition=Ready pod/grower --timeout=120s
kubectl patch pvc grow -p '{"spec":{"resources":{"requests":{"storage":"128Mi"}}}}'
sleep 20
kubectl get pvc grow -o jsonpath='spec: {.spec.resources.requests.storage} status: {.status.capacity.storage}{"\n"}'
echo "conditions: $(kubectl get pvc grow -o jsonpath='{.status.conditions}')"
kubectl get pvc grow -o jsonpath='{.status.allocatedResourceStatuses}{"\n"}'
kubectl describe pvc grow | tail -8
kubectl delete pod grower
kubectl delete pvc grow
# put the shared default class back the way it was
kubectl patch storageclass standard -p '{"allowVolumeExpansion":false}'outputcaptured 2026-09-12
$ kubectl get storageclass
NAME PROVISIONER RECLAIMPOLICY VOLUMEBINDINGMODE ALLOWVOLUMEEXPANSION AGE
standard (default) rancher.io/local-path Delete WaitForFirstConsumer false 20h
$ kubectl patch storageclass standard -p '{"allowVolumeExpansion":true}'
storageclass.storage.k8s.io/standard patched
$ kubectl apply -f - <<'EOF'
apiVersion: v1
kind: PersistentVolumeClaim
metadata: { name: grow, namespace: default }
spec:
accessModes: [ReadWriteOnce]
resources: { requests: { storage: 64Mi } }
EOF
persistentvolumeclaim/grow created
$ kubectl run grower --image=busybox:1.36 --restart=Never --overrides='{"spec":{"containers":[{"name":"c","image":"busybox:1.36","command":["sh","-c","sleep 86400"],"volumeMounts":[{"name":"d","mountPath":"/d"}]}],"volumes":[{"name":"d","persistentVolumeClaim":{"claimName":"grow"}}]}}'
pod/grower created
$ kubectl wait --for=condition=Ready pod/grower --timeout=120s
pod/grower condition met
$ kubectl patch pvc grow -p '{"spec":{"resources":{"requests":{"storage":"128Mi"}}}}'
persistentvolumeclaim/grow patched
$ sleep 20
$ kubectl get pvc grow -o jsonpath='spec: {.spec.resources.requests.storage} status: {.status.capacity.storage}{"\n"}'
spec: 128Mi status: 64Mi
$ echo "conditions: $(kubectl get pvc grow -o jsonpath='{.status.conditions}')"
conditions:
$ kubectl get pvc grow -o jsonpath='{.status.allocatedResourceStatuses}{"\n"}'
$ kubectl describe pvc grow | tail -8
Events:
Type Reason Age From Message
---- ------ ---- ---- -------
Normal WaitForFirstConsumer 32s persistentvolume-controller waiting for first consumer to be created before binding
Normal ExternalProvisioning 31s persistentvolume-controller Waiting for a volume to be created either by the external provisioner 'rancher.io/local-path' or manually by the system administrator. If volume creation is delayed, please verify that the provisioner is running and correctly registered.
Normal Provisioning 31s rancher.io/local-path_local-path-provisioner-855c7b7774-xsqw2_944b3c00-03a8-486a-896a-567844028b41 External provisioner is provisioning volume for claim "default/grow"
Normal ProvisioningSucceeded 25s rancher.io/local-path_local-path-provisioner-855c7b7774-xsqw2_944b3c00-03a8-486a-896a-567844028b41 Successfully provisioned volume pvc-f412d670-060c-4a9a-8fda-bd4cc5645a03
Normal ExternalExpanding 21s volume_expand waiting for an external controller to expand this PVC
$ kubectl delete pod grower
pod "grower" deleted from default namespace
$ kubectl delete pvc grow
persistentvolumeclaim "grow" deleted from default namespace
$ # put the shared default class back the way it was
$ kubectl patch storageclass standard -p '{"allowVolumeExpansion":false}'
storageclass.storage.k8s.io/standard patchedallowVolumeExpansion: true and the claim accepts the larger request, so spec reads 128Mi while status.capacity stays at 64Mi. status.conditions is empty and the only event is ExternalExpanding: waiting for an external controller to expand this PVC, which is the whole answer: local-path has no resize support, so no controller ever picks the request up and the claim waits forever with its spec and status disagreeing. There is no error to find, which is the failure mode worth recognizing.A generic ephemeral volume gives a pod a real PVC with a real class and a real lifetime of exactly one pod. The naming rule and the ownerReference are the whole trick.
kubectl apply -f - <<'EOF'
apiVersion: v1
kind: Pod
metadata: { name: eph, namespace: default }
spec:
containers:
- name: c
image: busybox:1.36
command: ["sh", "-c", "sleep 86400"]
volumeMounts: [{ name: scratch, mountPath: /scratch }]
volumes:
- name: scratch
ephemeral:
volumeClaimTemplate:
spec:
accessModes: [ReadWriteOnce]
resources: { requests: { storage: 64Mi } }
EOF
kubectl wait --for=condition=Ready pod/eph --timeout=120s
kubectl get pvc eph-scratch
kubectl get pvc eph-scratch -o jsonpath='{.metadata.ownerReferences}' | jq
kubectl delete pod eph
sleep 15
kubectl get pvc eph-scratchoutputcaptured 2026-09-13
$ kubectl apply -f - <<'EOF'
apiVersion: v1
kind: Pod
metadata: { name: eph, namespace: default }
spec:
containers:
- name: c
image: busybox:1.36
command: ["sh", "-c", "sleep 86400"]
volumeMounts: [{ name: scratch, mountPath: /scratch }]
volumes:
- name: scratch
ephemeral:
volumeClaimTemplate:
spec:
accessModes: [ReadWriteOnce]
resources: { requests: { storage: 64Mi } }
EOF
pod/eph created
$ kubectl wait --for=condition=Ready pod/eph --timeout=120s
pod/eph condition met
$ kubectl get pvc eph-scratch
NAME STATUS VOLUME CAPACITY ACCESS MODES STORAGECLASS VOLUMEATTRIBUTESCLASS AGE
eph-scratch Bound pvc-60e802fd-df09-43c4-90ba-64f0e90ad910 64Mi RWO standard <unset> 13s
$ kubectl get pvc eph-scratch -o jsonpath='{.metadata.ownerReferences}' | jq
[
{
"apiVersion": "v1",
"blockOwnerDeletion": true,
"controller": true,
"kind": "Pod",
"name": "eph",
"uid": "b7e5e146-ef47-42b5-afb7-8af422d2a318"
}
]
$ kubectl delete pod eph
pod "eph" deleted from default namespace
$ sleep 15
$ kubectl get pvc eph-scratch
Error from server (NotFound): persistentvolumeclaims "eph-scratch" not found<pod>-<volume>, owned by the pod, and gone once the pod is.Snapshots are three separate installs: the CRDs, the snapshot-controller, and driver support. kind with local-path has none of them, and knowing that is more useful than reciting the object names.
kubectl get crd | grep snapshot.storage.k8s.io || echo 'no snapshot CRDs installed'
kubectl api-resources | grep -i volumesnapshot || echo 'no volumesnapshot API on this cluster'
kubectl get csidrivers
kubectl get sc standard -o jsonpath='{.provisioner}{"\n"}'outputcaptured 2026-09-12
$ kubectl get crd | grep snapshot.storage.k8s.io || echo 'no snapshot CRDs installed'
no snapshot CRDs installed
$ kubectl api-resources | grep -i volumesnapshot || echo 'no volumesnapshot API on this cluster'
no volumesnapshot API on this cluster
$ kubectl get csidrivers
NAME ATTACHREQUIRED PODINFOONMOUNT STORAGECAPACITY TOKENREQUESTS REQUIRESREPUBLISH MODES AGE
csi.spiffe.io false true false <unset> false Ephemeral 21m
$ kubectl get sc standard -o jsonpath='{.provisioner}{"\n"}'
rancher.io/local-pathA PVC with no storageClassName is not bound to "the default forever": it records whatever default existed when admission saw it, and since retroactive default assignment it can be filled in later. Watch both halves.
kubectl get storageclass -o custom-columns='NAME:.metadata.name,DEFAULT:.metadata.annotations.storageclass\.kubernetes\.io/is-default-class'
kubectl annotate storageclass standard storageclass.kubernetes.io/is-default-class-
kubectl apply -f - <<'EOF'
apiVersion: v1
kind: PersistentVolumeClaim
metadata: { name: noclass, namespace: default }
spec:
accessModes: [ReadWriteOnce]
resources: { requests: { storage: 64Mi } }
EOF
kubectl get pvc noclass -o jsonpath='class=[{.spec.storageClassName}] phase={.status.phase}{"\n"}'
kubectl annotate storageclass standard storageclass.kubernetes.io/is-default-class=true
sleep 20
kubectl get pvc noclass -o jsonpath='class=[{.spec.storageClassName}] phase={.status.phase}{"\n"}'
kubectl delete pvc noclassoutputcaptured 2026-09-13
$ kubectl get storageclass -o custom-columns='NAME:.metadata.name,DEFAULT:.metadata.annotations.storageclass\.kubernetes\.io/is-default-class'
NAME DEFAULT
standard true
$ kubectl annotate storageclass standard storageclass.kubernetes.io/is-default-class-
storageclass.storage.k8s.io/standard annotated
$ kubectl apply -f - <<'EOF'
apiVersion: v1
kind: PersistentVolumeClaim
metadata: { name: noclass, namespace: default }
spec:
accessModes: [ReadWriteOnce]
resources: { requests: { storage: 64Mi } }
EOF
persistentvolumeclaim/noclass created
$ kubectl get pvc noclass -o jsonpath='class=[{.spec.storageClassName}] phase={.status.phase}{"\n"}'
class=[] phase=Pending
$ kubectl annotate storageclass standard storageclass.kubernetes.io/is-default-class=true
storageclass.storage.k8s.io/standard annotated
$ sleep 20
$ kubectl get pvc noclass -o jsonpath='class=[{.spec.storageClassName}] phase={.status.phase}{"\n"}'
class=[standard] phase=Pending
$ kubectl delete pvc noclass
persistentvolumeclaim "noclass" deleted from default namespacestandard written into its spec once the annotation is back. Restore the annotation before you leave this exercise or every later claim on the page hangs.Self-check
A PVC has been Pending for ten minutes and describe shows no events at all. Bug or not?
Not a bug if the class uses WaitForFirstConsumer and nothing mounts the claim yet: no scheduling decision exists, so there is nothing to provision and nothing to report. It becomes a finding the moment a pod references it and the claim stays Pending; then the events appear and name the real reason.
A PV sits in Released and your new, identical PVC will not bind to it. Why, and what is the fix?
Retain leaves spec.claimRef pointing at the deleted claim, and a PV with a claimRef is spoken for. Clear it (kubectl patch pv <name> -p '{"spec":{"claimRef":null}}') and it returns to Available. Deciding whether the old data should be wiped first is the actual judgment call.
You delete a StatefulSet and its PVCs remain. Accident or design?
Design. The claims outlive the set so scaling back up reattaches the same data, and so an accidental delete does not destroy the database. Newer clusters can opt in to cleanup via persistentVolumeClaimRetentionPolicy (whenDeleted / whenScaled).
Two pods on different nodes must share a directory. What do you need, and what will this lab do?
A class whose provisioner supports RWX (NFS, CephFS, a cloud file service). local-path cannot, so the claim pends and the events say so. The honest lab answer is "not possible here" plus the fix you would apply in a real cluster, which is also the right exam answer when the storage cannot do what the task implies.
Which quota lines cap tenant storage, and why is per-class quota interesting?
requests.storage and persistentvolumeclaims cap the total and the count; <class>.storageclass.storage.k8s.io/requests.storage caps a specific class. The per-class form is how a platform team offers a fast expensive tier without every tenant defaulting to it: a pricing decision expressed as a Kubernetes object.
You want scale-down to reclaim volumes but deleting the StatefulSet by mistake to keep them. Which fields, and since when is this GA?
persistentVolumeClaimRetentionPolicy: { whenScaled: Delete, whenDeleted: Retain }, GA since Kubernetes 1.32 (StatefulSetAutoDeletePVC). Delete works through an ownerReference on the PVC that the garbage collector honors after the pod is gone; a pod replaced after a node failure is not a scale event and keeps its claim.
An expansion to 500Gi failed because the backend has no capacity and the PVC keeps retrying. Can you back out?
Yes: RecoverVolumeExpansionFailure is GA since 1.34, so edit the request down to any value still above status.capacity (never below it; shrinking is impossible). Watch status.allocatedResourceStatuses and events. Before 1.34 the recovery was the manual dance of Retain, delete the PVC, clear claimRef, recreate.
Restore a database volume from last night's snapshot into a new PVC. What does the manifest look like and what must be true first?
A PVC whose dataSource is { kind: VolumeSnapshot, name: <snap>, apiGroup: snapshot.storage.k8s.io }, in the same namespace, with a StorageClass from the same CSI driver and a size at least the snapshot's restore size. The VolumeSnapshot must show status.readyToUse: true; a not-ready snapshot leaves the PVC Pending with a provisioner event saying so.
How do you change a volume's IOPS tier without recreating it?
VolumeAttributesClass (GA since 1.34): create a class with the driver's performance parameters and set spec.volumeAttributesClassName on the PVC; the field is mutable and the external-resizer applies the change. StorageClass parameters, by contrast, are fixed at provisioning time.
Docs to know your way around
- kubernetes.io: Persistent Volumes (the access-modes and reclaim tables), Storage Classes, Volume Snapshots, StatefulSet
volumeClaimTemplates. - Offline:
kubectl explain pvc.spec,kubectl explain sc,kubectl get sc -o wide, and PVCstatus.conditionswhen a resize misbehaves. - kubernetes.io: Persistent Volumes, "Expanding Persistent Volumes Claims" and "Recovering from Failure when Expanding Volumes": the exact rules for grow-only, filesystem-on-mount and the smaller-retry path.
- kubernetes.io: Ephemeral Volumes and Volume Attributes Classes: generic ephemeral volume naming and ownership; the mutable volumeAttributesClassName field.
- velero.io/docs: "How Velero Works" and "CSI snapshot support": the Backup/Schedule/Restore objects and the CSI versus file-system backup choice.
make down-api