The lab has no storage layer to install because kind ships one: the standard StorageClass backed by the local-path provisioner. That is enough to exercise every concept the exam touches, and its limitations are themselves instructive.

needsmake upmake api

Orientation

competency 1.1 · architecture best practices

Storage tasks on a performance exam are rarely "install a CSI driver". They are: this claim is Pending, say why; this pod lost its data, say why; make this workload survive a reschedule; grow this volume. All four are answered from four fields you can recite.

Exam angle

The graders can only see objects and their status. So the tell for a storage task is almost always a status: PVC Pending, PV Released, pod stuck ContainerCreating with a mount error in events. Learn to map those three states to their causes and you have the competency.

The model: three objects, one relationship

PV · PVC · StorageClass

A PersistentVolume is a piece of real storage. A PersistentVolumeClaim is a request for one. A StorageClass is the recipe for making PVs on demand, and its provisioner field names the code that does it. Static provisioning (an admin pre-creates PVs) still exists but dynamic is the assumed default: the PVC references a class, the provisioner makes the PV, the two bind one-to-one and exclusively.

pod ──mounts──▶ PVC ──binds 1:1──▶ PV ──backed by──▶ real disk / host path / cloud volume
                 │                  ▲
                 └── storageClassName ──▶ StorageClass ──▶ provisioner creates the PV
                                          (binding mode, reclaim policy, expansion, params)

Underneath, CSI is the plugin interface every modern driver implements: a controller component (provision, attach, snapshot) and a node component (mount). You will not be asked to write one, but knowing the split explains error locations: "failed to provision" is a controller-side message, "failed to mount" is node-side, and they point at different logs.

The fields that decide exam tasks

FieldValuesWhat it changes
accessModesRWO · ROX · RWX · RWOPA claim binds only to a PV offering what it asks. local-path only does RWO, so an RWX claim here pends forever; recognizing why is the skill.
volumeBindingModeImmediate · WaitForFirstConsumerWFFC keeps the PVC Pending until a pod uses it, so topology can be considered. The standard class here uses it, so you meet this "problem" immediately, and it is not a problem.
persistentVolumeReclaimPolicyDelete · RetainDelete throws data away with the claim; Retain keeps the PV in Released, and it will not rebind until someone clears spec.claimRef. Released-but-unusable PVs are a classic troubleshooting scenario.
allowVolumeExpansiontrue · falseOnly classes that set it let you grow a PVC by editing spec.resources.requests.storage. Shrinking is never allowed, anywhere.
volumeModeFilesystem · BlockBlock hands the raw device to the container. Databases sometimes want it; nothing else does.
Access modes are a claim, not a guarantee

RWX means "this PV supports many nodes mounting it read-write", and it is a property of the backing storage, not a wish you can express. Nothing in Kubernetes enforces that two pods writing an RWO volume from the same node behave sensibly, and requesting RWX does not turn a local disk into a shared filesystem. RWOP (ReadWriteOncePod) is the strict one: exactly one pod, enforced by the scheduler (the kubelet backstops it at mount), which is how you stop two replicas corrupting a single-writer database.

Lifecycle and the states you will be asked to explain

Pending → Bound → Released
SymptomUsual causeCheck
PVC Pending, no eventsWaitForFirstConsumer, no pod yetnormal: schedule a consumer
PVC Pending, provisioner eventsunsupported access mode, no capacity, bad class namekubectl describe pvc
Pod ContainerCreating forevermount/attach failure, node-sidekubectl describe pod → events
PV Released, not rebindingRetain policy leaves claimRef populatedkubectl patch pv … claimRef=null
PVC Terminating foreverkubernetes.io/pvc-protection finalizer: a pod still uses itkubectl get pods -o json | grep claimName
Resize stuckclass lacks expansion, or filesystem resize needs a pod restartstatus.conditions on the PVC

Two protection finalizers exist for good reasons and both look like bugs the first time: pvc-protection blocks deleting a claim that a pod mounts, and pv-protection blocks deleting a bound PV. The fix is never to strip the finalizer first; it is to remove the consumer, then let the controller clean up. Stripping finalizers to make an object disappear is the storage equivalent of pulling the disk out.

StatefulSets, where storage semantics become visible

volumeClaimTemplates gives every replica its own PVC, named <template>-<sts>-<ordinal>. Deleting the StatefulSet does not delete those PVCs (unless you set persistentVolumeClaimRetentionPolicy, which newer versions offer), so scaling back up reattaches the old data. That retention is deliberate and it is the entire reason StatefulSets exist rather than "a Deployment with a volume": stable identity, stable storage, ordered rollout.

And the part GitOps cannot do for you: data. "Delete the namespace and the controller rebuilds it" restores manifests, never the contents of a PV. Three things fill that gap. CSI snapshots give point-in-time copies within a cluster. A backup tool (Velero is the common one) snapshots volumes and exports object state for anything that must survive the cluster itself. Database operators bring their own backup story where one exists. Being able to say which of those three you are relying on is the whole DR answer.

Snapshots round out the vocabulary: VolumeSnapshotClass, VolumeSnapshot, VolumeSnapshotContent. They follow the same three-object pattern as class/claim/volume, and they are how a CSI driver exposes point-in-time copies. local-path has no snapshotter, so this lab teaches the nouns and CloudNativePG teaches the backup story instead.

What a platform engineer actually decides

classes as product

Section 3.1's "APIs as products" idea lands here concretely: StorageClasses are a platform API. You are choosing the menu your tenants order from, and each entry encodes a durability, performance and cost decision they should not have to make.

  • Name classes for intent (fast-ssd, cheap-hdd, shared-rwx), never for the implementation that happens to back them today.
  • Set exactly one default class (storageclass.kubernetes.io/is-default-class) and know what it costs, because every PVC without an explicit class silently buys it. If two are marked default, the DefaultStorageClass admission plugin, not the scheduler, picks the most recently created one. That is worse than an error: it is silent, and it changes the next time someone adds a class.
  • Prefer WaitForFirstConsumer in any topology-aware environment, or you will provision volumes in zones your pods cannot reach.
  • Delete for ephemeral tenant workloads, Retain for anything whose loss ends up in a postmortem.
  • Cap storage per tenant with quota: requests.storage for the total, <class>.storageclass.storage.k8s.io/requests.storage per class. Per-class quota is how you stop everyone ordering the expensive one (section 1.4).
Command reflex

kubectl get sc first, every time: it shows provisioner, reclaim policy, binding mode and expansion in one line, which is four of the five fields above. Then kubectl get pvc -A and look for anything not Bound.

Backup and restore as objects: Velero

When the DR answer is "a backup tool", Velero is the one the exam pool knows, and its model is four CRDs. A BackupStorageLocation names the object store (bucket, prefix, provider plugin); a VolumeSnapshotLocation names where disk snapshots go; a Backup selects what to copy (includedNamespaces, includedResources, labelSelector, ttl, default 30 days) and a Schedule is a Backup template with a cron expression, producing backups named <schedule>-<timestamp>. Volumes are copied either by CSI snapshot (the snapshot is retained only for the life of the backup, then moved or deleted) or by file-system backup, which reads the mounted files from a node agent and works for any volume type at the cost of consistency. Backup hooks (pre and post exec commands, such as flushing a database) are the honest answer to "is this backup consistent". A Restore replays a backup, optionally mapping namespaces (namespaceMapping: {abc: def}), and restored objects carry the label velero.io/restore-name. In a GitOps cluster the split to state out loud is: manifests come back from git, PV contents come back from Velero (or the operator's own backup), and neither replaces the other.

Lifecycle features by version, and copies of data

retention · expansion · snapshots · clones · ephemeral volumes · VolumeAttributesClass

Every item here is a field on a PVC, StatefulSet or StorageClass that a task can ask you to set, plus the version at which it stopped being optional. The cluster is 1.36, so everything marked GA is simply on.

StatefulSet PVC retention (GA since 1.32)

spec.persistentVolumeClaimRetentionPolicy has two knobs, each Retain (default) or Delete: whenDeleted applies when the StatefulSet is deleted, whenScaled when replicas are reduced. Delete works by putting an ownerReference on the PVC so the garbage collector removes it after the pod is gone; a pod replaced after a node failure keeps its PVC regardless. The common production choice is whenDeleted: Retain, whenScaled: Delete: an accidental delete keeps the data, a deliberate scale-down does not leave orphans.

Expansion and its failure path

  • Only with allowVolumeExpansion: true on the class and a CSI driver that supports it; you edit spec.resources.requests.storage on the PVC, never the PV (editing the PV first makes the controller think the resize already happened).
  • The block device grows immediately; the filesystem (ext3, ext4, xfs) is resized only when a pod mounts the claim, online if the driver supports it. Until then the PVC shows a FileSystemResizePending condition. "Resize stuck" with no consumer is therefore normal.
  • Shrinking is never allowed, but a failed grow can be retried smaller: RecoverVolumeExpansionFailure is GA since 1.34, so you may lower the request to any value above status.capacity. Progress lives in status.allocatedResources and status.allocatedResourceStatuses (for example ControllerResizeInProgress, NodeResizePending).

Snapshots, clones and populators

WantWriteNeeds
point-in-time copyVolumeSnapshot { source: { persistentVolumeClaimName } , volumeSnapshotClassName }snapshot CRDs + snapshot-controller (installed by the distro), csi-snapshotter sidecar in the driver; a VolumeSnapshotClass with deletionPolicy: Delete or Retain
new PVC from a snapshotdataSource: { kind: VolumeSnapshot, name, apiGroup: snapshot.storage.k8s.io }same namespace and a class from the same driver; wait for status.readyToUse: true on the snapshot first
copy an existing PVCdataSource: { kind: PersistentVolumeClaim, name }CSI clone support; same namespace; requested size at least the source size
populate from anything elsedataSourceRef: { apiGroup, kind, name }a volume populator controller for that kind (GA since 1.33); cross-namespace sources are still alpha

Two protections you will meet as "stuck": a PVC being snapshotted cannot be deleted until the snapshot is readyToUse, and a PV bound to a claim carries kubernetes.io/pv-protection. Both clear themselves; do not strip them.

Ephemeral volumes

  • Generic ephemeral volume (GA since 1.23): volumes[].ephemeral.volumeClaimTemplate in the pod. The controller creates a real PVC named <pod>-<volume>, owned by the pod, so it is deleted with the pod, yet while it exists it can be snapshotted, cloned or expanded like any PVC. Use WaitForFirstConsumer classes so the scheduler picks the node first.
  • CSI ephemeral volume: volumes[].csi inline, no PVC object, for drivers that declare the Ephemeral mode (secrets-store style drivers). Data disappears with the pod and never touches a StorageClass.
  • emptyDir counts against ephemeral-storage requests and limits; with medium: Memory it counts against memory instead and can be resized in place on cgroup v2 nodes.

VolumeAttributesClass (GA since 1.34)

A StorageClass decides how a volume is created; a VolumeAttributesClass (driverName plus parameters such as IOPS and throughput) decides how it performs, and pvc.spec.volumeAttributesClassName is mutable, so moving a claim from silver to gold is an edit rather than a migration. Quota can cap it per class (<vac>.volumeattributesclass.storage.k8s.io/...). It is the answer to "change the disk tier without recreating the PVC".

Smaller facts with a version

  • Retroactive default class (GA since 1.28): a PVC created with no class while no default exists is patched when a default appears; a PVC with storageClassName: "" is left alone and binds only to classless PVs.
  • ReadWriteOncePod is GA since 1.29 and CSI-only; the scheduler and kubelet enforce it, unlike the other three modes which are matching hints.
  • HonorPVReclaimPolicy adds an external-provisioner.volume.kubernetes.io/finalizer so a Delete PV whose PVC outlives it is still cleaned up in the backend when the PV is deleted first.
  • StatefulSet updateStrategy.rollingUpdate.maxUnavailable is beta since 1.35 (default-on in 1.35.0 to 1.35.3, off again from 1.35.4 through 1.36, on by default in 1.37); a Recreate update strategy is also feature-gated. Do not assume either in a task; the default is one pod at a time, highest ordinal first.
How this gets tested

"Grow the database volume to 20Gi" is graded on the PVC's status.capacity; the edit alone earns nothing. If the class lacks allowVolumeExpansion the edit is rejected outright ("only dynamically provisioned pvc can be resized and the storageclass that provisions the pvc must support resize"); if no pod mounts the claim, capacity stays at the old value with FileSystemResizePending, and the task is not done until a consumer runs.

Exercises

tick the dot when its check passes
kubectl get storageclass standard -o yaml   # read provisioner, bindingMode, reclaimPolicy
cat <<'EOF' | kubectl apply -f -
apiVersion: v1
kind: PersistentVolumeClaim
metadata: { name: scratch, namespace: default }
spec:
  accessModes: [ReadWriteOnce]
  resources: { requests: { storage: 1Gi } }
  storageClassName: standard
EOF
kubectl get pvc scratch    # Pending, and that is CORRECT
outputcaptured 2026-08-26
$ kubectl get storageclass standard -o yaml   # read provisioner, bindingMode, reclaimPolicy
apiVersion: storage.k8s.io/v1
kind: StorageClass
metadata:
  annotations:
    kubectl.kubernetes.io/last-applied-configuration: |
      {"apiVersion":"storage.k8s.io/v1","kind":"StorageClass","metadata":{"annotations":{"storageclass.kubernetes.io/is-default-class":"true"},"name":"standard"},"provisioner":"rancher.io/local-path","reclaimPolicy":"Delete","volumeBindingMode":"WaitForFirstConsumer"}
    storageclass.kubernetes.io/is-default-class: "true"
  creationTimestamp: "2026-08-27T01:42:26Z"
  name: standard
  resourceVersion: "317"
  uid: c527a04d-959b-4479-90bc-c1e5de4612ad
provisioner: rancher.io/local-path
reclaimPolicy: Delete
volumeBindingMode: WaitForFirstConsumer
$ cat <<'EOF' | kubectl apply -f -
apiVersion: v1
kind: PersistentVolumeClaim
metadata: { name: scratch, namespace: default }
spec:
  accessModes: [ReadWriteOnce]
  resources: { requests: { storage: 1Gi } }
  storageClassName: standard
EOF
persistentvolumeclaim/scratch created
$ kubectl get pvc scratch    # Pending, and that is CORRECT
NAME      STATUS    VOLUME   CAPACITY   ACCESS MODES   STORAGECLASS   VOLUMEATTRIBUTESCLASS   AGE
scratch   Pending                                      standard       <unset>                 0s

Now consume it:

kubectl run writer --image=busybox:1.37 --restart=Never \
  --overrides='{"spec":{"containers":[{"name":"writer","image":"busybox:1.37","command":["sh","-c","echo survived > /data/proof && sleep 3600"],"volumeMounts":[{"name":"d","mountPath":"/data"}]}],"volumes":[{"name":"d","persistentVolumeClaim":{"claimName":"scratch"}}]}}'
kubectl get pvc scratch    # Bound, seconds after the pod scheduled
outputcaptured 2026-08-26
$ kubectl run writer --image=busybox:1.37 --restart=Never \
  --overrides='{"spec":{"containers":[{"name":"writer","image":"busybox:1.37","command":["sh","-c","echo survived > /data/proof && sleep 3600"],"volumeMounts":[{"name":"d","mountPath":"/data"}]}],"volumes":[{"name":"d","persistentVolumeClaim":{"claimName":"scratch"}}]}}'
pod/writer created
$ kubectl get pvc scratch    # Bound, seconds after the pod scheduled
NAME      STATUS   VOLUME                                     CAPACITY   ACCESS MODES   STORAGECLASS   VOLUMEATTRIBUTESCLASS   AGE
scratch   Bound    pvc-5e6b3db5-7d24-4e9f-82bf-8cdad4dec78c   1Gi        RWO            standard       <unset>                 11s
verify: delete the pod, recreate it with a command of cat /data/proof, and kubectl logs writer must print survived.

Create a PVC identical to the above but accessModes: [ReadWriteMany], plus a pod that mounts it (without a consumer, WaitForFirstConsumer keeps the events silent and you learn nothing). Now it pends with a reason: kubectl describe pvc events show the provisioner refusing, because local-path only does RWO.

verify: you can state the fix (RWO, or a class whose provisioner supports RWX) and then delete pod and claim. The general lesson: a Pending PVC under WaitForFirstConsumer is normal until a pod consumes it; only then does silence become a finding.

make api installs the CloudNativePG operator; a Postgres exists once you create one (section 3.3 does, and it is worth jumping ahead for its first exercise). With one running:

kubectl get pvc -A | grep -v Bound        # unbound + consumed = a finding (your own scratch claims excepted)
kubectl get pvc -A -o custom-columns='NS:.metadata.namespace,NAME:.metadata.name,SC:.spec.storageClassName,MODE:.spec.accessModes[0],SIZE:.spec.resources.requests.storage'
outputcaptured 2026-08-26
$ kubectl get pvc -A | grep -v Bound        # unbound + consumed = a finding (your own scratch claims excepted)
NAMESPACE    NAME                        STATUS   VOLUME                                     CAPACITY   ACCESS MODES   STORAGECLASS   VOLUMEATTRIBUTESCLASS   AGE
$ kubectl get pvc -A -o custom-columns='NS:.metadata.namespace,NAME:.metadata.name,SC:.spec.storageClassName,MODE:.spec.accessModes[0],SIZE:.spec.resources.requests.storage'
NS           NAME                        SC         MODE            SIZE
default      pg-1                        standard   ReadWriteOnce   1Gi
default      pg-2                        standard   ReadWriteOnce   1Gi
default      scratch                     standard   ReadWriteOnce   1Gi
monitoring   storage-loki-0              standard   ReadWriteOnce   10Gi
spire        spire-data-spire-server-0   standard   ReadWriteOnce   1Gi
verify: you can point at each PVC and name what created it (a volumeClaimTemplate, an operator, a human). Note that a CNPG instance's PVC survives kubectl delete pod of the instance, and the new pod mounts the same data. That is the operator relying on exactly the PVC semantics above.

Create a PV of type hostPath (1Gi, RWO, storageClassName: manual), a PVC requesting it by the same class name, and show they bind with no provisioner involved.

verify: kubectl get pv shows STATUS Bound and CLAIM pointing at your PVC. Then delete the PVC and explain what the PV's new status means given its reclaim policy.

A StatefulSet's claims outlive it by default, which is either the safety you wanted or the leak you get billed for. The retention policy splits the two cases apart, and the ownerReference the controller writes on each claim is how the deletion half is carried out.

kubectl apply -f - <<'EOF'
apiVersion: apps/v1
kind: StatefulSet
metadata: { name: keep, namespace: default }
spec:
  serviceName: keep
  replicas: 3
  persistentVolumeClaimRetentionPolicy: { whenScaled: Delete, whenDeleted: Retain }
  selector: { matchLabels: { app: keep } }
  template:
    metadata: { labels: { app: keep } }
    spec:
      containers:
        - name: c
          image: busybox:1.36
          command: ["sh", "-c", "sleep 86400"]
          volumeMounts: [{ name: data, mountPath: /data }]
  volumeClaimTemplates:
    - metadata: { name: data }
      spec:
        accessModes: [ReadWriteOnce]
        resources: { requests: { storage: 64Mi } }
EOF
kubectl rollout status statefulset/keep --timeout=180s
kubectl get pvc -l app=keep
kubectl get pvc data-keep-2 -o jsonpath='{.metadata.ownerReferences}' | jq   # nothing owns it yet
kubectl scale statefulset keep --replicas=1
sleep 15
kubectl get pvc data-keep-2 -o jsonpath='{.metadata.ownerReferences}' | jq   # now the condemned pod does
sleep 75
kubectl get pvc -l app=keep
kubectl delete statefulset keep
sleep 15
kubectl get pvc -l app=keep
kubectl delete pvc data-keep-0
kubectl delete pvc -l app=keep --ignore-not-found
outputcaptured 2026-09-13
$ kubectl apply -f - <<'EOF'
apiVersion: apps/v1
kind: StatefulSet
metadata: { name: keep, namespace: default }
spec:
  serviceName: keep
  replicas: 3
  persistentVolumeClaimRetentionPolicy: { whenScaled: Delete, whenDeleted: Retain }
  selector: { matchLabels: { app: keep } }
  template:
    metadata: { labels: { app: keep } }
    spec:
      containers:
        - name: c
          image: busybox:1.36
          command: ["sh", "-c", "sleep 86400"]
          volumeMounts: [{ name: data, mountPath: /data }]
  volumeClaimTemplates:
    - metadata: { name: data }
      spec:
        accessModes: [ReadWriteOnce]
        resources: { requests: { storage: 64Mi } }
EOF
statefulset.apps/keep created
$ kubectl rollout status statefulset/keep --timeout=180s
Waiting for statefulset spec update to be observed...
Waiting for 3 pods to be ready...
Waiting for 3 pods to be ready...
Waiting for 2 pods to be ready...
Waiting for 2 pods to be ready...
Waiting for 1 pods to be ready...
Waiting for 1 pods to be ready...
partitioned roll out complete: 3 new pods have been updated...
$ kubectl get pvc -l app=keep
NAME          STATUS   VOLUME                                     CAPACITY   ACCESS MODES   STORAGECLASS   VOLUMEATTRIBUTESCLASS   AGE
data-keep-0   Bound    pvc-5d487b38-6213-435f-89fa-e84c958cc0f6   64Mi       RWO            standard       <unset>                 24s
data-keep-1   Bound    pvc-7e38cb51-e9b5-4d16-9f05-27bc88315451   64Mi       RWO            standard       <unset>                 17s
data-keep-2   Bound    pvc-cde3b56b-bd6d-4cfb-8233-ad4868f74623   64Mi       RWO            standard       <unset>                 9s
$ kubectl get pvc data-keep-2 -o jsonpath='{.metadata.ownerReferences}' | jq   # nothing owns it yet
$ kubectl scale statefulset keep --replicas=1
statefulset.apps/keep scaled
$ sleep 15
$ kubectl get pvc data-keep-2 -o jsonpath='{.metadata.ownerReferences}' | jq   # now the condemned pod does
[
  {
    "apiVersion": "v1",
    "blockOwnerDeletion": true,
    "controller": true,
    "kind": "Pod",
    "name": "keep-2",
    "uid": "a92afd0d-1914-4e04-8ba2-85d6a507a65a"
  }
]
$ sleep 75
$ kubectl get pvc -l app=keep
NAME          STATUS   VOLUME                                     CAPACITY   ACCESS MODES   STORAGECLASS   VOLUMEATTRIBUTESCLASS   AGE
data-keep-0   Bound    pvc-5d487b38-6213-435f-89fa-e84c958cc0f6   64Mi       RWO            standard       <unset>                 114s
$ kubectl delete statefulset keep
statefulset.apps "keep" deleted from default namespace
$ sleep 15
$ kubectl get pvc -l app=keep
NAME          STATUS   VOLUME                                     CAPACITY   ACCESS MODES   STORAGECLASS   VOLUMEATTRIBUTESCLASS   AGE
data-keep-0   Bound    pvc-5d487b38-6213-435f-89fa-e84c958cc0f6   64Mi       RWO            standard       <unset>                 2m10s
$ kubectl delete pvc data-keep-0
persistentvolumeclaim "data-keep-0" deleted from default namespace
$ kubectl delete pvc -l app=keep --ignore-not-found
No resources found
verify: while all three replicas are up the claims carry no ownerReferences at all, because whenDeleted: Retain means the set never owns them. Scale down and the two condemned claims each gain an ownerReference to their own pod, which is what deletes them once that pod is gone; the last claim then survives deletion of the set itself. Two halves of one field, opposite outcomes.

Expansion is a property of the driver, not of the StorageClass field that advertises it. Setting allowVolumeExpansion on a class whose provisioner cannot expand produces one of the most quoted error strings in storage tasks, so meet it here rather than in a task.

kubectl get storageclass
kubectl patch storageclass standard -p '{"allowVolumeExpansion":true}'
kubectl apply -f - <<'EOF'
apiVersion: v1
kind: PersistentVolumeClaim
metadata: { name: grow, namespace: default }
spec:
  accessModes: [ReadWriteOnce]
  resources: { requests: { storage: 64Mi } }
EOF
kubectl run grower --image=busybox:1.36 --restart=Never --overrides='{"spec":{"containers":[{"name":"c","image":"busybox:1.36","command":["sh","-c","sleep 86400"],"volumeMounts":[{"name":"d","mountPath":"/d"}]}],"volumes":[{"name":"d","persistentVolumeClaim":{"claimName":"grow"}}]}}'
kubectl wait --for=condition=Ready pod/grower --timeout=120s
kubectl patch pvc grow -p '{"spec":{"resources":{"requests":{"storage":"128Mi"}}}}'
sleep 20
kubectl get pvc grow -o jsonpath='spec: {.spec.resources.requests.storage}  status: {.status.capacity.storage}{"\n"}'
echo "conditions: $(kubectl get pvc grow -o jsonpath='{.status.conditions}')"
kubectl get pvc grow -o jsonpath='{.status.allocatedResourceStatuses}{"\n"}'
kubectl describe pvc grow | tail -8
kubectl delete pod grower
kubectl delete pvc grow
# put the shared default class back the way it was
kubectl patch storageclass standard -p '{"allowVolumeExpansion":false}'
outputcaptured 2026-09-12
$ kubectl get storageclass
NAME                 PROVISIONER             RECLAIMPOLICY   VOLUMEBINDINGMODE      ALLOWVOLUMEEXPANSION   AGE
standard (default)   rancher.io/local-path   Delete          WaitForFirstConsumer   false                  20h
$ kubectl patch storageclass standard -p '{"allowVolumeExpansion":true}'
storageclass.storage.k8s.io/standard patched
$ kubectl apply -f - <<'EOF'
apiVersion: v1
kind: PersistentVolumeClaim
metadata: { name: grow, namespace: default }
spec:
  accessModes: [ReadWriteOnce]
  resources: { requests: { storage: 64Mi } }
EOF
persistentvolumeclaim/grow created
$ kubectl run grower --image=busybox:1.36 --restart=Never --overrides='{"spec":{"containers":[{"name":"c","image":"busybox:1.36","command":["sh","-c","sleep 86400"],"volumeMounts":[{"name":"d","mountPath":"/d"}]}],"volumes":[{"name":"d","persistentVolumeClaim":{"claimName":"grow"}}]}}'
pod/grower created
$ kubectl wait --for=condition=Ready pod/grower --timeout=120s
pod/grower condition met
$ kubectl patch pvc grow -p '{"spec":{"resources":{"requests":{"storage":"128Mi"}}}}'
persistentvolumeclaim/grow patched
$ sleep 20
$ kubectl get pvc grow -o jsonpath='spec: {.spec.resources.requests.storage}  status: {.status.capacity.storage}{"\n"}'
spec: 128Mi  status: 64Mi
$ echo "conditions: $(kubectl get pvc grow -o jsonpath='{.status.conditions}')"
conditions: 
$ kubectl get pvc grow -o jsonpath='{.status.allocatedResourceStatuses}{"\n"}'
$ kubectl describe pvc grow | tail -8
Events:
  Type    Reason                 Age   From                                                                                                Message
  ----    ------                 ----  ----                                                                                                -------
  Normal  WaitForFirstConsumer   32s   persistentvolume-controller                                                                         waiting for first consumer to be created before binding
  Normal  ExternalProvisioning   31s   persistentvolume-controller                                                                         Waiting for a volume to be created either by the external provisioner 'rancher.io/local-path' or manually by the system administrator. If volume creation is delayed, please verify that the provisioner is running and correctly registered.
  Normal  Provisioning           31s   rancher.io/local-path_local-path-provisioner-855c7b7774-xsqw2_944b3c00-03a8-486a-896a-567844028b41  External provisioner is provisioning volume for claim "default/grow"
  Normal  ProvisioningSucceeded  25s   rancher.io/local-path_local-path-provisioner-855c7b7774-xsqw2_944b3c00-03a8-486a-896a-567844028b41  Successfully provisioned volume pvc-f412d670-060c-4a9a-8fda-bd4cc5645a03
  Normal  ExternalExpanding      21s   volume_expand                                                                                       waiting for an external controller to expand this PVC
$ kubectl delete pod grower
pod "grower" deleted from default namespace
$ kubectl delete pvc grow
persistentvolumeclaim "grow" deleted from default namespace
$ # put the shared default class back the way it was
$ kubectl patch storageclass standard -p '{"allowVolumeExpansion":false}'
storageclass.storage.k8s.io/standard patched
verify: the class accepts allowVolumeExpansion: true and the claim accepts the larger request, so spec reads 128Mi while status.capacity stays at 64Mi. status.conditions is empty and the only event is ExternalExpanding: waiting for an external controller to expand this PVC, which is the whole answer: local-path has no resize support, so no controller ever picks the request up and the claim waits forever with its spec and status disagreeing. There is no error to find, which is the failure mode worth recognizing.

A generic ephemeral volume gives a pod a real PVC with a real class and a real lifetime of exactly one pod. The naming rule and the ownerReference are the whole trick.

kubectl apply -f - <<'EOF'
apiVersion: v1
kind: Pod
metadata: { name: eph, namespace: default }
spec:
  containers:
    - name: c
      image: busybox:1.36
      command: ["sh", "-c", "sleep 86400"]
      volumeMounts: [{ name: scratch, mountPath: /scratch }]
  volumes:
    - name: scratch
      ephemeral:
        volumeClaimTemplate:
          spec:
            accessModes: [ReadWriteOnce]
            resources: { requests: { storage: 64Mi } }
EOF
kubectl wait --for=condition=Ready pod/eph --timeout=120s
kubectl get pvc eph-scratch
kubectl get pvc eph-scratch -o jsonpath='{.metadata.ownerReferences}' | jq
kubectl delete pod eph
sleep 15
kubectl get pvc eph-scratch
outputcaptured 2026-09-13
$ kubectl apply -f - <<'EOF'
apiVersion: v1
kind: Pod
metadata: { name: eph, namespace: default }
spec:
  containers:
    - name: c
      image: busybox:1.36
      command: ["sh", "-c", "sleep 86400"]
      volumeMounts: [{ name: scratch, mountPath: /scratch }]
  volumes:
    - name: scratch
      ephemeral:
        volumeClaimTemplate:
          spec:
            accessModes: [ReadWriteOnce]
            resources: { requests: { storage: 64Mi } }
EOF
pod/eph created
$ kubectl wait --for=condition=Ready pod/eph --timeout=120s
pod/eph condition met
$ kubectl get pvc eph-scratch
NAME          STATUS   VOLUME                                     CAPACITY   ACCESS MODES   STORAGECLASS   VOLUMEATTRIBUTESCLASS   AGE
eph-scratch   Bound    pvc-60e802fd-df09-43c4-90ba-64f0e90ad910   64Mi       RWO            standard       <unset>                 13s
$ kubectl get pvc eph-scratch -o jsonpath='{.metadata.ownerReferences}' | jq
[
  {
    "apiVersion": "v1",
    "blockOwnerDeletion": true,
    "controller": true,
    "kind": "Pod",
    "name": "eph",
    "uid": "b7e5e146-ef47-42b5-afb7-8af422d2a318"
  }
]
$ kubectl delete pod eph
pod "eph" deleted from default namespace
$ sleep 15
$ kubectl get pvc eph-scratch
Error from server (NotFound): persistentvolumeclaims "eph-scratch" not found
verify: the claim is named <pod>-<volume>, owned by the pod, and gone once the pod is.

Snapshots are three separate installs: the CRDs, the snapshot-controller, and driver support. kind with local-path has none of them, and knowing that is more useful than reciting the object names.

kubectl get crd | grep snapshot.storage.k8s.io || echo 'no snapshot CRDs installed'
kubectl api-resources | grep -i volumesnapshot || echo 'no volumesnapshot API on this cluster'
kubectl get csidrivers
kubectl get sc standard -o jsonpath='{.provisioner}{"\n"}'
outputcaptured 2026-09-12
$ kubectl get crd | grep snapshot.storage.k8s.io || echo 'no snapshot CRDs installed'
no snapshot CRDs installed
$ kubectl api-resources | grep -i volumesnapshot || echo 'no volumesnapshot API on this cluster'
no volumesnapshot API on this cluster
$ kubectl get csidrivers
NAME            ATTACHREQUIRED   PODINFOONMOUNT   STORAGECAPACITY   TOKENREQUESTS   REQUIRESREPUBLISH   MODES       AGE
csi.spiffe.io   false            true             false             <unset>         false               Ephemeral   21m
$ kubectl get sc standard -o jsonpath='{.provisioner}{"\n"}'
rancher.io/local-path
verify: you can state the three pieces a working VolumeSnapshot needs and which of them this cluster has. On the exam cluster, run the same two commands before you write a VolumeSnapshot manifest at all.

A PVC with no storageClassName is not bound to "the default forever": it records whatever default existed when admission saw it, and since retroactive default assignment it can be filled in later. Watch both halves.

kubectl get storageclass -o custom-columns='NAME:.metadata.name,DEFAULT:.metadata.annotations.storageclass\.kubernetes\.io/is-default-class'
kubectl annotate storageclass standard storageclass.kubernetes.io/is-default-class-
kubectl apply -f - <<'EOF'
apiVersion: v1
kind: PersistentVolumeClaim
metadata: { name: noclass, namespace: default }
spec:
  accessModes: [ReadWriteOnce]
  resources: { requests: { storage: 64Mi } }
EOF
kubectl get pvc noclass -o jsonpath='class=[{.spec.storageClassName}] phase={.status.phase}{"\n"}'
kubectl annotate storageclass standard storageclass.kubernetes.io/is-default-class=true
sleep 20
kubectl get pvc noclass -o jsonpath='class=[{.spec.storageClassName}] phase={.status.phase}{"\n"}'
kubectl delete pvc noclass
outputcaptured 2026-09-13
$ kubectl get storageclass -o custom-columns='NAME:.metadata.name,DEFAULT:.metadata.annotations.storageclass\.kubernetes\.io/is-default-class'
NAME       DEFAULT
standard   true
$ kubectl annotate storageclass standard storageclass.kubernetes.io/is-default-class-
storageclass.storage.k8s.io/standard annotated
$ kubectl apply -f - <<'EOF'
apiVersion: v1
kind: PersistentVolumeClaim
metadata: { name: noclass, namespace: default }
spec:
  accessModes: [ReadWriteOnce]
  resources: { requests: { storage: 64Mi } }
EOF
persistentvolumeclaim/noclass created
$ kubectl get pvc noclass -o jsonpath='class=[{.spec.storageClassName}] phase={.status.phase}{"\n"}'
class=[] phase=Pending
$ kubectl annotate storageclass standard storageclass.kubernetes.io/is-default-class=true
storageclass.storage.k8s.io/standard annotated
$ sleep 20
$ kubectl get pvc noclass -o jsonpath='class=[{.spec.storageClassName}] phase={.status.phase}{"\n"}'
class=[standard] phase=Pending
$ kubectl delete pvc noclass
persistentvolumeclaim "noclass" deleted from default namespace
verify: the claim sits Pending with an empty class while no default exists, then has standard written into its spec once the annotation is back. Restore the annotation before you leave this exercise or every later claim on the page hangs.

Self-check

answer before opening
A PVC has been Pending for ten minutes and describe shows no events at all. Bug or not?

Not a bug if the class uses WaitForFirstConsumer and nothing mounts the claim yet: no scheduling decision exists, so there is nothing to provision and nothing to report. It becomes a finding the moment a pod references it and the claim stays Pending; then the events appear and name the real reason.

A PV sits in Released and your new, identical PVC will not bind to it. Why, and what is the fix?

Retain leaves spec.claimRef pointing at the deleted claim, and a PV with a claimRef is spoken for. Clear it (kubectl patch pv <name> -p '{"spec":{"claimRef":null}}') and it returns to Available. Deciding whether the old data should be wiped first is the actual judgment call.

You delete a StatefulSet and its PVCs remain. Accident or design?

Design. The claims outlive the set so scaling back up reattaches the same data, and so an accidental delete does not destroy the database. Newer clusters can opt in to cleanup via persistentVolumeClaimRetentionPolicy (whenDeleted / whenScaled).

Two pods on different nodes must share a directory. What do you need, and what will this lab do?

A class whose provisioner supports RWX (NFS, CephFS, a cloud file service). local-path cannot, so the claim pends and the events say so. The honest lab answer is "not possible here" plus the fix you would apply in a real cluster, which is also the right exam answer when the storage cannot do what the task implies.

Which quota lines cap tenant storage, and why is per-class quota interesting?

requests.storage and persistentvolumeclaims cap the total and the count; <class>.storageclass.storage.k8s.io/requests.storage caps a specific class. The per-class form is how a platform team offers a fast expensive tier without every tenant defaulting to it: a pricing decision expressed as a Kubernetes object.

You want scale-down to reclaim volumes but deleting the StatefulSet by mistake to keep them. Which fields, and since when is this GA?

persistentVolumeClaimRetentionPolicy: { whenScaled: Delete, whenDeleted: Retain }, GA since Kubernetes 1.32 (StatefulSetAutoDeletePVC). Delete works through an ownerReference on the PVC that the garbage collector honors after the pod is gone; a pod replaced after a node failure is not a scale event and keeps its claim.

An expansion to 500Gi failed because the backend has no capacity and the PVC keeps retrying. Can you back out?

Yes: RecoverVolumeExpansionFailure is GA since 1.34, so edit the request down to any value still above status.capacity (never below it; shrinking is impossible). Watch status.allocatedResourceStatuses and events. Before 1.34 the recovery was the manual dance of Retain, delete the PVC, clear claimRef, recreate.

Restore a database volume from last night's snapshot into a new PVC. What does the manifest look like and what must be true first?

A PVC whose dataSource is { kind: VolumeSnapshot, name: <snap>, apiGroup: snapshot.storage.k8s.io }, in the same namespace, with a StorageClass from the same CSI driver and a size at least the snapshot's restore size. The VolumeSnapshot must show status.readyToUse: true; a not-ready snapshot leaves the PVC Pending with a provisioner event saying so.

How do you change a volume's IOPS tier without recreating it?

VolumeAttributesClass (GA since 1.34): create a class with the driver's performance parameters and set spec.volumeAttributesClassName on the PVC; the field is mutable and the external-resizer applies the change. StorageClass parameters, by contrast, are fixed at provisioning time.

Docs to know your way around

study time, not exam time
  • kubernetes.io: Persistent Volumes (the access-modes and reclaim tables), Storage Classes, Volume Snapshots, StatefulSet volumeClaimTemplates.
  • Offline: kubectl explain pvc.spec, kubectl explain sc, kubectl get sc -o wide, and PVC status.conditions when a resize misbehaves.
  • kubernetes.io: Persistent Volumes, "Expanding Persistent Volumes Claims" and "Recovering from Failure when Expanding Volumes": the exact rules for grow-only, filesystem-on-mount and the smaller-retry path.
  • kubernetes.io: Ephemeral Volumes and Volume Attributes Classes: generic ephemeral volume naming and ownership; the mutable volumeAttributesClassName field.
  • velero.io/docs: "How Velero Works" and "CSI snapshot support": the Backup/Schedule/Restore objects and the CSI versus file-system backup choice.
free before 1.4make down-api