The admission control you get without installing anything: three named profiles enforced by the built-in Pod Security admission controller, driven entirely by namespace labels. Less expressive than Kyverno or Gatekeeper, and that is its virtue.

needsmake up sec

Orientation

competency 5.4 · Pod Security Standards

Know exactly what each profile means and the label grammar, and these become the fastest points on the paper. PSS replaced PodSecurityPolicy, which was removed; if a scenario mentions PSP, the expected answer is "removed; use PSS, or a policy engine for anything PSS cannot express".

Profiles and modes

three profiles × three modes
ProfileBlocks / requiresUse for
privilegednothing: unrestrictedsystem namespaces, CNI, storage drivers
baselineblocks the known-bad: privileged containers, hostNetwork/hostPID/hostIPC, hostPath volumes, added capabilities beyond a safe list, unconfined seccomp/AppArmor, host portsthe sane tenant default
restricteddemands actively-good: runAsNonRoot, allowPrivilegeEscalation: false, all capabilities dropped (ALL, NET_BIND_SERVICE permitted back), seccompProfile: RuntimeDefault, no hostPath, restricted volume typesanything you can make comply

The one-liner: baseline stops you being dangerous, restricted forces you to be safe.

The modes are set independently per namespace: enforce rejects at admission, audit annotates the audit log, warn prints warnings to the client. Label grammar:

pod-security.kubernetes.io/enforce: baseline
pod-security.kubernetes.io/enforce-version: v1.31     # pin, or "latest"
pod-security.kubernetes.io/warn: restricted
pod-security.kubernetes.io/audit: restricted

Version pinning matters more than it looks: profiles gain checks across releases, so latest means an upgrade can start rejecting pods that were fine yesterday. Pinning trades that surprise for a deliberate migration.

Build a pod spec and watch the three profiles decide:

The lab's stance is the recommended one

team-a runs enforce baseline, warn and audit restricted: nothing overtly dangerous gets in, and every restricted violation is visible to both the user (a warning in their terminal) and the platform (the audit log) before anyone tightens the screw. That staged posture mirrors the Audit→Enforce choreography from 5.2; the pattern generalizes to every control in this domain.

The profiles at field level

what baseline forbids, what restricted demands, path by path

The rejection message names the control (privileged, hostPath volumes, allowPrivilegeEscalation != false, unrestricted capabilities, runAsNonRoot != true, seccompProfile). Knowing which field each control reads turns that message into a one-line edit.

ControlField(s), on pod and every container (containers, initContainers, ephemeralContainers)BaselineRestricted adds
Host namespacesspec.hostNetwork, hostPID, hostIPCmust be false/unsetsame
PrivilegedsecurityContext.privilegedfalse/unsetsame
CapabilitiessecurityContext.capabilities.add / dropadd only from: AUDIT_WRITE, CHOWN, DAC_OVERRIDE, FOWNER, FSETID, KILL, MKNOD, NET_BIND_SERVICE, SETFCAP, SETGID, SETPCAP, SETUID, SYS_CHROOTdrop: ["ALL"] required; add may contain only NET_BIND_SERVICE
HostPath volumesspec.volumes[*].hostPathforbiddensame
Host portscontainers[*].ports[*].hostPortunset or 0same
AppArmorsecurityContext.appArmorProfile.typeRuntimeDefault, Localhost, or unset; not Unconfinedsame
SELinuxsecurityContext.seLinuxOptions.type / user / roletype in container_t, container_init_t, container_kvm_t, container_engine_t or unset; user and role must be unsetsame
/proc mountsecurityContext.procMountDefault or unsetsame
SeccompsecurityContext.seccompProfile.typenot Unconfined (unset allowed)must be set: RuntimeDefault or Localhost, at pod level or on every container
Sysctlsspec.securityContext.sysctls[*].nameonly the safe set (kernel.shm_rmid_forced, net.ipv4.ip_local_port_range, net.ipv4.tcp_syncookies, net.ipv4.ping_group_range, net.ipv4.ip_unprivileged_port_start and a few more)same
Volume typesspec.volumes[*]anything but hostPathonly configMap, csi, downwardAPI, emptyDir, ephemeral, persistentVolumeClaim, projected, secret
Privilege escalationcontainers[*].securityContext.allowPrivilegeEscalationn/amust be false on every container (Linux pods)
Non-rootsecurityContext.runAsNonRoot; runAsUsern/arunAsNonRoot: true at pod level or on every container; runAsUser must not be 0 where set
Windows HostProcesssecurityContext.windowsOptions.hostProcessfalse/unsetsame
  • Pod level vs container level. runAsNonRoot, runAsUser, seccompProfile, appArmorProfile and seLinuxOptions can be set on spec.securityContext and inherited, with a container able to override; allowPrivilegeEscalation, capabilities, privileged, procMount and readOnlyRootFilesystem exist only per container. A restricted failure that says "container x" wants a per-container field; one that names the pod wants the pod-level block.
  • Not required by restricted. readOnlyRootFilesystem, resource limits, a non-root fsGroup. Tasks that ask for them are asking for a policy engine rule or a hardened manifest, not PSS.
  • Images that run as root. runAsNonRoot: true with an image whose default user is root fails at container start with container has runAsNonRoot and image will run as root (a CreateContainerConfigError), not at admission. Add runAsUser: 1000 (and usually runAsGroup), or use an image built for a non-root user.
  • Ephemeral containers count. kubectl debug into a restricted namespace fails unless you pass --profile=restricted, because the injected container also has to satisfy the profile.
Reflex

Reading a rejection: split the message on commas; each clause is one row of the table and names the container. Fix them all in one edit; PSS reports every violation at once, so a second apply should be clean.

Two operational facts that decide troubleshooting tasks

and one preview command
  1. PSS evaluates pods. A Deployment whose template violates the profile is created successfully and then fails to make pods. The evidence is on the ReplicaSet, one level below where you were looking. Same indirection as quota (1.4). Recognizing the pattern is worth more than any single profile detail.
  2. Enforcement is admission-time. Tightening a namespace label does not evict existing violators; they run until their next write. Which is why the preview command exists:
kubectl label --dry-run=server --overwrite ns <ns> pod-security.kubernetes.io/enforce=restricted
outputcaptured 2026-08-26
(a merely-baseline pod, lazy from the first exercise, is running in team-b for this preview)
$ kubectl label --dry-run=server --overwrite ns team-b pod-security.kubernetes.io/enforce=restricted
Warning: existing pods in namespace "team-b" violate the new PodSecurity enforce level "restricted:latest"
Warning: lazy: allowPrivilegeEscalation != false, unrestricted capabilities, runAsNonRoot != true, seccompProfile
namespace/team-b labeled (server dry run)

The API server evaluates every existing pod against the proposed profile and warns about each one that would break, without changing anything. It is the single most useful PSS command to know exists, and it turns a risky tightening into a planned migration.

What a compliant pod looks like

securityContext:                 # pod level
  runAsNonRoot: true
  runAsUser: 1000
  seccompProfile: { type: RuntimeDefault }
containers:
  - name: app
    securityContext:             # container level
      allowPrivilegeEscalation: false
      readOnlyRootFilesystem: true          # not required by restricted, but good practice
      capabilities: { drop: ["ALL"] }

The demo app's manifest (examples/demo-app/base/deployment.yaml) is a worked answer to "make this pass restricted": read its securityContext blocks line by line, because writing exactly that block from memory is a plausible task. Note the level split: runAsNonRoot and seccompProfile sit naturally at pod level, capabilities and allowPrivilegeEscalation are per container.

PSS versus a policy engine

PSS is fixed, built in, free, and covers exactly the pod-hardening dimension. A policy engine is arbitrary, installable, and covers everything else (image registries, labels, ownership, cross-object rules) at the cost of running a webhook. The professional answer to "which should we use" is both: PSS as the always-on floor, an engine for organization-specific rules. The engine can also enforce PSS-like rules in namespaces where you need exceptions PSS cannot express.

Cluster-wide defaults, exemptions and the newer knobs

AdmissionConfiguration, workload resources, user namespaces, AppArmor

Defaults and exemptions live on the API server

Namespace labels are the per-namespace layer. The cluster-wide layer is the PodSecurity admission plugin's configuration, passed to kube-apiserver with --admission-control-config-file:

apiVersion: apiserver.config.k8s.io/v1
kind: AdmissionConfiguration
plugins:
- name: PodSecurity
  configuration:
    apiVersion: pod-security.admission.config.k8s.io/v1
    kind: PodSecurityConfiguration
    defaults:                       # applied to namespaces with no label for that mode
      enforce: baseline
      enforce-version: latest
      audit: restricted
      audit-version: latest
      warn: restricted
      warn-version: latest
    exemptions:
      usernames: []                 # authenticated or impersonated users
      runtimeClasses: [kata]        # pods with this runtimeClassName
      namespaces: [kube-system]     # whole namespaces
  • A namespace label always overrides the default for that mode; the default only fills gaps. So "every new namespace must be at least baseline" is this file, not a Kyverno policy that labels namespaces (though that works too and is the answer when you cannot touch the API server, which on managed clusters you usually cannot).
  • Exempt requests skip all three modes. Exempting a username exempts what that user creates directly, not what a controller creates on their behalf, which is why the docs say not to exempt controller ServiceAccounts such as the ReplicaSet controller: that would exempt everyone who can create a Deployment.
  • pod_security_evaluations_total and pod_security_exemptions_total on the API server tell you how often each mode fires; a spike in decision="deny" after a label change is your rollout signal.

Workload resources and updates

warn and audit are applied to Deployments, StatefulSets, Jobs and other pod-template resources, so the person applying a Deployment sees the warning; enforce is applied only to the resulting pods, which is the two-level failure from the exercises. Pod updates that touch only metadata (other than seccomp or AppArmor annotations), activeDeadlineSeconds or tolerations are exempt from re-evaluation; anything that creates a new pod, including a Deployment rollout, is evaluated in full, which is when a tightened label first bites a running workload. kubectl label --dry-run=server is the preview for the label change; kubectl apply --dry-run=server on a workload previews the warnings without creating it.

Newer knobs that change what "compliant" means

  • User namespaces (GA in 1.36). spec.hostUsers: false maps container UIDs to unprivileged host UIDs, so root in the container is not root on the node. Needs a Linux 6.3+ kernel with idmap mounts on every filesystem the pod uses (tmpfs included), containerd 2.0+ or CRI-O, and runc 1.2+ or crun 1.9+. It is Linux-only and does not apply to hostPath or host namespaces. PSS still evaluates the pod normally; user namespaces are defense in depth under it.
  • AppArmor (GA in 1.30). The field is securityContext.appArmorProfile: { type: RuntimeDefault | Localhost | Unconfined, localhostProfile: name }; the old container.apparmor.security.beta.kubernetes.io/<container> annotation is deprecated and rejected when it disagrees with the field. Baseline reads the field.
  • Seccomp. RuntimeDefault uses the container runtime's profile; Localhost needs localhostProfile relative to the kubelet's /var/lib/kubelet/seccomp/ directory on every node. The kubelet flag --seccomp-default (or seccompDefault: true in KubeletConfiguration) makes RuntimeDefault the default for pods that set nothing, which is how you satisfy restricted's seccomp requirement fleet-wide without editing manifests; PSS, though, checks the pod spec, not the kubelet default, so restricted still wants the field written.
  • PodSecurityPolicy was removed in 1.25; a policy/v1beta1 PodSecurityPolicy manifest in a task is a trick, and the answer is labels plus an engine.
How this gets tested

"Make namespace X enforce restricted without breaking its current workloads" is: preview with --dry-run=server, fix the manifests the warnings name (usually allowPrivilegeEscalation, capabilities.drop, runAsNonRoot, seccompProfile), then set the label with a pinned enforce-version. "Allow one system workload that needs hostPath in an otherwise restricted cluster" is an exemption by namespace or a separate privileged namespace, never a per-pod exception, because PSS has none.

Exercises

tick the dot when its check passes

In team-a (enforce: baseline):

kubectl -n team-a run priv --image=busybox:1.37 --restart=Never \
  --overrides='{"spec":{"containers":[{"name":"priv","image":"busybox:1.37","command":["sleep","300"],"securityContext":{"privileged":true}}]}}'
outputcaptured 2026-08-26
$ kubectl -n team-a run priv --image=busybox:1.37 --restart=Never \
  --overrides='{"spec":{"containers":[{"name":"priv","image":"busybox:1.37","command":["sleep","300"],"securityContext":{"privileged":true}}]}}'
Error from server (Forbidden): pods "priv" is forbidden: violates PodSecurity "baseline:latest": privileged (container "priv" must not set securityContext.privileged=true)

Then run a merely-lazy pod (kubectl -n team-a run lazy --image=busybox:1.37 --restart=Never -- sleep 300) and read the warnings it prints: admitted under baseline, flagged against restricted by the warn label.

verify: the first is rejected at admission with an error naming the violated control; the second is admitted with warnings. Two different outcomes, one namespace, and you can name which label produced each line of output.

Create namespace hardened with enforce restricted. Take the lazy pod spec and make it pass: runAsNonRoot: true (busybox needs runAsUser too, e.g. 1000), allowPrivilegeEscalation: false, capabilities drop ALL, seccompProfile RuntimeDefault. Iterate against the live error messages; they name the missing control each time, which makes PSS self-documenting under exam conditions.

verify: the pod runs in hardened, and your final securityContext agrees with the demo app's.

In hardened, kubectl -n hardened create deploy sneaky --image=busybox:1.37 -- sleep 300. The deploy is created; no pod appears. Find the rejection where it actually lives:

kubectl -n hardened get deploy sneaky   # READY 0/1, no error here
kubectl -n hardened describe rs -l app=sneaky | tail -5
outputcaptured 2026-08-26
$ kubectl -n hardened get deploy sneaky   # READY 0/1, no error here
NAME     READY   UP-TO-DATE   AVAILABLE   AGE
sneaky   0/1     0            0           11s
$ kubectl -n hardened describe rs -l app=sneaky | tail -5
  Warning  FailedCreate  11s              replicaset-controller  Error creating: pods "sneaky-595c4bc9c7-bs46b" is forbidden: violates PodSecurity "restricted:latest": allowPrivilegeEscalation != false (container "busybox" must set securityContext.allowPrivilegeEscalation=false), unrestricted capabilities (container "busybox" must set securityContext.capabilities.drop=["ALL"]), runAsNonRoot != true (pod or container "busybox" must set securityContext.runAsNonRoot=true), seccompProfile (pod or container "busybox" must set securityContext.seccompProfile.type to "RuntimeDefault" or "Localhost")
  Warning  FailedCreate  11s              replicaset-controller  Error creating: pods "sneaky-595c4bc9c7-df7n2" is forbidden: violates PodSecurity "restricted:latest": allowPrivilegeEscalation != false (container "busybox" must set securityContext.allowPrivilegeEscalation=false), unrestricted capabilities (container "busybox" must set securityContext.capabilities.drop=["ALL"]), runAsNonRoot != true (pod or container "busybox" must set securityContext.runAsNonRoot=true), seccompProfile (pod or container "busybox" must set securityContext.seccompProfile.type to "RuntimeDefault" or "Localhost")
  Warning  FailedCreate  11s              replicaset-controller  Error creating: pods "sneaky-595c4bc9c7-v8xwt" is forbidden: violates PodSecurity "restricted:latest": allowPrivilegeEscalation != false (container "busybox" must set securityContext.allowPrivilegeEscalation=false), unrestricted capabilities (container "busybox" must set securityContext.capabilities.drop=["ALL"]), runAsNonRoot != true (pod or container "busybox" must set securityContext.runAsNonRoot=true), seccompProfile (pod or container "busybox" must set securityContext.seccompProfile.type to "RuntimeDefault" or "Localhost")
  Warning  FailedCreate  10s              replicaset-controller  Error creating: pods "sneaky-595c4bc9c7-z48jz" is forbidden: violates PodSecurity "restricted:latest": allowPrivilegeEscalation != false (container "busybox" must set securityContext.allowPrivilegeEscalation=false), unrestricted capabilities (container "busybox" must set securityContext.capabilities.drop=["ALL"]), runAsNonRoot != true (pod or container "busybox" must set securityContext.runAsNonRoot=true), seccompProfile (pod or container "busybox" must set securityContext.seccompProfile.type to "RuntimeDefault" or "Localhost")
  Warning  FailedCreate  1s (x3 over 9s)  replicaset-controller  (combined from similar events): Error creating: pods "sneaky-595c4bc9c7-pf2pl" is forbidden: violates PodSecurity "restricted:latest": allowPrivilegeEscalation != false (container "busybox" must set securityContext.allowPrivilegeEscalation=false), unrestricted capabilities (container "busybox" must set securityContext.capabilities.drop=["ALL"]), runAsNonRoot != true (pod or container "busybox" must set securityContext.runAsNonRoot=true), seccompProfile (pod or container "busybox" must set securityContext.seccompProfile.type to "RuntimeDefault" or "Localhost")
verify: the ReplicaSet events carry the PSS denial. This indirection is a stock exam scenario for every admission mechanism (5.2's engines included); PSS is just where it is cheapest to practice.

kubectl label --dry-run=server --overwrite ns team-b pod-security.kubernetes.io/enforce=restricted and read which existing pods would violate.

verify: a warning per non-compliant pod, zero mutations made. Pair with the audit log (section 5.4) to see what the audit label has been quietly recording all along.

Pod Security admits this pod happily: the spec is compliant. The kubelet then refuses to start it, because the image's user is root and the spec said not root. The refusal therefore lands in the container status, not at admission, which is why the pod is created and never runs.

kubectl create ns hardened --dry-run=client -o yaml | kubectl apply -f -
kubectl label ns hardened pod-security.kubernetes.io/enforce=restricted --overwrite
kubectl -n hardened apply -f - <<'EOF'
apiVersion: v1
kind: Pod
metadata: { name: rootimage }
spec:
  securityContext:
    runAsNonRoot: true
    seccompProfile: { type: RuntimeDefault }
  containers:
    - name: c
      image: nginx:1.27-alpine
      securityContext:
        allowPrivilegeEscalation: false
        capabilities: { drop: ["ALL"] }
EOF
sleep 20
kubectl -n hardened get pod rootimage -o jsonpath='{.status.containerStatuses[0].state}' | jq
kubectl -n hardened describe pod rootimage | sed -n '/Events/,$p'
kubectl -n hardened delete pod rootimage
outputcaptured 2026-09-12
$ kubectl create ns hardened --dry-run=client -o yaml | kubectl apply -f -
namespace/hardened created
$ kubectl label ns hardened pod-security.kubernetes.io/enforce=restricted --overwrite
namespace/hardened labeled
$ kubectl -n hardened apply -f - <<'EOF'
apiVersion: v1
kind: Pod
metadata: { name: rootimage }
spec:
  securityContext:
    runAsNonRoot: true
    seccompProfile: { type: RuntimeDefault }
  containers:
    - name: c
      image: nginx:1.27-alpine
      securityContext:
        allowPrivilegeEscalation: false
        capabilities: { drop: ["ALL"] }
EOF
pod/rootimage created
$ sleep 20
$ kubectl -n hardened get pod rootimage -o jsonpath='{.status.containerStatuses[0].state}' | jq
{
  "waiting": {
    "message": "container has runAsNonRoot and image will run as root (pod: \"rootimage_hardened(b1353a96-63bb-4895-b6f2-838249d72469)\", container: c)",
    "reason": "CreateContainerConfigError"
  }
}
$ kubectl -n hardened describe pod rootimage | sed -n '/Events/,$p'
Events:
  Type     Reason           Age               From               Message
  ----     ------           ----              ----               -------
  Warning  PolicyViolation  21s               kyverno-admission  policy require-resource-requests/ fail: every container must set cpu and memory requests
  Normal   Scheduled        21s               default-scheduler  Successfully assigned hardened/rootimage to cnpe-worker
  Normal   Pulled           7s (x3 over 19s)  kubelet            spec.containers{c}: Container image "nginx:1.27-alpine" already present on machine and can be accessed by the pod
  Warning  Failed           7s (x3 over 19s)  kubelet            spec.containers{c}: Error: container has runAsNonRoot and image will run as root (pod: "rootimage_hardened(b1353a96-63bb-4895-b6f2-838249d72469)", container: c)
$ kubectl -n hardened delete pod rootimage
pod "rootimage" deleted from hardened namespace
verify: admission accepts the pod and the container status is CreateContainerConfigError with an event saying the container has runAsNonRoot and the image will run as root.

kubectl debug adds an ephemeral container, and that container goes through admission like any other. In a restricted namespace the default debug container is refused, and the profile flag is the fix people do not know exists.

kubectl -n hardened run target --image=ghcr.io/nginxinc/nginx-unprivileged:1.27-alpine --restart=Never --overrides='{"spec":{"securityContext":{"runAsNonRoot":true,"seccompProfile":{"type":"RuntimeDefault"}},"containers":[{"name":"target","image":"ghcr.io/nginxinc/nginx-unprivileged:1.27-alpine","securityContext":{"allowPrivilegeEscalation":false,"capabilities":{"drop":["ALL"]}}}]}}'
kubectl -n hardened wait --for=condition=Ready pod/target --timeout=120s
kubectl -n hardened debug target --image=busybox:1.36 --target=target -- true
kubectl -n hardened debug target --image=busybox:1.36 --target=target --profile=restricted -- true
kubectl -n hardened get pod target -o jsonpath='{range .spec.ephemeralContainers[*]}{.name} {.securityContext.allowPrivilegeEscalation} {.securityContext.capabilities.drop}{"\n"}{end}'
sleep 10
kubectl -n hardened get pod target -o json | jq -r '.status.ephemeralContainerStatuses[]? | "\(.name) \(.state|keys[0])"'
kubectl -n hardened delete pod target
outputcaptured 2026-09-13
$ kubectl -n hardened run target --image=ghcr.io/nginxinc/nginx-unprivileged:1.27-alpine --restart=Never --overrides='{"spec":{"securityContext":{"runAsNonRoot":true,"seccompProfile":{"type":"RuntimeDefault"}},"containers":[{"name":"target","image":"ghcr.io/nginxinc/nginx-unprivileged:1.27-alpine","securityContext":{"allowPrivilegeEscalation":false,"capabilities":{"drop":["ALL"]}}}]}}'
pod/target created
$ kubectl -n hardened wait --for=condition=Ready pod/target --timeout=120s
pod/target condition met
$ kubectl -n hardened debug target --image=busybox:1.36 --target=target -- true
Targeting container "target". If you don't see processes from this container it may be because the container runtime doesn't support this feature.
Defaulting debug container name to debugger-s56nq.
Error from server (Forbidden): pods "target" is forbidden: violates PodSecurity "restricted:latest": allowPrivilegeEscalation != false (container "debugger-s56nq" must set securityContext.allowPrivilegeEscalation=false), unrestricted capabilities (container "debugger-s56nq" must set securityContext.capabilities.drop=["ALL"]; container "debugger-s56nq" must not include "SYS_PTRACE" in securityContext.capabilities.add)
$ kubectl -n hardened debug target --image=busybox:1.36 --target=target --profile=restricted -- true
Targeting container "target". If you don't see processes from this container it may be because the container runtime doesn't support this feature.
Defaulting debug container name to debugger-67shh.
$ kubectl -n hardened get pod target -o jsonpath='{range .spec.ephemeralContainers[*]}{.name} {.securityContext.allowPrivilegeEscalation} {.securityContext.capabilities.drop}{"\n"}{end}'
debugger-67shh false ["ALL"]
$ sleep 10
$ kubectl -n hardened get pod target -o json | jq -r '.status.ephemeralContainerStatuses[]? | "\(.name) \(.state|keys[0])"'
debugger-67shh waiting
$ kubectl -n hardened delete pod target
pod "target" deleted from hardened namespace
verify: the first attempt is rejected by Pod Security naming the fields the ephemeral container is missing, and --profile=restricted is accepted. The spec listing then shows the accepted debug container with allowPrivilegeEscalation false and capabilities drop ["ALL"], which is what the profile built for you. Read the spec rather than the status: the kubelet writes ephemeralContainerStatuses a beat later, so a status read this soon shows waiting, or nothing at all. The profile changes the container it builds, not the namespace's policy.

hostUsers: false maps the container's root to an unprivileged uid on the host, which defangs a whole class of escapes. Whether it works depends on the kernel and the runtime, and in this lab it does not: a kind node is itself a container, and the runtime will not give a nested user namespace the mounts it needs. The refusal is the exercise, because it is the same message you will read on any cluster where the feature is unavailable.

kubectl -n default apply -f - <<'EOF'
apiVersion: v1
kind: Pod
metadata: { name: userns }
spec:
  hostUsers: false
  containers:
    - name: c
      image: busybox:1.36
      command: ["sh", "-c", "sleep 600"]
EOF
sleep 30
kubectl -n default get pod userns -o jsonpath='{.status.phase} {.status.conditions[*].message}{"\n"}'
kubectl -n default get events --field-selector involvedObject.name=userns -o jsonpath='{range .items[*]}{.reason}: {.message}{"\n"}{end}' | tail -1
kubectl -n default exec userns -- cat /proc/self/uid_map
kubectl -n default exec userns -- id
kubectl -n default delete pod userns
outputcaptured 2026-09-13
$ kubectl -n default apply -f - <<'EOF'
apiVersion: v1
kind: Pod
metadata: { name: userns }
spec:
  hostUsers: false
  containers:
    - name: c
      image: busybox:1.36
      command: ["sh", "-c", "sleep 600"]
EOF
pod/userns created
$ sleep 30
$ kubectl -n default get pod userns -o jsonpath='{.status.phase} {.status.conditions[*].message}{"\n"}'
Pending containers with unready status: [c] containers with unready status: [c]
$ kubectl -n default get events --field-selector involvedObject.name=userns -o jsonpath='{range .items[*]}{.reason}: {.message}{"\n"}{end}' | tail -1
FailedCreatePodSandBox: Failed to create pod sandbox: rpc error: code = Unknown desc = failed to start sandbox "8f8262b522a1e5b8f11a084559e731dd00390acf1c8d8f4ebddd62f0452dc38d": failed to create containerd task: failed to create shim task: OCI runtime create failed: runc create failed: unable to start container process: error during container init: error mounting "sysfs" to rootfs at "/sys": mount src=sysfs, dst=/sys, dstFd=/proc/thread-self/fd/11, flags=MS_RDONLY|MS_NOSUID|MS_NODEV|MS_NOEXEC: operation not permitted
$ kubectl -n default exec userns -- cat /proc/self/uid_map
error: unable to upgrade connection: container not found ("c")
$ kubectl -n default exec userns -- id
error: unable to upgrade connection: container not found ("c")
$ kubectl -n default delete pod userns
pod "userns" deleted from default namespace
verify: the pod stays Pending with containers with unready status: [c], and the event is FailedCreatePodSandBox ... error mounting "sysfs" to rootfs at "/sys": ... operation not permitted. That is the runtime, not the API, refusing: admission accepted hostUsers: false quite happily, so the exec that follows fails with container not found rather than with anything about user namespaces. On a cluster where it does work, cat /proc/self/uid_map inside the pod shows container uid 0 mapped to a high host uid rather than to 0, which is the check worth remembering.

Warn and audit labels let you find out what restricted would reject before you enforce it. A server-side dry run is the same question asked about one workload, and it costs nothing.

kubectl create ns warnzone --dry-run=client -o yaml | kubectl apply -f -
kubectl label ns warnzone pod-security.kubernetes.io/warn=restricted pod-security.kubernetes.io/audit=restricted --overwrite
kubectl -n warnzone apply --dry-run=server -f - <<'EOF'
apiVersion: apps/v1
kind: Deployment
metadata: { name: legacy-app }
spec:
  replicas: 1
  selector: { matchLabels: { app: legacy-app } }
  template:
    metadata: { labels: { app: legacy-app } }
    spec:
      containers:
        - name: c
          image: nginx:1.27-alpine
EOF
kubectl -n warnzone apply -f - <<'EOF'
apiVersion: apps/v1
kind: Deployment
metadata: { name: legacy-app }
spec:
  replicas: 1
  selector: { matchLabels: { app: legacy-app } }
  template:
    metadata: { labels: { app: legacy-app } }
    spec:
      containers:
        - name: c
          image: nginx:1.27-alpine
EOF
kubectl -n warnzone get deploy legacy-app
kubectl delete ns warnzone
outputcaptured 2026-09-12
$ kubectl create ns warnzone --dry-run=client -o yaml | kubectl apply -f -
namespace/warnzone created
$ kubectl label ns warnzone pod-security.kubernetes.io/warn=restricted pod-security.kubernetes.io/audit=restricted --overwrite
namespace/warnzone labeled
$ kubectl -n warnzone apply --dry-run=server -f - <<'EOF'
apiVersion: apps/v1
kind: Deployment
metadata: { name: legacy-app }
spec:
  replicas: 1
  selector: { matchLabels: { app: legacy-app } }
  template:
    metadata: { labels: { app: legacy-app } }
    spec:
      containers:
        - name: c
          image: nginx:1.27-alpine
EOF
Warning: would violate PodSecurity "restricted:latest": allowPrivilegeEscalation != false (container "c" must set securityContext.allowPrivilegeEscalation=false), unrestricted capabilities (container "c" must set securityContext.capabilities.drop=["ALL"]), runAsNonRoot != true (pod or container "c" must set securityContext.runAsNonRoot=true), seccompProfile (pod or container "c" must set securityContext.seccompProfile.type to "RuntimeDefault" or "Localhost")
deployment.apps/legacy-app created (server dry run)
$ kubectl -n warnzone apply -f - <<'EOF'
apiVersion: apps/v1
kind: Deployment
metadata: { name: legacy-app }
spec:
  replicas: 1
  selector: { matchLabels: { app: legacy-app } }
  template:
    metadata: { labels: { app: legacy-app } }
    spec:
      containers:
        - name: c
          image: nginx:1.27-alpine
EOF
Warning: would violate PodSecurity "restricted:latest": allowPrivilegeEscalation != false (container "c" must set securityContext.allowPrivilegeEscalation=false), unrestricted capabilities (container "c" must set securityContext.capabilities.drop=["ALL"]), runAsNonRoot != true (pod or container "c" must set securityContext.runAsNonRoot=true), seccompProfile (pod or container "c" must set securityContext.seccompProfile.type to "RuntimeDefault" or "Localhost")
deployment.apps/legacy-app created
$ kubectl -n warnzone get deploy legacy-app
NAME         READY   UP-TO-DATE   AVAILABLE   AGE
legacy-app   0/1     0            0           0s
$ kubectl delete ns warnzone
namespace "warnzone" deleted
verify: the dry run prints the restricted warning listing every field the workload is missing, and the real apply prints the same warning and creates it anyway. That list is the migration plan; paste it into the ticket.

Self-check

answer before opening
Name the three profiles in one clause each.

Privileged: no restrictions. Baseline: blocks known privilege escalations (privileged containers, host namespaces, hostPath, extra capabilities). Restricted: additionally requires non-root, no privilege escalation, all capabilities dropped, and RuntimeDefault seccomp.

A Deployment is created and no pods appear. Where is the PSS error?

On the ReplicaSet: PSS is a pod-level admission decision, so the Deployment controller records the rejection one level down. kubectl describe rs -l … or namespace events.

You want to know who breaks if you tighten a namespace. Command?

kubectl label --dry-run=server --overwrite ns <ns> pod-security.kubernetes.io/enforce=restricted: the server evaluates existing pods and warns for each violator, changing nothing.

Why pin enforce-version?

Because profiles gain checks in new Kubernetes releases: with latest, a cluster upgrade can start rejecting pods that were compliant before. Pinning turns that into a deliberate migration you schedule rather than a surprise you discover.

PSS or Kyverno for "images must come from our registry"?

Kyverno (or Gatekeeper): PSS covers only the fixed pod-hardening dimension and has no notion of registries. PSS is the floor; the engine is for organization-specific rules, and it is fine, indeed normal, to run both.

A pod passes restricted admission and then sits in CreateContainerConfigError with container has runAsNonRoot and image will run as root. What happened?

You set runAsNonRoot: true, which satisfies PSS at admission, but the image's default user is root, so the kubelet refuses to start it. Set runAsUser (and runAsGroup) to a non-zero id, or use an image whose USER is non-root. PSS checks the spec; the kubelet checks the image.

How do you make every namespace without labels enforce baseline and warn on restricted, cluster-wide, and what does a namespace label do to that default?

An AdmissionConfiguration with a PodSecurityConfiguration whose defaults set enforce: baseline, warn: restricted (plus versions), passed to the API server with --admission-control-config-file. A namespace label for a mode overrides the default for that mode only. On a managed control plane you cannot set this, so the fallback is a policy engine that labels namespaces at creation.

Restricted rejects a container for seccompProfile, yet the kubelet runs with seccompDefault: true. Why is that not enough?

PSS evaluates the pod spec at admission; the kubelet default is applied later, at container creation, and is invisible to the admission controller. Restricted requires seccompProfile.type to be RuntimeDefault or Localhost in the spec, at pod level or on every container. The kubelet flag is still worth setting: it protects pods in namespaces that are not restricted.

Docs to know your way around

study time, not exam time
  • kubernetes.io: Pod Security Standards (the profile tables; do not memorize the capability lists, know where they are) and Enforce Pod Security Standards with Namespace Labels.
  • Offline: the rejection messages themselves; PSS names every violated control, which makes it the most self-documenting admission mechanism in the cluster; plus kubectl explain pod.spec.securityContext.
  • kubernetes.io/docs/concepts/security/pod-security-standards: the field-by-field tables (use them, do not memorize the capability list); /docs/concepts/security/pod-security-admission for workload-resource behavior and exemptions.
  • kubernetes.io/docs/tasks/configure-pod-container/enforce-standards-admission-controller: the AdmissionConfiguration file; /docs/concepts/workloads/pods/user-namespaces for hostUsers and its runtime requirements.