The admission control you get without installing anything: three named profiles enforced by the built-in Pod Security admission controller, driven entirely by namespace labels. Less expressive than Kyverno or Gatekeeper, and that is its virtue.
make up secOrientation
Know exactly what each profile means and the label grammar, and these become the fastest points on the paper. PSS replaced PodSecurityPolicy, which was removed; if a scenario mentions PSP, the expected answer is "removed; use PSS, or a policy engine for anything PSS cannot express".
Profiles and modes
| Profile | Blocks / requires | Use for |
|---|---|---|
| privileged | nothing: unrestricted | system namespaces, CNI, storage drivers |
| baseline | blocks the known-bad: privileged containers, hostNetwork/hostPID/hostIPC, hostPath volumes, added capabilities beyond a safe list, unconfined seccomp/AppArmor, host ports | the sane tenant default |
| restricted | demands actively-good: runAsNonRoot, allowPrivilegeEscalation: false, all capabilities dropped (ALL, NET_BIND_SERVICE permitted back), seccompProfile: RuntimeDefault, no hostPath, restricted volume types | anything you can make comply |
The one-liner: baseline stops you being dangerous, restricted forces you to be safe.
The modes are set independently per namespace: enforce rejects at admission, audit annotates the audit log, warn prints warnings to the client. Label grammar:
pod-security.kubernetes.io/enforce: baseline
pod-security.kubernetes.io/enforce-version: v1.31 # pin, or "latest"
pod-security.kubernetes.io/warn: restricted
pod-security.kubernetes.io/audit: restrictedVersion pinning matters more than it looks: profiles gain checks across releases, so latest means an upgrade can start rejecting pods that were fine yesterday. Pinning trades that surprise for a deliberate migration.
Build a pod spec and watch the three profiles decide:
team-a runs enforce baseline, warn and audit restricted: nothing overtly dangerous gets in, and every restricted violation is visible to both the user (a warning in their terminal) and the platform (the audit log) before anyone tightens the screw. That staged posture mirrors the Audit→Enforce choreography from 5.2; the pattern generalizes to every control in this domain.
The profiles at field level
The rejection message names the control (privileged, hostPath volumes, allowPrivilegeEscalation != false, unrestricted capabilities, runAsNonRoot != true, seccompProfile). Knowing which field each control reads turns that message into a one-line edit.
| Control | Field(s), on pod and every container (containers, initContainers, ephemeralContainers) | Baseline | Restricted adds |
|---|---|---|---|
| Host namespaces | spec.hostNetwork, hostPID, hostIPC | must be false/unset | same |
| Privileged | securityContext.privileged | false/unset | same |
| Capabilities | securityContext.capabilities.add / drop | add only from: AUDIT_WRITE, CHOWN, DAC_OVERRIDE, FOWNER, FSETID, KILL, MKNOD, NET_BIND_SERVICE, SETFCAP, SETGID, SETPCAP, SETUID, SYS_CHROOT | drop: ["ALL"] required; add may contain only NET_BIND_SERVICE |
| HostPath volumes | spec.volumes[*].hostPath | forbidden | same |
| Host ports | containers[*].ports[*].hostPort | unset or 0 | same |
| AppArmor | securityContext.appArmorProfile.type | RuntimeDefault, Localhost, or unset; not Unconfined | same |
| SELinux | securityContext.seLinuxOptions.type / user / role | type in container_t, container_init_t, container_kvm_t, container_engine_t or unset; user and role must be unset | same |
| /proc mount | securityContext.procMount | Default or unset | same |
| Seccomp | securityContext.seccompProfile.type | not Unconfined (unset allowed) | must be set: RuntimeDefault or Localhost, at pod level or on every container |
| Sysctls | spec.securityContext.sysctls[*].name | only the safe set (kernel.shm_rmid_forced, net.ipv4.ip_local_port_range, net.ipv4.tcp_syncookies, net.ipv4.ping_group_range, net.ipv4.ip_unprivileged_port_start and a few more) | same |
| Volume types | spec.volumes[*] | anything but hostPath | only configMap, csi, downwardAPI, emptyDir, ephemeral, persistentVolumeClaim, projected, secret |
| Privilege escalation | containers[*].securityContext.allowPrivilegeEscalation | n/a | must be false on every container (Linux pods) |
| Non-root | securityContext.runAsNonRoot; runAsUser | n/a | runAsNonRoot: true at pod level or on every container; runAsUser must not be 0 where set |
| Windows HostProcess | securityContext.windowsOptions.hostProcess | false/unset | same |
- Pod level vs container level.
runAsNonRoot,runAsUser,seccompProfile,appArmorProfileandseLinuxOptionscan be set onspec.securityContextand inherited, with a container able to override;allowPrivilegeEscalation,capabilities,privileged,procMountandreadOnlyRootFilesystemexist only per container. A restricted failure that says "container x" wants a per-container field; one that names the pod wants the pod-level block. - Not required by restricted.
readOnlyRootFilesystem, resource limits, a non-rootfsGroup. Tasks that ask for them are asking for a policy engine rule or a hardened manifest, not PSS. - Images that run as root.
runAsNonRoot: truewith an image whose default user is root fails at container start withcontainer has runAsNonRoot and image will run as root(aCreateContainerConfigError), not at admission. AddrunAsUser: 1000(and usuallyrunAsGroup), or use an image built for a non-root user. - Ephemeral containers count.
kubectl debuginto a restricted namespace fails unless you pass--profile=restricted, because the injected container also has to satisfy the profile.
Reading a rejection: split the message on commas; each clause is one row of the table and names the container. Fix them all in one edit; PSS reports every violation at once, so a second apply should be clean.
Two operational facts that decide troubleshooting tasks
- PSS evaluates pods. A Deployment whose template violates the profile is created successfully and then fails to make pods. The evidence is on the ReplicaSet, one level below where you were looking. Same indirection as quota (1.4). Recognizing the pattern is worth more than any single profile detail.
- Enforcement is admission-time. Tightening a namespace label does not evict existing violators; they run until their next write. Which is why the preview command exists:
kubectl label --dry-run=server --overwrite ns <ns> pod-security.kubernetes.io/enforce=restrictedoutputcaptured 2026-08-26
(a merely-baseline pod, lazy from the first exercise, is running in team-b for this preview)
$ kubectl label --dry-run=server --overwrite ns team-b pod-security.kubernetes.io/enforce=restricted
Warning: existing pods in namespace "team-b" violate the new PodSecurity enforce level "restricted:latest"
Warning: lazy: allowPrivilegeEscalation != false, unrestricted capabilities, runAsNonRoot != true, seccompProfile
namespace/team-b labeled (server dry run)The API server evaluates every existing pod against the proposed profile and warns about each one that would break, without changing anything. It is the single most useful PSS command to know exists, and it turns a risky tightening into a planned migration.
What a compliant pod looks like
securityContext: # pod level
runAsNonRoot: true
runAsUser: 1000
seccompProfile: { type: RuntimeDefault }
containers:
- name: app
securityContext: # container level
allowPrivilegeEscalation: false
readOnlyRootFilesystem: true # not required by restricted, but good practice
capabilities: { drop: ["ALL"] }The demo app's manifest (examples/demo-app/base/deployment.yaml) is a worked answer to "make this pass restricted": read its securityContext blocks line by line, because writing exactly that block from memory is a plausible task. Note the level split: runAsNonRoot and seccompProfile sit naturally at pod level, capabilities and allowPrivilegeEscalation are per container.
PSS is fixed, built in, free, and covers exactly the pod-hardening dimension. A policy engine is arbitrary, installable, and covers everything else (image registries, labels, ownership, cross-object rules) at the cost of running a webhook. The professional answer to "which should we use" is both: PSS as the always-on floor, an engine for organization-specific rules. The engine can also enforce PSS-like rules in namespaces where you need exceptions PSS cannot express.
Cluster-wide defaults, exemptions and the newer knobs
Defaults and exemptions live on the API server
Namespace labels are the per-namespace layer. The cluster-wide layer is the PodSecurity admission plugin's configuration, passed to kube-apiserver with --admission-control-config-file:
apiVersion: apiserver.config.k8s.io/v1
kind: AdmissionConfiguration
plugins:
- name: PodSecurity
configuration:
apiVersion: pod-security.admission.config.k8s.io/v1
kind: PodSecurityConfiguration
defaults: # applied to namespaces with no label for that mode
enforce: baseline
enforce-version: latest
audit: restricted
audit-version: latest
warn: restricted
warn-version: latest
exemptions:
usernames: [] # authenticated or impersonated users
runtimeClasses: [kata] # pods with this runtimeClassName
namespaces: [kube-system] # whole namespaces- A namespace label always overrides the default for that mode; the default only fills gaps. So "every new namespace must be at least baseline" is this file, not a Kyverno policy that labels namespaces (though that works too and is the answer when you cannot touch the API server, which on managed clusters you usually cannot).
- Exempt requests skip all three modes. Exempting a username exempts what that user creates directly, not what a controller creates on their behalf, which is why the docs say not to exempt controller ServiceAccounts such as the ReplicaSet controller: that would exempt everyone who can create a Deployment.
pod_security_evaluations_totalandpod_security_exemptions_totalon the API server tell you how often each mode fires; a spike indecision="deny"after a label change is your rollout signal.
Workload resources and updates
warn and audit are applied to Deployments, StatefulSets, Jobs and other pod-template resources, so the person applying a Deployment sees the warning; enforce is applied only to the resulting pods, which is the two-level failure from the exercises. Pod updates that touch only metadata (other than seccomp or AppArmor annotations), activeDeadlineSeconds or tolerations are exempt from re-evaluation; anything that creates a new pod, including a Deployment rollout, is evaluated in full, which is when a tightened label first bites a running workload. kubectl label --dry-run=server is the preview for the label change; kubectl apply --dry-run=server on a workload previews the warnings without creating it.
Newer knobs that change what "compliant" means
- User namespaces (GA in 1.36).
spec.hostUsers: falsemaps container UIDs to unprivileged host UIDs, so root in the container is not root on the node. Needs a Linux 6.3+ kernel with idmap mounts on every filesystem the pod uses (tmpfs included), containerd 2.0+ or CRI-O, and runc 1.2+ or crun 1.9+. It is Linux-only and does not apply tohostPathor host namespaces. PSS still evaluates the pod normally; user namespaces are defense in depth under it. - AppArmor (GA in 1.30). The field is
securityContext.appArmorProfile: { type: RuntimeDefault | Localhost | Unconfined, localhostProfile: name }; the oldcontainer.apparmor.security.beta.kubernetes.io/<container>annotation is deprecated and rejected when it disagrees with the field. Baseline reads the field. - Seccomp.
RuntimeDefaultuses the container runtime's profile;LocalhostneedslocalhostProfilerelative to the kubelet's/var/lib/kubelet/seccomp/directory on every node. The kubelet flag--seccomp-default(orseccompDefault: trueinKubeletConfiguration) makesRuntimeDefaultthe default for pods that set nothing, which is how you satisfy restricted's seccomp requirement fleet-wide without editing manifests; PSS, though, checks the pod spec, not the kubelet default, so restricted still wants the field written. - PodSecurityPolicy was removed in 1.25; a
policy/v1beta1 PodSecurityPolicymanifest in a task is a trick, and the answer is labels plus an engine.
"Make namespace X enforce restricted without breaking its current workloads" is: preview with --dry-run=server, fix the manifests the warnings name (usually allowPrivilegeEscalation, capabilities.drop, runAsNonRoot, seccompProfile), then set the label with a pinned enforce-version. "Allow one system workload that needs hostPath in an otherwise restricted cluster" is an exemption by namespace or a separate privileged namespace, never a per-pod exception, because PSS has none.
Exercises
In team-a (enforce: baseline):
kubectl -n team-a run priv --image=busybox:1.37 --restart=Never \
--overrides='{"spec":{"containers":[{"name":"priv","image":"busybox:1.37","command":["sleep","300"],"securityContext":{"privileged":true}}]}}'outputcaptured 2026-08-26
$ kubectl -n team-a run priv --image=busybox:1.37 --restart=Never \
--overrides='{"spec":{"containers":[{"name":"priv","image":"busybox:1.37","command":["sleep","300"],"securityContext":{"privileged":true}}]}}'
Error from server (Forbidden): pods "priv" is forbidden: violates PodSecurity "baseline:latest": privileged (container "priv" must not set securityContext.privileged=true)Then run a merely-lazy pod (kubectl -n team-a run lazy --image=busybox:1.37 --restart=Never -- sleep 300) and read the warnings it prints: admitted under baseline, flagged against restricted by the warn label.
Create namespace hardened with enforce restricted. Take the lazy pod spec and make it pass: runAsNonRoot: true (busybox needs runAsUser too, e.g. 1000), allowPrivilegeEscalation: false, capabilities drop ALL, seccompProfile RuntimeDefault. Iterate against the live error messages; they name the missing control each time, which makes PSS self-documenting under exam conditions.
hardened, and your final securityContext agrees with the demo app's.In hardened, kubectl -n hardened create deploy sneaky --image=busybox:1.37 -- sleep 300. The deploy is created; no pod appears. Find the rejection where it actually lives:
kubectl -n hardened get deploy sneaky # READY 0/1, no error here
kubectl -n hardened describe rs -l app=sneaky | tail -5outputcaptured 2026-08-26
$ kubectl -n hardened get deploy sneaky # READY 0/1, no error here
NAME READY UP-TO-DATE AVAILABLE AGE
sneaky 0/1 0 0 11s
$ kubectl -n hardened describe rs -l app=sneaky | tail -5
Warning FailedCreate 11s replicaset-controller Error creating: pods "sneaky-595c4bc9c7-bs46b" is forbidden: violates PodSecurity "restricted:latest": allowPrivilegeEscalation != false (container "busybox" must set securityContext.allowPrivilegeEscalation=false), unrestricted capabilities (container "busybox" must set securityContext.capabilities.drop=["ALL"]), runAsNonRoot != true (pod or container "busybox" must set securityContext.runAsNonRoot=true), seccompProfile (pod or container "busybox" must set securityContext.seccompProfile.type to "RuntimeDefault" or "Localhost")
Warning FailedCreate 11s replicaset-controller Error creating: pods "sneaky-595c4bc9c7-df7n2" is forbidden: violates PodSecurity "restricted:latest": allowPrivilegeEscalation != false (container "busybox" must set securityContext.allowPrivilegeEscalation=false), unrestricted capabilities (container "busybox" must set securityContext.capabilities.drop=["ALL"]), runAsNonRoot != true (pod or container "busybox" must set securityContext.runAsNonRoot=true), seccompProfile (pod or container "busybox" must set securityContext.seccompProfile.type to "RuntimeDefault" or "Localhost")
Warning FailedCreate 11s replicaset-controller Error creating: pods "sneaky-595c4bc9c7-v8xwt" is forbidden: violates PodSecurity "restricted:latest": allowPrivilegeEscalation != false (container "busybox" must set securityContext.allowPrivilegeEscalation=false), unrestricted capabilities (container "busybox" must set securityContext.capabilities.drop=["ALL"]), runAsNonRoot != true (pod or container "busybox" must set securityContext.runAsNonRoot=true), seccompProfile (pod or container "busybox" must set securityContext.seccompProfile.type to "RuntimeDefault" or "Localhost")
Warning FailedCreate 10s replicaset-controller Error creating: pods "sneaky-595c4bc9c7-z48jz" is forbidden: violates PodSecurity "restricted:latest": allowPrivilegeEscalation != false (container "busybox" must set securityContext.allowPrivilegeEscalation=false), unrestricted capabilities (container "busybox" must set securityContext.capabilities.drop=["ALL"]), runAsNonRoot != true (pod or container "busybox" must set securityContext.runAsNonRoot=true), seccompProfile (pod or container "busybox" must set securityContext.seccompProfile.type to "RuntimeDefault" or "Localhost")
Warning FailedCreate 1s (x3 over 9s) replicaset-controller (combined from similar events): Error creating: pods "sneaky-595c4bc9c7-pf2pl" is forbidden: violates PodSecurity "restricted:latest": allowPrivilegeEscalation != false (container "busybox" must set securityContext.allowPrivilegeEscalation=false), unrestricted capabilities (container "busybox" must set securityContext.capabilities.drop=["ALL"]), runAsNonRoot != true (pod or container "busybox" must set securityContext.runAsNonRoot=true), seccompProfile (pod or container "busybox" must set securityContext.seccompProfile.type to "RuntimeDefault" or "Localhost")kubectl label --dry-run=server --overwrite ns team-b pod-security.kubernetes.io/enforce=restricted and read which existing pods would violate.
audit label has been quietly recording all along.Pod Security admits this pod happily: the spec is compliant. The kubelet then refuses to start it, because the image's user is root and the spec said not root. The refusal therefore lands in the container status, not at admission, which is why the pod is created and never runs.
kubectl create ns hardened --dry-run=client -o yaml | kubectl apply -f -
kubectl label ns hardened pod-security.kubernetes.io/enforce=restricted --overwrite
kubectl -n hardened apply -f - <<'EOF'
apiVersion: v1
kind: Pod
metadata: { name: rootimage }
spec:
securityContext:
runAsNonRoot: true
seccompProfile: { type: RuntimeDefault }
containers:
- name: c
image: nginx:1.27-alpine
securityContext:
allowPrivilegeEscalation: false
capabilities: { drop: ["ALL"] }
EOF
sleep 20
kubectl -n hardened get pod rootimage -o jsonpath='{.status.containerStatuses[0].state}' | jq
kubectl -n hardened describe pod rootimage | sed -n '/Events/,$p'
kubectl -n hardened delete pod rootimageoutputcaptured 2026-09-12
$ kubectl create ns hardened --dry-run=client -o yaml | kubectl apply -f -
namespace/hardened created
$ kubectl label ns hardened pod-security.kubernetes.io/enforce=restricted --overwrite
namespace/hardened labeled
$ kubectl -n hardened apply -f - <<'EOF'
apiVersion: v1
kind: Pod
metadata: { name: rootimage }
spec:
securityContext:
runAsNonRoot: true
seccompProfile: { type: RuntimeDefault }
containers:
- name: c
image: nginx:1.27-alpine
securityContext:
allowPrivilegeEscalation: false
capabilities: { drop: ["ALL"] }
EOF
pod/rootimage created
$ sleep 20
$ kubectl -n hardened get pod rootimage -o jsonpath='{.status.containerStatuses[0].state}' | jq
{
"waiting": {
"message": "container has runAsNonRoot and image will run as root (pod: \"rootimage_hardened(b1353a96-63bb-4895-b6f2-838249d72469)\", container: c)",
"reason": "CreateContainerConfigError"
}
}
$ kubectl -n hardened describe pod rootimage | sed -n '/Events/,$p'
Events:
Type Reason Age From Message
---- ------ ---- ---- -------
Warning PolicyViolation 21s kyverno-admission policy require-resource-requests/ fail: every container must set cpu and memory requests
Normal Scheduled 21s default-scheduler Successfully assigned hardened/rootimage to cnpe-worker
Normal Pulled 7s (x3 over 19s) kubelet spec.containers{c}: Container image "nginx:1.27-alpine" already present on machine and can be accessed by the pod
Warning Failed 7s (x3 over 19s) kubelet spec.containers{c}: Error: container has runAsNonRoot and image will run as root (pod: "rootimage_hardened(b1353a96-63bb-4895-b6f2-838249d72469)", container: c)
$ kubectl -n hardened delete pod rootimage
pod "rootimage" deleted from hardened namespaceCreateContainerConfigError with an event saying the container has runAsNonRoot and the image will run as root.kubectl debug adds an ephemeral container, and that container goes through admission like any other. In a restricted namespace the default debug container is refused, and the profile flag is the fix people do not know exists.
kubectl -n hardened run target --image=ghcr.io/nginxinc/nginx-unprivileged:1.27-alpine --restart=Never --overrides='{"spec":{"securityContext":{"runAsNonRoot":true,"seccompProfile":{"type":"RuntimeDefault"}},"containers":[{"name":"target","image":"ghcr.io/nginxinc/nginx-unprivileged:1.27-alpine","securityContext":{"allowPrivilegeEscalation":false,"capabilities":{"drop":["ALL"]}}}]}}'
kubectl -n hardened wait --for=condition=Ready pod/target --timeout=120s
kubectl -n hardened debug target --image=busybox:1.36 --target=target -- true
kubectl -n hardened debug target --image=busybox:1.36 --target=target --profile=restricted -- true
kubectl -n hardened get pod target -o jsonpath='{range .spec.ephemeralContainers[*]}{.name} {.securityContext.allowPrivilegeEscalation} {.securityContext.capabilities.drop}{"\n"}{end}'
sleep 10
kubectl -n hardened get pod target -o json | jq -r '.status.ephemeralContainerStatuses[]? | "\(.name) \(.state|keys[0])"'
kubectl -n hardened delete pod targetoutputcaptured 2026-09-13
$ kubectl -n hardened run target --image=ghcr.io/nginxinc/nginx-unprivileged:1.27-alpine --restart=Never --overrides='{"spec":{"securityContext":{"runAsNonRoot":true,"seccompProfile":{"type":"RuntimeDefault"}},"containers":[{"name":"target","image":"ghcr.io/nginxinc/nginx-unprivileged:1.27-alpine","securityContext":{"allowPrivilegeEscalation":false,"capabilities":{"drop":["ALL"]}}}]}}'
pod/target created
$ kubectl -n hardened wait --for=condition=Ready pod/target --timeout=120s
pod/target condition met
$ kubectl -n hardened debug target --image=busybox:1.36 --target=target -- true
Targeting container "target". If you don't see processes from this container it may be because the container runtime doesn't support this feature.
Defaulting debug container name to debugger-s56nq.
Error from server (Forbidden): pods "target" is forbidden: violates PodSecurity "restricted:latest": allowPrivilegeEscalation != false (container "debugger-s56nq" must set securityContext.allowPrivilegeEscalation=false), unrestricted capabilities (container "debugger-s56nq" must set securityContext.capabilities.drop=["ALL"]; container "debugger-s56nq" must not include "SYS_PTRACE" in securityContext.capabilities.add)
$ kubectl -n hardened debug target --image=busybox:1.36 --target=target --profile=restricted -- true
Targeting container "target". If you don't see processes from this container it may be because the container runtime doesn't support this feature.
Defaulting debug container name to debugger-67shh.
$ kubectl -n hardened get pod target -o jsonpath='{range .spec.ephemeralContainers[*]}{.name} {.securityContext.allowPrivilegeEscalation} {.securityContext.capabilities.drop}{"\n"}{end}'
debugger-67shh false ["ALL"]
$ sleep 10
$ kubectl -n hardened get pod target -o json | jq -r '.status.ephemeralContainerStatuses[]? | "\(.name) \(.state|keys[0])"'
debugger-67shh waiting
$ kubectl -n hardened delete pod target
pod "target" deleted from hardened namespace--profile=restricted is accepted. The spec listing then shows the accepted debug container with allowPrivilegeEscalation false and capabilities drop ["ALL"], which is what the profile built for you. Read the spec rather than the status: the kubelet writes ephemeralContainerStatuses a beat later, so a status read this soon shows waiting, or nothing at all. The profile changes the container it builds, not the namespace's policy.hostUsers: false maps the container's root to an unprivileged uid on the host, which defangs a whole class of escapes. Whether it works depends on the kernel and the runtime, and in this lab it does not: a kind node is itself a container, and the runtime will not give a nested user namespace the mounts it needs. The refusal is the exercise, because it is the same message you will read on any cluster where the feature is unavailable.
kubectl -n default apply -f - <<'EOF'
apiVersion: v1
kind: Pod
metadata: { name: userns }
spec:
hostUsers: false
containers:
- name: c
image: busybox:1.36
command: ["sh", "-c", "sleep 600"]
EOF
sleep 30
kubectl -n default get pod userns -o jsonpath='{.status.phase} {.status.conditions[*].message}{"\n"}'
kubectl -n default get events --field-selector involvedObject.name=userns -o jsonpath='{range .items[*]}{.reason}: {.message}{"\n"}{end}' | tail -1
kubectl -n default exec userns -- cat /proc/self/uid_map
kubectl -n default exec userns -- id
kubectl -n default delete pod usernsoutputcaptured 2026-09-13
$ kubectl -n default apply -f - <<'EOF'
apiVersion: v1
kind: Pod
metadata: { name: userns }
spec:
hostUsers: false
containers:
- name: c
image: busybox:1.36
command: ["sh", "-c", "sleep 600"]
EOF
pod/userns created
$ sleep 30
$ kubectl -n default get pod userns -o jsonpath='{.status.phase} {.status.conditions[*].message}{"\n"}'
Pending containers with unready status: [c] containers with unready status: [c]
$ kubectl -n default get events --field-selector involvedObject.name=userns -o jsonpath='{range .items[*]}{.reason}: {.message}{"\n"}{end}' | tail -1
FailedCreatePodSandBox: Failed to create pod sandbox: rpc error: code = Unknown desc = failed to start sandbox "8f8262b522a1e5b8f11a084559e731dd00390acf1c8d8f4ebddd62f0452dc38d": failed to create containerd task: failed to create shim task: OCI runtime create failed: runc create failed: unable to start container process: error during container init: error mounting "sysfs" to rootfs at "/sys": mount src=sysfs, dst=/sys, dstFd=/proc/thread-self/fd/11, flags=MS_RDONLY|MS_NOSUID|MS_NODEV|MS_NOEXEC: operation not permitted
$ kubectl -n default exec userns -- cat /proc/self/uid_map
error: unable to upgrade connection: container not found ("c")
$ kubectl -n default exec userns -- id
error: unable to upgrade connection: container not found ("c")
$ kubectl -n default delete pod userns
pod "userns" deleted from default namespacePending with containers with unready status: [c], and the event is FailedCreatePodSandBox ... error mounting "sysfs" to rootfs at "/sys": ... operation not permitted. That is the runtime, not the API, refusing: admission accepted hostUsers: false quite happily, so the exec that follows fails with container not found rather than with anything about user namespaces. On a cluster where it does work, cat /proc/self/uid_map inside the pod shows container uid 0 mapped to a high host uid rather than to 0, which is the check worth remembering.Warn and audit labels let you find out what restricted would reject before you enforce it. A server-side dry run is the same question asked about one workload, and it costs nothing.
kubectl create ns warnzone --dry-run=client -o yaml | kubectl apply -f -
kubectl label ns warnzone pod-security.kubernetes.io/warn=restricted pod-security.kubernetes.io/audit=restricted --overwrite
kubectl -n warnzone apply --dry-run=server -f - <<'EOF'
apiVersion: apps/v1
kind: Deployment
metadata: { name: legacy-app }
spec:
replicas: 1
selector: { matchLabels: { app: legacy-app } }
template:
metadata: { labels: { app: legacy-app } }
spec:
containers:
- name: c
image: nginx:1.27-alpine
EOF
kubectl -n warnzone apply -f - <<'EOF'
apiVersion: apps/v1
kind: Deployment
metadata: { name: legacy-app }
spec:
replicas: 1
selector: { matchLabels: { app: legacy-app } }
template:
metadata: { labels: { app: legacy-app } }
spec:
containers:
- name: c
image: nginx:1.27-alpine
EOF
kubectl -n warnzone get deploy legacy-app
kubectl delete ns warnzoneoutputcaptured 2026-09-12
$ kubectl create ns warnzone --dry-run=client -o yaml | kubectl apply -f -
namespace/warnzone created
$ kubectl label ns warnzone pod-security.kubernetes.io/warn=restricted pod-security.kubernetes.io/audit=restricted --overwrite
namespace/warnzone labeled
$ kubectl -n warnzone apply --dry-run=server -f - <<'EOF'
apiVersion: apps/v1
kind: Deployment
metadata: { name: legacy-app }
spec:
replicas: 1
selector: { matchLabels: { app: legacy-app } }
template:
metadata: { labels: { app: legacy-app } }
spec:
containers:
- name: c
image: nginx:1.27-alpine
EOF
Warning: would violate PodSecurity "restricted:latest": allowPrivilegeEscalation != false (container "c" must set securityContext.allowPrivilegeEscalation=false), unrestricted capabilities (container "c" must set securityContext.capabilities.drop=["ALL"]), runAsNonRoot != true (pod or container "c" must set securityContext.runAsNonRoot=true), seccompProfile (pod or container "c" must set securityContext.seccompProfile.type to "RuntimeDefault" or "Localhost")
deployment.apps/legacy-app created (server dry run)
$ kubectl -n warnzone apply -f - <<'EOF'
apiVersion: apps/v1
kind: Deployment
metadata: { name: legacy-app }
spec:
replicas: 1
selector: { matchLabels: { app: legacy-app } }
template:
metadata: { labels: { app: legacy-app } }
spec:
containers:
- name: c
image: nginx:1.27-alpine
EOF
Warning: would violate PodSecurity "restricted:latest": allowPrivilegeEscalation != false (container "c" must set securityContext.allowPrivilegeEscalation=false), unrestricted capabilities (container "c" must set securityContext.capabilities.drop=["ALL"]), runAsNonRoot != true (pod or container "c" must set securityContext.runAsNonRoot=true), seccompProfile (pod or container "c" must set securityContext.seccompProfile.type to "RuntimeDefault" or "Localhost")
deployment.apps/legacy-app created
$ kubectl -n warnzone get deploy legacy-app
NAME READY UP-TO-DATE AVAILABLE AGE
legacy-app 0/1 0 0 0s
$ kubectl delete ns warnzone
namespace "warnzone" deletedSelf-check
Name the three profiles in one clause each.
Privileged: no restrictions. Baseline: blocks known privilege escalations (privileged containers, host namespaces, hostPath, extra capabilities). Restricted: additionally requires non-root, no privilege escalation, all capabilities dropped, and RuntimeDefault seccomp.
A Deployment is created and no pods appear. Where is the PSS error?
On the ReplicaSet: PSS is a pod-level admission decision, so the Deployment controller records the rejection one level down. kubectl describe rs -l … or namespace events.
You want to know who breaks if you tighten a namespace. Command?
kubectl label --dry-run=server --overwrite ns <ns> pod-security.kubernetes.io/enforce=restricted: the server evaluates existing pods and warns for each violator, changing nothing.
Why pin enforce-version?
Because profiles gain checks in new Kubernetes releases: with latest, a cluster upgrade can start rejecting pods that were compliant before. Pinning turns that into a deliberate migration you schedule rather than a surprise you discover.
PSS or Kyverno for "images must come from our registry"?
Kyverno (or Gatekeeper): PSS covers only the fixed pod-hardening dimension and has no notion of registries. PSS is the floor; the engine is for organization-specific rules, and it is fine, indeed normal, to run both.
A pod passes restricted admission and then sits in CreateContainerConfigError with container has runAsNonRoot and image will run as root. What happened?
You set runAsNonRoot: true, which satisfies PSS at admission, but the image's default user is root, so the kubelet refuses to start it. Set runAsUser (and runAsGroup) to a non-zero id, or use an image whose USER is non-root. PSS checks the spec; the kubelet checks the image.
How do you make every namespace without labels enforce baseline and warn on restricted, cluster-wide, and what does a namespace label do to that default?
An AdmissionConfiguration with a PodSecurityConfiguration whose defaults set enforce: baseline, warn: restricted (plus versions), passed to the API server with --admission-control-config-file. A namespace label for a mode overrides the default for that mode only. On a managed control plane you cannot set this, so the fallback is a policy engine that labels namespaces at creation.
Restricted rejects a container for seccompProfile, yet the kubelet runs with seccompDefault: true. Why is that not enough?
PSS evaluates the pod spec at admission; the kubelet default is applied later, at container creation, and is invisible to the admission controller. Restricted requires seccompProfile.type to be RuntimeDefault or Localhost in the spec, at pod level or on every container. The kubelet flag is still worth setting: it protects pods in namespaces that are not restricted.
Docs to know your way around
- kubernetes.io: Pod Security Standards (the profile tables; do not memorize the capability lists, know where they are) and Enforce Pod Security Standards with Namespace Labels.
- Offline: the rejection messages themselves; PSS names every violated control, which makes it the most self-documenting admission mechanism in the cluster; plus
kubectl explain pod.spec.securityContext. - kubernetes.io/docs/concepts/security/pod-security-standards: the field-by-field tables (use them, do not memorize the capability list); /docs/concepts/security/pod-security-admission for workload-resource behavior and exemptions.
- kubernetes.io/docs/tasks/configure-pod-container/enforce-standards-admission-controller: the
AdmissionConfigurationfile; /docs/concepts/workloads/pods/user-namespaces forhostUsersand its runtime requirements.