Requests and limits look like beginner material and then show up everywhere: quotas count requests, OpenCost bills requests, LimitRanges inject them, Kyverno policies demand them, HPAs divide by them, and the resources break drill corrupts them. Get the mechanics exact and four other sections get easier.
make upOrientation
Two numbers per container decide scheduling, cost, eviction order, throttling, and whether your workload survives a noisy neighbor. This section states precisely what each number does at each layer.
A request is a claim against the scheduler's arithmetic; a limit is a ceiling enforced by the kernel at runtime. Nothing reconciles the two after scheduling, which is how a node can be 30% used and 100% requested at the same time. Section 1.5 turns that gap into money.
How requests and limits are enforced
What the scheduler does with requests
The scheduler sums the requests of all non-terminated pods on a node and compares against the node's allocatable, not its capacity. Allocatable is capacity minus kube-reserved, system-reserved, and the eviction threshold. Live usage is never consulted. A node running at 5% CPU with every millicore requested is, to the scheduler, completely full.
node capacity ████████████████████████████████████████ 4000m
allocatable ██████████████████████████████████░░░░░░ 3800m (minus kubelet/system/eviction)
sum of requests ████████████████████████████████░░░░░░░░ 3600m ← what the scheduler sees
actual usage ██████████░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░ 1100m ← what you actually use; the rest is paid for anyway
└────────────────────┘
the right-sizing gap: 2500m reserved and never used
The four numbers, live: drag the requests up until the node refuses another pod, then drag usage down and watch what you are paying for:
What the kernel does with limits
| Resource | request becomes | limit becomes | over the limit |
|---|---|---|---|
| cpu | cpu.weight (relative share) | cpu.max (quota per 100ms period) | throttled: silently, no event, no restart |
| memory | eviction ranking only | memory.max | OOMKilled: instant, by the kernel |
| ephemeral-storage | scheduling claim | kubelet-enforced | pod evicted |
| hugepages / devices | requests must equal limits; no overcommit possible | ||
CPU is compressible: exceed the limit and the CFS scheduler simply stops giving you cycles until the next 100ms period. Latency-sensitive services with a tight CPU limit get periodic stalls that look like network problems; container_cpu_cfs_throttled_periods_total is where the truth lives, and a nonzero rate on a healthy-looking service is a real finding. Memory is incompressible: there is no "slow down", only the OOM killer.
For a pod with --restart=Never, an OOM kill lands in .status.containerStatuses[0].state.terminated.reason. In a crash-looping Deployment, the container has already restarted, so the same evidence is in lastState.terminated.reason and state shows waiting: CrashLoopBackOff. Reading the wrong field and concluding "no OOM here" is a classic self-inflicted wound.
QoS, derived not memorized
| Class | Condition | Eviction order | oom_score_adj |
|---|---|---|---|
| Guaranteed | every container: requests == limits for both cpu and memory | last | -997 |
| Burstable | at least one request or limit set, but not Guaranteed | middle: those exceeding requests first | 2–999, scaled |
| BestEffort | nothing set at all | first | 1000 |
Derive it rather than recall it. Set the four numbers and read the class, the eviction rank and the two runtime behaviors:
Under node memory pressure the kubelet ranks victims in a specific order: pods whose usage exceeds their requests first, then by Pod Priority (lowest goes first), then by how far over the request they are. QoS class is a consequence of that first test rather than an input: a BestEffort pod requested nothing, so it always exceeds. The practical corollary: a high-priority Burstable pod outlives a low-priority BestEffort one, which is why PriorityClass protects platform components from eviction, not just from preemption. The same ranking is mirrored to the kernel via oom_score_adj so a system-level OOM picks the same victim. CPU pressure never evicts anything; it only throttles. If someone tells you a pod was "evicted for CPU", they are describing something else.
Placement: the levers, in the order you reach for them
| Lever | Direction | Hard or soft | Reach for it when |
|---|---|---|---|
| nodeSelector | pod → node labels | hard only | trivially simple pinning; blunt |
| nodeAffinity | pod → node labels | required… / preferred… with weights | expressive pinning, operators (In, NotIn, Exists, Gt, Lt) |
| podAffinity / podAntiAffinity | pod → other pods, within a topologyKey | both | co-locate with a cache; spread replicas across nodes |
| taints + tolerations | node repels pods | NoSchedule, PreferNoSchedule, NoExecute | dedicated pools, control planes, GPU nodes |
| topologySpreadConstraints | even distribution over a label domain | DoNotSchedule / ScheduleAnyway | zone/node spread with a bounded skew, the modern default |
| priorityClass | who wins when the node is full | preemption | platform components must outrank tenant workloads |
Three details that decide tasks. Taints repel pods, tolerations do not attract: a toleration only says "I can live here", never "put me here". Pairing a taint with a matching nodeAffinity is how you actually dedicate a pool. NoExecute evicts already-running pods that lack the toleration, and tolerationSeconds is what makes the node-not-ready eviction delay configurable. And anti-affinity is expensive to evaluate at scale, which is exactly why topologySpreadConstraints exists: same intent, bounded cost, plus maxSkew to say how even is even enough.
The kind config labels its workers into two zones (topology.kubernetes.io/zone) precisely so spread constraints do something observable. kubectl get nodes -L topology.kubernetes.io/zone shows you the domains before you write the constraint. Read the labels first, or you will write a constraint against a key no node carries, which fails open under ScheduleAnyway and closed under DoNotSchedule.
One neighbor worth naming because availability tasks touch it: a PodDisruptionBudget constrains voluntary disruptions (drains, node upgrades) with minAvailable or maxUnavailable. It does nothing about crashes or node failures. A PDB of minAvailable: 100% is the classic way to make a cluster upgrade hang forever, and recognizing that stalled-drain symptom is worth a mark.
Autoscaling: three different scalers, three different jobs
| Scaler | Changes | Input | Gotcha |
|---|---|---|---|
| HPA | replica count | live metrics (Resource, Pods, Object, External) | needs metrics-server; percentage targets are relative to requests |
| VPA | the requests themselves | historical usage | fights an HPA on the same resource; updateMode: Off makes it advisory |
| Cluster Autoscaler / Karpenter | node count / node size | unschedulable pods | reacts to Pending pods, so it is downstream of requests too |
| KEDA | replicas, including 0 | queue depth, topic lag, cron, any external source | a ScaledObject generates an HPA underneath; it is the thing that produces those "External" metrics, and the only way to scale to zero |
The HPA formula, which explains every surprise
desiredReplicas = ceil( currentReplicas × ( currentMetricValue / desiredMetricValue ) )With one caveat that explains most "why didn't it scale" moments: the HPA does nothing while the ratio sits inside its tolerance (10% by default), and pods that are unready or missing metrics are left out of the average entirely. With --cpu-percent=20 and containers requesting 25m, the target is 5m of actual usage per pod. That is why a demo app with a tiny request scales up under a trivial load: the percentage is a fraction of the request, not of the node. Change the request and you change the autoscaler's behavior without touching the HPA.
Scaling up is fast; scaling down waits out a stabilization window (300s by default) so a brief dip cannot flap your fleet. Both directions are configurable per-HPA under spec.behavior with policies (pods or percent, per period) and selectPolicy. Knowing that downscale lag exists, and that it is deliberate, is worth a mark on its own.
VPA modes
Off: compute recommendations only. The exam-relevant one: a pure right-sizing oracle you can read without letting anything evict.Initial: apply recommendations at pod creation only.Auto/Recreate: evict and recreate pods to resize them. Disruptive by design, because changing a running pod's requests historically required a new pod.
Newer clusters can resize CPU and memory in place, which softens that trade-off. Do not test for it with kubectl explain pod.spec.containers.resizePolicy; that field has been in the schema since 1.27 whether or not the feature is on. The honest check is whether the subresource exists: kubectl get --raw /api/v1 | grep -o 'pods/resize', or simply try the patch and read the error.
kubectl top pod --containers for the instantaneous truth, kubectl describe node for the requests-vs-allocatable table at the bottom (the single most useful capacity view kubectl gives you), and kubectl get vpa -o jsonpath='{.items[*].status.recommendation…}' for what the numbers ought to be.
Node scaling and the disruption vocabulary
The autoscaling table above names the node scalers; a scenario question names their fields. Both react to the same input (Pending pods whose requests fit nowhere) and both consolidate on requests, never on live usage, which is why right-sizing requests (1.5) is the precondition for any node-level saving.
Cluster Autoscaler vs Karpenter
| Aspect | Cluster Autoscaler | Karpenter |
|---|---|---|
| Unit it manages | pre-configured node groups (cloud VM groups); picks the group via an expander (least-waste default, priority, most-pods, random) | NodePool constraints (instance families, zones, arch, karpenter.sh/capacity-type spot/on-demand) and per-node NodeClaims; auto-provisions the cheapest instance that fits |
| Scale-down | node unneeded for --scale-down-unneeded-time (10m default) below --scale-down-utilization-threshold (0.5) | spec.disruption.consolidationPolicy: WhenEmpty, WhenEmptyOrUnderutilized; consolidateAfter (or Never); can replace a node with a cheaper one |
| Lifecycle | autoscaling only | also drift (re-create nodes when NodePool or AMI changes) and expiry (expireAfter, default 720h) |
| Rate limiting | flags | spec.disruption.budgets (default nodes: 10%, optional cron schedule/duration and reasons); spec.limits caps total CPU/memory a pool may own |
| Pod opt-out | cluster-autoscaler.kubernetes.io/safe-to-evict: "false" | karpenter.sh/do-not-disrupt: "true" (or a duration) |
| What blocks removal | a restrictive PDB, pods with no controller, pods with local storage (CA), kube-system pods without a PDB (CA); both engines respect PDBs and both give up on a node whose eviction is refused | |
Karpenter posts Unconsolidatable events on the node that name the blocker ("pdb default/x prevents pod evictions", "can't replace with a lower-priced node"), which is where to look when a node you expected to disappear is still there.
PodDisruptionBudget
- One of
minAvailableormaxUnavailable, integer or percentage; the intended count comes from the owning workload's.spec.replicasvia ownerReferences, so a PDB over bare pods cannot use percentages. - It gates the Eviction API only (drain, autoscaler consolidation, descheduler). Node crashes, OOM kills and rolling updates are not evictions; they still count against the budget but are not blocked by it.
unhealthyPodEvictionPolicy(GA since 1.31):IfHealthyBudget(default) refuses to evict even unhealthy pods until the budget is met, which is how a crash-looping app stalls a drain;AlwaysAllowlets unhealthy pods go. Kubernetes' own recommendation isAlwaysAllowfor anything that may misbehave.- Scheduler preemption respects PDBs on a best-effort basis only: if no PDB-safe victim exists, it preempts anyway.
PriorityClass, the fields
- Cluster-scoped;
valueup to 1000000000 for user classes;system-cluster-critical(2000000000) andsystem-node-critical(2000001000) are reserved for the control plane and node agents. globalDefault: trueon at most one class applies to pods with nopriorityClassName; otherwise the default priority is 0. It does not retroactively change running pods.preemptionPolicy: Nevermakes a class queue ahead of lower priorities without evicting anyone: the batch-job pattern. The defaultPreemptLowerPriorityevicts.- A preempting pod gets
status.nominatedNodeNamewhile victims drain; victims receive their fullterminationGracePeriodSeconds. - Quota can fence a class: a ResourceQuota with
scopeSelectorscopeName: PriorityClass, operator: In, values: [high]caps how much a tenant may run at that priority (1.4).
Where allocatable comes from
allocatable = capacity - kubeReserved - systemReserved - evictionHard. The kubelet fields are kubeReserved (kubelet, container runtime), systemReserved (sshd, kernel, journald) and evictionHard, whose defaults are memory.available: 100Mi, nodefs.available: 10%, nodefs.inodesFree: 5%, imagefs.available: 15%. enforceNodeAllocatable: [pods] is the default enforcement; adding kube-reserved or system-reserved enforces those cgroups too and is where clusters get hurt when the estimate is wrong. Read all four numbers from kubectl describe node (Capacity vs Allocatable) before you reason about why a pod does not fit.
Node sizing rules of thumb
- Default cap is 110 pods per node (
maxPods); DaemonSets cost one pod and their requests on every node, so many small nodes multiply that overhead, while few huge nodes widen the blast radius of a node failure and make bin-packing lumpier. Two to three zones with at least one spare node's worth of headroom per zone is the usual middle. - Request-heavy, bursty workloads pack better on larger nodes; Guaranteed pods (requests equal limits) pack worst because nothing can borrow their slack.
- Spot or preemptible capacity belongs in its own NodePool with a taint, for replaceable workloads with a PDB, never for anything holding a single-writer volume.
"A node has been unschedulable and cordoned for an hour; the drain never finishes" is a PDB task: kubectl get pdb -A, find the one whose ALLOWED DISRUPTIONS is 0, and decide whether the right fix is scaling the workload up, relaxing minAvailable, or setting unhealthyPodEvictionPolicy: AlwaysAllow because the pods are unhealthy anyway. Deleting the PDB is the answer only when the task says so.
In-place resize and what else moved recently
In-place pod resize
InPlacePodVerticalScaling is GA since Kubernetes 1.35 (alpha 1.27, beta and on by default 1.33), so on a 1.36 cluster the pods/resize subresource exists and the schema check the panel above warns about is moot. The mechanics:
- You patch through the subresource:
kubectl patch pod x --subresource resize --patch '{...}'(alsokubectl edit pod x --subresource resize). A plain patch ofspec.containers[].resourcesis rejected as immutable. resizePolicyper container and per resource:restartPolicy: NotRequired(default, apply live) orRestartContainer. A pod withrestartPolicy: Nevermay only useNotRequired.- Progress is in pod conditions:
PodResizePendingwith reasonInfeasible(node can never fit it) orDeferred(retry later, for example when another pod leaves);PodResizeInProgresswhile the kubelet actuates, withreason: Errorif the runtime refused.status.containerStatuses[].allocatedResourcesshows what was actually granted, andstatus.observedGenerationtells you whether the kubelet has seen your latest spec. - Only CPU and memory; QoS class is fixed at creation (a Guaranteed pod must stay requests-equal-limits, a Burstable pod may not become Guaranteed, a BestEffort pod may not gain requests); you cannot remove a request once set; memory decrease below current usage is skipped and stays In Progress; non-restartable init containers and ephemeral containers cannot be resized; Windows and static CPU-manager pods are excluded.
- VPA caught up:
updateMode: InPlaceOrRecreateis GA in VPA 1.6 and resizes without eviction where the kubelet allows it, falling back to recreation. That removes the old objection to running VPA in an active mode on latency-sensitive services, though the HPA conflict on the same metric remains.
Pod-level resources
spec.resources on the Pod itself (beta since 1.34, on by default) sets a budget for the whole pod; containers may leave their own requests unset and share it. Pod-level values take precedence for scheduling and QoS. In-place resize of the pod-level budget is beta in 1.36. Quota and LimitRange count it like any other request. Say the version when you mention it; it is new enough to be a trap.
Topology spread, the optional fields
minDomains(only withDoNotSchedule): treat fewer eligible domains than this as skew, forcing the autoscaler to open a new zone rather than piling into two.nodeAffinityPolicyandnodeTaintsPolicy(GA since 1.33):HonororIgnorewhether nodes excluded by affinity or taints count as domains. Default honors affinity and ignores taints, which is why tainted spare nodes can distort skew.matchLabelKeys(typicallypod-template-hash): spread each ReplicaSet revision independently so a rolling update does not fight the constraint.
HPA details that decide behavior
- Tolerance is 10% globally; per-HPA
spec.behavior.scaleUp.tolerance/scaleDown.toleranceis beta in 1.35 and 1.36 (off by default until GA in 1.37), so do not rely on it in a task unless the cluster shows the field. - Not-ready pods and pods with missing metrics are set aside; on scale-down they are assumed to be at 100% of target, on scale-up at 0%, which dampens both directions. CPU metrics from a pod that became ready in the last 30 s (
--horizontal-pod-autoscaler-initial-readiness-delay) are ignored. behavior.scaleDown.stabilizationWindowSecondsdefaults to 300, scale-up to 0;policiescombine byselectPolicy: Max(default),Min, orDisabledto freeze a direction.ContainerResourcemetrics target one container's usage rather than the pod sum, the fix for sidecar-heavy pods.
Before resizing anything, kubectl get pod x -o jsonpath='{.status.qosClass}'. The class cannot change, so a request that would flip it is refused with a validation error rather than deferred, and the message names the field.
Exercises
Create three pods in default: one with requests==limits, one with only requests, one with nothing. Predict each class before checking:
kubectl get pod <name> -o jsonpath='{.status.qosClass}{"\n"}'outputcaptured 2026-08-26
$ kubectl get pod qos-guaranteed -o jsonpath='{.status.qosClass}{"\n"}'
Guaranteed
$ kubectl get pod qos-burstable -o jsonpath='{.status.qosClass}{"\n"}'
Burstable
$ kubectl get pod qos-besteffort -o jsonpath='{.status.qosClass}{"\n"}'
BestEffortMemory first. kubectl run lost its resource flags a while back, so this is also a rep for the --overrides escape hatch:
kubectl run oom --image=polinux/stress --restart=Never \
--overrides='{"spec":{"containers":[{"name":"oom","image":"polinux/stress","command":["stress","--vm","1","--vm-bytes","128M","--vm-hang","0"],"resources":{"requests":{"memory":"64Mi"},"limits":{"memory":"64Mi"}}}]}}'
kubectl get pod oom -w # until STATUS shows OOMKilledoutputcaptured 2026-08-26
$ kubectl run oom --image=polinux/stress --restart=Never \
--overrides='{"spec":{"containers":[{"name":"oom","image":"polinux/stress","command":["stress","--vm","1","--vm-bytes","128M","--vm-hang","0"],"resources":{"requests":{"memory":"64Mi"},"limits":{"memory":"64Mi"}}}]}}'
pod/oom created
$ kubectl get pod oom -w # until STATUS shows OOMKilled
NAME READY STATUS RESTARTS AGE
oom 0/1 Pending 0 1s
oom 0/1 ContainerCreating 0 1s
oom 0/1 ContainerCreating 0 3s
oom 1/1 Running 0 12s
oom 0/1 OOMKilled 0 12s
oom 0/1 OOMKilled 0 14s
^Ckubectl get pod oom -o jsonpath='{.status.containerStatuses[0].state.terminated.reason}' prints OOMKilled. With --restart=Never nothing restarts, so the evidence sits in state, not lastState; in a crash-looping Deployment it is the other way round. CPU, by contrast, would have throttled silently.Deploy 4 replicas of anything with:
topologySpreadConstraints:
- maxSkew: 1
topologyKey: topology.kubernetes.io/zone
whenUnsatisfiable: DoNotSchedule
labelSelector: { matchLabels: { app: spread-demo } }kubectl get pods -l app=spread-demo -o wide shows 2+2 across the two worker zones, and kubectl get nodes -L topology.kubernetes.io/zone confirms which zone each node carries.Deploy examples/demo-app/base into default (kubectl apply -k examples/demo-app/base), then:
kubectl autoscale deploy demo --min=2 --max=6 --cpu-percent=20
kubectl run load --image=busybox:1.37 --restart=Never -- \
sh -c 'while true; do wget -qO- http://demo.default.svc:80 >/dev/null; done'
kubectl get hpa demo -woutputcaptured 2026-08-26
$ kubectl autoscale deploy demo --min=2 --max=6 --cpu-percent=20
Flag --cpu-percent has been deprecated, Use --cpu with percentage or resource quantity format (e.g., '70%' for utilization or '500m' for milliCPU).
horizontalpodautoscaler.autoscaling/demo autoscaled
$ kubectl run load --image=busybox:1.37 --restart=Never -- \
sh -c 'while true; do wget -qO- http://demo.default.svc:80 >/dev/null; done'
pod/load created
$ kubectl get hpa demo -w
NAME REFERENCE TARGETS MINPODS MAXPODS REPLICAS AGE
demo Deployment/demo cpu: <unknown>/20% 2 6 0 0s
demo Deployment/demo cpu: <unknown>/20% 2 6 2 15s
demo Deployment/demo cpu: 114%/20% 2 6 2 30s
demo Deployment/demo cpu: 98%/20% 2 6 4 45s
demo Deployment/demo cpu: 98%/20% 2 6 6 60s
demo Deployment/demo cpu: 34%/20% 2 6 6 75s
demo Deployment/demo cpu: 36%/20% 2 6 6 91s
demo Deployment/demo cpu: 41%/20% 2 6 6 106s
demo Deployment/demo cpu: 30%/20% 2 6 6 2m1s
demo Deployment/demo cpu: 25%/20% 2 6 6 2m16s
demo Deployment/demo cpu: 30%/20% 2 6 6 2m31s
demo Deployment/demo cpu: 39%/20% 2 6 6 2m46s
demo Deployment/demo cpu: 38%/20% 2 6 6 3m1s
demo Deployment/demo cpu: 25%/20% 2 6 6 3m16s
demo Deployment/demo cpu: 26%/20% 2 6 6 3m31s
demo Deployment/demo cpu: 38%/20% 2 6 6 3m46s
demo Deployment/demo cpu: 39%/20% 2 6 6 4m1s
demo Deployment/demo cpu: 30%/20% 2 6 6 4m16s
^Cload and watch it settle back after the stabilization window (about 5 minutes). Note the demo container requests 25m CPU, which is why 20% is reachable at all.With demo still running:
apiVersion: autoscaling.k8s.io/v1
kind: VerticalPodAutoscaler
metadata: { name: demo, namespace: default }
spec:
targetRef: { apiVersion: apps/v1, kind: Deployment, name: demo }
updatePolicy: { updateMode: "Off" }Give it a few minutes, then:
kubectl get vpa demo -o jsonpath='{.status.recommendation.containerRecommendations[0]}' | jqoutputcaptured 2026-08-26
$ kubectl get vpa demo -o jsonpath='{.status.recommendation.containerRecommendations[0]}' | jq
{
"containerName": "web",
"lowerBound": {
"cpu": "15m",
"memory": "100Mi"
},
"target": {
"cpu": "23m",
"memory": "100Mi"
},
"uncappedTarget": {
"cpu": "23m",
"memory": "100Mi"
},
"upperBound": {
"cpu": "12479m",
"memory": "8405796509"
}
}In-place resize changed what "change the resources" means: a CPU request can move without the container restarting. The subresource is resize, not a normal patch, and the evidence lives in status.containerStatuses[].allocatedResources. The second half of this exercise asks for more CPU than the node has, to find out where a resize you cannot have gets refused.
kubectl run burst --image=nginx:1.27-alpine --restart=Never --overrides='{"spec":{"containers":[{"name":"burst","image":"nginx:1.27-alpine","resources":{"requests":{"cpu":"100m","memory":"64Mi"},"limits":{"cpu":"200m","memory":"128Mi"}}}]}}'
kubectl wait --for=condition=Ready pod/burst --timeout=90s
kubectl get pod burst -o jsonpath='{.status.qosClass}{"\n"}'
kubectl patch pod burst --subresource resize --patch '{"spec":{"containers":[{"name":"burst","resources":{"requests":{"cpu":"200m","memory":"64Mi"},"limits":{"cpu":"400m","memory":"128Mi"}}}]}}'
kubectl get pod burst -o jsonpath='{.status.containerStatuses[0].allocatedResources}{"\n"}{.status.containerStatuses[0].restartCount}{"\n"}'
kubectl get pod burst -o jsonpath='{.status.conditions}' | jq '.[] | {type, status, reason}'
kubectl patch pod burst --subresource resize --patch '{"spec":{"containers":[{"name":"burst","resources":{"requests":{"cpu":"64","memory":"64Mi"},"limits":{"cpu":"80","memory":"128Mi"}}}]}}'
kubectl get pod burst -o jsonpath='{.status.conditions}' | jq '.[] | {type, status, reason, message}'
kubectl delete pod burstoutputcaptured 2026-09-12
$ kubectl run burst --image=nginx:1.27-alpine --restart=Never --overrides='{"spec":{"containers":[{"name":"burst","image":"nginx:1.27-alpine","resources":{"requests":{"cpu":"100m","memory":"64Mi"},"limits":{"cpu":"200m","memory":"128Mi"}}}]}}'
pod/burst created
$ kubectl wait --for=condition=Ready pod/burst --timeout=90s
pod/burst condition met
$ kubectl get pod burst -o jsonpath='{.status.qosClass}{"\n"}'
Burstable
$ kubectl patch pod burst --subresource resize --patch '{"spec":{"containers":[{"name":"burst","resources":{"requests":{"cpu":"200m","memory":"64Mi"},"limits":{"cpu":"400m","memory":"128Mi"}}}]}}'
pod/burst patched
$ kubectl get pod burst -o jsonpath='{.status.containerStatuses[0].allocatedResources}{"\n"}{.status.containerStatuses[0].restartCount}{"\n"}'
{"cpu":"200m","memory":"64Mi"}
0
$ kubectl get pod burst -o jsonpath='{.status.conditions}' | jq '.[] | {type, status, reason}'
{
"type": "PodReadyToStartContainers",
"status": "True",
"reason": null
}
{
"type": "Initialized",
"status": "True",
"reason": null
}
{
"type": "Ready",
"status": "True",
"reason": null
}
{
"type": "ContainersReady",
"status": "True",
"reason": null
}
{
"type": "PodScheduled",
"status": "True",
"reason": null
}
$ kubectl patch pod burst --subresource resize --patch '{"spec":{"containers":[{"name":"burst","resources":{"requests":{"cpu":"64","memory":"64Mi"},"limits":{"cpu":"80","memory":"128Mi"}}}]}}'
Error from server (Forbidden): pods "burst" is forbidden: node didn't have enough allocatable resources: cpu, requested: 64000, allocatable: 16000
$ kubectl get pod burst -o jsonpath='{.status.conditions}' | jq '.[] | {type, status, reason, message}'
{
"type": "PodReadyToStartContainers",
"status": "True",
"reason": null,
"message": null
}
{
"type": "Initialized",
"status": "True",
"reason": null,
"message": null
}
{
"type": "Ready",
"status": "True",
"reason": null,
"message": null
}
{
"type": "ContainersReady",
"status": "True",
"reason": null,
"message": null
}
{
"type": "PodScheduled",
"status": "True",
"reason": null,
"message": null
}
$ kubectl delete pod burst
pod "burst" deleted from default namespaceallocatedResources at {"cpu":"200m","memory":"64Mi"} and restartCount still 0. The impossible one does not park anywhere: the API refuses the patch outright with Error from server (Forbidden): pods "burst" is forbidden: node didn't have enough allocatable resources: cpu, requested: 64000, allocatable: 16000, and the pod's conditions are unchanged, with no PodResizePending among them.QoS class is computed from requests and limits and it is immutable for the life of the pod. A resize that would move a Guaranteed pod into Burstable is therefore not a scheduling problem, it is a validation error, and the message says so.
kubectl run guar --image=nginx:1.27-alpine --restart=Never --overrides='{"spec":{"containers":[{"name":"guar","image":"nginx:1.27-alpine","resources":{"requests":{"cpu":"100m","memory":"64Mi"},"limits":{"cpu":"100m","memory":"64Mi"}}}]}}'
kubectl wait --for=condition=Ready pod/guar --timeout=90s
kubectl get pod guar -o jsonpath='{.status.qosClass}{"\n"}'
kubectl patch pod guar --subresource resize --patch '{"spec":{"containers":[{"name":"guar","resources":{"requests":{"cpu":"50m","memory":"64Mi"},"limits":{"cpu":"100m","memory":"64Mi"}}}]}}'
kubectl delete pod guaroutputcaptured 2026-09-12
$ kubectl run guar --image=nginx:1.27-alpine --restart=Never --overrides='{"spec":{"containers":[{"name":"guar","image":"nginx:1.27-alpine","resources":{"requests":{"cpu":"100m","memory":"64Mi"},"limits":{"cpu":"100m","memory":"64Mi"}}}]}}'
pod/guar created
$ kubectl wait --for=condition=Ready pod/guar --timeout=90s
pod/guar condition met
$ kubectl get pod guar -o jsonpath='{.status.qosClass}{"\n"}'
Guaranteed
$ kubectl patch pod guar --subresource resize --patch '{"spec":{"containers":[{"name":"guar","resources":{"requests":{"cpu":"50m","memory":"64Mi"},"limits":{"cpu":"100m","memory":"64Mi"}}}]}}'
The Pod "guar" is invalid: spec: Invalid value: "Guaranteed": Pod QOS Class may not change as a result of resizing
$ kubectl delete pod guar
pod "guar" deleted from default namespaceA PodDisruptionBudget that can never be satisfied turns a routine node drain into a stuck terminal. The interesting part is that unhealthy pods count against the budget unless you tell the API otherwise.
kubectl get pdb -A -o custom-columns=NS:.metadata.namespace,NAME:.metadata.name,ALLOWED:.status.disruptionsAllowed
kubectl apply -f - <<'EOF'
apiVersion: apps/v1
kind: Deployment
metadata: { name: crasher, namespace: default }
spec:
replicas: 2
selector: { matchLabels: { app: crasher } }
template:
metadata: { labels: { app: crasher } }
spec:
nodeSelector: { kubernetes.io/hostname: cnpe-worker2 }
containers:
- name: c
image: busybox:1.36
command: ["sh", "-c", "exit 1"]
---
apiVersion: policy/v1
kind: PodDisruptionBudget
metadata: { name: crasher, namespace: default }
spec:
minAvailable: 100%
selector: { matchLabels: { app: crasher } }
EOF
sleep 30
# --pod-selector keeps the drain to this exercise's pods; without it you evict the lab off this node
kubectl get pdb crasher -o jsonpath='{.status.disruptionsAllowed}{"\n"}'
kubectl drain cnpe-worker2 --pod-selector app=crasher --ignore-daemonsets --delete-emptydir-data --timeout=60s
kubectl patch pdb crasher --type merge -p '{"spec":{"unhealthyPodEvictionPolicy":"AlwaysAllow"}}'
kubectl drain cnpe-worker2 --pod-selector app=crasher --ignore-daemonsets --delete-emptydir-data --timeout=60s
kubectl uncordon cnpe-worker2
kubectl delete deployment crasher
kubectl delete pdb crasheroutputcaptured 2026-09-13
$ kubectl get pdb -A -o custom-columns=NS:.metadata.namespace,NAME:.metadata.name,ALLOWED:.status.disruptionsAllowed
NS NAME ALLOWED
gatekeeper-system gatekeeper-controller-manager 0
team-a pg-primary 0
tracing otel-collector 1
$ kubectl apply -f - <<'EOF'
apiVersion: apps/v1
kind: Deployment
metadata: { name: crasher, namespace: default }
spec:
replicas: 2
selector: { matchLabels: { app: crasher } }
template:
metadata: { labels: { app: crasher } }
spec:
nodeSelector: { kubernetes.io/hostname: cnpe-worker2 }
containers:
- name: c
image: busybox:1.36
command: ["sh", "-c", "exit 1"]
---
apiVersion: policy/v1
kind: PodDisruptionBudget
metadata: { name: crasher, namespace: default }
spec:
minAvailable: 100%
selector: { matchLabels: { app: crasher } }
EOF
deployment.apps/crasher created
poddisruptionbudget.policy/crasher created
$ sleep 30
$ # --pod-selector keeps the drain to this exercise's pods; without it you evict the lab off this node
$ kubectl get pdb crasher -o jsonpath='{.status.disruptionsAllowed}{"\n"}'
0
$ kubectl drain cnpe-worker2 --pod-selector app=crasher --ignore-daemonsets --delete-emptydir-data --timeout=60s
node/cnpe-worker2 cordoned
evicting pod default/crasher-7857b795db-hbpzj
evicting pod default/crasher-7857b795db-854th
error when evicting pods/"crasher-7857b795db-hbpzj" -n "default" (will retry after 5s): Cannot evict pod as it would violate the pod's disruption budget.
error when evicting pods/"crasher-7857b795db-854th" -n "default" (will retry after 5s): Cannot evict pod as it would violate the pod's disruption budget.
evicting pod default/crasher-7857b795db-hbpzj
error when evicting pods/"crasher-7857b795db-hbpzj" -n "default" (will retry after 5s): Cannot evict pod as it would violate the pod's disruption budget.
evicting pod default/crasher-7857b795db-854th
error when evicting pods/"crasher-7857b795db-854th" -n "default" (will retry after 5s): Cannot evict pod as it would violate the pod's disruption budget.
evicting pod default/crasher-7857b795db-hbpzj
error when evicting pods/"crasher-7857b795db-hbpzj" -n "default" (will retry after 5s): Cannot evict pod as it would violate the pod's disruption budget.
evicting pod default/crasher-7857b795db-854th
error when evicting pods/"crasher-7857b795db-854th" -n "default" (will retry after 5s): Cannot evict pod as it would violate the pod's disruption budget.
evicting pod default/crasher-7857b795db-hbpzj
error when evicting pods/"crasher-7857b795db-hbpzj" -n "default" (will retry after 5s): Cannot evict pod as it would violate the pod's disruption budget.
evicting pod default/crasher-7857b795db-854th
error when evicting pods/"crasher-7857b795db-854th" -n "default" (will retry after 5s): Cannot evict pod as it would violate the pod's disruption budget.
evicting pod default/crasher-7857b795db-hbpzj
evicting pod default/crasher-7857b795db-854th
error when evicting pods/"crasher-7857b795db-hbpzj" -n "default" (will retry after 5s): Cannot evict pod as it would violate the pod's disruption budget.
error when evicting pods/"crasher-7857b795db-854th" -n "default" (will retry after 5s): Cannot evict pod as it would violate the pod's disruption budget.
evicting pod default/crasher-7857b795db-hbpzj
evicting pod default/crasher-7857b795db-854th
error when evicting pods/"crasher-7857b795db-hbpzj" -n "default" (will retry after 5s): Cannot evict pod as it would violate the pod's disruption budget.
error when evicting pods/"crasher-7857b795db-854th" -n "default" (will retry after 5s): Cannot evict pod as it would violate the pod's disruption budget.
evicting pod default/crasher-7857b795db-hbpzj
evicting pod default/crasher-7857b795db-854th
error when evicting pods/"crasher-7857b795db-854th" -n "default" (will retry after 5s): Cannot evict pod as it would violate the pod's disruption budget.
error when evicting pods/"crasher-7857b795db-hbpzj" -n "default" (will retry after 5s): Cannot evict pod as it would violate the pod's disruption budget.
evicting pod default/crasher-7857b795db-854th
evicting pod default/crasher-7857b795db-hbpzj
error when evicting pods/"crasher-7857b795db-854th" -n "default" (will retry after 5s): Cannot evict pod as it would violate the pod's disruption budget.
error when evicting pods/"crasher-7857b795db-hbpzj" -n "default" (will retry after 5s): Cannot evict pod as it would violate the pod's disruption budget.
evicting pod default/crasher-7857b795db-854th
evicting pod default/crasher-7857b795db-hbpzj
error when evicting pods/"crasher-7857b795db-854th" -n "default" (will retry after 5s): Cannot evict pod as it would violate the pod's disruption budget.
error when evicting pods/"crasher-7857b795db-hbpzj" -n "default" (will retry after 5s): Cannot evict pod as it would violate the pod's disruption budget.
evicting pod default/crasher-7857b795db-854th
evicting pod default/crasher-7857b795db-hbpzj
error when evicting pods/"crasher-7857b795db-hbpzj" -n "default" (will retry after 5s): Cannot evict pod as it would violate the pod's disruption budget.
... 19 more lines
$ kubectl patch pdb crasher --type merge -p '{"spec":{"unhealthyPodEvictionPolicy":"AlwaysAllow"}}'
poddisruptionbudget.policy/crasher patched
$ kubectl drain cnpe-worker2 --pod-selector app=crasher --ignore-daemonsets --delete-emptydir-data --timeout=60s
node/cnpe-worker2 already cordoned
evicting pod default/crasher-7857b795db-hbpzj
evicting pod default/crasher-7857b795db-854th
pod/crasher-7857b795db-854th evicted
pod/crasher-7857b795db-hbpzj evicted
node/cnpe-worker2 drained
$ kubectl uncordon cnpe-worker2
node/cnpe-worker2 uncordoned
$ kubectl delete deployment crasher
deployment.apps "crasher" deleted from default namespace
$ kubectl delete pdb crasher
poddisruptionbudget.policy "crasher" deleted from default namespacedisruptionsAllowed is 0, and the first drain refuses both pods with Cannot evict pod as it would violate the pod's disruption budget until it times out. After unhealthyPodEvictionPolicy: AlwaysAllow the same command evicts both and prints node/cnpe-worker2 drained. Note the --pod-selector: without it this drain empties the node, and half the lab is on it.Allocatable is capacity minus what the kubelet reserved, and the reservation is a kubelet flag rather than anything in the API. Read both halves on the same node so the subtraction is yours, not a slide's.
kubectl get node cnpe-worker -o jsonpath='{"capacity: cpu "}{.status.capacity.cpu}{", memory "}{.status.capacity.memory}{"\n"}{"allocatable: cpu "}{.status.allocatable.cpu}{", memory "}{.status.allocatable.memory}{"\n"}'
kubectl get --raw /api/v1/nodes/cnpe-worker/proxy/configz | jq '.kubeletconfig | {kubeReserved, systemReserved, evictionHard}'
kubectl describe node cnpe-worker | sed -n '/Allocated resources/,/Events/p'outputcaptured 2026-09-13
$ kubectl get node cnpe-worker -o jsonpath='{"capacity: cpu "}{.status.capacity.cpu}{", memory "}{.status.capacity.memory}{"\n"}{"allocatable: cpu "}{.status.allocatable.cpu}{", memory "}{.status.allocatable.memory}{"\n"}'
capacity: cpu 16, memory 32747836Ki
allocatable: cpu 16, memory 32747836Ki
$ kubectl get --raw /api/v1/nodes/cnpe-worker/proxy/configz | jq '.kubeletconfig | {kubeReserved, systemReserved, evictionHard}'
{
"kubeReserved": null,
"systemReserved": null,
"evictionHard": {
"imagefs.available": "0%",
"nodefs.available": "0%",
"nodefs.inodesFree": "0%"
}
}
$ kubectl describe node cnpe-worker | sed -n '/Allocated resources/,/Events/p'
Allocated resources:
(Total limits may be over 100 percent, i.e., overcommitted.)
Resource Requests Limits
-------- -------- ------
cpu 1811m (11%) 9450m (59%)
memory 3714Mi (11%) 12800Mi (40%)
ephemeral-storage 0 (0%) 0 (0%)
hugepages-1Gi 0 (0%) 0 (0%)
hugepages-2Mi 0 (0%) 0 (0%)
Events:kubeReserved and systemReserved both read null and evictionHard is 0% on every signal, so there is no gap to subtract. The scheduler adds pod requests up against Allocatable, which is the number the Allocated resources block takes its percentages from.Two priority classes at the same value behave completely differently if one sets preemptionPolicy: Never. Batch work usually wants that: high enough to jump the scheduling queue, never rude enough to evict a neighbor.
kubectl create priorityclass batch --value=1000 --preemption-policy=Never
kubectl create priorityclass urgent --value=1000
kubectl run filler --image=nginx:1.27-alpine --restart=Never --overrides='{"spec":{"nodeSelector":{"kubernetes.io/hostname":"cnpe-worker2"},"containers":[{"name":"filler","image":"nginx:1.27-alpine","resources":{"requests":{"cpu":"12"}}}]}}'
kubectl wait --for=condition=Ready pod/filler --timeout=120s
kubectl run batchpod --image=nginx:1.27-alpine --restart=Never --overrides='{"spec":{"priorityClassName":"batch","nodeSelector":{"kubernetes.io/hostname":"cnpe-worker2"},"containers":[{"name":"batchpod","image":"nginx:1.27-alpine","resources":{"requests":{"cpu":"8"}}}]}}'
sleep 20
kubectl get pod batchpod -o wide
kubectl describe pod batchpod | tail -4
kubectl run urgentpod --image=nginx:1.27-alpine --restart=Never --overrides='{"spec":{"priorityClassName":"urgent","nodeSelector":{"kubernetes.io/hostname":"cnpe-worker2"},"containers":[{"name":"urgentpod","image":"nginx:1.27-alpine","resources":{"requests":{"cpu":"8"}}}]}}'
sleep 30
kubectl get pods filler batchpod urgentpod -o wide
kubectl delete pod filler batchpod urgentpod --ignore-not-found
kubectl delete priorityclass batch urgentoutputcaptured 2026-09-12
$ kubectl create priorityclass batch --value=1000 --preemption-policy=Never
priorityclass.scheduling.k8s.io/batch created
$ kubectl create priorityclass urgent --value=1000
priorityclass.scheduling.k8s.io/urgent created
$ kubectl run filler --image=nginx:1.27-alpine --restart=Never --overrides='{"spec":{"nodeSelector":{"kubernetes.io/hostname":"cnpe-worker2"},"containers":[{"name":"filler","image":"nginx:1.27-alpine","resources":{"requests":{"cpu":"12"}}}]}}'
pod/filler created
$ kubectl wait --for=condition=Ready pod/filler --timeout=120s
pod/filler condition met
$ kubectl run batchpod --image=nginx:1.27-alpine --restart=Never --overrides='{"spec":{"priorityClassName":"batch","nodeSelector":{"kubernetes.io/hostname":"cnpe-worker2"},"containers":[{"name":"batchpod","image":"nginx:1.27-alpine","resources":{"requests":{"cpu":"8"}}}]}}'
pod/batchpod created
$ sleep 20
$ kubectl get pod batchpod -o wide
NAME READY STATUS RESTARTS AGE IP NODE NOMINATED NODE READINESS GATES
batchpod 0/1 Pending 0 20s <none> <none> <none> <none>
$ kubectl describe pod batchpod | tail -4
Type Reason Age From Message
---- ------ ---- ---- -------
Warning PolicyViolation 20s kyverno-admission policy require-resource-requests/ fail: every container must set cpu and memory requests
Warning FailedScheduling 20s default-scheduler 0/3 nodes are available: 1 Insufficient cpu, 1 node(s) didn't match Pod's node affinity/selector, 1 node(s) had untolerated taint(s). no new claims to deallocate, preemption: not eligible due to preemptionPolicy=Never.
$ kubectl run urgentpod --image=nginx:1.27-alpine --restart=Never --overrides='{"spec":{"priorityClassName":"urgent","nodeSelector":{"kubernetes.io/hostname":"cnpe-worker2"},"containers":[{"name":"urgentpod","image":"nginx:1.27-alpine","resources":{"requests":{"cpu":"8"}}}]}}'
pod/urgentpod created
$ sleep 30
$ kubectl get pods filler batchpod urgentpod -o wide
NAME READY STATUS RESTARTS AGE IP NODE NOMINATED NODE READINESS GATES
batchpod 0/1 Pending 0 51s <none> <none> <none> <none>
urgentpod 1/1 Running 0 31s 10.244.2.42 cnpe-worker2 <none> <none>
Error from server (NotFound): pods "filler" not found
$ kubectl delete pod filler batchpod urgentpod --ignore-not-found
pod "batchpod" deleted from default namespace
pod "urgentpod" deleted from default namespace
$ kubectl delete priorityclass batch urgent
priorityclass.scheduling.k8s.io "batch" deleted
priorityclass.scheduling.k8s.io "urgent" deletedbatchpod stays Pending with an Insufficient cpu event and nothing is evicted, while urgentpod at the same value takes the room by pushing filler out. The two classes carry the same value; only preemptionPolicy differs.An HPA created with one command still has a full scale-up and scale-down policy; you just did not type it. Know where the defaults live before a task asks you to slow a flapping scale-down.
kubectl create deployment hpademo --image=nginx:1.27-alpine
kubectl set resources deployment hpademo --requests=cpu=50m
kubectl autoscale deployment hpademo --cpu-percent=80 --min=1 --max=5
kubectl get hpa hpademo -o jsonpath='{.spec.behavior}{"\n"}'
kubectl get hpa hpademo -o yaml | sed -n '/^spec:/,/^status:/p'
kubectl explain hpa.spec.behavior.scaleDown
kubectl delete hpa hpademo
kubectl delete deployment hpademooutputcaptured 2026-09-12
$ kubectl create deployment hpademo --image=nginx:1.27-alpine
deployment.apps/hpademo created
$ kubectl set resources deployment hpademo --requests=cpu=50m
deployment.apps/hpademo resource requirements updated
$ kubectl autoscale deployment hpademo --cpu-percent=80 --min=1 --max=5
Flag --cpu-percent has been deprecated, Use --cpu with percentage or resource quantity format (e.g., '70%' for utilization or '500m' for milliCPU).
horizontalpodautoscaler.autoscaling/hpademo autoscaled
$ kubectl get hpa hpademo -o jsonpath='{.spec.behavior}{"\n"}'
$ kubectl get hpa hpademo -o yaml | sed -n '/^spec:/,/^status:/p'
spec:
maxReplicas: 5
metrics:
- resource:
name: cpu
target:
averageUtilization: 80
type: Utilization
type: Resource
minReplicas: 1
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: hpademo
status:
$ kubectl explain hpa.spec.behavior.scaleDown
GROUP: autoscaling
KIND: HorizontalPodAutoscaler
VERSION: v2
FIELD: scaleDown <HPAScalingRules>
DESCRIPTION:
scaleDown is scaling policy for scaling Down. If not set, the default value
is to allow to scale down to minReplicas pods, with a 300 second
stabilization window (i.e., the highest recommendation for the last 300sec
is used).
HPAScalingRules configures the scaling behavior for one direction via
scaling Policy Rules and a configurable metric tolerance.
Scaling Policy Rules are applied after calculating DesiredReplicas from
metrics for the HPA. They can limit the scaling velocity by specifying
scaling policies. They can prevent flapping by specifying the stabilization
window, so that the number of replicas is not set instantly, instead, the
safest value from the stabilization window is chosen.
The tolerance is applied to the metric values and prevents scaling too
eagerly for small metric variations. (Note that setting a tolerance requires
the beta HPAConfigurableTolerance feature gate to be enabled.)
FIELDS:
policies <[]HPAScalingPolicy>
policies is a list of potential scaling polices which can be used during
scaling. If not set, use the default values: - For scale up: allow doubling
the number of pods, or an absolute change of 4 pods in a 15s window. - For
scale down: allow all pods to be removed in a 15s window.
selectPolicy <string>
selectPolicy is used to specify which policy should be used. If not set, the
default value Max is used.
stabilizationWindowSeconds <integer>
stabilizationWindowSeconds is the number of seconds for which past
recommendations should be considered while scaling up or scaling down.
StabilizationWindowSeconds must be greater than or equal to zero and less
... 17 more lines
$ kubectl delete hpa hpademo
horizontalpodautoscaler.autoscaling "hpademo" deleted from default namespace
$ kubectl delete deployment hpademo
deployment.apps "hpademo" deleted from default namespacespec.behavior reads back empty, the defaults are the controller's rather than the object's, and that is the answer to give.Self-check
A node shows 25% CPU usage and refuses to schedule a pod requesting 200m. Explain, in one sentence.
The scheduler counts requests against allocatable, not usage: the node's requests are already at allocatable even though the processes are idle. Fix by right-sizing the existing requests (or adding capacity), not by adding CPU headroom that already exists.
Which pod does the kubelet evict first under memory pressure, and why?
BestEffort, because it declared no memory request at all and therefore ranks worst. Then Burstable pods exceeding their requests, ordered by how far over they are. Guaranteed last. The kernel's oom_score_adj mirrors that ranking so a system-level OOM picks the same victim.
Your service has p99 latency spikes every few seconds; CPU usage sits at 60% of the limit. First metric you check?
rate(container_cpu_cfs_throttled_periods_total[5m]). Average usage below the limit hides per-period throttling: the container burns its 100ms quota early and stalls until the next period. Average utilization is the wrong lens for CPU limits.
Why can a VPA in Auto mode and an HPA on CPU not coexist on the same workload?
They form a loop: the HPA scales replicas based on usage-against-requests while the VPA rewrites those same requests, so each one keeps invalidating the other's denominator. Standard resolution: HPA on CPU with VPA in Off (advisory), or HPA on a custom/external metric while VPA owns the resources.
A drain hangs forever on one node. What are the two usual causes?
A PodDisruptionBudget that cannot be satisfied (often minAvailable equal to the replica count), or unmanaged pods: bare pods with no controller, which kubectl drain refuses to evict without --force. Both are voluntary-disruption mechanics; neither has anything to do with node health.
Is in-place pod resize usable on a Kubernetes 1.36 cluster, and what does the request look like?
Yes: InPlacePodVerticalScaling has been GA since 1.35 (beta and default-on since 1.33). You patch the resize subresource (kubectl patch pod x --subresource resize ...); a plain spec patch is rejected as immutable. Watch PodResizePending (Infeasible or Deferred) and PodResizeInProgress conditions, and remember the QoS class cannot change.
Karpenter refuses to consolidate a node you consider empty enough. Where is the reason, and name two things it could be.
Events on the node with reason Unconsolidatable. Typical messages: a PDB prevents pod evictions, a pod carries karpenter.sh/do-not-disrupt, no cheaper replacement exists, or the NodePool's disruption budget is exhausted. Cluster Autoscaler's equivalents are the safe-to-evict annotation, uncontrolled pods and local storage.
A PDB has minAvailable 2 for a 2-replica Deployment whose pods are CrashLoopBackOff. Drain hangs. Fix without deleting the PDB?
Set unhealthyPodEvictionPolicy: AlwaysAllow on the PDB: the default IfHealthyBudget refuses to evict unhealthy pods until the budget is met, which can never happen here. Alternatively scale the Deployment so the budget is satisfiable. Deleting the PDB removes the protection permanently rather than for the drain.
What is allocatable, and which four kubelet settings subtract from capacity?
Allocatable is what the scheduler may hand to pods: capacity minus kubeReserved, systemReserved and the evictionHard thresholds (defaults memory.available 100Mi, nodefs.available 10%, nodefs.inodesFree 5%, imagefs.available 15%). enforceNodeAllocatable decides which of those cgroups are actually enforced; pods is the default.
Docs to know your way around
- kubernetes.io: Resource Management for Pods and Containers; Pod Quality of Service Classes; Node-pressure Eviction; Assigning Pods to Nodes; Pod Topology Spread Constraints; the HorizontalPodAutoscaler walkthrough (the algorithm section especially).
- github.com/kubernetes/autoscaler: the VPA README, in particular the
updateModetable. - Offline:
kubectl explain pod.spec.containers.resources,kubectl explain hpa.spec.behavior --recursive, and the requests table at the bottom ofkubectl describe node. - kubernetes.io/docs/tasks/configure-pod-container/resize-container-resources: the resize subresource, resizePolicy, and the PodResizePending / PodResizeInProgress conditions with their reasons.
- kubernetes.io/docs/concepts/cluster-administration/node-autoscaling: the neutral Cluster Autoscaler vs Karpenter comparison; karpenter.sh/docs/concepts/disruption for consolidationPolicy, budgets and do-not-disrupt.
- kubernetes.io/docs/tasks/administer-cluster/reserve-compute-resources: kubeReserved, systemReserved, evictionHard and enforceNodeAllocatable with the worked allocatable example.