Requests and limits look like beginner material and then show up everywhere: quotas count requests, OpenCost bills requests, LimitRanges inject them, Kyverno policies demand them, HPAs divide by them, and the resources break drill corrupts them. Get the mechanics exact and four other sections get easier.

needsmake up

Orientation

competency 1.1 + 1.2 · compute and scaling

Two numbers per container decide scheduling, cost, eviction order, throttling, and whether your workload survives a noisy neighbor. This section states precisely what each number does at each layer.

One sentence to hold everything together

A request is a claim against the scheduler's arithmetic; a limit is a ceiling enforced by the kernel at runtime. Nothing reconciles the two after scheduling, which is how a node can be 30% used and 100% requested at the same time. Section 1.5 turns that gap into money.

How requests and limits are enforced

scheduler arithmetic · cgroup enforcement

What the scheduler does with requests

The scheduler sums the requests of all non-terminated pods on a node and compares against the node's allocatable, not its capacity. Allocatable is capacity minus kube-reserved, system-reserved, and the eviction threshold. Live usage is never consulted. A node running at 5% CPU with every millicore requested is, to the scheduler, completely full.

node capacity          ████████████████████████████████████████  4000m
allocatable            ██████████████████████████████████░░░░░░  3800m  (minus kubelet/system/eviction)
sum of requests        ████████████████████████████████░░░░░░░░  3600m  ← what the scheduler sees
actual usage           ██████████░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░  1100m  ← what you actually use; the rest is paid for anyway
                                 └────────────────────┘
                                  the right-sizing gap: 2500m reserved and never used

The four numbers, live: drag the requests up until the node refuses another pod, then drag usage down and watch what you are paying for:

What the kernel does with limits

Resourcerequest becomeslimit becomesover the limit
cpucpu.weight (relative share)cpu.max (quota per 100ms period)throttled: silently, no event, no restart
memoryeviction ranking onlymemory.maxOOMKilled: instant, by the kernel
ephemeral-storagescheduling claimkubelet-enforcedpod evicted
hugepages / devicesrequests must equal limits; no overcommit possible

CPU is compressible: exceed the limit and the CFS scheduler simply stops giving you cycles until the next 100ms period. Latency-sensitive services with a tight CPU limit get periodic stalls that look like network problems; container_cpu_cfs_throttled_periods_total is where the truth lives, and a nonzero rate on a healthy-looking service is a real finding. Memory is incompressible: there is no "slow down", only the OOM killer.

Know which field holds the evidence

For a pod with --restart=Never, an OOM kill lands in .status.containerStatuses[0].state.terminated.reason. In a crash-looping Deployment, the container has already restarted, so the same evidence is in lastState.terminated.reason and state shows waiting: CrashLoopBackOff. Reading the wrong field and concluding "no OOM here" is a classic self-inflicted wound.

QoS, derived not memorized

ClassConditionEviction orderoom_score_adj
Guaranteedevery container: requests == limits for both cpu and memorylast-997
Burstableat least one request or limit set, but not Guaranteedmiddle: those exceeding requests first2–999, scaled
BestEffortnothing set at allfirst1000

Derive it rather than recall it. Set the four numbers and read the class, the eviction rank and the two runtime behaviors:

Under node memory pressure the kubelet ranks victims in a specific order: pods whose usage exceeds their requests first, then by Pod Priority (lowest goes first), then by how far over the request they are. QoS class is a consequence of that first test rather than an input: a BestEffort pod requested nothing, so it always exceeds. The practical corollary: a high-priority Burstable pod outlives a low-priority BestEffort one, which is why PriorityClass protects platform components from eviction, not just from preemption. The same ranking is mirrored to the kernel via oom_score_adj so a system-level OOM picks the same victim. CPU pressure never evicts anything; it only throttles. If someone tells you a pod was "evicted for CPU", they are describing something else.

Placement: the levers, in the order you reach for them

and what each one costs
LeverDirectionHard or softReach for it when
nodeSelectorpod → node labelshard onlytrivially simple pinning; blunt
nodeAffinitypod → node labelsrequired… / preferred… with weightsexpressive pinning, operators (In, NotIn, Exists, Gt, Lt)
podAffinity / podAntiAffinitypod → other pods, within a topologyKeybothco-locate with a cache; spread replicas across nodes
taints + tolerationsnode repels podsNoSchedule, PreferNoSchedule, NoExecutededicated pools, control planes, GPU nodes
topologySpreadConstraintseven distribution over a label domainDoNotSchedule / ScheduleAnywayzone/node spread with a bounded skew, the modern default
priorityClasswho wins when the node is fullpreemptionplatform components must outrank tenant workloads

Three details that decide tasks. Taints repel pods, tolerations do not attract: a toleration only says "I can live here", never "put me here". Pairing a taint with a matching nodeAffinity is how you actually dedicate a pool. NoExecute evicts already-running pods that lack the toleration, and tolerationSeconds is what makes the node-not-ready eviction delay configurable. And anti-affinity is expensive to evaluate at scale, which is exactly why topologySpreadConstraints exists: same intent, bounded cost, plus maxSkew to say how even is even enough.

This cluster is built for it

The kind config labels its workers into two zones (topology.kubernetes.io/zone) precisely so spread constraints do something observable. kubectl get nodes -L topology.kubernetes.io/zone shows you the domains before you write the constraint. Read the labels first, or you will write a constraint against a key no node carries, which fails open under ScheduleAnyway and closed under DoNotSchedule.

One neighbor worth naming because availability tasks touch it: a PodDisruptionBudget constrains voluntary disruptions (drains, node upgrades) with minAvailable or maxUnavailable. It does nothing about crashes or node failures. A PDB of minAvailable: 100% is the classic way to make a cluster upgrade hang forever, and recognizing that stalled-drain symptom is worth a mark.

Autoscaling: three different scalers, three different jobs

HPA · VPA · node scaling
ScalerChangesInputGotcha
HPAreplica countlive metrics (Resource, Pods, Object, External)needs metrics-server; percentage targets are relative to requests
VPAthe requests themselveshistorical usagefights an HPA on the same resource; updateMode: Off makes it advisory
Cluster Autoscaler / Karpenternode count / node sizeunschedulable podsreacts to Pending pods, so it is downstream of requests too
KEDAreplicas, including 0queue depth, topic lag, cron, any external sourcea ScaledObject generates an HPA underneath; it is the thing that produces those "External" metrics, and the only way to scale to zero

The HPA formula, which explains every surprise

desiredReplicas = ceil( currentReplicas × ( currentMetricValue / desiredMetricValue ) )

With one caveat that explains most "why didn't it scale" moments: the HPA does nothing while the ratio sits inside its tolerance (10% by default), and pods that are unready or missing metrics are left out of the average entirely. With --cpu-percent=20 and containers requesting 25m, the target is 5m of actual usage per pod. That is why a demo app with a tiny request scales up under a trivial load: the percentage is a fraction of the request, not of the node. Change the request and you change the autoscaler's behavior without touching the HPA.

Scaling up is fast; scaling down waits out a stabilization window (300s by default) so a brief dip cannot flap your fleet. Both directions are configurable per-HPA under spec.behavior with policies (pods or percent, per period) and selectPolicy. Knowing that downscale lag exists, and that it is deliberate, is worth a mark on its own.

VPA modes

  • Off: compute recommendations only. The exam-relevant one: a pure right-sizing oracle you can read without letting anything evict.
  • Initial: apply recommendations at pod creation only.
  • Auto / Recreate: evict and recreate pods to resize them. Disruptive by design, because changing a running pod's requests historically required a new pod.

Newer clusters can resize CPU and memory in place, which softens that trade-off. Do not test for it with kubectl explain pod.spec.containers.resizePolicy; that field has been in the schema since 1.27 whether or not the feature is on. The honest check is whether the subresource exists: kubectl get --raw /api/v1 | grep -o 'pods/resize', or simply try the patch and read the error.

Command reflex

kubectl top pod --containers for the instantaneous truth, kubectl describe node for the requests-vs-allocatable table at the bottom (the single most useful capacity view kubectl gives you), and kubectl get vpa -o jsonpath='{.items[*].status.recommendation…}' for what the numbers ought to be.

Node scaling and the disruption vocabulary

Cluster Autoscaler vs Karpenter · PDB semantics · PriorityClass · kubelet reservations

The autoscaling table above names the node scalers; a scenario question names their fields. Both react to the same input (Pending pods whose requests fit nowhere) and both consolidate on requests, never on live usage, which is why right-sizing requests (1.5) is the precondition for any node-level saving.

Cluster Autoscaler vs Karpenter

AspectCluster AutoscalerKarpenter
Unit it managespre-configured node groups (cloud VM groups); picks the group via an expander (least-waste default, priority, most-pods, random)NodePool constraints (instance families, zones, arch, karpenter.sh/capacity-type spot/on-demand) and per-node NodeClaims; auto-provisions the cheapest instance that fits
Scale-downnode unneeded for --scale-down-unneeded-time (10m default) below --scale-down-utilization-threshold (0.5)spec.disruption.consolidationPolicy: WhenEmpty, WhenEmptyOrUnderutilized; consolidateAfter (or Never); can replace a node with a cheaper one
Lifecycleautoscaling onlyalso drift (re-create nodes when NodePool or AMI changes) and expiry (expireAfter, default 720h)
Rate limitingflagsspec.disruption.budgets (default nodes: 10%, optional cron schedule/duration and reasons); spec.limits caps total CPU/memory a pool may own
Pod opt-outcluster-autoscaler.kubernetes.io/safe-to-evict: "false"karpenter.sh/do-not-disrupt: "true" (or a duration)
What blocks removala restrictive PDB, pods with no controller, pods with local storage (CA), kube-system pods without a PDB (CA); both engines respect PDBs and both give up on a node whose eviction is refused

Karpenter posts Unconsolidatable events on the node that name the blocker ("pdb default/x prevents pod evictions", "can't replace with a lower-priced node"), which is where to look when a node you expected to disappear is still there.

PodDisruptionBudget

  • One of minAvailable or maxUnavailable, integer or percentage; the intended count comes from the owning workload's .spec.replicas via ownerReferences, so a PDB over bare pods cannot use percentages.
  • It gates the Eviction API only (drain, autoscaler consolidation, descheduler). Node crashes, OOM kills and rolling updates are not evictions; they still count against the budget but are not blocked by it.
  • unhealthyPodEvictionPolicy (GA since 1.31): IfHealthyBudget (default) refuses to evict even unhealthy pods until the budget is met, which is how a crash-looping app stalls a drain; AlwaysAllow lets unhealthy pods go. Kubernetes' own recommendation is AlwaysAllow for anything that may misbehave.
  • Scheduler preemption respects PDBs on a best-effort basis only: if no PDB-safe victim exists, it preempts anyway.

PriorityClass, the fields

  • Cluster-scoped; value up to 1000000000 for user classes; system-cluster-critical (2000000000) and system-node-critical (2000001000) are reserved for the control plane and node agents.
  • globalDefault: true on at most one class applies to pods with no priorityClassName; otherwise the default priority is 0. It does not retroactively change running pods.
  • preemptionPolicy: Never makes a class queue ahead of lower priorities without evicting anyone: the batch-job pattern. The default PreemptLowerPriority evicts.
  • A preempting pod gets status.nominatedNodeName while victims drain; victims receive their full terminationGracePeriodSeconds.
  • Quota can fence a class: a ResourceQuota with scopeSelector scopeName: PriorityClass, operator: In, values: [high] caps how much a tenant may run at that priority (1.4).

Where allocatable comes from

allocatable = capacity - kubeReserved - systemReserved - evictionHard. The kubelet fields are kubeReserved (kubelet, container runtime), systemReserved (sshd, kernel, journald) and evictionHard, whose defaults are memory.available: 100Mi, nodefs.available: 10%, nodefs.inodesFree: 5%, imagefs.available: 15%. enforceNodeAllocatable: [pods] is the default enforcement; adding kube-reserved or system-reserved enforces those cgroups too and is where clusters get hurt when the estimate is wrong. Read all four numbers from kubectl describe node (Capacity vs Allocatable) before you reason about why a pod does not fit.

Node sizing rules of thumb

  • Default cap is 110 pods per node (maxPods); DaemonSets cost one pod and their requests on every node, so many small nodes multiply that overhead, while few huge nodes widen the blast radius of a node failure and make bin-packing lumpier. Two to three zones with at least one spare node's worth of headroom per zone is the usual middle.
  • Request-heavy, bursty workloads pack better on larger nodes; Guaranteed pods (requests equal limits) pack worst because nothing can borrow their slack.
  • Spot or preemptible capacity belongs in its own NodePool with a taint, for replaceable workloads with a PDB, never for anything holding a single-writer volume.
How this gets tested

"A node has been unschedulable and cordoned for an hour; the drain never finishes" is a PDB task: kubectl get pdb -A, find the one whose ALLOWED DISRUPTIONS is 0, and decide whether the right fix is scaling the workload up, relaxing minAvailable, or setting unhealthyPodEvictionPolicy: AlwaysAllow because the pods are unhealthy anyway. Deleting the PDB is the answer only when the task says so.

In-place resize and what else moved recently

version status stated precisely · Kubernetes 1.36 cluster

In-place pod resize

InPlacePodVerticalScaling is GA since Kubernetes 1.35 (alpha 1.27, beta and on by default 1.33), so on a 1.36 cluster the pods/resize subresource exists and the schema check the panel above warns about is moot. The mechanics:

  • You patch through the subresource: kubectl patch pod x --subresource resize --patch '{...}' (also kubectl edit pod x --subresource resize). A plain patch of spec.containers[].resources is rejected as immutable.
  • resizePolicy per container and per resource: restartPolicy: NotRequired (default, apply live) or RestartContainer. A pod with restartPolicy: Never may only use NotRequired.
  • Progress is in pod conditions: PodResizePending with reason Infeasible (node can never fit it) or Deferred (retry later, for example when another pod leaves); PodResizeInProgress while the kubelet actuates, with reason: Error if the runtime refused. status.containerStatuses[].allocatedResources shows what was actually granted, and status.observedGeneration tells you whether the kubelet has seen your latest spec.
  • Only CPU and memory; QoS class is fixed at creation (a Guaranteed pod must stay requests-equal-limits, a Burstable pod may not become Guaranteed, a BestEffort pod may not gain requests); you cannot remove a request once set; memory decrease below current usage is skipped and stays In Progress; non-restartable init containers and ephemeral containers cannot be resized; Windows and static CPU-manager pods are excluded.
  • VPA caught up: updateMode: InPlaceOrRecreate is GA in VPA 1.6 and resizes without eviction where the kubelet allows it, falling back to recreation. That removes the old objection to running VPA in an active mode on latency-sensitive services, though the HPA conflict on the same metric remains.

Pod-level resources

spec.resources on the Pod itself (beta since 1.34, on by default) sets a budget for the whole pod; containers may leave their own requests unset and share it. Pod-level values take precedence for scheduling and QoS. In-place resize of the pod-level budget is beta in 1.36. Quota and LimitRange count it like any other request. Say the version when you mention it; it is new enough to be a trap.

Topology spread, the optional fields

  • minDomains (only with DoNotSchedule): treat fewer eligible domains than this as skew, forcing the autoscaler to open a new zone rather than piling into two.
  • nodeAffinityPolicy and nodeTaintsPolicy (GA since 1.33): Honor or Ignore whether nodes excluded by affinity or taints count as domains. Default honors affinity and ignores taints, which is why tainted spare nodes can distort skew.
  • matchLabelKeys (typically pod-template-hash): spread each ReplicaSet revision independently so a rolling update does not fight the constraint.

HPA details that decide behavior

  • Tolerance is 10% globally; per-HPA spec.behavior.scaleUp.tolerance / scaleDown.tolerance is beta in 1.35 and 1.36 (off by default until GA in 1.37), so do not rely on it in a task unless the cluster shows the field.
  • Not-ready pods and pods with missing metrics are set aside; on scale-down they are assumed to be at 100% of target, on scale-up at 0%, which dampens both directions. CPU metrics from a pod that became ready in the last 30 s (--horizontal-pod-autoscaler-initial-readiness-delay) are ignored.
  • behavior.scaleDown.stabilizationWindowSeconds defaults to 300, scale-up to 0; policies combine by selectPolicy: Max (default), Min, or Disabled to freeze a direction.
  • ContainerResource metrics target one container's usage rather than the pod sum, the fix for sidecar-heavy pods.
Reflex

Before resizing anything, kubectl get pod x -o jsonpath='{.status.qosClass}'. The class cannot change, so a request that would flip it is refused with a validation error rather than deferred, and the message names the field.

Exercises

tick the dot when its check passes

Create three pods in default: one with requests==limits, one with only requests, one with nothing. Predict each class before checking:

kubectl get pod <name> -o jsonpath='{.status.qosClass}{"\n"}'
outputcaptured 2026-08-26
$ kubectl get pod qos-guaranteed -o jsonpath='{.status.qosClass}{"\n"}'
Guaranteed
$ kubectl get pod qos-burstable -o jsonpath='{.status.qosClass}{"\n"}'
Burstable
$ kubectl get pod qos-besteffort -o jsonpath='{.status.qosClass}{"\n"}'
BestEffort
verify: three different answers, all matching your prediction. Delete them.

Memory first. kubectl run lost its resource flags a while back, so this is also a rep for the --overrides escape hatch:

kubectl run oom --image=polinux/stress --restart=Never \
  --overrides='{"spec":{"containers":[{"name":"oom","image":"polinux/stress","command":["stress","--vm","1","--vm-bytes","128M","--vm-hang","0"],"resources":{"requests":{"memory":"64Mi"},"limits":{"memory":"64Mi"}}}]}}'
kubectl get pod oom -w    # until STATUS shows OOMKilled
outputcaptured 2026-08-26
$ kubectl run oom --image=polinux/stress --restart=Never \
  --overrides='{"spec":{"containers":[{"name":"oom","image":"polinux/stress","command":["stress","--vm","1","--vm-bytes","128M","--vm-hang","0"],"resources":{"requests":{"memory":"64Mi"},"limits":{"memory":"64Mi"}}}]}}'
pod/oom created
$ kubectl get pod oom -w    # until STATUS shows OOMKilled
NAME   READY   STATUS    RESTARTS   AGE
oom    0/1     Pending   0          1s
oom    0/1     ContainerCreating   0          1s
oom    0/1     ContainerCreating   0          3s
oom    1/1     Running             0          12s
oom    0/1     OOMKilled           0          12s
oom    0/1     OOMKilled           0          14s
^C
verify: kubectl get pod oom -o jsonpath='{.status.containerStatuses[0].state.terminated.reason}' prints OOMKilled. With --restart=Never nothing restarts, so the evidence sits in state, not lastState; in a crash-looping Deployment it is the other way round. CPU, by contrast, would have throttled silently.

Deploy 4 replicas of anything with:

topologySpreadConstraints:
  - maxSkew: 1
    topologyKey: topology.kubernetes.io/zone
    whenUnsatisfiable: DoNotSchedule
    labelSelector: { matchLabels: { app: spread-demo } }
verify: kubectl get pods -l app=spread-demo -o wide shows 2+2 across the two worker zones, and kubectl get nodes -L topology.kubernetes.io/zone confirms which zone each node carries.

Deploy examples/demo-app/base into default (kubectl apply -k examples/demo-app/base), then:

kubectl autoscale deploy demo --min=2 --max=6 --cpu-percent=20
kubectl run load --image=busybox:1.37 --restart=Never -- \
  sh -c 'while true; do wget -qO- http://demo.default.svc:80 >/dev/null; done'
kubectl get hpa demo -w
outputcaptured 2026-08-26
$ kubectl autoscale deploy demo --min=2 --max=6 --cpu-percent=20
Flag --cpu-percent has been deprecated, Use --cpu with percentage or resource quantity format (e.g., '70%' for utilization or '500m' for milliCPU).
horizontalpodautoscaler.autoscaling/demo autoscaled
$ kubectl run load --image=busybox:1.37 --restart=Never -- \
  sh -c 'while true; do wget -qO- http://demo.default.svc:80 >/dev/null; done'
pod/load created
$ kubectl get hpa demo -w
NAME   REFERENCE         TARGETS              MINPODS   MAXPODS   REPLICAS   AGE
demo   Deployment/demo   cpu: <unknown>/20%   2         6         0          0s
demo   Deployment/demo   cpu: <unknown>/20%   2         6         2          15s
demo   Deployment/demo   cpu: 114%/20%        2         6         2          30s
demo   Deployment/demo   cpu: 98%/20%         2         6         4          45s
demo   Deployment/demo   cpu: 98%/20%         2         6         6          60s
demo   Deployment/demo   cpu: 34%/20%         2         6         6          75s
demo   Deployment/demo   cpu: 36%/20%         2         6         6          91s
demo   Deployment/demo   cpu: 41%/20%         2         6         6          106s
demo   Deployment/demo   cpu: 30%/20%         2         6         6          2m1s
demo   Deployment/demo   cpu: 25%/20%         2         6         6          2m16s
demo   Deployment/demo   cpu: 30%/20%         2         6         6          2m31s
demo   Deployment/demo   cpu: 39%/20%         2         6         6          2m46s
demo   Deployment/demo   cpu: 38%/20%         2         6         6          3m1s
demo   Deployment/demo   cpu: 25%/20%         2         6         6          3m16s
demo   Deployment/demo   cpu: 26%/20%         2         6         6          3m31s
demo   Deployment/demo   cpu: 38%/20%         2         6         6          3m46s
demo   Deployment/demo   cpu: 39%/20%         2         6         6          4m1s
demo   Deployment/demo   cpu: 30%/20%         2         6         6          4m16s
^C
verify: REPLICAS climbs above 2 within a couple of minutes. Kill load and watch it settle back after the stabilization window (about 5 minutes). Note the demo container requests 25m CPU, which is why 20% is reachable at all.

With demo still running:

apiVersion: autoscaling.k8s.io/v1
kind: VerticalPodAutoscaler
metadata: { name: demo, namespace: default }
spec:
  targetRef: { apiVersion: apps/v1, kind: Deployment, name: demo }
  updatePolicy: { updateMode: "Off" }

Give it a few minutes, then:

kubectl get vpa demo -o jsonpath='{.status.recommendation.containerRecommendations[0]}' | jq
outputcaptured 2026-08-26
$ kubectl get vpa demo -o jsonpath='{.status.recommendation.containerRecommendations[0]}' | jq
{
  "containerName": "web",
  "lowerBound": {
    "cpu": "15m",
    "memory": "100Mi"
  },
  "target": {
    "cpu": "23m",
    "memory": "100Mi"
  },
  "uncappedTarget": {
    "cpu": "23m",
    "memory": "100Mi"
  },
  "upperBound": {
    "cpu": "12479m",
    "memory": "8405796509"
  }
}
verify: a target/lowerBound/upperBound block. Compare target against the 25m request in the manifest and say, out loud, whether this workload is over- or under-provisioned. Section 1.5 turns that comparison into money.

In-place resize changed what "change the resources" means: a CPU request can move without the container restarting. The subresource is resize, not a normal patch, and the evidence lives in status.containerStatuses[].allocatedResources. The second half of this exercise asks for more CPU than the node has, to find out where a resize you cannot have gets refused.

kubectl run burst --image=nginx:1.27-alpine --restart=Never --overrides='{"spec":{"containers":[{"name":"burst","image":"nginx:1.27-alpine","resources":{"requests":{"cpu":"100m","memory":"64Mi"},"limits":{"cpu":"200m","memory":"128Mi"}}}]}}'
kubectl wait --for=condition=Ready pod/burst --timeout=90s
kubectl get pod burst -o jsonpath='{.status.qosClass}{"\n"}'
kubectl patch pod burst --subresource resize --patch '{"spec":{"containers":[{"name":"burst","resources":{"requests":{"cpu":"200m","memory":"64Mi"},"limits":{"cpu":"400m","memory":"128Mi"}}}]}}'
kubectl get pod burst -o jsonpath='{.status.containerStatuses[0].allocatedResources}{"\n"}{.status.containerStatuses[0].restartCount}{"\n"}'
kubectl get pod burst -o jsonpath='{.status.conditions}' | jq '.[] | {type, status, reason}'
kubectl patch pod burst --subresource resize --patch '{"spec":{"containers":[{"name":"burst","resources":{"requests":{"cpu":"64","memory":"64Mi"},"limits":{"cpu":"80","memory":"128Mi"}}}]}}'
kubectl get pod burst -o jsonpath='{.status.conditions}' | jq '.[] | {type, status, reason, message}'
kubectl delete pod burst
outputcaptured 2026-09-12
$ kubectl run burst --image=nginx:1.27-alpine --restart=Never --overrides='{"spec":{"containers":[{"name":"burst","image":"nginx:1.27-alpine","resources":{"requests":{"cpu":"100m","memory":"64Mi"},"limits":{"cpu":"200m","memory":"128Mi"}}}]}}'
pod/burst created
$ kubectl wait --for=condition=Ready pod/burst --timeout=90s
pod/burst condition met
$ kubectl get pod burst -o jsonpath='{.status.qosClass}{"\n"}'
Burstable
$ kubectl patch pod burst --subresource resize --patch '{"spec":{"containers":[{"name":"burst","resources":{"requests":{"cpu":"200m","memory":"64Mi"},"limits":{"cpu":"400m","memory":"128Mi"}}}]}}'
pod/burst patched
$ kubectl get pod burst -o jsonpath='{.status.containerStatuses[0].allocatedResources}{"\n"}{.status.containerStatuses[0].restartCount}{"\n"}'
{"cpu":"200m","memory":"64Mi"}
0
$ kubectl get pod burst -o jsonpath='{.status.conditions}' | jq '.[] | {type, status, reason}'
{
  "type": "PodReadyToStartContainers",
  "status": "True",
  "reason": null
}
{
  "type": "Initialized",
  "status": "True",
  "reason": null
}
{
  "type": "Ready",
  "status": "True",
  "reason": null
}
{
  "type": "ContainersReady",
  "status": "True",
  "reason": null
}
{
  "type": "PodScheduled",
  "status": "True",
  "reason": null
}
$ kubectl patch pod burst --subresource resize --patch '{"spec":{"containers":[{"name":"burst","resources":{"requests":{"cpu":"64","memory":"64Mi"},"limits":{"cpu":"80","memory":"128Mi"}}}]}}'
Error from server (Forbidden): pods "burst" is forbidden: node didn't have enough allocatable resources: cpu, requested: 64000, allocatable: 16000
$ kubectl get pod burst -o jsonpath='{.status.conditions}' | jq '.[] | {type, status, reason, message}'
{
  "type": "PodReadyToStartContainers",
  "status": "True",
  "reason": null,
  "message": null
}
{
  "type": "Initialized",
  "status": "True",
  "reason": null,
  "message": null
}
{
  "type": "Ready",
  "status": "True",
  "reason": null,
  "message": null
}
{
  "type": "ContainersReady",
  "status": "True",
  "reason": null,
  "message": null
}
{
  "type": "PodScheduled",
  "status": "True",
  "reason": null,
  "message": null
}
$ kubectl delete pod burst
pod "burst" deleted from default namespace
verify: the reachable resize lands with allocatedResources at {"cpu":"200m","memory":"64Mi"} and restartCount still 0. The impossible one does not park anywhere: the API refuses the patch outright with Error from server (Forbidden): pods "burst" is forbidden: node didn't have enough allocatable resources: cpu, requested: 64000, allocatable: 16000, and the pod's conditions are unchanged, with no PodResizePending among them.

QoS class is computed from requests and limits and it is immutable for the life of the pod. A resize that would move a Guaranteed pod into Burstable is therefore not a scheduling problem, it is a validation error, and the message says so.

kubectl run guar --image=nginx:1.27-alpine --restart=Never --overrides='{"spec":{"containers":[{"name":"guar","image":"nginx:1.27-alpine","resources":{"requests":{"cpu":"100m","memory":"64Mi"},"limits":{"cpu":"100m","memory":"64Mi"}}}]}}'
kubectl wait --for=condition=Ready pod/guar --timeout=90s
kubectl get pod guar -o jsonpath='{.status.qosClass}{"\n"}'
kubectl patch pod guar --subresource resize --patch '{"spec":{"containers":[{"name":"guar","resources":{"requests":{"cpu":"50m","memory":"64Mi"},"limits":{"cpu":"100m","memory":"64Mi"}}}]}}'
kubectl delete pod guar
outputcaptured 2026-09-12
$ kubectl run guar --image=nginx:1.27-alpine --restart=Never --overrides='{"spec":{"containers":[{"name":"guar","image":"nginx:1.27-alpine","resources":{"requests":{"cpu":"100m","memory":"64Mi"},"limits":{"cpu":"100m","memory":"64Mi"}}}]}}'
pod/guar created
$ kubectl wait --for=condition=Ready pod/guar --timeout=90s
pod/guar condition met
$ kubectl get pod guar -o jsonpath='{.status.qosClass}{"\n"}'
Guaranteed
$ kubectl patch pod guar --subresource resize --patch '{"spec":{"containers":[{"name":"guar","resources":{"requests":{"cpu":"50m","memory":"64Mi"},"limits":{"cpu":"100m","memory":"64Mi"}}}]}}'
The Pod "guar" is invalid: spec: Invalid value: "Guaranteed": Pod QOS Class may not change as a result of resizing
$ kubectl delete pod guar
pod "guar" deleted from default namespace
verify: the patch is refused with a message naming the QoS class, and the pod is untouched.

A PodDisruptionBudget that can never be satisfied turns a routine node drain into a stuck terminal. The interesting part is that unhealthy pods count against the budget unless you tell the API otherwise.

kubectl get pdb -A -o custom-columns=NS:.metadata.namespace,NAME:.metadata.name,ALLOWED:.status.disruptionsAllowed
kubectl apply -f - <<'EOF'
apiVersion: apps/v1
kind: Deployment
metadata: { name: crasher, namespace: default }
spec:
  replicas: 2
  selector: { matchLabels: { app: crasher } }
  template:
    metadata: { labels: { app: crasher } }
    spec:
      nodeSelector: { kubernetes.io/hostname: cnpe-worker2 }
      containers:
        - name: c
          image: busybox:1.36
          command: ["sh", "-c", "exit 1"]
---
apiVersion: policy/v1
kind: PodDisruptionBudget
metadata: { name: crasher, namespace: default }
spec:
  minAvailable: 100%
  selector: { matchLabels: { app: crasher } }
EOF
sleep 30
# --pod-selector keeps the drain to this exercise's pods; without it you evict the lab off this node
kubectl get pdb crasher -o jsonpath='{.status.disruptionsAllowed}{"\n"}'
kubectl drain cnpe-worker2 --pod-selector app=crasher --ignore-daemonsets --delete-emptydir-data --timeout=60s
kubectl patch pdb crasher --type merge -p '{"spec":{"unhealthyPodEvictionPolicy":"AlwaysAllow"}}'
kubectl drain cnpe-worker2 --pod-selector app=crasher --ignore-daemonsets --delete-emptydir-data --timeout=60s
kubectl uncordon cnpe-worker2
kubectl delete deployment crasher
kubectl delete pdb crasher
outputcaptured 2026-09-13
$ kubectl get pdb -A -o custom-columns=NS:.metadata.namespace,NAME:.metadata.name,ALLOWED:.status.disruptionsAllowed
NS                  NAME                            ALLOWED
gatekeeper-system   gatekeeper-controller-manager   0
team-a              pg-primary                      0
tracing             otel-collector                  1
$ kubectl apply -f - <<'EOF'
apiVersion: apps/v1
kind: Deployment
metadata: { name: crasher, namespace: default }
spec:
  replicas: 2
  selector: { matchLabels: { app: crasher } }
  template:
    metadata: { labels: { app: crasher } }
    spec:
      nodeSelector: { kubernetes.io/hostname: cnpe-worker2 }
      containers:
        - name: c
          image: busybox:1.36
          command: ["sh", "-c", "exit 1"]
---
apiVersion: policy/v1
kind: PodDisruptionBudget
metadata: { name: crasher, namespace: default }
spec:
  minAvailable: 100%
  selector: { matchLabels: { app: crasher } }
EOF
deployment.apps/crasher created
poddisruptionbudget.policy/crasher created
$ sleep 30
$ # --pod-selector keeps the drain to this exercise's pods; without it you evict the lab off this node
$ kubectl get pdb crasher -o jsonpath='{.status.disruptionsAllowed}{"\n"}'
0
$ kubectl drain cnpe-worker2 --pod-selector app=crasher --ignore-daemonsets --delete-emptydir-data --timeout=60s
node/cnpe-worker2 cordoned
evicting pod default/crasher-7857b795db-hbpzj
evicting pod default/crasher-7857b795db-854th
error when evicting pods/"crasher-7857b795db-hbpzj" -n "default" (will retry after 5s): Cannot evict pod as it would violate the pod's disruption budget.
error when evicting pods/"crasher-7857b795db-854th" -n "default" (will retry after 5s): Cannot evict pod as it would violate the pod's disruption budget.
evicting pod default/crasher-7857b795db-hbpzj
error when evicting pods/"crasher-7857b795db-hbpzj" -n "default" (will retry after 5s): Cannot evict pod as it would violate the pod's disruption budget.
evicting pod default/crasher-7857b795db-854th
error when evicting pods/"crasher-7857b795db-854th" -n "default" (will retry after 5s): Cannot evict pod as it would violate the pod's disruption budget.
evicting pod default/crasher-7857b795db-hbpzj
error when evicting pods/"crasher-7857b795db-hbpzj" -n "default" (will retry after 5s): Cannot evict pod as it would violate the pod's disruption budget.
evicting pod default/crasher-7857b795db-854th
error when evicting pods/"crasher-7857b795db-854th" -n "default" (will retry after 5s): Cannot evict pod as it would violate the pod's disruption budget.
evicting pod default/crasher-7857b795db-hbpzj
error when evicting pods/"crasher-7857b795db-hbpzj" -n "default" (will retry after 5s): Cannot evict pod as it would violate the pod's disruption budget.
evicting pod default/crasher-7857b795db-854th
error when evicting pods/"crasher-7857b795db-854th" -n "default" (will retry after 5s): Cannot evict pod as it would violate the pod's disruption budget.
evicting pod default/crasher-7857b795db-hbpzj
evicting pod default/crasher-7857b795db-854th
error when evicting pods/"crasher-7857b795db-hbpzj" -n "default" (will retry after 5s): Cannot evict pod as it would violate the pod's disruption budget.
error when evicting pods/"crasher-7857b795db-854th" -n "default" (will retry after 5s): Cannot evict pod as it would violate the pod's disruption budget.
evicting pod default/crasher-7857b795db-hbpzj
evicting pod default/crasher-7857b795db-854th
error when evicting pods/"crasher-7857b795db-hbpzj" -n "default" (will retry after 5s): Cannot evict pod as it would violate the pod's disruption budget.
error when evicting pods/"crasher-7857b795db-854th" -n "default" (will retry after 5s): Cannot evict pod as it would violate the pod's disruption budget.
evicting pod default/crasher-7857b795db-hbpzj
evicting pod default/crasher-7857b795db-854th
error when evicting pods/"crasher-7857b795db-854th" -n "default" (will retry after 5s): Cannot evict pod as it would violate the pod's disruption budget.
error when evicting pods/"crasher-7857b795db-hbpzj" -n "default" (will retry after 5s): Cannot evict pod as it would violate the pod's disruption budget.
evicting pod default/crasher-7857b795db-854th
evicting pod default/crasher-7857b795db-hbpzj
error when evicting pods/"crasher-7857b795db-854th" -n "default" (will retry after 5s): Cannot evict pod as it would violate the pod's disruption budget.
error when evicting pods/"crasher-7857b795db-hbpzj" -n "default" (will retry after 5s): Cannot evict pod as it would violate the pod's disruption budget.
evicting pod default/crasher-7857b795db-854th
evicting pod default/crasher-7857b795db-hbpzj
error when evicting pods/"crasher-7857b795db-854th" -n "default" (will retry after 5s): Cannot evict pod as it would violate the pod's disruption budget.
error when evicting pods/"crasher-7857b795db-hbpzj" -n "default" (will retry after 5s): Cannot evict pod as it would violate the pod's disruption budget.
evicting pod default/crasher-7857b795db-854th
evicting pod default/crasher-7857b795db-hbpzj
error when evicting pods/"crasher-7857b795db-hbpzj" -n "default" (will retry after 5s): Cannot evict pod as it would violate the pod's disruption budget.
... 19 more lines
$ kubectl patch pdb crasher --type merge -p '{"spec":{"unhealthyPodEvictionPolicy":"AlwaysAllow"}}'
poddisruptionbudget.policy/crasher patched
$ kubectl drain cnpe-worker2 --pod-selector app=crasher --ignore-daemonsets --delete-emptydir-data --timeout=60s
node/cnpe-worker2 already cordoned
evicting pod default/crasher-7857b795db-hbpzj
evicting pod default/crasher-7857b795db-854th
pod/crasher-7857b795db-854th evicted
pod/crasher-7857b795db-hbpzj evicted
node/cnpe-worker2 drained
$ kubectl uncordon cnpe-worker2
node/cnpe-worker2 uncordoned
$ kubectl delete deployment crasher
deployment.apps "crasher" deleted from default namespace
$ kubectl delete pdb crasher
poddisruptionbudget.policy "crasher" deleted from default namespace
verify: disruptionsAllowed is 0, and the first drain refuses both pods with Cannot evict pod as it would violate the pod's disruption budget until it times out. After unhealthyPodEvictionPolicy: AlwaysAllow the same command evicts both and prints node/cnpe-worker2 drained. Note the --pod-selector: without it this drain empties the node, and half the lab is on it.

Allocatable is capacity minus what the kubelet reserved, and the reservation is a kubelet flag rather than anything in the API. Read both halves on the same node so the subtraction is yours, not a slide's.

kubectl get node cnpe-worker -o jsonpath='{"capacity:    cpu "}{.status.capacity.cpu}{", memory "}{.status.capacity.memory}{"\n"}{"allocatable: cpu "}{.status.allocatable.cpu}{", memory "}{.status.allocatable.memory}{"\n"}'
kubectl get --raw /api/v1/nodes/cnpe-worker/proxy/configz | jq '.kubeletconfig | {kubeReserved, systemReserved, evictionHard}'
kubectl describe node cnpe-worker | sed -n '/Allocated resources/,/Events/p'
outputcaptured 2026-09-13
$ kubectl get node cnpe-worker -o jsonpath='{"capacity:    cpu "}{.status.capacity.cpu}{", memory "}{.status.capacity.memory}{"\n"}{"allocatable: cpu "}{.status.allocatable.cpu}{", memory "}{.status.allocatable.memory}{"\n"}'
capacity:    cpu 16, memory 32747836Ki
allocatable: cpu 16, memory 32747836Ki
$ kubectl get --raw /api/v1/nodes/cnpe-worker/proxy/configz | jq '.kubeletconfig | {kubeReserved, systemReserved, evictionHard}'
{
  "kubeReserved": null,
  "systemReserved": null,
  "evictionHard": {
    "imagefs.available": "0%",
    "nodefs.available": "0%",
    "nodefs.inodesFree": "0%"
  }
}
$ kubectl describe node cnpe-worker | sed -n '/Allocated resources/,/Events/p'
Allocated resources:
  (Total limits may be over 100 percent, i.e., overcommitted.)
  Resource           Requests      Limits
  --------           --------      ------
  cpu                1811m (11%)   9450m (59%)
  memory             3714Mi (11%)  12800Mi (40%)
  ephemeral-storage  0 (0%)        0 (0%)
  hugepages-1Gi      0 (0%)        0 (0%)
  hugepages-2Mi      0 (0%)        0 (0%)
Events:
verify: Capacity and Allocatable are identical on this node. The kubelet reserves nothing here, kubeReserved and systemReserved both read null and evictionHard is 0% on every signal, so there is no gap to subtract. The scheduler adds pod requests up against Allocatable, which is the number the Allocated resources block takes its percentages from.

Two priority classes at the same value behave completely differently if one sets preemptionPolicy: Never. Batch work usually wants that: high enough to jump the scheduling queue, never rude enough to evict a neighbor.

kubectl create priorityclass batch --value=1000 --preemption-policy=Never
kubectl create priorityclass urgent --value=1000
kubectl run filler --image=nginx:1.27-alpine --restart=Never --overrides='{"spec":{"nodeSelector":{"kubernetes.io/hostname":"cnpe-worker2"},"containers":[{"name":"filler","image":"nginx:1.27-alpine","resources":{"requests":{"cpu":"12"}}}]}}'
kubectl wait --for=condition=Ready pod/filler --timeout=120s
kubectl run batchpod --image=nginx:1.27-alpine --restart=Never --overrides='{"spec":{"priorityClassName":"batch","nodeSelector":{"kubernetes.io/hostname":"cnpe-worker2"},"containers":[{"name":"batchpod","image":"nginx:1.27-alpine","resources":{"requests":{"cpu":"8"}}}]}}'
sleep 20
kubectl get pod batchpod -o wide
kubectl describe pod batchpod | tail -4
kubectl run urgentpod --image=nginx:1.27-alpine --restart=Never --overrides='{"spec":{"priorityClassName":"urgent","nodeSelector":{"kubernetes.io/hostname":"cnpe-worker2"},"containers":[{"name":"urgentpod","image":"nginx:1.27-alpine","resources":{"requests":{"cpu":"8"}}}]}}'
sleep 30
kubectl get pods filler batchpod urgentpod -o wide
kubectl delete pod filler batchpod urgentpod --ignore-not-found
kubectl delete priorityclass batch urgent
outputcaptured 2026-09-12
$ kubectl create priorityclass batch --value=1000 --preemption-policy=Never
priorityclass.scheduling.k8s.io/batch created
$ kubectl create priorityclass urgent --value=1000
priorityclass.scheduling.k8s.io/urgent created
$ kubectl run filler --image=nginx:1.27-alpine --restart=Never --overrides='{"spec":{"nodeSelector":{"kubernetes.io/hostname":"cnpe-worker2"},"containers":[{"name":"filler","image":"nginx:1.27-alpine","resources":{"requests":{"cpu":"12"}}}]}}'
pod/filler created
$ kubectl wait --for=condition=Ready pod/filler --timeout=120s
pod/filler condition met
$ kubectl run batchpod --image=nginx:1.27-alpine --restart=Never --overrides='{"spec":{"priorityClassName":"batch","nodeSelector":{"kubernetes.io/hostname":"cnpe-worker2"},"containers":[{"name":"batchpod","image":"nginx:1.27-alpine","resources":{"requests":{"cpu":"8"}}}]}}'
pod/batchpod created
$ sleep 20
$ kubectl get pod batchpod -o wide
NAME       READY   STATUS    RESTARTS   AGE   IP       NODE     NOMINATED NODE   READINESS GATES
batchpod   0/1     Pending   0          20s   <none>   <none>   <none>           <none>
$ kubectl describe pod batchpod | tail -4
  Type     Reason            Age   From               Message
  ----     ------            ----  ----               -------
  Warning  PolicyViolation   20s   kyverno-admission  policy require-resource-requests/ fail: every container must set cpu and memory requests
  Warning  FailedScheduling  20s   default-scheduler  0/3 nodes are available: 1 Insufficient cpu, 1 node(s) didn't match Pod's node affinity/selector, 1 node(s) had untolerated taint(s). no new claims to deallocate, preemption: not eligible due to preemptionPolicy=Never.
$ kubectl run urgentpod --image=nginx:1.27-alpine --restart=Never --overrides='{"spec":{"priorityClassName":"urgent","nodeSelector":{"kubernetes.io/hostname":"cnpe-worker2"},"containers":[{"name":"urgentpod","image":"nginx:1.27-alpine","resources":{"requests":{"cpu":"8"}}}]}}'
pod/urgentpod created
$ sleep 30
$ kubectl get pods filler batchpod urgentpod -o wide
NAME        READY   STATUS    RESTARTS   AGE   IP            NODE           NOMINATED NODE   READINESS GATES
batchpod    0/1     Pending   0          51s   <none>        <none>         <none>           <none>
urgentpod   1/1     Running   0          31s   10.244.2.42   cnpe-worker2   <none>           <none>
Error from server (NotFound): pods "filler" not found
$ kubectl delete pod filler batchpod urgentpod --ignore-not-found
pod "batchpod" deleted from default namespace
pod "urgentpod" deleted from default namespace
$ kubectl delete priorityclass batch urgent
priorityclass.scheduling.k8s.io "batch" deleted
priorityclass.scheduling.k8s.io "urgent" deleted
verify: batchpod stays Pending with an Insufficient cpu event and nothing is evicted, while urgentpod at the same value takes the room by pushing filler out. The two classes carry the same value; only preemptionPolicy differs.

An HPA created with one command still has a full scale-up and scale-down policy; you just did not type it. Know where the defaults live before a task asks you to slow a flapping scale-down.

kubectl create deployment hpademo --image=nginx:1.27-alpine
kubectl set resources deployment hpademo --requests=cpu=50m
kubectl autoscale deployment hpademo --cpu-percent=80 --min=1 --max=5
kubectl get hpa hpademo -o jsonpath='{.spec.behavior}{"\n"}'
kubectl get hpa hpademo -o yaml | sed -n '/^spec:/,/^status:/p'
kubectl explain hpa.spec.behavior.scaleDown
kubectl delete hpa hpademo
kubectl delete deployment hpademo
outputcaptured 2026-09-12
$ kubectl create deployment hpademo --image=nginx:1.27-alpine
deployment.apps/hpademo created
$ kubectl set resources deployment hpademo --requests=cpu=50m
deployment.apps/hpademo resource requirements updated
$ kubectl autoscale deployment hpademo --cpu-percent=80 --min=1 --max=5
Flag --cpu-percent has been deprecated, Use --cpu with percentage or resource quantity format (e.g., '70%' for utilization or '500m' for milliCPU).
horizontalpodautoscaler.autoscaling/hpademo autoscaled
$ kubectl get hpa hpademo -o jsonpath='{.spec.behavior}{"\n"}'
$ kubectl get hpa hpademo -o yaml | sed -n '/^spec:/,/^status:/p'
spec:
  maxReplicas: 5
  metrics:
  - resource:
      name: cpu
      target:
        averageUtilization: 80
        type: Utilization
    type: Resource
  minReplicas: 1
  scaleTargetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: hpademo
status:
$ kubectl explain hpa.spec.behavior.scaleDown
GROUP:      autoscaling
KIND:       HorizontalPodAutoscaler
VERSION:    v2

FIELD: scaleDown <HPAScalingRules>


DESCRIPTION:
    scaleDown is scaling policy for scaling Down. If not set, the default value
    is to allow to scale down to minReplicas pods, with a 300 second
    stabilization window (i.e., the highest recommendation for the last 300sec
    is used).
    HPAScalingRules configures the scaling behavior for one direction via
    scaling Policy Rules and a configurable metric tolerance.
    
    Scaling Policy Rules are applied after calculating DesiredReplicas from
    metrics for the HPA. They can limit the scaling velocity by specifying
    scaling policies. They can prevent flapping by specifying the stabilization
    window, so that the number of replicas is not set instantly, instead, the
    safest value from the stabilization window is chosen.
    
    The tolerance is applied to the metric values and prevents scaling too
    eagerly for small metric variations. (Note that setting a tolerance requires
    the beta HPAConfigurableTolerance feature gate to be enabled.)
    
FIELDS:
  policies	<[]HPAScalingPolicy>
    policies is a list of potential scaling polices which can be used during
    scaling. If not set, use the default values: - For scale up: allow doubling
    the number of pods, or an absolute change of 4 pods in a 15s window. - For
    scale down: allow all pods to be removed in a 15s window.

  selectPolicy	<string>
    selectPolicy is used to specify which policy should be used. If not set, the
    default value Max is used.

  stabilizationWindowSeconds	<integer>
    stabilizationWindowSeconds is the number of seconds for which past
    recommendations should be considered while scaling up or scaling down.
    StabilizationWindowSeconds must be greater than or equal to zero and less
... 17 more lines
$ kubectl delete hpa hpademo
horizontalpodautoscaler.autoscaling "hpademo" deleted from default namespace
$ kubectl delete deployment hpademo
deployment.apps "hpademo" deleted from default namespace
verify: you can state the stabilization window that applies to scale-down on this HPA and where that number came from. If spec.behavior reads back empty, the defaults are the controller's rather than the object's, and that is the answer to give.

Self-check

answer before opening
A node shows 25% CPU usage and refuses to schedule a pod requesting 200m. Explain, in one sentence.

The scheduler counts requests against allocatable, not usage: the node's requests are already at allocatable even though the processes are idle. Fix by right-sizing the existing requests (or adding capacity), not by adding CPU headroom that already exists.

Which pod does the kubelet evict first under memory pressure, and why?

BestEffort, because it declared no memory request at all and therefore ranks worst. Then Burstable pods exceeding their requests, ordered by how far over they are. Guaranteed last. The kernel's oom_score_adj mirrors that ranking so a system-level OOM picks the same victim.

Your service has p99 latency spikes every few seconds; CPU usage sits at 60% of the limit. First metric you check?

rate(container_cpu_cfs_throttled_periods_total[5m]). Average usage below the limit hides per-period throttling: the container burns its 100ms quota early and stalls until the next period. Average utilization is the wrong lens for CPU limits.

Why can a VPA in Auto mode and an HPA on CPU not coexist on the same workload?

They form a loop: the HPA scales replicas based on usage-against-requests while the VPA rewrites those same requests, so each one keeps invalidating the other's denominator. Standard resolution: HPA on CPU with VPA in Off (advisory), or HPA on a custom/external metric while VPA owns the resources.

A drain hangs forever on one node. What are the two usual causes?

A PodDisruptionBudget that cannot be satisfied (often minAvailable equal to the replica count), or unmanaged pods: bare pods with no controller, which kubectl drain refuses to evict without --force. Both are voluntary-disruption mechanics; neither has anything to do with node health.

Is in-place pod resize usable on a Kubernetes 1.36 cluster, and what does the request look like?

Yes: InPlacePodVerticalScaling has been GA since 1.35 (beta and default-on since 1.33). You patch the resize subresource (kubectl patch pod x --subresource resize ...); a plain spec patch is rejected as immutable. Watch PodResizePending (Infeasible or Deferred) and PodResizeInProgress conditions, and remember the QoS class cannot change.

Karpenter refuses to consolidate a node you consider empty enough. Where is the reason, and name two things it could be.

Events on the node with reason Unconsolidatable. Typical messages: a PDB prevents pod evictions, a pod carries karpenter.sh/do-not-disrupt, no cheaper replacement exists, or the NodePool's disruption budget is exhausted. Cluster Autoscaler's equivalents are the safe-to-evict annotation, uncontrolled pods and local storage.

A PDB has minAvailable 2 for a 2-replica Deployment whose pods are CrashLoopBackOff. Drain hangs. Fix without deleting the PDB?

Set unhealthyPodEvictionPolicy: AlwaysAllow on the PDB: the default IfHealthyBudget refuses to evict unhealthy pods until the budget is met, which can never happen here. Alternatively scale the Deployment so the budget is satisfiable. Deleting the PDB removes the protection permanently rather than for the drain.

What is allocatable, and which four kubelet settings subtract from capacity?

Allocatable is what the scheduler may hand to pods: capacity minus kubeReserved, systemReserved and the evictionHard thresholds (defaults memory.available 100Mi, nodefs.available 10%, nodefs.inodesFree 5%, imagefs.available 15%). enforceNodeAllocatable decides which of those cgroups are actually enforced; pods is the default.

Docs to know your way around

study time, not exam time
  • kubernetes.io: Resource Management for Pods and Containers; Pod Quality of Service Classes; Node-pressure Eviction; Assigning Pods to Nodes; Pod Topology Spread Constraints; the HorizontalPodAutoscaler walkthrough (the algorithm section especially).
  • github.com/kubernetes/autoscaler: the VPA README, in particular the updateMode table.
  • Offline: kubectl explain pod.spec.containers.resources, kubectl explain hpa.spec.behavior --recursive, and the requests table at the bottom of kubectl describe node.
  • kubernetes.io/docs/tasks/configure-pod-container/resize-container-resources: the resize subresource, resizePolicy, and the PodResizePending / PodResizeInProgress conditions with their reasons.
  • kubernetes.io/docs/concepts/cluster-administration/node-autoscaling: the neutral Cluster Autoscaler vs Karpenter comparison; karpenter.sh/docs/concepts/disruption for consolidationPolicy, budgets and do-not-disrupt.
  • kubernetes.io/docs/tasks/administer-cluster/reserve-compute-resources: kubeReserved, systemReserved, evictionHard and enforceNodeAllocatable with the worked allocatable example.