Being able to say, for any packet, what touches it and in what order: pod, CNI, Service, kube-proxy or its eBPF replacement, DNS, policy, ingress. Most delivery and incident tasks in the other four domains eventually collapse into this one.

needs make upmake gitea gitopsmake sec

Orientation

competency 1.1 · architecture best practices

Networking is the substrate every other domain stands on. A Crossplane XR that never goes Ready, an Argo CD app stuck Progressing, a Prometheus target that will not come up, a canary that never receives traffic: a surprising share of those are one Service selector, one missing DNS egress rule, or one unready endpoint.

What the exam actually asks

Not "explain CNI". It asks you to make traffic work or make traffic stop, then prove it: expose a workload, restrict a namespace, author a Gateway and an HTTPRoute, or explain why a Service resolves and then refuses connections. Every one of those is a five-minute task if the model in your head is exact, and twenty minutes of guessing if it is fuzzy.

The four rules of the Kubernetes network model

  1. Every pod gets its own IP, cluster-routable, no NAT between pods.
  2. Pods on a node can reach all pods on all nodes without NAT.
  3. Agents on a node (kubelet, system daemons) can reach all pods on that node.
  4. A pod sees its own IP as the same address other pods use to reach it.

The CNI plugin is whatever implements those rules: routing, overlay, or eBPF datapath. Everything above (Services, DNS, policy, Gateways) is built on the assumption that the four rules already hold. When they do not, nothing above them behaves sanely, which is why "is it a CNI problem or a Service problem" is the first fork in any network diagnosis.

Services and the thing that actually makes them work

ClusterIP · NodePort · LoadBalancer · headless

ClusterIP is a virtual IP that exists only as translation rules on each node; nothing listens on it, no interface owns it, you cannot ping it in any meaningful sense. NodePort opens the same Service on a high port (30000–32767 by default) of every node. LoadBalancer is NodePort plus something external handing out a real IP (in this lab, cloud-provider-kind). Headless (clusterIP: None) skips the VIP entirely and returns pod IPs straight from DNS, which is what StatefulSets need for stable peer discovery. ExternalName is a CNAME with no proxying at all.

Typespec bitsReaches it fromUse it when
ClusterIPdefaultinside the clusterthe 90% case; every internal call
NodePorttype: NodePortany node IPbootstrapping, bare metal, demos
LoadBalancertype: LoadBalanceroutsidereal external entry point (loadBalancerClass only picks between implementations)
HeadlessclusterIP: NoneDNS returns pod IPsStatefulSet peers, client-side LB
ExternalNameexternalName: hostDNS CNAME onlyaliasing an out-of-cluster host
The idea everything else rests on

The thing that makes a Service work is not the Service. It is the EndpointSlice behind it, and the selector that fills it. A Service with no ready endpoints resolves fine and then refuses connections, which is one of the failures the exam leans on most. Check order, every time: selector matches pod labels → pods are Ready (readiness gates endpoint membership) → kubectl get endpointslices -l kubernetes.io/service-name=<svc>.

Two related fields decide whether traffic reaches a pod at all. readinessProbe failure pulls the pod out of the slice, which is the intended way to drain a pod. publishNotReadyAddresses: true overrides that and is how headless Services for clustered databases let peers find each other before they are serving. And terminationGracePeriodSeconds plus a preStop sleep is the standard trick for the race where a pod is deleted but nodes have not yet removed its rules; connections refused during rollouts almost always trace here.

Traffic policies

  • externalTrafficPolicy: Cluster (default) SNATs and may hop to another node, so the backend sees the node's IP, not the client's. Local preserves the client source IP and only routes to pods on the receiving node, with the trade that a node holding no pod blackholes the traffic (health checks are what stop the external LB from sending there).
  • internalTrafficPolicy: Local is the same idea for in-cluster traffic: node-local endpoints only. Used for node-local caches and log shippers.
  • sessionAffinity: ClientIP is the only affinity a Service offers. Anything richer belongs in a mesh or gateway.

Five conditions, one request. Turn any of them off and see which hop drops it, what the user reports, and the command that proves it:

Command reflex

kubectl get endpointslices -l kubernetes.io/service-name=X beats describe svc, because it shows you conditions per address (ready, serving, terminating) rather than a summarized list. When a rollout half-breaks, those three booleans tell the whole story.

The datapath: kube-proxy, eBPF, and this cluster's choice

Cilium · Hubble · identities

Classic kube-proxy watches Services and EndpointSlices and programs the node: iptables mode writes DNAT chains (simple, and linear-ish to rule count), IPVS mode uses kernel load balancing with real scheduling algorithms. eBPF datapaths (Cilium, Calico) replace kube-proxy entirely, doing the translation in the socket or TC layer, which removes the rule-table scaling problem and unlocks flow visibility.

client pod
   │ connect 10.96.0.42:80          ← ClusterIP, exists only as a rule
   ▼
[ datapath ]  kube-proxy iptables/IPVS  or  eBPF (Cilium, this lab)
   │ DNAT → picks one endpoint from the EndpointSlice
   ▼
backend pod IP 10.244.2.17:8080
   │
   ▼ policy is evaluated on the pod identity, not the Service VIP

Note the last line. This is why an ipBlock rule naming a ClusterIP never matches: by the time policy is evaluated, the destination has already been rewritten to a pod IP. That single fact explains a whole family of "my NetworkPolicy does nothing" tasks; section 3.3 sets the same trap with CloudNativePG.

This cluster runs Cilium instead of kindnet plus kube-proxy for one decisive reason: kindnet does not enforce NetworkPolicy at all. A lab where policies apply cleanly and change nothing teaches you the wrong lesson permanently. Cilium also brings identity-based policy (labels are compiled into numeric identities, so policy survives pod IP churn) and Hubble, which shows flows with their verdict: forwarded or dropped, and by which policy.

Cilium extras worth knowing by name

CiliumNetworkPolicy adds what upstream NetworkPolicy cannot express: toFQDNs (allow api.github.com by name, enforced at DNS), toEntities (kube-apiserver, world, host, remote-node), and L7 rules for HTTP/DNS/Kafka. The exam will not test CRD field names, but "which policy engine can express egress to the API server" is a fair scenario, and the answer here is toEntities: [kube-apiserver].

DNS: the layer that fails quietly

CoreDNS · search domains · ndots

Names are <svc>.<ns>.svc.cluster.local, served by CoreDNS in kube-system, backed by a Service called kube-dns (the name is left over from kube-dns; the pods are CoreDNS). Pods get a /etc/resolv.conf with a search list and ndots:5, which means any name with fewer than five dots is tried against every search domain first. backend.default.svc.cluster.local (four dots, still under five) becomes four or five queries before the right one lands; a trailing dot (backend.default.svc.cluster.local.) skips the search list entirely, and that is the cheap fix for DNS-heavy workloads.

RecordResolves toNotes
svc.ns.svc.cluster.localClusterIPthe normal A record
svc.nsClusterIPworks via search domains
headless.ns.svc…all ready pod IPsmultiple A records, client picks
pod-0.headless.ns.svc…one podStatefulSet stable identity
_port._tcp.svc.ns.svc…SRV recordnamed ports; how peers discover ports

spec.dnsPolicy and dnsConfig on a pod let you override all of this: ClusterFirst (default), None plus explicit nameservers, Default (inherit the node's). It comes up when a workload must resolve an external private zone.

The failure mode this lab teaches on purpose

team-a allows DNS egress through one explicit NetworkPolicy rule (UDP/TCP 53 to kube-dns). The netpol break drill removes exactly that rule. Result: every pod stays Running, every probe stays green, and nothing can resolve anything. No pod listing will ever show it. You find it by making something try (nslookup from inside) and then reading Hubble's DROPPED verdicts on port 53. The same break returns in section 4.6.

NetworkPolicy semantics

additive allow-lists · both ends must open
  • Policies are namespaced and select pods, never Services.
  • A pod is "isolated" for a direction the moment any policy selects it for that direction. Until then, everything is allowed.
  • Rules are additive allow-lists. There is no deny rule. You loosen by adding and tighten only by removing. The break drill exploits that asymmetry: a missing rule looks exactly like a rule that was never there.
  • Ingress and egress are independent. A cross-namespace call needs an egress allow on the caller's side and an ingress allow on the callee's. Discovering that empirically once (exercise 3 below) saves you re-deriving it under pressure.
  • Default-deny is itself a policy: empty podSelector, policyTypes: [Ingress, Egress], no rules.
  • Selector scoping inside a rule matters: namespaceSelector and podSelector in the same list item is an AND (those pods in those namespaces); as two separate items it is an OR. This is the single most common authoring bug.
Exam angle

A task that says "team-a must reach only team-b's web service and DNS" is asking for three objects, not one: default-deny in team-a, an egress allow in team-a, an ingress allow in team-b. Grade yourself the way a grader would: a curl that returns 200 and a second curl to something else that times out.

Ingress vs Gateway API

the exam-era answer is Gateway API

Ingress was one object owned by nobody in particular, extended by controller-specific annotations. Gateway API replaces it with a role-split model, and the roles are the point:

KindOwned bySays
GatewayClassinfrastructure providerwhich controller implements Gateways of this class
Gatewayplatform teamlisteners: ports, protocols, TLS, and which routes may attach
HTTPRoute / GRPCRoute / TCPRouteapplication teammatch rules, filters, weighted backends
ReferenceGrantthe namespace being referencedconsent for a cross-namespace backend reference

Route attachment is two-sided, in the same spirit as NetworkPolicy: the route names a parentRef, and the Gateway's listener declares allowedRoutes (same namespace, selected namespaces, or all). Both must agree, and a route whose attachment is refused reports it in status rather than failing to apply.

Status conditions are how you grade your own work: Accepted (the controller understood it), Programmed (the data plane is configured), ResolvedRefs (backends and secrets exist and are permitted). Reading those three beats guessing every time.

Honest lab caveat

This cluster installs the standard-channel Gateway API CRDs so you can author and schema-validate all of it, but no controller programs a data plane for them on the main cluster. Your Gateway will sit unprogrammed; do not mistake a clean apply for a working Gateway. On the mesh cluster (make mesh), Istio serves Gateway API for real: the same manifests, an actual Programmed condition.

One neighbor the Gateway needs and Kubernetes does not provide: certificates. A TLS listener references a Secret, and something has to keep that Secret valid. In practice that something is cert-manager: an Issuer/ClusterIssuer (ACME, CA, or Vault) plus a Certificate, or the cert-manager.io/cluster-issuer annotation on the Gateway, which issues into the Secret and renews before expiry. Not installed in this lab, but "who renews that certificate" is a fair question about any Gateway you design, and ResolvedRefs: False on a listener is usually its absence.

Progressive delivery ties in here (section 2.5): weighted backendRefs on an HTTPRoute are how a mesh shifts traffic by route weight instead of by replica count, and Flagger drives exactly those weights.

Gateway API, field by field

listeners · matches · filters · weights · ReferenceGrant · status reasons

The role table above is the model. A task hands you the fields. Gateway API v1.5 (April 2026) is the current release; everything below is Standard channel unless marked. On the exam cluster, kubectl explain gateway.spec.listeners --recursive and kubectl explain httproute.spec.rules --recursive print the same fields, version-correct.

Gateway listeners

  • name, port, protocol (HTTP, HTTPS, TLS, TCP, UDP) and an optional hostname. Two listeners with the same port, protocol and hostname conflict; the loser carries Conflicted: True with reason HostnameConflict or ProtocolConflict.
  • tls.mode: Terminate (the Gateway decrypts; needs tls.certificateRefs pointing at a kubernetes.io/tls Secret) or Passthrough (SNI routing only, used with TLSRoute). A Secret in another namespace needs a ReferenceGrant in the Secret's namespace, otherwise ResolvedRefs: False with RefNotPermitted.
  • allowedRoutes.namespaces.from: Same (default), All, or Selector with a label selector; allowedRoutes.kinds restricts route kinds. This is the Gateway's half of attachment consent.
  • Gateway conditions: Accepted (reasons include ListenersNotValid, Pending), Programmed (AddressNotAssigned, NoResources), per-listener Accepted, ResolvedRefs, Programmed. A Gateway whose class has no controller sits at Accepted: Unknown forever.

HTTPRoute rules

  • parentRefs name the Gateway; add sectionName (listener name) or port to bind to one listener rather than all of them. hostnames on the route must intersect the listener hostname or you get Accepted: False, NoMatchingListenerHostname.
  • matches: a list where each item is ANDed inside (path AND headers AND queryParams AND method) and items are ORed against each other. Path types: PathPrefix (default), Exact, RegularExpression (implementation-specific). No matches at all means PathPrefix /, which matches everything.
  • filters: RequestHeaderModifier, ResponseHeaderModifier, RequestRedirect, URLRewrite (these two cannot be combined), RequestMirror, CORS (Standard since v1.5) and ExtensionRef for vendor CRDs. Filters replace almost every Ingress annotation you used to write.
  • backendRefs: name, port, optional namespace (needs a ReferenceGrant) and weight (default 1). Weights are shares of the total: 90 and 10 split 90/10, 3 and 1 split 75/25. A rule with no backendRefs and no redirect answers 500. A missing Service gives ResolvedRefs: False, BackendNotFound.
  • timeouts (Standard since v1.2): request for the whole exchange, backendRequest for one attempt; backendRequest may not exceed request; 0s disables.
apiVersion: gateway.networking.k8s.io/v1
kind: Gateway
metadata: { name: web, namespace: infra }
spec:
  gatewayClassName: istio
  listeners:
    - name: https
      port: 443
      protocol: HTTPS
      hostname: "*.example.com"
      tls: { mode: Terminate, certificateRefs: [{ kind: Secret, name: wildcard-tls }] }
      allowedRoutes: { namespaces: { from: Selector, selector: { matchLabels: { tenant: "true" } } } }
---
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata: { name: demo, namespace: team-a }
spec:
  parentRefs: [{ name: web, namespace: infra, sectionName: https }]
  hostnames: ["demo.example.com"]
  rules:
    - matches: [{ path: { type: PathPrefix, value: /api } }]
      filters: [{ type: RequestHeaderModifier, requestHeaderModifier: { add: [{ name: X-Env, value: staging }] } }]
      backendRefs:
        - { name: demo-stable, port: 80, weight: 90 }
        - { name: demo-canary, port: 80, weight: 10 }

The neighbors that arrived recently

  • ReferenceGrant is v1 since v1.5: from (group, kind, namespace of the referrer) and to (group, kind, optionally name) in the namespace being referenced. It is consent, never a route: it grants nothing by itself.
  • BackendTLSPolicy (Standard since v1.4) configures TLS from the Gateway to the backend pods: targetRefs a Service, validation.hostname as SNI, and either caCertificateRefs or wellKnownCACertificates: System. Same namespace as the Service, by design.
  • ListenerSet (Standard since v1.5) lets a team add listeners to a shared Gateway without editing it; the Gateway must opt in with spec.allowedListeners.namespaces.from (default None). Parent listeners win conflicts. It also lifts the 64-listener ceiling.
  • TLSRoute is v1 since v1.5; GRPCRoute has been Standard since v1.1.

Ingress to Gateway, mapped

IngressGateway API
IngressClassGatewayClass (controllerName instead of controller)
spec.tls[].secretName + hostsGateway listener protocol HTTPS, tls.certificateRefs, hostname
spec.rules[].host + pathsHTTPRoute hostnames + rules[].matches[].path
pathType Prefix / Exact / ImplementationSpecificPathPrefix / Exact / RegularExpression
backend.service.name/portbackendRefs[].name/port (+ weight)
controller annotations (rewrite, redirect, canary weight, CORS)filters and backendRefs weights, portable across controllers
one object, one ownerGateway owned by platform, HTTPRoute owned by the app team, ReferenceGrant by the referenced namespace

The community ingress2gateway tool does the mechanical conversion for common controllers; the judgment it cannot do for you is deciding who owns the Gateway and which namespaces its listeners admit.

How this gets tested

"Expose service X on host Y through the existing Gateway in namespace infra" is three checks: does the Gateway's listener admit routes from your namespace (allowedRoutes), does your hostname fit the listener hostname, and does status.parents[].conditions read Accepted: True and ResolvedRefs: True. A cross-namespace backend adds a fourth object, the ReferenceGrant, in the backend's namespace, not yours.

Datapath and architecture details a scenario can name

kube-proxy modes · Cilium replacement · traffic distribution · Corefile knobs · HA and zones

kube-proxy modes and their replacement

ModeStatusWhat to remember
iptablesdefaultDNAT chains; rule count grows with services; minSyncPeriod and syncPeriod tune programming latency
ipvsdeprecated since v1.35kernel load balancing with real algorithms (rr, lc, sh); needs ipvs kernel modules
nftablesGA since v1.33kernel 5.13+; faster endpoint updates; NodePorts default to --nodeport-addresses primary (the node primary IP only)
kernelspaceWindows onlythe Windows equivalent
Cilium kube-proxy replacementHelm kubeProxyReplacement=trueClusterIP, NodePort, LoadBalancer, hostPort all in eBPF; socket-level LB for east-west (the connect() is rewritten, so there is no per-packet DNAT to trace); DSR and Maglev options; cilium status reports the mode

Two Service fields that ride on this layer. spec.trafficDistribution (GA since v1.33) takes PreferSameZone or PreferSameNode (PreferClose is the deprecated older alias); it is a preference, and externalTrafficPolicy / internalTrafficPolicy: Local override it. The old Endpoints API is deprecated since v1.33 in favor of EndpointSlice, and spec.externalIPs is deprecated since v1.36; when a task says "expose", the modern answers are LoadBalancer or a Gateway.

CoreDNS: the Corefile you will be asked to read or edit

The ConfigMap is coredns in kube-system; the reload plugin picks up edits within about two minutes, so do not restart pods to make a change land. The stock server block and what each line buys you:

LineEffectKnob a task turns
kubernetes cluster.local in-addr.arpa ip6.arpa { pods insecure; fallthrough in-addr.arpa ip6.arpa; ttl 30 }serves Service and Pod records for the cluster domainpods verified returns pod A records only for pods that exist; ttl 0-3600 s (default 5)
forward . /etc/resolv.confeverything not in the cluster domain goes upstreamforward . 172.16.0.1 pins the upstream; a second server block corp.example:53 { forward . 10.0.0.53 } is a stub domain
cache 30response cacheraise it for DNS-heavy tenants, or add NodeLocal DNSCache (a DaemonSet cache on each node that also upgrades upstream queries to TCP and avoids conntrack races)
loop · reload · loadbalance · errors · health · ready · prometheus :9153loop detection, live config reload, A-record shuffling, logging, probes, metricsmostly leave alone; log is the debugging plugin you add temporarily

Pod side: dnsPolicy is ClusterFirst unless you say otherwise; a hostNetwork: true pod needs ClusterFirstWithHostNet or it silently falls back to the node's resolver; None plus dnsConfig (at most 3 nameservers, up to 32 search domains, options: [{name: ndots, value: "2"}]) is how you cut the search-list storm without a trailing dot. The autopath plugin does the same server-side but requires pods verified.

Policy details that differ between engines

  • Target a namespace by name with the immutable label kubernetes.io/metadata.name in a namespaceSelector; do not use an empty {} selector, which matches every namespace.
  • In Cilium, ipBlock rules do not match pod or node IPs by default (policyCIDRMatchMode changes it); pods are selected by labels and identities. That is a second reason, beyond DNAT, why CIDR rules "do nothing".
  • CiliumClusterwideNetworkPolicy is the cluster-scoped twin of CiliumNetworkPolicy: same spec, no namespace, for baseline rules such as "everyone may reach DNS and the API server". Cilium 1.20 also implements the upstream ClusterNetworkPolicy API with admin and baseline tiers that ordinary NetworkPolicy cannot override.

Architecture decisions a scenario can name

  • HA control plane: three control plane nodes minimum. Stacked etcd (etcd on the control plane nodes, kubeadm's default) is simpler; external etcd needs six hosts but losing a node no longer removes both an API server and an etcd member. Either way the API server sits behind a load balancer; Kubernetes itself does not fail over the endpoint.
  • Zones: nodes carry topology.kubernetes.io/zone and region; spread control plane replicas over at least three zones, spread workloads with topologySpreadConstraints, and use WaitForFirstConsumer storage so a volume is not provisioned in a zone the pod cannot reach.
  • Scale limits worth quoting: 110 pods per node by default, 5000 nodes, 150000 pods per cluster.
  • Multi-cluster: a hub Argo CD with cluster Secrets and an ApplicationSet clusters generator, or one Flux per cluster each reconciling clusters/<name> from a shared repo. Hub-and-spoke centralizes credentials and UI; agent-per-cluster survives hub loss and keeps credentials local. Section 2.2 covers the objects.
The hostNetwork DNS trap

A monitoring or CNI-adjacent pod on hostNetwork: true that "cannot resolve services" is almost always dnsPolicy: ClusterFirst behaving as Default. Set ClusterFirstWithHostNet. The symptom looks exactly like a NetworkPolicy drop, and Hubble shows nothing because the query never went to CoreDNS.

Exercises

click a title to collapse · tick the dot when its check passes

With gitops up, Argo CD sits behind a LoadBalancer (running server.insecure, so plain http):

kubectl -n argocd get svc argocd-server -o wide      # note the EXTERNAL-IP
kubectl -n argocd get endpointslices -l kubernetes.io/service-name=argocd-server
curl -s -o /dev/null -w '%{http_code}\n' http://<EXTERNAL-IP>
outputcaptured 2026-08-26
$ kubectl -n argocd get svc argocd-server -o wide      # note the EXTERNAL-IP
NAME            TYPE           CLUSTER-IP     EXTERNAL-IP   PORT(S)                      AGE   SELECTOR
argocd-server   LoadBalancer   10.96.20.112   172.18.0.9    80:30153/TCP,443:30860/TCP   25m   app.kubernetes.io/instance=argocd,app.kubernetes.io/name=argocd-server
$ kubectl -n argocd get endpointslices -l kubernetes.io/service-name=argocd-server
NAME                  ADDRESSTYPE   PORTS       ENDPOINTS      AGE
argocd-server-lm5nc   IPv4          8080,8080   10.244.1.181   25m
$ curl -s -o /dev/null -w '%{http_code}\n' http://172.18.0.9
200

The EXTERNAL-IP is real only while cloud-provider-kind runs; without it, kubectl -n argocd port-forward svc/argocd-server 8080:80 and curl localhost:8080 instead, and the rest of the exercise is unchanged. Now break it in a way you can explain:

kubectl -n argocd patch svc argocd-server --type=merge -p '{"spec":{"selector":{"app":"nope"}}}'
outputcaptured 2026-08-26
$ kubectl -n argocd patch svc argocd-server --type=merge -p '{"spec":{"selector":{"app":"nope"}}}'
service/argocd-server patched

Watch the endpointslice lose its endpoints and curl start failing while DNS still resolves. Then look at what your patch actually did before undoing it: kubectl -n argocd get svc argocd-server -o jsonpath='{.spec.selector}' shows the original two labels plus app: nope, because a JSON merge patch merges maps rather than replacing them. So the honest restore is removing your key, not re-adding theirs: --type=merge -p '{"spec":{"selector":{"app":null}}}'.

verify: the selector is back to exactly app.kubernetes.io/name: argocd-server and app.kubernetes.io/instance: argocd, and curl returns 200. The merge-patch semantics come up on their own exam tasks.

Run a probe that policy must kill:

cilium hubble port-forward &
kubectl -n team-a run probe --image=curlimages/curl:8.11.1 --restart=Never \
  -- curl -s -m 5 http://example.com
hubble observe --namespace team-a --verdict DROPPED --last 20
outputcaptured 2026-08-26
$ cilium hubble port-forward &
[1] 954968
$ kubectl -n team-a run probe --image=curlimages/curl:8.11.1 --restart=Never \
  -- curl -s -m 5 http://example.com
Warning: would violate PodSecurity "restricted:latest": allowPrivilegeEscalation != false (container "probe" must set securityContext.allowPrivilegeEscalation=false), unrestricted capabilities (container "probe" must set securityContext.capabilities.drop=["ALL"]), runAsNonRoot != true (pod or container "probe" must set securityContext.runAsNonRoot=true), seccompProfile (pod or container "probe" must set securityContext.seccompProfile.type to "RuntimeDefault" or "Localhost")
pod/probe created
$ hubble observe --namespace team-a --verdict DROPPED --last 20
time=2026-08-26T22:13:56.514-04:00 level=WARN msg="Hubble CLI version is lower than Hubble Relay, API compatibility is not guaranteed, updating to a matching or higher version is recommended" hubble-cli-version=1.19.4 hubble-relay-version=1.20.1+g7d68cfb3
Aug 27 02:13:45.229: team-a/probe:39428 (ID:2901) <> 104.20.23.154:80 (world) policy-verdict:none TRAFFIC_DIRECTION_UNKNOWN DENIED (TCP Flags: SYN)
Aug 27 02:13:45.229: team-a/probe:39428 (ID:2901) <> 104.20.23.154:80 (world) Policy denied DROPPED (TCP Flags: SYN)
Aug 27 02:13:46.270: team-a/probe:39428 (ID:2901) <> 104.20.23.154:80 (world) policy-verdict:none TRAFFIC_DIRECTION_UNKNOWN DENIED (TCP Flags: SYN)
Aug 27 02:13:46.270: team-a/probe:39428 (ID:2901) <> 104.20.23.154:80 (world) Policy denied DROPPED (TCP Flags: SYN)
Aug 27 02:13:47.294: team-a/probe:39428 (ID:2901) <> 104.20.23.154:80 (world) policy-verdict:none TRAFFIC_DIRECTION_UNKNOWN DENIED (TCP Flags: SYN)
Aug 27 02:13:47.294: team-a/probe:39428 (ID:2901) <> 104.20.23.154:80 (world) Policy denied DROPPED (TCP Flags: SYN)
Aug 27 02:13:47.629: team-a/probe:48966 (ID:2901) <> 172.66.147.243:80 (world) policy-verdict:none TRAFFIC_DIRECTION_UNKNOWN DENIED (TCP Flags: SYN)
Aug 27 02:13:47.629: team-a/probe:48966 (ID:2901) <> 172.66.147.243:80 (world) Policy denied DROPPED (TCP Flags: SYN)
Aug 27 02:13:48.639: team-a/probe:48966 (ID:2901) <> 172.66.147.243:80 (world) policy-verdict:none TRAFFIC_DIRECTION_UNKNOWN DENIED (TCP Flags: SYN)
Aug 27 02:13:48.639: team-a/probe:48966 (ID:2901) <> 172.66.147.243:80 (world) Policy denied DROPPED (TCP Flags: SYN)
Aug 27 02:13:49.663: team-a/probe:48966 (ID:2901) <> 172.66.147.243:80 (world) policy-verdict:none TRAFFIC_DIRECTION_UNKNOWN DENIED (TCP Flags: SYN)
Aug 27 02:13:49.663: team-a/probe:48966 (ID:2901) <> 172.66.147.243:80 (world) Policy denied DROPPED (TCP Flags: SYN)
verify: DROPPED flows from the probe pod, and the pod's curl exits non-zero. This is the difference between "policy exists" and "policy enforces".

Write a Gateway named web using listener port 80, protocol HTTP, plus an HTTPRoute that matches path /demo and backends a Service demo:80. Don't copy from docs; build it from kubectl explain gateway.spec.listeners and kubectl explain httproute.spec.rules.

verify: both apply cleanly (schema-valid), and kubectl get gateway web -o yaml shows the spec you meant. Status stays unprogrammed here; on the mesh cluster Istio would accept the same manifests.

From a pod in default, resolve short and long names:

kubectl run dnsprobe --image=busybox:1.28 --restart=Never -it --rm -- \
  sh -c 'nslookup argocd-server.argocd && nslookup argocd-server.argocd.svc.cluster.local'
outputcaptured 2026-08-26
$ kubectl run dnsprobe --image=busybox:1.28 --restart=Never -it --rm -- \
  sh -c 'nslookup argocd-server.argocd && nslookup argocd-server.argocd.svc.cluster.local'
Server:    10.96.0.10
Address 1: 10.96.0.10 kube-dns.kube-system.svc.cluster.local

Name:      argocd-server.argocd
Address 1: 10.96.20.112 argocd-server.argocd.svc.cluster.local
Server:    10.96.0.10
Address 1: 10.96.0.10 kube-dns.kube-system.svc.cluster.local

Name:      argocd-server.argocd.svc.cluster.local
Address 1: 10.96.20.112 argocd-server.argocd.svc.cluster.local
All commands and output from this session will be recorded in container logs, including credentials and sensitive information passed through the command prompt.
If you don't see a command prompt, try pressing enter.
pod "dnsprobe" deleted from default namespace

(busybox:1.28 on purpose: newer busybox nslookup ignores the search path and returns NXDOMAIN for the short name even though the pod's resolver handles it fine. A famous trap, worth meeting here rather than in an exam.)

Then repeat inside team-a and explain why it still works (the tenant policy allows port 53 to kube-dns explicitly; find that rule in examples/multitenancy/team-a.yaml).

verify: both names resolve from both namespaces, and you can point at the exact policy rule that permits it.

There is no browser in the exam, and the Gateway API field names are the kind you half-remember. The CRDs carry their own documentation, so kubectl explain is the reference, and it is correct for the channel this cluster actually installed.

kubectl explain gateway.spec.listeners.allowedRoutes.namespaces.from
kubectl explain gateway.spec.listeners.allowedRoutes --recursive | head -12
kubectl explain httproute.spec.rules.backendRefs
outputcaptured 2026-09-13
$ kubectl explain gateway.spec.listeners.allowedRoutes.namespaces.from
GROUP:      gateway.networking.k8s.io
KIND:       Gateway
VERSION:    v1

FIELD: from <string>
ENUM:
    All
    Selector
    Same

DESCRIPTION:
    From indicates where Routes will be selected for this Gateway. Possible
    values are:
    
    * All: Routes in all namespaces may be used by this Gateway.
    * Selector: Routes in namespaces selected by the selector may be used by
      this Gateway.
    * Same: Only Routes in the same namespace may be used by this Gateway.
    
    Support: Core
    
$ kubectl explain gateway.spec.listeners.allowedRoutes --recursive | head -12
GROUP:      gateway.networking.k8s.io
KIND:       Gateway
VERSION:    v1

FIELD: allowedRoutes <Object>


DESCRIPTION:
    AllowedRoutes defines the types of routes that MAY be attached to a
    Listener and the trusted namespaces where those Route resources MAY be
    present.
    
$ kubectl explain httproute.spec.rules.backendRefs
GROUP:      gateway.networking.k8s.io
KIND:       HTTPRoute
VERSION:    v1

FIELD: backendRefs <[]Object>


DESCRIPTION:
    BackendRefs defines the backend(s) where matching requests should be
    sent.
    
    Failure behavior here depends on how many BackendRefs are specified and
    how many are invalid.
    
    If *all* entries in BackendRefs are invalid, and there are also no filters
    specified in this route rule, *all* traffic which matches this rule MUST
    receive a 500 status code.
    
    See the HTTPBackendRef definition for the rules about what makes a single
    HTTPBackendRef invalid.
    
    When a HTTPBackendRef is invalid, 500 status codes MUST be returned for
    requests that would have otherwise been routed to an invalid backend. If
    multiple backends are specified, and some are invalid, the proportion of
    requests that would otherwise have been routed to an invalid backend
    MUST receive a 500 status code.
    
    For example, if two backends are specified with equal weights, and one is
    invalid, 50 percent of traffic must receive a 500. Implementations may
    choose how that 50 percent is determined.
    
    When a HTTPBackendRef refers to a Service that has no ready endpoints,
    implementations SHOULD return a 503 for requests to that backend instead.
    If an implementation chooses to do this, all of the above rules for 500
    responses
    MUST also apply for responses that return a 503.
    
    Support: Core for Kubernetes Service
    
    Support: Extended for Kubernetes ServiceImport
... 76 more lines
verify: allowedRoutes.namespaces.from documents All, Selector and Same, and backendRefs lists weight next to name, namespace and port. A field you expected and cannot find is a field the standard channel does not ship.

A route in one namespace may not send traffic to a Service in another until the owner of that Service says so. The permission travels the other way from the reference, which is the part people get wrong under time pressure.

Run this on kind-mesh: the main cluster has the CRDs but no controller, so the route's status stays empty and there is nothing to read.

kubectl --context kind-mesh create ns team-a --dry-run=client -o yaml | kubectl --context kind-mesh apply -f -
kubectl --context kind-mesh create ns team-b --dry-run=client -o yaml | kubectl --context kind-mesh apply -f -
kubectl --context kind-mesh -n team-b create deployment web --image=nginx:1.27-alpine
kubectl --context kind-mesh -n team-b expose deployment web --port=80
kubectl --context kind-mesh apply -f - <<'EOF'
apiVersion: gateway.networking.k8s.io/v1
kind: Gateway
metadata: { name: web, namespace: team-a }
spec:
  gatewayClassName: istio
  listeners:
    - name: http
      port: 80
      protocol: HTTP
      allowedRoutes: { namespaces: { from: Same } }
---
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata: { name: demo, namespace: team-a }
spec:
  parentRefs: [{ name: web }]
  rules:
    - backendRefs: [{ name: web, namespace: team-b, port: 80 }]
EOF
sleep 10
kubectl --context kind-mesh -n team-a get httproute demo -o jsonpath='{.status.parents[*].conditions}' | jq '.[] | {type, status, reason}'
kubectl --context kind-mesh apply -f - <<'EOF'
apiVersion: gateway.networking.k8s.io/v1beta1
kind: ReferenceGrant
metadata: { name: from-team-a, namespace: team-b }
spec:
  from: [{ group: gateway.networking.k8s.io, kind: HTTPRoute, namespace: team-a }]
  to: [{ group: "", kind: Service }]
EOF
sleep 10
kubectl --context kind-mesh -n team-a get httproute demo -o jsonpath='{.status.parents[*].conditions}' | jq '.[] | {type, status, reason}'
outputcaptured 2026-09-13
$ kubectl --context kind-mesh create ns team-a --dry-run=client -o yaml | kubectl --context kind-mesh apply -f -
namespace/team-a created
$ kubectl --context kind-mesh create ns team-b --dry-run=client -o yaml | kubectl --context kind-mesh apply -f -
namespace/team-b created
$ kubectl --context kind-mesh -n team-b create deployment web --image=nginx:1.27-alpine
deployment.apps/web created
$ kubectl --context kind-mesh -n team-b expose deployment web --port=80
service/web exposed
$ kubectl --context kind-mesh apply -f - <<'EOF'
apiVersion: gateway.networking.k8s.io/v1
kind: Gateway
metadata: { name: web, namespace: team-a }
spec:
  gatewayClassName: istio
  listeners:
    - name: http
      port: 80
      protocol: HTTP
      allowedRoutes: { namespaces: { from: Same } }
---
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata: { name: demo, namespace: team-a }
spec:
  parentRefs: [{ name: web }]
  rules:
    - backendRefs: [{ name: web, namespace: team-b, port: 80 }]
EOF
gateway.gateway.networking.k8s.io/web created
httproute.gateway.networking.k8s.io/demo created
$ sleep 10
$ kubectl --context kind-mesh -n team-a get httproute demo -o jsonpath='{.status.parents[*].conditions}' | jq '.[] | {type, status, reason}'
{
  "type": "Accepted",
  "status": "True",
  "reason": "Accepted"
}
{
  "type": "ResolvedRefs",
  "status": "False",
  "reason": "RefNotPermitted"
}
{
  "type": "ResolvedWaypoints",
  "status": "True",
  "reason": "ResolvedWaypoints"
}
$ kubectl --context kind-mesh apply -f - <<'EOF'
apiVersion: gateway.networking.k8s.io/v1beta1
kind: ReferenceGrant
metadata: { name: from-team-a, namespace: team-b }
spec:
  from: [{ group: gateway.networking.k8s.io, kind: HTTPRoute, namespace: team-a }]
  to: [{ group: "", kind: Service }]
EOF
referencegrant.gateway.networking.k8s.io/from-team-a created
$ sleep 10
$ kubectl --context kind-mesh -n team-a get httproute demo -o jsonpath='{.status.parents[*].conditions}' | jq '.[] | {type, status, reason}'
{
  "type": "Accepted",
  "status": "True",
  "reason": "Accepted"
}
{
  "type": "ResolvedRefs",
  "status": "True",
  "reason": "ResolvedRefs"
}
{
  "type": "ResolvedWaypoints",
  "status": "True",
  "reason": "ResolvedWaypoints"
}
verify: ResolvedRefs is False with reason RefNotPermitted before the grant and True after it, with nothing about the route itself changed. The grant lives with the backend, not with the route.

Every cluster you inherit has a Corefile, and half the DNS incidents you will be handed are a plugin block someone added to it. Read the stock one, add a stub domain, and watch the reload land.

kubectl -n kube-system get cm coredns -o jsonpath='{.data.Corefile}' | tee /tmp/Corefile.orig
printf '%s\nconsul.local:53 {\n    errors\n    cache 30\n    forward . 10.150.0.1\n}\n' "$(cat /tmp/Corefile.orig)" > /tmp/Corefile.new
kubectl -n kube-system create configmap coredns --from-file=Corefile=/tmp/Corefile.new --dry-run=client -o yaml | kubectl -n kube-system apply -f -
sleep 120
kubectl -n kube-system logs deploy/coredns --tail=10
kubectl -n kube-system create configmap coredns --from-file=Corefile=/tmp/Corefile.orig --dry-run=client -o yaml | kubectl -n kube-system apply -f -
outputcaptured 2026-09-13
$ kubectl -n kube-system get cm coredns -o jsonpath='{.data.Corefile}' | tee /tmp/Corefile.orig
.:53 {
    errors
    health {
       lameduck 5s
    }
    ready
    hosts {
        172.18.0.6 gitea.lab
        172.18.0.2 kind-registry
        fallthrough
    }
    kubernetes cluster.local in-addr.arpa ip6.arpa {
       pods insecure
       fallthrough in-addr.arpa ip6.arpa
       ttl 30
    }
    prometheus :9153
    forward . /etc/resolv.conf {
       max_concurrent 1000
    }
    cache 30 {
       disable success cluster.local
       disable denial cluster.local
    }
    loop
    reload
    loadbalance
}
$ printf '%s\nconsul.local:53 {\n    errors\n    cache 30\n    forward . 10.150.0.1\n}\n' "$(cat /tmp/Corefile.orig)" > /tmp/Corefile.new
$ kubectl -n kube-system create configmap coredns --from-file=Corefile=/tmp/Corefile.new --dry-run=client -o yaml | kubectl -n kube-system apply -f -
configmap/coredns configured
$ sleep 120
$ kubectl -n kube-system logs deploy/coredns --tail=10
Found 2 pods, using pod/coredns-754f9cbf9f-9kjvm
[INFO] plugin/reload: Running configuration SHA512 = 21ac5a0b49ecc87479896d161ce04e13d5464755235ff622730e3dfae4608be276f4a5f8b10043b699e176d468200893ab00c7ce401fe4e331b9c0e19f77eb6d
CoreDNS-1.14.2
linux/amd64, go1.26.1, dd1df4f
[ERROR] plugin/errors: 2 _grpclb._tcp.argocd-repo-server. SRV: read udp 10.244.2.138:45773->172.18.0.1:53: i/o timeout
[ERROR] plugin/errors: 2 _grpclb._tcp.spire-server.spire. SRV: read udp 10.244.2.138:41440->172.18.0.1:53: i/o timeout
[ERROR] plugin/errors: 2 _grpclb._tcp.spire-server.spire. SRV: read udp 10.244.2.138:45718->172.18.0.1:53: i/o timeout
[ERROR] plugin/errors: 2 _grpclb._tcp.spire-server.spire. SRV: read udp 10.244.2.138:60045->172.18.0.1:53: i/o timeout
[INFO] Reloading
[INFO] plugin/reload: Running configuration SHA512 = fa9459fcfd404334de10bf2f59ce4f79c08316d97dff8ba20f46ce959b71a49c67999e92a94381f4cab42431a2af64dcc71ec6d1681cd790c23927cc3c07ac13
[INFO] Reloading complete
$ kubectl -n kube-system create configmap coredns --from-file=Corefile=/tmp/Corefile.orig --dry-run=client -o yaml | kubectl -n kube-system apply -f -
configmap/coredns configured
verify: the log shows a reload line naming the new plugin set, and resolution of everything else is unaffected. Break the block on purpose (drop a brace) and the same log line becomes a parse error with the offending line number; that error is the one you will be shown in a task.

A pod on the host network gets the node's resolver unless you say otherwise, so cluster names stop resolving and nothing in the pod spec looks wrong. It is a two-word fix and a favorite of exam authors.

kubectl run hn --image=busybox:1.28 --restart=Never --overrides='{"spec":{"hostNetwork":true}}' -- nslookup kubernetes.default
sleep 15
kubectl logs hn
kubectl run hn2 --image=busybox:1.28 --restart=Never --overrides='{"spec":{"hostNetwork":true,"dnsPolicy":"ClusterFirstWithHostNet"}}' -- nslookup kubernetes.default
sleep 15
kubectl logs hn2
kubectl delete pod hn hn2
outputcaptured 2026-09-12
$ kubectl run hn --image=busybox:1.28 --restart=Never --overrides='{"spec":{"hostNetwork":true}}' -- nslookup kubernetes.default
pod/hn created
$ sleep 15
$ kubectl logs hn
nslookup: can't resolve 'kubernetes.default'
Server:    172.18.0.1
Address 1: 172.18.0.1 omarchy
$ kubectl run hn2 --image=busybox:1.28 --restart=Never --overrides='{"spec":{"hostNetwork":true,"dnsPolicy":"ClusterFirstWithHostNet"}}' -- nslookup kubernetes.default
pod/hn2 created
$ sleep 15
$ kubectl logs hn2
Server:    10.96.0.10
Address 1: 10.96.0.10 kube-dns.kube-system.svc.cluster.local

Name:      kubernetes.default
Address 1: 10.96.0.1 kubernetes.default.svc.cluster.local
$ kubectl delete pod hn hn2
pod "hn" deleted from default namespace
pod "hn2" deleted from default namespace
verify: the first pod cannot resolve kubernetes.default and the second can, with dnsPolicy: ClusterFirstWithHostNet the only difference between them.

Cilium may or may not be doing kube-proxy's job, and the answer changes what you look at when a Service stops working. Ask the agent rather than guessing from the pod list.

kubectl -n kube-system exec ds/cilium -c cilium-agent -- cilium-dbg status | grep -i -A2 'KubeProxyReplacement'
kubectl -n kube-system get cm cilium-config -o jsonpath='{.data.kube-proxy-replacement}{"\n"}'
kubectl -n kube-system get ds -l k8s-app=kube-proxy
outputcaptured 2026-09-13
$ kubectl -n kube-system exec ds/cilium -c cilium-agent -- cilium-dbg status | grep -i -A2 'KubeProxyReplacement'
KubeProxyReplacement:    False   
Host firewall:           Disabled
SRv6:                    Disabled
$ kubectl -n kube-system get cm cilium-config -o jsonpath='{.data.kube-proxy-replacement}{"\n"}'
false
$ kubectl -n kube-system get ds -l k8s-app=kube-proxy
NAME         DESIRED   CURRENT   READY   UP-TO-DATE   AVAILABLE   NODE SELECTOR            AGE
kube-proxy   3         3         3       3            3           kubernetes.io/os=linux   16h
verify: the agent's own report and the ConfigMap key agree, and you can say whether iptables or eBPF is programming Service VIPs on this cluster.

trafficDistribution is a hint, not a rule: the field is accepted anywhere, but the EndpointSlice controller only writes zone hints when the endpoints actually carry zones. On kind the nodes have no zone label until you add one, which makes the mechanism easy to see.

kubectl create deployment demo --image=nginx:1.27-alpine --replicas=4
kubectl expose deployment demo --port=80
kubectl patch svc demo -p '{"spec":{"trafficDistribution":"PreferSameZone"}}'
kubectl get svc demo -o jsonpath='{.spec.trafficDistribution}{"\n"}'
kubectl get endpointslices -l kubernetes.io/service-name=demo -o yaml | grep -B2 -A3 hints
kubectl label node cnpe-worker topology.kubernetes.io/zone=a --overwrite
kubectl label node cnpe-worker2 topology.kubernetes.io/zone=b --overwrite
kubectl rollout restart deployment demo
kubectl rollout status deployment demo
kubectl get endpointslices -l kubernetes.io/service-name=demo -o yaml | grep -B2 -A3 hints
# the scheduling exercises on this page assume unzoned nodes
kubectl delete svc demo --ignore-not-found
kubectl delete deployment demo --ignore-not-found
kubectl label node cnpe-worker topology.kubernetes.io/zone-
kubectl label node cnpe-worker2 topology.kubernetes.io/zone-
outputcaptured 2026-09-12
$ kubectl create deployment demo --image=nginx:1.27-alpine --replicas=4
deployment.apps/demo created
$ kubectl expose deployment demo --port=80
service/demo exposed
$ kubectl patch svc demo -p '{"spec":{"trafficDistribution":"PreferSameZone"}}'
service/demo patched
$ kubectl get svc demo -o jsonpath='{.spec.trafficDistribution}{"\n"}'
PreferSameZone
$ kubectl get endpointslices -l kubernetes.io/service-name=demo -o yaml | grep -B2 -A3 hints
$ kubectl label node cnpe-worker topology.kubernetes.io/zone=a --overwrite
node/cnpe-worker labeled
$ kubectl label node cnpe-worker2 topology.kubernetes.io/zone=b --overwrite
node/cnpe-worker2 labeled
$ kubectl rollout restart deployment demo
deployment.apps/demo restarted
$ kubectl rollout status deployment demo
Waiting for deployment spec update to be observed...
Waiting for deployment "demo" rollout to finish: 1 out of 4 new replicas have been updated...
Waiting for deployment "demo" rollout to finish: 1 out of 4 new replicas have been updated...
Waiting for deployment "demo" rollout to finish: 2 out of 4 new replicas have been updated...
Waiting for deployment "demo" rollout to finish: 2 out of 4 new replicas have been updated...
Waiting for deployment "demo" rollout to finish: 2 out of 4 new replicas have been updated...
Waiting for deployment "demo" rollout to finish: 2 out of 4 new replicas have been updated...
Waiting for deployment "demo" rollout to finish: 2 out of 4 new replicas have been updated...
Waiting for deployment "demo" rollout to finish: 2 out of 4 new replicas have been updated...
Waiting for deployment "demo" rollout to finish: 2 out of 4 new replicas have been updated...
Waiting for deployment "demo" rollout to finish: 2 out of 4 new replicas have been updated...
Waiting for deployment "demo" rollout to finish: 3 out of 4 new replicas have been updated...
Waiting for deployment "demo" rollout to finish: 3 out of 4 new replicas have been updated...
Waiting for deployment "demo" rollout to finish: 1 old replicas are pending termination...
Waiting for deployment "demo" rollout to finish: 1 old replicas are pending termination...
Waiting for deployment "demo" rollout to finish: 1 old replicas are pending termination...
Waiting for deployment "demo" rollout to finish: 1 old replicas are pending termination...
Waiting for deployment "demo" rollout to finish: 3 of 4 updated replicas are available...
deployment "demo" successfully rolled out
$ kubectl get endpointslices -l kubernetes.io/service-name=demo -o yaml | grep -B2 -A3 hints
      serving: true
      terminating: false
    hints:
      forZones:
      - name: a
    nodeName: cnpe-worker
--
      serving: true
      terminating: false
    hints:
      forZones:
      - name: b
    nodeName: cnpe-worker2
--
      serving: true
      terminating: false
    hints:
      forZones:
      - name: a
    nodeName: cnpe-worker
--
      serving: true
      terminating: false
    hints:
      forZones:
      - name: b
    nodeName: cnpe-worker2
$ # the scheduling exercises on this page assume unzoned nodes
$ kubectl delete svc demo --ignore-not-found
service "demo" deleted from default namespace
$ kubectl delete deployment demo --ignore-not-found
deployment.apps "demo" deleted from default namespace
$ kubectl label node cnpe-worker topology.kubernetes.io/zone-
node/cnpe-worker unlabeled
$ kubectl label node cnpe-worker2 topology.kubernetes.io/zone-
node/cnpe-worker2 unlabeled
verify: hints appear only once the nodes carry topology.kubernetes.io/zone, and each endpoint's hint names its own zone. Clean up the labels afterwards so the scheduling exercises still behave.

Self-check

answer out loud before opening
A Service resolves in DNS but every connection is refused. Name the three things you check, in order.

Selector against actual pod labels; pod readiness (an unready pod is not in the slice); the EndpointSlice itself. If the slice has addresses with ready: false, it is a probe problem, not a Service problem. If the slice is empty, it is a selector or a scheduling problem.

Why does a NetworkPolicy ipBlock naming a Service's ClusterIP never work?

Because DNAT happens before policy evaluation: by the time the packet is checked, its destination is a pod IP. Policy is expressed in pod selectors and identities, not virtual IPs. For the API server specifically, Cilium's toEntities: [kube-apiserver] is the expressible form.

Your backend needs the real client IP. What do you change, and what breaks?

externalTrafficPolicy: Local on the Service. The trade: nodes with no backing pod stop serving that traffic, so you now depend on the load balancer's health checks to avoid them, and you lose the even spread that Cluster gave you.

What is the difference between a headless Service and a ClusterIP Service with one endpoint?

Headless has no VIP and no datapath translation: DNS returns pod IPs directly and the client chooses. A one-endpoint ClusterIP still goes through the VIP and its rules, so the client never learns the pod IP. StatefulSets need the former for stable per-pod names (pod-0.svc…).

Team-a can curl team-b's Service after you added an egress allow, but it still fails. What did you forget?

Team-b's ingress side. Both ends must open; policies are additive allow-lists evaluated independently at each pod. Also check DNS: if egress is now restricted in team-a, resolving web.team-b.svc needs its own port-53 allow.

An HTTPRoute applies cleanly but nothing routes. Where do you look first?

status.parents[].conditions on the route: Accepted false usually means attachment was refused by the Gateway's allowedRoutes; ResolvedRefs false means a backend Service or a TLS secret is missing (or needs a ReferenceGrant across namespaces). On the Gateway, Programmed false means no data plane, which is the permanent state on this lab's main cluster.

An HTTPRoute in team-a references a Service in team-b. Which object, in which namespace, and what does status say until it exists?

A ReferenceGrant in team-b (the namespace being referenced) with from group gateway.networking.k8s.io kind HTTPRoute namespace team-a and to kind Service. Until it exists the route reports ResolvedRefs: False with reason RefNotPermitted, and the implementation answers 500 for that backend.

Two backendRefs carry weight 3 and weight 1. What split is that, and what happens if one backend Service is missing?

75/25: each weight is a share of the total, so 3 and 1 make four shares. A missing backend produces ResolvedRefs: False, BackendNotFound; the implementation must still route the valid share and return 500 for the share that points at the invalid backend, so users see intermittent errors proportional to the weight.

A stub domain for corp.example must resolve via 10.0.0.53. Where does the change go and how long before it applies?

A second server block in the coredns ConfigMap in kube-system: corp.example:53 { errors; cache 30; forward . 10.0.0.53 }. The reload plugin applies it within about two minutes without restarting pods. Verify from a pod with nslookup, and check kubectl -n kube-system logs deploy/coredns for a Corefile parse error if nothing changes.

Which kube-proxy mode is GA and newest, and what changes for NodePort when you switch to it?

nftables, GA since Kubernetes 1.33 (kernel 5.13+). NodePorts default to --nodeport-addresses primary, so they answer on the node's primary IP only; other interfaces and localhost stop answering, and kube-proxy no longer punches firewall holes for the NodePort range. IPVS is deprecated since 1.35; iptables remains the default.

Docs to know your way around

study time, not exam time
  • kubernetes.io: Services, DNS for Services and Pods, Network Policies. The DNS page's search-domain section is the one people never read.
  • gateway-api.sigs.k8s.io: the API model page with the role diagram; the "route attachment" section.
  • docs.cilium.io: Hubble observe reference; CiliumNetworkPolicy entities and FQDN rules.
  • Offline, in the exam: kubectl explain service.spec, kubectl explain networkpolicy.spec.egress, kubectl explain httproute.spec.rules --recursive. Faster than docs and always version-correct for the cluster in front of you.
  • gateway-api.sigs.k8s.io/reference/spec: the field reference for Gateway listeners, HTTPRoute matches/filters/backendRefs, ReferenceGrant, BackendTLSPolicy and ListenerSet; the condition reasons (NotAllowedByListeners, RefNotPermitted, BackendNotFound) are listed there verbatim.
  • kubernetes.io/docs/reference/networking/virtual-ips: kube-proxy modes, traffic policies, terminating endpoints and trafficDistribution on one page.
  • kubernetes.io/docs/tasks/administer-cluster/dns-custom-nameservers: the stock Corefile with every plugin explained, and the stub-domain example you will copy under time pressure.
free before 1.2make down-gitopsmake down-sec