Being able to say, for any packet, what touches it and in what order: pod, CNI, Service, kube-proxy or its eBPF replacement, DNS, policy, ingress. Most delivery and incident tasks in the other four domains eventually collapse into this one.
make upmake gitea gitopsmake sec
Orientation
Networking is the substrate every other domain stands on. A Crossplane XR that never goes Ready, an Argo CD app stuck Progressing, a Prometheus target that will not come up, a canary that never receives traffic: a surprising share of those are one Service selector, one missing DNS egress rule, or one unready endpoint.
Not "explain CNI". It asks you to make traffic work or make traffic stop, then prove it: expose a workload, restrict a namespace, author a Gateway and an HTTPRoute, or explain why a Service resolves and then refuses connections. Every one of those is a five-minute task if the model in your head is exact, and twenty minutes of guessing if it is fuzzy.
The four rules of the Kubernetes network model
- Every pod gets its own IP, cluster-routable, no NAT between pods.
- Pods on a node can reach all pods on all nodes without NAT.
- Agents on a node (kubelet, system daemons) can reach all pods on that node.
- A pod sees its own IP as the same address other pods use to reach it.
The CNI plugin is whatever implements those rules: routing, overlay, or eBPF datapath. Everything above (Services, DNS, policy, Gateways) is built on the assumption that the four rules already hold. When they do not, nothing above them behaves sanely, which is why "is it a CNI problem or a Service problem" is the first fork in any network diagnosis.
Services and the thing that actually makes them work
ClusterIP is a virtual IP that exists only as translation rules on each node; nothing listens on it, no interface owns it, you cannot ping it in any meaningful sense. NodePort opens the same Service on a high port (30000–32767 by default) of every node. LoadBalancer is NodePort plus something external handing out a real IP (in this lab, cloud-provider-kind). Headless (clusterIP: None) skips the VIP entirely and returns pod IPs straight from DNS, which is what StatefulSets need for stable peer discovery. ExternalName is a CNAME with no proxying at all.
| Type | spec bits | Reaches it from | Use it when |
|---|---|---|---|
| ClusterIP | default | inside the cluster | the 90% case; every internal call |
| NodePort | type: NodePort | any node IP | bootstrapping, bare metal, demos |
| LoadBalancer | type: LoadBalancer | outside | real external entry point (loadBalancerClass only picks between implementations) |
| Headless | clusterIP: None | DNS returns pod IPs | StatefulSet peers, client-side LB |
| ExternalName | externalName: host | DNS CNAME only | aliasing an out-of-cluster host |
The thing that makes a Service work is not the Service. It is the EndpointSlice behind it, and the selector that fills it. A Service with no ready endpoints resolves fine and then refuses connections, which is one of the failures the exam leans on most. Check order, every time: selector matches pod labels → pods are Ready (readiness gates endpoint membership) → kubectl get endpointslices -l kubernetes.io/service-name=<svc>.
Two related fields decide whether traffic reaches a pod at all. readinessProbe failure pulls the pod out of the slice, which is the intended way to drain a pod. publishNotReadyAddresses: true overrides that and is how headless Services for clustered databases let peers find each other before they are serving. And terminationGracePeriodSeconds plus a preStop sleep is the standard trick for the race where a pod is deleted but nodes have not yet removed its rules; connections refused during rollouts almost always trace here.
Traffic policies
externalTrafficPolicy: Cluster(default) SNATs and may hop to another node, so the backend sees the node's IP, not the client's.Localpreserves the client source IP and only routes to pods on the receiving node, with the trade that a node holding no pod blackholes the traffic (health checks are what stop the external LB from sending there).internalTrafficPolicy: Localis the same idea for in-cluster traffic: node-local endpoints only. Used for node-local caches and log shippers.sessionAffinity: ClientIPis the only affinity a Service offers. Anything richer belongs in a mesh or gateway.
Five conditions, one request. Turn any of them off and see which hop drops it, what the user reports, and the command that proves it:
kubectl get endpointslices -l kubernetes.io/service-name=X beats describe svc, because it shows you conditions per address (ready, serving, terminating) rather than a summarized list. When a rollout half-breaks, those three booleans tell the whole story.
The datapath: kube-proxy, eBPF, and this cluster's choice
Classic kube-proxy watches Services and EndpointSlices and programs the node: iptables mode writes DNAT chains (simple, and linear-ish to rule count), IPVS mode uses kernel load balancing with real scheduling algorithms. eBPF datapaths (Cilium, Calico) replace kube-proxy entirely, doing the translation in the socket or TC layer, which removes the rule-table scaling problem and unlocks flow visibility.
client pod │ connect 10.96.0.42:80 ← ClusterIP, exists only as a rule ▼ [ datapath ] kube-proxy iptables/IPVS or eBPF (Cilium, this lab) │ DNAT → picks one endpoint from the EndpointSlice ▼ backend pod IP 10.244.2.17:8080 │ ▼ policy is evaluated on thepod identity, not the Service VIP
Note the last line. This is why an ipBlock rule naming a ClusterIP never matches: by the time policy is evaluated, the destination has already been rewritten to a pod IP. That single fact explains a whole family of "my NetworkPolicy does nothing" tasks; section 3.3 sets the same trap with CloudNativePG.
This cluster runs Cilium instead of kindnet plus kube-proxy for one decisive reason: kindnet does not enforce NetworkPolicy at all. A lab where policies apply cleanly and change nothing teaches you the wrong lesson permanently. Cilium also brings identity-based policy (labels are compiled into numeric identities, so policy survives pod IP churn) and Hubble, which shows flows with their verdict: forwarded or dropped, and by which policy.
CiliumNetworkPolicy adds what upstream NetworkPolicy cannot express: toFQDNs (allow api.github.com by name, enforced at DNS), toEntities (kube-apiserver, world, host, remote-node), and L7 rules for HTTP/DNS/Kafka. The exam will not test CRD field names, but "which policy engine can express egress to the API server" is a fair scenario, and the answer here is toEntities: [kube-apiserver].
DNS: the layer that fails quietly
Names are <svc>.<ns>.svc.cluster.local, served by CoreDNS in kube-system, backed by a Service called kube-dns (the name is left over from kube-dns; the pods are CoreDNS). Pods get a /etc/resolv.conf with a search list and ndots:5, which means any name with fewer than five dots is tried against every search domain first. backend.default.svc.cluster.local (four dots, still under five) becomes four or five queries before the right one lands; a trailing dot (backend.default.svc.cluster.local.) skips the search list entirely, and that is the cheap fix for DNS-heavy workloads.
| Record | Resolves to | Notes |
|---|---|---|
| svc.ns.svc.cluster.local | ClusterIP | the normal A record |
| svc.ns | ClusterIP | works via search domains |
| headless.ns.svc… | all ready pod IPs | multiple A records, client picks |
| pod-0.headless.ns.svc… | one pod | StatefulSet stable identity |
| _port._tcp.svc.ns.svc… | SRV record | named ports; how peers discover ports |
spec.dnsPolicy and dnsConfig on a pod let you override all of this: ClusterFirst (default), None plus explicit nameservers, Default (inherit the node's). It comes up when a workload must resolve an external private zone.
team-a allows DNS egress through one explicit NetworkPolicy rule (UDP/TCP 53 to kube-dns). The netpol break drill removes exactly that rule. Result: every pod stays Running, every probe stays green, and nothing can resolve anything. No pod listing will ever show it. You find it by making something try (nslookup from inside) and then reading Hubble's DROPPED verdicts on port 53. The same break returns in section 4.6.
NetworkPolicy semantics
- Policies are namespaced and select pods, never Services.
- A pod is "isolated" for a direction the moment any policy selects it for that direction. Until then, everything is allowed.
- Rules are additive allow-lists. There is no deny rule. You loosen by adding and tighten only by removing. The break drill exploits that asymmetry: a missing rule looks exactly like a rule that was never there.
- Ingress and egress are independent. A cross-namespace call needs an egress allow on the caller's side and an ingress allow on the callee's. Discovering that empirically once (exercise 3 below) saves you re-deriving it under pressure.
- Default-deny is itself a policy: empty
podSelector,policyTypes: [Ingress, Egress], no rules. - Selector scoping inside a rule matters:
namespaceSelectorandpodSelectorin the same list item is an AND (those pods in those namespaces); as two separate items it is an OR. This is the single most common authoring bug.
A task that says "team-a must reach only team-b's web service and DNS" is asking for three objects, not one: default-deny in team-a, an egress allow in team-a, an ingress allow in team-b. Grade yourself the way a grader would: a curl that returns 200 and a second curl to something else that times out.
Ingress vs Gateway API
Ingress was one object owned by nobody in particular, extended by controller-specific annotations. Gateway API replaces it with a role-split model, and the roles are the point:
| Kind | Owned by | Says |
|---|---|---|
| GatewayClass | infrastructure provider | which controller implements Gateways of this class |
| Gateway | platform team | listeners: ports, protocols, TLS, and which routes may attach |
| HTTPRoute / GRPCRoute / TCPRoute | application team | match rules, filters, weighted backends |
| ReferenceGrant | the namespace being referenced | consent for a cross-namespace backend reference |
Route attachment is two-sided, in the same spirit as NetworkPolicy: the route names a parentRef, and the Gateway's listener declares allowedRoutes (same namespace, selected namespaces, or all). Both must agree, and a route whose attachment is refused reports it in status rather than failing to apply.
Status conditions are how you grade your own work: Accepted (the controller understood it), Programmed (the data plane is configured), ResolvedRefs (backends and secrets exist and are permitted). Reading those three beats guessing every time.
This cluster installs the standard-channel Gateway API CRDs so you can author and schema-validate all of it, but no controller programs a data plane for them on the main cluster. Your Gateway will sit unprogrammed; do not mistake a clean apply for a working Gateway. On the mesh cluster (make mesh), Istio serves Gateway API for real: the same manifests, an actual Programmed condition.
One neighbor the Gateway needs and Kubernetes does not provide: certificates. A TLS listener references a Secret, and something has to keep that Secret valid. In practice that something is cert-manager: an Issuer/ClusterIssuer (ACME, CA, or Vault) plus a Certificate, or the cert-manager.io/cluster-issuer annotation on the Gateway, which issues into the Secret and renews before expiry. Not installed in this lab, but "who renews that certificate" is a fair question about any Gateway you design, and ResolvedRefs: False on a listener is usually its absence.
Progressive delivery ties in here (section 2.5): weighted backendRefs on an HTTPRoute are how a mesh shifts traffic by route weight instead of by replica count, and Flagger drives exactly those weights.
Gateway API, field by field
The role table above is the model. A task hands you the fields. Gateway API v1.5 (April 2026) is the current release; everything below is Standard channel unless marked. On the exam cluster, kubectl explain gateway.spec.listeners --recursive and kubectl explain httproute.spec.rules --recursive print the same fields, version-correct.
Gateway listeners
name,port,protocol(HTTP, HTTPS, TLS, TCP, UDP) and an optionalhostname. Two listeners with the same port, protocol and hostname conflict; the loser carriesConflicted: Truewith reasonHostnameConflictorProtocolConflict.tls.mode:Terminate(the Gateway decrypts; needstls.certificateRefspointing at akubernetes.io/tlsSecret) orPassthrough(SNI routing only, used with TLSRoute). A Secret in another namespace needs a ReferenceGrant in the Secret's namespace, otherwiseResolvedRefs: FalsewithRefNotPermitted.allowedRoutes.namespaces.from:Same(default),All, orSelectorwith a label selector;allowedRoutes.kindsrestricts route kinds. This is the Gateway's half of attachment consent.- Gateway conditions:
Accepted(reasons includeListenersNotValid,Pending),Programmed(AddressNotAssigned,NoResources), per-listenerAccepted,ResolvedRefs,Programmed. A Gateway whose class has no controller sits atAccepted: Unknownforever.
HTTPRoute rules
parentRefsname the Gateway; addsectionName(listener name) orportto bind to one listener rather than all of them.hostnameson the route must intersect the listener hostname or you getAccepted: False,NoMatchingListenerHostname.matches: a list where each item is ANDed inside (path AND headers AND queryParams AND method) and items are ORed against each other. Path types:PathPrefix(default),Exact,RegularExpression(implementation-specific). No matches at all meansPathPrefix /, which matches everything.filters:RequestHeaderModifier,ResponseHeaderModifier,RequestRedirect,URLRewrite(these two cannot be combined),RequestMirror,CORS(Standard since v1.5) andExtensionReffor vendor CRDs. Filters replace almost every Ingress annotation you used to write.backendRefs:name,port, optionalnamespace(needs a ReferenceGrant) andweight(default 1). Weights are shares of the total: 90 and 10 split 90/10, 3 and 1 split 75/25. A rule with no backendRefs and no redirect answers 500. A missing Service givesResolvedRefs: False,BackendNotFound.timeouts(Standard since v1.2):requestfor the whole exchange,backendRequestfor one attempt;backendRequestmay not exceedrequest;0sdisables.
apiVersion: gateway.networking.k8s.io/v1
kind: Gateway
metadata: { name: web, namespace: infra }
spec:
gatewayClassName: istio
listeners:
- name: https
port: 443
protocol: HTTPS
hostname: "*.example.com"
tls: { mode: Terminate, certificateRefs: [{ kind: Secret, name: wildcard-tls }] }
allowedRoutes: { namespaces: { from: Selector, selector: { matchLabels: { tenant: "true" } } } }
---
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata: { name: demo, namespace: team-a }
spec:
parentRefs: [{ name: web, namespace: infra, sectionName: https }]
hostnames: ["demo.example.com"]
rules:
- matches: [{ path: { type: PathPrefix, value: /api } }]
filters: [{ type: RequestHeaderModifier, requestHeaderModifier: { add: [{ name: X-Env, value: staging }] } }]
backendRefs:
- { name: demo-stable, port: 80, weight: 90 }
- { name: demo-canary, port: 80, weight: 10 }The neighbors that arrived recently
- ReferenceGrant is
v1since v1.5:from(group, kind, namespace of the referrer) andto(group, kind, optionally name) in the namespace being referenced. It is consent, never a route: it grants nothing by itself. - BackendTLSPolicy (Standard since v1.4) configures TLS from the Gateway to the backend pods:
targetRefsa Service,validation.hostnameas SNI, and eithercaCertificateRefsorwellKnownCACertificates: System. Same namespace as the Service, by design. - ListenerSet (Standard since v1.5) lets a team add listeners to a shared Gateway without editing it; the Gateway must opt in with
spec.allowedListeners.namespaces.from(defaultNone). Parent listeners win conflicts. It also lifts the 64-listener ceiling. - TLSRoute is
v1since v1.5; GRPCRoute has been Standard since v1.1.
Ingress to Gateway, mapped
| Ingress | Gateway API |
|---|---|
| IngressClass | GatewayClass (controllerName instead of controller) |
| spec.tls[].secretName + hosts | Gateway listener protocol HTTPS, tls.certificateRefs, hostname |
| spec.rules[].host + paths | HTTPRoute hostnames + rules[].matches[].path |
| pathType Prefix / Exact / ImplementationSpecific | PathPrefix / Exact / RegularExpression |
| backend.service.name/port | backendRefs[].name/port (+ weight) |
| controller annotations (rewrite, redirect, canary weight, CORS) | filters and backendRefs weights, portable across controllers |
| one object, one owner | Gateway owned by platform, HTTPRoute owned by the app team, ReferenceGrant by the referenced namespace |
The community ingress2gateway tool does the mechanical conversion for common controllers; the judgment it cannot do for you is deciding who owns the Gateway and which namespaces its listeners admit.
"Expose service X on host Y through the existing Gateway in namespace infra" is three checks: does the Gateway's listener admit routes from your namespace (allowedRoutes), does your hostname fit the listener hostname, and does status.parents[].conditions read Accepted: True and ResolvedRefs: True. A cross-namespace backend adds a fourth object, the ReferenceGrant, in the backend's namespace, not yours.
Datapath and architecture details a scenario can name
kube-proxy modes and their replacement
| Mode | Status | What to remember |
|---|---|---|
| iptables | default | DNAT chains; rule count grows with services; minSyncPeriod and syncPeriod tune programming latency |
| ipvs | deprecated since v1.35 | kernel load balancing with real algorithms (rr, lc, sh); needs ipvs kernel modules |
| nftables | GA since v1.33 | kernel 5.13+; faster endpoint updates; NodePorts default to --nodeport-addresses primary (the node primary IP only) |
| kernelspace | Windows only | the Windows equivalent |
| Cilium kube-proxy replacement | Helm kubeProxyReplacement=true | ClusterIP, NodePort, LoadBalancer, hostPort all in eBPF; socket-level LB for east-west (the connect() is rewritten, so there is no per-packet DNAT to trace); DSR and Maglev options; cilium status reports the mode |
Two Service fields that ride on this layer. spec.trafficDistribution (GA since v1.33) takes PreferSameZone or PreferSameNode (PreferClose is the deprecated older alias); it is a preference, and externalTrafficPolicy / internalTrafficPolicy: Local override it. The old Endpoints API is deprecated since v1.33 in favor of EndpointSlice, and spec.externalIPs is deprecated since v1.36; when a task says "expose", the modern answers are LoadBalancer or a Gateway.
CoreDNS: the Corefile you will be asked to read or edit
The ConfigMap is coredns in kube-system; the reload plugin picks up edits within about two minutes, so do not restart pods to make a change land. The stock server block and what each line buys you:
| Line | Effect | Knob a task turns |
|---|---|---|
| kubernetes cluster.local in-addr.arpa ip6.arpa { pods insecure; fallthrough in-addr.arpa ip6.arpa; ttl 30 } | serves Service and Pod records for the cluster domain | pods verified returns pod A records only for pods that exist; ttl 0-3600 s (default 5) |
| forward . /etc/resolv.conf | everything not in the cluster domain goes upstream | forward . 172.16.0.1 pins the upstream; a second server block corp.example:53 { forward . 10.0.0.53 } is a stub domain |
| cache 30 | response cache | raise it for DNS-heavy tenants, or add NodeLocal DNSCache (a DaemonSet cache on each node that also upgrades upstream queries to TCP and avoids conntrack races) |
| loop · reload · loadbalance · errors · health · ready · prometheus :9153 | loop detection, live config reload, A-record shuffling, logging, probes, metrics | mostly leave alone; log is the debugging plugin you add temporarily |
Pod side: dnsPolicy is ClusterFirst unless you say otherwise; a hostNetwork: true pod needs ClusterFirstWithHostNet or it silently falls back to the node's resolver; None plus dnsConfig (at most 3 nameservers, up to 32 search domains, options: [{name: ndots, value: "2"}]) is how you cut the search-list storm without a trailing dot. The autopath plugin does the same server-side but requires pods verified.
Policy details that differ between engines
- Target a namespace by name with the immutable label
kubernetes.io/metadata.namein anamespaceSelector; do not use an empty{}selector, which matches every namespace. - In Cilium,
ipBlockrules do not match pod or node IPs by default (policyCIDRMatchModechanges it); pods are selected by labels and identities. That is a second reason, beyond DNAT, why CIDR rules "do nothing". CiliumClusterwideNetworkPolicyis the cluster-scoped twin ofCiliumNetworkPolicy: same spec, no namespace, for baseline rules such as "everyone may reach DNS and the API server". Cilium 1.20 also implements the upstreamClusterNetworkPolicyAPI with admin and baseline tiers that ordinary NetworkPolicy cannot override.
Architecture decisions a scenario can name
- HA control plane: three control plane nodes minimum. Stacked etcd (etcd on the control plane nodes, kubeadm's default) is simpler; external etcd needs six hosts but losing a node no longer removes both an API server and an etcd member. Either way the API server sits behind a load balancer; Kubernetes itself does not fail over the endpoint.
- Zones: nodes carry
topology.kubernetes.io/zoneandregion; spread control plane replicas over at least three zones, spread workloads with topologySpreadConstraints, and useWaitForFirstConsumerstorage so a volume is not provisioned in a zone the pod cannot reach. - Scale limits worth quoting: 110 pods per node by default, 5000 nodes, 150000 pods per cluster.
- Multi-cluster: a hub Argo CD with cluster Secrets and an ApplicationSet clusters generator, or one Flux per cluster each reconciling
clusters/<name>from a shared repo. Hub-and-spoke centralizes credentials and UI; agent-per-cluster survives hub loss and keeps credentials local. Section 2.2 covers the objects.
A monitoring or CNI-adjacent pod on hostNetwork: true that "cannot resolve services" is almost always dnsPolicy: ClusterFirst behaving as Default. Set ClusterFirstWithHostNet. The symptom looks exactly like a NetworkPolicy drop, and Hubble shows nothing because the query never went to CoreDNS.
Exercises
With gitops up, Argo CD sits behind a LoadBalancer (running server.insecure, so plain http):
kubectl -n argocd get svc argocd-server -o wide # note the EXTERNAL-IP
kubectl -n argocd get endpointslices -l kubernetes.io/service-name=argocd-server
curl -s -o /dev/null -w '%{http_code}\n' http://<EXTERNAL-IP>outputcaptured 2026-08-26
$ kubectl -n argocd get svc argocd-server -o wide # note the EXTERNAL-IP
NAME TYPE CLUSTER-IP EXTERNAL-IP PORT(S) AGE SELECTOR
argocd-server LoadBalancer 10.96.20.112 172.18.0.9 80:30153/TCP,443:30860/TCP 25m app.kubernetes.io/instance=argocd,app.kubernetes.io/name=argocd-server
$ kubectl -n argocd get endpointslices -l kubernetes.io/service-name=argocd-server
NAME ADDRESSTYPE PORTS ENDPOINTS AGE
argocd-server-lm5nc IPv4 8080,8080 10.244.1.181 25m
$ curl -s -o /dev/null -w '%{http_code}\n' http://172.18.0.9
200The EXTERNAL-IP is real only while cloud-provider-kind runs; without it, kubectl -n argocd port-forward svc/argocd-server 8080:80 and curl localhost:8080 instead, and the rest of the exercise is unchanged. Now break it in a way you can explain:
kubectl -n argocd patch svc argocd-server --type=merge -p '{"spec":{"selector":{"app":"nope"}}}'outputcaptured 2026-08-26
$ kubectl -n argocd patch svc argocd-server --type=merge -p '{"spec":{"selector":{"app":"nope"}}}'
service/argocd-server patchedWatch the endpointslice lose its endpoints and curl start failing while DNS still resolves. Then look at what your patch actually did before undoing it: kubectl -n argocd get svc argocd-server -o jsonpath='{.spec.selector}' shows the original two labels plus app: nope, because a JSON merge patch merges maps rather than replacing them. So the honest restore is removing your key, not re-adding theirs: --type=merge -p '{"spec":{"selector":{"app":null}}}'.
app.kubernetes.io/name: argocd-server and app.kubernetes.io/instance: argocd, and curl returns 200. The merge-patch semantics come up on their own exam tasks.Run a probe that policy must kill:
cilium hubble port-forward &
kubectl -n team-a run probe --image=curlimages/curl:8.11.1 --restart=Never \
-- curl -s -m 5 http://example.com
hubble observe --namespace team-a --verdict DROPPED --last 20outputcaptured 2026-08-26
$ cilium hubble port-forward &
[1] 954968
$ kubectl -n team-a run probe --image=curlimages/curl:8.11.1 --restart=Never \
-- curl -s -m 5 http://example.com
Warning: would violate PodSecurity "restricted:latest": allowPrivilegeEscalation != false (container "probe" must set securityContext.allowPrivilegeEscalation=false), unrestricted capabilities (container "probe" must set securityContext.capabilities.drop=["ALL"]), runAsNonRoot != true (pod or container "probe" must set securityContext.runAsNonRoot=true), seccompProfile (pod or container "probe" must set securityContext.seccompProfile.type to "RuntimeDefault" or "Localhost")
pod/probe created
$ hubble observe --namespace team-a --verdict DROPPED --last 20
time=2026-08-26T22:13:56.514-04:00 level=WARN msg="Hubble CLI version is lower than Hubble Relay, API compatibility is not guaranteed, updating to a matching or higher version is recommended" hubble-cli-version=1.19.4 hubble-relay-version=1.20.1+g7d68cfb3
Aug 27 02:13:45.229: team-a/probe:39428 (ID:2901) <> 104.20.23.154:80 (world) policy-verdict:none TRAFFIC_DIRECTION_UNKNOWN DENIED (TCP Flags: SYN)
Aug 27 02:13:45.229: team-a/probe:39428 (ID:2901) <> 104.20.23.154:80 (world) Policy denied DROPPED (TCP Flags: SYN)
Aug 27 02:13:46.270: team-a/probe:39428 (ID:2901) <> 104.20.23.154:80 (world) policy-verdict:none TRAFFIC_DIRECTION_UNKNOWN DENIED (TCP Flags: SYN)
Aug 27 02:13:46.270: team-a/probe:39428 (ID:2901) <> 104.20.23.154:80 (world) Policy denied DROPPED (TCP Flags: SYN)
Aug 27 02:13:47.294: team-a/probe:39428 (ID:2901) <> 104.20.23.154:80 (world) policy-verdict:none TRAFFIC_DIRECTION_UNKNOWN DENIED (TCP Flags: SYN)
Aug 27 02:13:47.294: team-a/probe:39428 (ID:2901) <> 104.20.23.154:80 (world) Policy denied DROPPED (TCP Flags: SYN)
Aug 27 02:13:47.629: team-a/probe:48966 (ID:2901) <> 172.66.147.243:80 (world) policy-verdict:none TRAFFIC_DIRECTION_UNKNOWN DENIED (TCP Flags: SYN)
Aug 27 02:13:47.629: team-a/probe:48966 (ID:2901) <> 172.66.147.243:80 (world) Policy denied DROPPED (TCP Flags: SYN)
Aug 27 02:13:48.639: team-a/probe:48966 (ID:2901) <> 172.66.147.243:80 (world) policy-verdict:none TRAFFIC_DIRECTION_UNKNOWN DENIED (TCP Flags: SYN)
Aug 27 02:13:48.639: team-a/probe:48966 (ID:2901) <> 172.66.147.243:80 (world) Policy denied DROPPED (TCP Flags: SYN)
Aug 27 02:13:49.663: team-a/probe:48966 (ID:2901) <> 172.66.147.243:80 (world) policy-verdict:none TRAFFIC_DIRECTION_UNKNOWN DENIED (TCP Flags: SYN)
Aug 27 02:13:49.663: team-a/probe:48966 (ID:2901) <> 172.66.147.243:80 (world) Policy denied DROPPED (TCP Flags: SYN)Write a Gateway named web using listener port 80, protocol HTTP, plus an HTTPRoute that matches path /demo and backends a Service demo:80. Don't copy from docs; build it from kubectl explain gateway.spec.listeners and kubectl explain httproute.spec.rules.
kubectl get gateway web -o yaml shows the spec you meant. Status stays unprogrammed here; on the mesh cluster Istio would accept the same manifests.From a pod in default, resolve short and long names:
kubectl run dnsprobe --image=busybox:1.28 --restart=Never -it --rm -- \
sh -c 'nslookup argocd-server.argocd && nslookup argocd-server.argocd.svc.cluster.local'outputcaptured 2026-08-26
$ kubectl run dnsprobe --image=busybox:1.28 --restart=Never -it --rm -- \
sh -c 'nslookup argocd-server.argocd && nslookup argocd-server.argocd.svc.cluster.local'
Server: 10.96.0.10
Address 1: 10.96.0.10 kube-dns.kube-system.svc.cluster.local
Name: argocd-server.argocd
Address 1: 10.96.20.112 argocd-server.argocd.svc.cluster.local
Server: 10.96.0.10
Address 1: 10.96.0.10 kube-dns.kube-system.svc.cluster.local
Name: argocd-server.argocd.svc.cluster.local
Address 1: 10.96.20.112 argocd-server.argocd.svc.cluster.local
All commands and output from this session will be recorded in container logs, including credentials and sensitive information passed through the command prompt.
If you don't see a command prompt, try pressing enter.
pod "dnsprobe" deleted from default namespace(busybox:1.28 on purpose: newer busybox nslookup ignores the search path and returns NXDOMAIN for the short name even though the pod's resolver handles it fine. A famous trap, worth meeting here rather than in an exam.)
Then repeat inside team-a and explain why it still works (the tenant policy allows port 53 to kube-dns explicitly; find that rule in examples/multitenancy/team-a.yaml).
There is no browser in the exam, and the Gateway API field names are the kind you half-remember. The CRDs carry their own documentation, so kubectl explain is the reference, and it is correct for the channel this cluster actually installed.
kubectl explain gateway.spec.listeners.allowedRoutes.namespaces.from
kubectl explain gateway.spec.listeners.allowedRoutes --recursive | head -12
kubectl explain httproute.spec.rules.backendRefsoutputcaptured 2026-09-13
$ kubectl explain gateway.spec.listeners.allowedRoutes.namespaces.from
GROUP: gateway.networking.k8s.io
KIND: Gateway
VERSION: v1
FIELD: from <string>
ENUM:
All
Selector
Same
DESCRIPTION:
From indicates where Routes will be selected for this Gateway. Possible
values are:
* All: Routes in all namespaces may be used by this Gateway.
* Selector: Routes in namespaces selected by the selector may be used by
this Gateway.
* Same: Only Routes in the same namespace may be used by this Gateway.
Support: Core
$ kubectl explain gateway.spec.listeners.allowedRoutes --recursive | head -12
GROUP: gateway.networking.k8s.io
KIND: Gateway
VERSION: v1
FIELD: allowedRoutes <Object>
DESCRIPTION:
AllowedRoutes defines the types of routes that MAY be attached to a
Listener and the trusted namespaces where those Route resources MAY be
present.
$ kubectl explain httproute.spec.rules.backendRefs
GROUP: gateway.networking.k8s.io
KIND: HTTPRoute
VERSION: v1
FIELD: backendRefs <[]Object>
DESCRIPTION:
BackendRefs defines the backend(s) where matching requests should be
sent.
Failure behavior here depends on how many BackendRefs are specified and
how many are invalid.
If *all* entries in BackendRefs are invalid, and there are also no filters
specified in this route rule, *all* traffic which matches this rule MUST
receive a 500 status code.
See the HTTPBackendRef definition for the rules about what makes a single
HTTPBackendRef invalid.
When a HTTPBackendRef is invalid, 500 status codes MUST be returned for
requests that would have otherwise been routed to an invalid backend. If
multiple backends are specified, and some are invalid, the proportion of
requests that would otherwise have been routed to an invalid backend
MUST receive a 500 status code.
For example, if two backends are specified with equal weights, and one is
invalid, 50 percent of traffic must receive a 500. Implementations may
choose how that 50 percent is determined.
When a HTTPBackendRef refers to a Service that has no ready endpoints,
implementations SHOULD return a 503 for requests to that backend instead.
If an implementation chooses to do this, all of the above rules for 500
responses
MUST also apply for responses that return a 503.
Support: Core for Kubernetes Service
Support: Extended for Kubernetes ServiceImport
... 76 more linesallowedRoutes.namespaces.from documents All, Selector and Same, and backendRefs lists weight next to name, namespace and port. A field you expected and cannot find is a field the standard channel does not ship.A route in one namespace may not send traffic to a Service in another until the owner of that Service says so. The permission travels the other way from the reference, which is the part people get wrong under time pressure.
Run this on kind-mesh: the main cluster has the CRDs but no controller, so the route's status stays empty and there is nothing to read.
kubectl --context kind-mesh create ns team-a --dry-run=client -o yaml | kubectl --context kind-mesh apply -f -
kubectl --context kind-mesh create ns team-b --dry-run=client -o yaml | kubectl --context kind-mesh apply -f -
kubectl --context kind-mesh -n team-b create deployment web --image=nginx:1.27-alpine
kubectl --context kind-mesh -n team-b expose deployment web --port=80
kubectl --context kind-mesh apply -f - <<'EOF'
apiVersion: gateway.networking.k8s.io/v1
kind: Gateway
metadata: { name: web, namespace: team-a }
spec:
gatewayClassName: istio
listeners:
- name: http
port: 80
protocol: HTTP
allowedRoutes: { namespaces: { from: Same } }
---
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata: { name: demo, namespace: team-a }
spec:
parentRefs: [{ name: web }]
rules:
- backendRefs: [{ name: web, namespace: team-b, port: 80 }]
EOF
sleep 10
kubectl --context kind-mesh -n team-a get httproute demo -o jsonpath='{.status.parents[*].conditions}' | jq '.[] | {type, status, reason}'
kubectl --context kind-mesh apply -f - <<'EOF'
apiVersion: gateway.networking.k8s.io/v1beta1
kind: ReferenceGrant
metadata: { name: from-team-a, namespace: team-b }
spec:
from: [{ group: gateway.networking.k8s.io, kind: HTTPRoute, namespace: team-a }]
to: [{ group: "", kind: Service }]
EOF
sleep 10
kubectl --context kind-mesh -n team-a get httproute demo -o jsonpath='{.status.parents[*].conditions}' | jq '.[] | {type, status, reason}'outputcaptured 2026-09-13
$ kubectl --context kind-mesh create ns team-a --dry-run=client -o yaml | kubectl --context kind-mesh apply -f -
namespace/team-a created
$ kubectl --context kind-mesh create ns team-b --dry-run=client -o yaml | kubectl --context kind-mesh apply -f -
namespace/team-b created
$ kubectl --context kind-mesh -n team-b create deployment web --image=nginx:1.27-alpine
deployment.apps/web created
$ kubectl --context kind-mesh -n team-b expose deployment web --port=80
service/web exposed
$ kubectl --context kind-mesh apply -f - <<'EOF'
apiVersion: gateway.networking.k8s.io/v1
kind: Gateway
metadata: { name: web, namespace: team-a }
spec:
gatewayClassName: istio
listeners:
- name: http
port: 80
protocol: HTTP
allowedRoutes: { namespaces: { from: Same } }
---
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata: { name: demo, namespace: team-a }
spec:
parentRefs: [{ name: web }]
rules:
- backendRefs: [{ name: web, namespace: team-b, port: 80 }]
EOF
gateway.gateway.networking.k8s.io/web created
httproute.gateway.networking.k8s.io/demo created
$ sleep 10
$ kubectl --context kind-mesh -n team-a get httproute demo -o jsonpath='{.status.parents[*].conditions}' | jq '.[] | {type, status, reason}'
{
"type": "Accepted",
"status": "True",
"reason": "Accepted"
}
{
"type": "ResolvedRefs",
"status": "False",
"reason": "RefNotPermitted"
}
{
"type": "ResolvedWaypoints",
"status": "True",
"reason": "ResolvedWaypoints"
}
$ kubectl --context kind-mesh apply -f - <<'EOF'
apiVersion: gateway.networking.k8s.io/v1beta1
kind: ReferenceGrant
metadata: { name: from-team-a, namespace: team-b }
spec:
from: [{ group: gateway.networking.k8s.io, kind: HTTPRoute, namespace: team-a }]
to: [{ group: "", kind: Service }]
EOF
referencegrant.gateway.networking.k8s.io/from-team-a created
$ sleep 10
$ kubectl --context kind-mesh -n team-a get httproute demo -o jsonpath='{.status.parents[*].conditions}' | jq '.[] | {type, status, reason}'
{
"type": "Accepted",
"status": "True",
"reason": "Accepted"
}
{
"type": "ResolvedRefs",
"status": "True",
"reason": "ResolvedRefs"
}
{
"type": "ResolvedWaypoints",
"status": "True",
"reason": "ResolvedWaypoints"
}ResolvedRefs is False with reason RefNotPermitted before the grant and True after it, with nothing about the route itself changed. The grant lives with the backend, not with the route.Every cluster you inherit has a Corefile, and half the DNS incidents you will be handed are a plugin block someone added to it. Read the stock one, add a stub domain, and watch the reload land.
kubectl -n kube-system get cm coredns -o jsonpath='{.data.Corefile}' | tee /tmp/Corefile.orig
printf '%s\nconsul.local:53 {\n errors\n cache 30\n forward . 10.150.0.1\n}\n' "$(cat /tmp/Corefile.orig)" > /tmp/Corefile.new
kubectl -n kube-system create configmap coredns --from-file=Corefile=/tmp/Corefile.new --dry-run=client -o yaml | kubectl -n kube-system apply -f -
sleep 120
kubectl -n kube-system logs deploy/coredns --tail=10
kubectl -n kube-system create configmap coredns --from-file=Corefile=/tmp/Corefile.orig --dry-run=client -o yaml | kubectl -n kube-system apply -f -outputcaptured 2026-09-13
$ kubectl -n kube-system get cm coredns -o jsonpath='{.data.Corefile}' | tee /tmp/Corefile.orig
.:53 {
errors
health {
lameduck 5s
}
ready
hosts {
172.18.0.6 gitea.lab
172.18.0.2 kind-registry
fallthrough
}
kubernetes cluster.local in-addr.arpa ip6.arpa {
pods insecure
fallthrough in-addr.arpa ip6.arpa
ttl 30
}
prometheus :9153
forward . /etc/resolv.conf {
max_concurrent 1000
}
cache 30 {
disable success cluster.local
disable denial cluster.local
}
loop
reload
loadbalance
}
$ printf '%s\nconsul.local:53 {\n errors\n cache 30\n forward . 10.150.0.1\n}\n' "$(cat /tmp/Corefile.orig)" > /tmp/Corefile.new
$ kubectl -n kube-system create configmap coredns --from-file=Corefile=/tmp/Corefile.new --dry-run=client -o yaml | kubectl -n kube-system apply -f -
configmap/coredns configured
$ sleep 120
$ kubectl -n kube-system logs deploy/coredns --tail=10
Found 2 pods, using pod/coredns-754f9cbf9f-9kjvm
[INFO] plugin/reload: Running configuration SHA512 = 21ac5a0b49ecc87479896d161ce04e13d5464755235ff622730e3dfae4608be276f4a5f8b10043b699e176d468200893ab00c7ce401fe4e331b9c0e19f77eb6d
CoreDNS-1.14.2
linux/amd64, go1.26.1, dd1df4f
[ERROR] plugin/errors: 2 _grpclb._tcp.argocd-repo-server. SRV: read udp 10.244.2.138:45773->172.18.0.1:53: i/o timeout
[ERROR] plugin/errors: 2 _grpclb._tcp.spire-server.spire. SRV: read udp 10.244.2.138:41440->172.18.0.1:53: i/o timeout
[ERROR] plugin/errors: 2 _grpclb._tcp.spire-server.spire. SRV: read udp 10.244.2.138:45718->172.18.0.1:53: i/o timeout
[ERROR] plugin/errors: 2 _grpclb._tcp.spire-server.spire. SRV: read udp 10.244.2.138:60045->172.18.0.1:53: i/o timeout
[INFO] Reloading
[INFO] plugin/reload: Running configuration SHA512 = fa9459fcfd404334de10bf2f59ce4f79c08316d97dff8ba20f46ce959b71a49c67999e92a94381f4cab42431a2af64dcc71ec6d1681cd790c23927cc3c07ac13
[INFO] Reloading complete
$ kubectl -n kube-system create configmap coredns --from-file=Corefile=/tmp/Corefile.orig --dry-run=client -o yaml | kubectl -n kube-system apply -f -
configmap/coredns configuredA pod on the host network gets the node's resolver unless you say otherwise, so cluster names stop resolving and nothing in the pod spec looks wrong. It is a two-word fix and a favorite of exam authors.
kubectl run hn --image=busybox:1.28 --restart=Never --overrides='{"spec":{"hostNetwork":true}}' -- nslookup kubernetes.default
sleep 15
kubectl logs hn
kubectl run hn2 --image=busybox:1.28 --restart=Never --overrides='{"spec":{"hostNetwork":true,"dnsPolicy":"ClusterFirstWithHostNet"}}' -- nslookup kubernetes.default
sleep 15
kubectl logs hn2
kubectl delete pod hn hn2outputcaptured 2026-09-12
$ kubectl run hn --image=busybox:1.28 --restart=Never --overrides='{"spec":{"hostNetwork":true}}' -- nslookup kubernetes.default
pod/hn created
$ sleep 15
$ kubectl logs hn
nslookup: can't resolve 'kubernetes.default'
Server: 172.18.0.1
Address 1: 172.18.0.1 omarchy
$ kubectl run hn2 --image=busybox:1.28 --restart=Never --overrides='{"spec":{"hostNetwork":true,"dnsPolicy":"ClusterFirstWithHostNet"}}' -- nslookup kubernetes.default
pod/hn2 created
$ sleep 15
$ kubectl logs hn2
Server: 10.96.0.10
Address 1: 10.96.0.10 kube-dns.kube-system.svc.cluster.local
Name: kubernetes.default
Address 1: 10.96.0.1 kubernetes.default.svc.cluster.local
$ kubectl delete pod hn hn2
pod "hn" deleted from default namespace
pod "hn2" deleted from default namespacekubernetes.default and the second can, with dnsPolicy: ClusterFirstWithHostNet the only difference between them.Cilium may or may not be doing kube-proxy's job, and the answer changes what you look at when a Service stops working. Ask the agent rather than guessing from the pod list.
kubectl -n kube-system exec ds/cilium -c cilium-agent -- cilium-dbg status | grep -i -A2 'KubeProxyReplacement'
kubectl -n kube-system get cm cilium-config -o jsonpath='{.data.kube-proxy-replacement}{"\n"}'
kubectl -n kube-system get ds -l k8s-app=kube-proxyoutputcaptured 2026-09-13
$ kubectl -n kube-system exec ds/cilium -c cilium-agent -- cilium-dbg status | grep -i -A2 'KubeProxyReplacement'
KubeProxyReplacement: False
Host firewall: Disabled
SRv6: Disabled
$ kubectl -n kube-system get cm cilium-config -o jsonpath='{.data.kube-proxy-replacement}{"\n"}'
false
$ kubectl -n kube-system get ds -l k8s-app=kube-proxy
NAME DESIRED CURRENT READY UP-TO-DATE AVAILABLE NODE SELECTOR AGE
kube-proxy 3 3 3 3 3 kubernetes.io/os=linux 16htrafficDistribution is a hint, not a rule: the field is accepted anywhere, but the EndpointSlice controller only writes zone hints when the endpoints actually carry zones. On kind the nodes have no zone label until you add one, which makes the mechanism easy to see.
kubectl create deployment demo --image=nginx:1.27-alpine --replicas=4
kubectl expose deployment demo --port=80
kubectl patch svc demo -p '{"spec":{"trafficDistribution":"PreferSameZone"}}'
kubectl get svc demo -o jsonpath='{.spec.trafficDistribution}{"\n"}'
kubectl get endpointslices -l kubernetes.io/service-name=demo -o yaml | grep -B2 -A3 hints
kubectl label node cnpe-worker topology.kubernetes.io/zone=a --overwrite
kubectl label node cnpe-worker2 topology.kubernetes.io/zone=b --overwrite
kubectl rollout restart deployment demo
kubectl rollout status deployment demo
kubectl get endpointslices -l kubernetes.io/service-name=demo -o yaml | grep -B2 -A3 hints
# the scheduling exercises on this page assume unzoned nodes
kubectl delete svc demo --ignore-not-found
kubectl delete deployment demo --ignore-not-found
kubectl label node cnpe-worker topology.kubernetes.io/zone-
kubectl label node cnpe-worker2 topology.kubernetes.io/zone-outputcaptured 2026-09-12
$ kubectl create deployment demo --image=nginx:1.27-alpine --replicas=4
deployment.apps/demo created
$ kubectl expose deployment demo --port=80
service/demo exposed
$ kubectl patch svc demo -p '{"spec":{"trafficDistribution":"PreferSameZone"}}'
service/demo patched
$ kubectl get svc demo -o jsonpath='{.spec.trafficDistribution}{"\n"}'
PreferSameZone
$ kubectl get endpointslices -l kubernetes.io/service-name=demo -o yaml | grep -B2 -A3 hints
$ kubectl label node cnpe-worker topology.kubernetes.io/zone=a --overwrite
node/cnpe-worker labeled
$ kubectl label node cnpe-worker2 topology.kubernetes.io/zone=b --overwrite
node/cnpe-worker2 labeled
$ kubectl rollout restart deployment demo
deployment.apps/demo restarted
$ kubectl rollout status deployment demo
Waiting for deployment spec update to be observed...
Waiting for deployment "demo" rollout to finish: 1 out of 4 new replicas have been updated...
Waiting for deployment "demo" rollout to finish: 1 out of 4 new replicas have been updated...
Waiting for deployment "demo" rollout to finish: 2 out of 4 new replicas have been updated...
Waiting for deployment "demo" rollout to finish: 2 out of 4 new replicas have been updated...
Waiting for deployment "demo" rollout to finish: 2 out of 4 new replicas have been updated...
Waiting for deployment "demo" rollout to finish: 2 out of 4 new replicas have been updated...
Waiting for deployment "demo" rollout to finish: 2 out of 4 new replicas have been updated...
Waiting for deployment "demo" rollout to finish: 2 out of 4 new replicas have been updated...
Waiting for deployment "demo" rollout to finish: 2 out of 4 new replicas have been updated...
Waiting for deployment "demo" rollout to finish: 2 out of 4 new replicas have been updated...
Waiting for deployment "demo" rollout to finish: 3 out of 4 new replicas have been updated...
Waiting for deployment "demo" rollout to finish: 3 out of 4 new replicas have been updated...
Waiting for deployment "demo" rollout to finish: 1 old replicas are pending termination...
Waiting for deployment "demo" rollout to finish: 1 old replicas are pending termination...
Waiting for deployment "demo" rollout to finish: 1 old replicas are pending termination...
Waiting for deployment "demo" rollout to finish: 1 old replicas are pending termination...
Waiting for deployment "demo" rollout to finish: 3 of 4 updated replicas are available...
deployment "demo" successfully rolled out
$ kubectl get endpointslices -l kubernetes.io/service-name=demo -o yaml | grep -B2 -A3 hints
serving: true
terminating: false
hints:
forZones:
- name: a
nodeName: cnpe-worker
--
serving: true
terminating: false
hints:
forZones:
- name: b
nodeName: cnpe-worker2
--
serving: true
terminating: false
hints:
forZones:
- name: a
nodeName: cnpe-worker
--
serving: true
terminating: false
hints:
forZones:
- name: b
nodeName: cnpe-worker2
$ # the scheduling exercises on this page assume unzoned nodes
$ kubectl delete svc demo --ignore-not-found
service "demo" deleted from default namespace
$ kubectl delete deployment demo --ignore-not-found
deployment.apps "demo" deleted from default namespace
$ kubectl label node cnpe-worker topology.kubernetes.io/zone-
node/cnpe-worker unlabeled
$ kubectl label node cnpe-worker2 topology.kubernetes.io/zone-
node/cnpe-worker2 unlabeledtopology.kubernetes.io/zone, and each endpoint's hint names its own zone. Clean up the labels afterwards so the scheduling exercises still behave.Self-check
A Service resolves in DNS but every connection is refused. Name the three things you check, in order.
Selector against actual pod labels; pod readiness (an unready pod is not in the slice); the EndpointSlice itself. If the slice has addresses with ready: false, it is a probe problem, not a Service problem. If the slice is empty, it is a selector or a scheduling problem.
Why does a NetworkPolicy ipBlock naming a Service's ClusterIP never work?
Because DNAT happens before policy evaluation: by the time the packet is checked, its destination is a pod IP. Policy is expressed in pod selectors and identities, not virtual IPs. For the API server specifically, Cilium's toEntities: [kube-apiserver] is the expressible form.
Your backend needs the real client IP. What do you change, and what breaks?
externalTrafficPolicy: Local on the Service. The trade: nodes with no backing pod stop serving that traffic, so you now depend on the load balancer's health checks to avoid them, and you lose the even spread that Cluster gave you.
What is the difference between a headless Service and a ClusterIP Service with one endpoint?
Headless has no VIP and no datapath translation: DNS returns pod IPs directly and the client chooses. A one-endpoint ClusterIP still goes through the VIP and its rules, so the client never learns the pod IP. StatefulSets need the former for stable per-pod names (pod-0.svc…).
Team-a can curl team-b's Service after you added an egress allow, but it still fails. What did you forget?
Team-b's ingress side. Both ends must open; policies are additive allow-lists evaluated independently at each pod. Also check DNS: if egress is now restricted in team-a, resolving web.team-b.svc needs its own port-53 allow.
An HTTPRoute applies cleanly but nothing routes. Where do you look first?
status.parents[].conditions on the route: Accepted false usually means attachment was refused by the Gateway's allowedRoutes; ResolvedRefs false means a backend Service or a TLS secret is missing (or needs a ReferenceGrant across namespaces). On the Gateway, Programmed false means no data plane, which is the permanent state on this lab's main cluster.
An HTTPRoute in team-a references a Service in team-b. Which object, in which namespace, and what does status say until it exists?
A ReferenceGrant in team-b (the namespace being referenced) with from group gateway.networking.k8s.io kind HTTPRoute namespace team-a and to kind Service. Until it exists the route reports ResolvedRefs: False with reason RefNotPermitted, and the implementation answers 500 for that backend.
Two backendRefs carry weight 3 and weight 1. What split is that, and what happens if one backend Service is missing?
75/25: each weight is a share of the total, so 3 and 1 make four shares. A missing backend produces ResolvedRefs: False, BackendNotFound; the implementation must still route the valid share and return 500 for the share that points at the invalid backend, so users see intermittent errors proportional to the weight.
A stub domain for corp.example must resolve via 10.0.0.53. Where does the change go and how long before it applies?
A second server block in the coredns ConfigMap in kube-system: corp.example:53 { errors; cache 30; forward . 10.0.0.53 }. The reload plugin applies it within about two minutes without restarting pods. Verify from a pod with nslookup, and check kubectl -n kube-system logs deploy/coredns for a Corefile parse error if nothing changes.
Which kube-proxy mode is GA and newest, and what changes for NodePort when you switch to it?
nftables, GA since Kubernetes 1.33 (kernel 5.13+). NodePorts default to --nodeport-addresses primary, so they answer on the node's primary IP only; other interfaces and localhost stop answering, and kube-proxy no longer punches firewall holes for the NodePort range. IPVS is deprecated since 1.35; iptables remains the default.
Docs to know your way around
- kubernetes.io: Services, DNS for Services and Pods, Network Policies. The DNS page's search-domain section is the one people never read.
- gateway-api.sigs.k8s.io: the API model page with the role diagram; the "route attachment" section.
- docs.cilium.io: Hubble observe reference; CiliumNetworkPolicy entities and FQDN rules.
- Offline, in the exam:
kubectl explain service.spec,kubectl explain networkpolicy.spec.egress,kubectl explain httproute.spec.rules --recursive. Faster than docs and always version-correct for the cluster in front of you. - gateway-api.sigs.k8s.io/reference/spec: the field reference for Gateway listeners, HTTPRoute matches/filters/backendRefs, ReferenceGrant, BackendTLSPolicy and ListenerSet; the condition reasons (NotAllowedByListeners, RefNotPermitted, BackendNotFound) are listed there verbatim.
- kubernetes.io/docs/reference/networking/virtual-ips: kube-proxy modes, traffic policies, terminating endpoints and trafficDistribution on one page.
- kubernetes.io/docs/tasks/administer-cluster/dns-custom-nameservers: the stock Corefile with every plugin explained, and the stub-domain example you will copy under time pressure.
make down-gitopsmake down-sec