Broken delivery is the most likely domain 2 exam task. Failures sort into four buckets, and the first diagnostic move is deciding which bucket you are in. Everything here is practice for that decision.

needsmake core

Orientation

the diagnostic half of every domain 2 competency

Under time pressure, the expensive mistake is not a wrong fix; it is fixing in the wrong layer. Bucket first, then descend. Say the bucket out loud before you touch anything; it costs three seconds and it stops you editing live state when the problem is a commit.

The 30-second triage
argocd app get <app>            # sync status + health + last sync time + conditions
flux get kustomizations -A       # Ready / Suspended / revision
kubectl -n <ns> get pods         # is anything even running
kubectl -n <ns> get events --sort-by=.lastTimestamp | tail -20
outputcaptured 2026-08-26
$ argocd app get demo-staging     # sync status + health + last sync time + conditions
Name:               argocd/demo-staging
Project:            default
Server:             https://kubernetes.default.svc
Namespace:          staging
URL:                https://argocd.example.com/applications/demo-staging
Source:
- Repo:             http://gitea.lab:3000/lab/platform.git
  Target:           main
  Path:             demo-app/overlays/staging
SyncWindow:         Sync Allowed
Sync Policy:        Automated (Prune)
Sync Status:        Synced to main (c7cb9a1)
Health Status:      Healthy

GROUP  KIND        NAMESPACE  NAME          STATUS  HEALTH   HOOK  MESSAGE
apps   Deployment  team-a     staging-demo  Synced  Healthy        deployment.apps/staging-demo configured
       Service     team-a     staging-demo  Synced  Healthy        
$ flux get kustomizations -A       # Ready / Suspended / revision
NAMESPACE  	NAME     	REVISION          	SUSPENDED	READY	MESSAGE                              
flux-system	demo-flux	main@sha1:c7cb9a1e	False    	True 	Applied revision: main@sha1:c7cb9a1e	
$ kubectl -n team-a get pods       # is anything even running
NAME                            READY   STATUS    RESTARTS   AGE
staging-demo-66955f8974-2xxd2   1/1     Running   0          4m21s
staging-demo-66955f8974-pzxdh   1/1     Running   0          4m21s
staging-demo-66955f8974-tzpqx   1/1     Running   0          4m22s
$ kubectl -n team-a get events --sort-by=.lastTimestamp | tail -20
4m16s       Normal    Created             pod/staging-demo-66955f8974-2xxd2    Container created
4m15s       Normal    Created             pod/staging-demo-66955f8974-tzpqx    Container created
4m15s       Normal    Started             pod/staging-demo-66955f8974-pzxdh    Container started
4m15s       Normal    Started             pod/staging-demo-66955f8974-tzpqx    Container started
4m1s        Normal    SuccessfulCreate    replicaset/staging-demo-66955f8974   Created pod: staging-demo-66955f8974-7xnfb
4m1s        Normal    ScalingReplicaSet   deployment/staging-demo              Scaled up replica set staging-demo-66955f8974 from 3 to 5
4m1s        Normal    Scheduled           pod/staging-demo-66955f8974-7xnfb    Successfully assigned team-a/staging-demo-66955f8974-7xnfb to cnpe-worker
4m1s        Normal    SuccessfulCreate    replicaset/staging-demo-66955f8974   Created pod: staging-demo-66955f8974-jh7rr
4m          Normal    Scheduled           pod/staging-demo-66955f8974-jh7rr    Successfully assigned team-a/staging-demo-66955f8974-jh7rr to cnpe-worker2
3m59s       Normal    SuccessfulDelete    replicaset/staging-demo-66955f8974   Deleted pod: staging-demo-66955f8974-7xnfb
3m59s       Normal    ScalingReplicaSet   deployment/staging-demo              Scaled down replica set staging-demo-66955f8974 from 5 to 3
3m59s       Normal    SuccessfulDelete    replicaset/staging-demo-66955f8974   Deleted pod: staging-demo-66955f8974-jh7rr
3m58s       Normal    Pulled              pod/staging-demo-66955f8974-7xnfb    Container image "ghcr.io/nginxinc/nginx-unprivileged:1.27-alpine" already present on machine and can be accessed by the pod
3m58s       Normal    Pulled              pod/staging-demo-66955f8974-jh7rr    Container image "ghcr.io/nginxinc/nginx-unprivileged:1.27-alpine" already present on machine and can be accessed by the pod
3m57s       Normal    Killing             pod/staging-demo-66955f8974-jh7rr    Stopping container web
3m57s       Normal    Started             pod/staging-demo-66955f8974-jh7rr    Container started
3m57s       Normal    Created             pod/staging-demo-66955f8974-jh7rr    Container created
3m57s       Normal    Killing             pod/staging-demo-66955f8974-7xnfb    Stopping container web
3m57s       Normal    Started             pod/staging-demo-66955f8974-7xnfb    Container started
3m57s       Normal    Created             pod/staging-demo-66955f8974-7xnfb    Container created

Those four commands place you in one of the four buckets below almost every time.

The four buckets

symptom → evidence → who fixes it

1 · Git is wrong, the cluster is faithful

Synced + Degraded. Bad image tag, impossible resource request, missing ConfigMap key, a probe that can never pass. The controller did its job; fix the commit. Evidence: app health Degraded (or Progressing on its way there) while sync status is green, and pod events carry the real error: ImagePullBackOff, CreateContainerConfigError, CrashLoopBackOff.

2 · The cluster refuses what git says

Sync fails outright. An admission policy denies the manifest, the API version does not exist on this cluster, a field is immutable (Service clusterIP, label selectors, PVC shrink), the CRD is not installed yet. Evidence: the sync operation's error message, which quotes the API server's rejection verbatim. Immutable-field errors mean delete-and-recreate or Replace=true.

3 · The controller lacks permission or access

Repo unreachable, credentials rotated, an Application targeting a namespace the controller's RBAC cannot touch, an ApplicationSet token missing a scope. Evidence: errors mention the controller's own identity, appear in controller logs (kubectl -n argocd logs deploy/argocd-repo-server, flux logs), and say nothing about your workload. No pod is ever involved; that absence is the tell.

4 · Nobody is wrong, the state is stale

Webhook lost, refresh interval not elapsed, reconciliation suspended and forgotten, a branch that moved while the app points at a tag. Evidence: everything green but old. flux get kustomizations shows a Suspended row; argocd app get shows a last-sync timestamp from an hour ago. Fix with argocd app get --refresh / flux reconcile … --with-source, and check for suspension first.

BucketSyncHealthWhere the text comes fromFix lives in
1 git wrongSyncedDegradedpod eventsa commit
2 cluster refusesFailed / OutOfSyncn/aAPI server, quoted by the sync opa commit + a sync option
3 accessUnknown / errorn/acontroller logsa Secret, RBAC, or a token
4 staleSyncedHealthytimestamps, Suspended flaga reconcile or a resume
Drift deserves one more sentence

With self-heal on, drift self-corrects and the interesting question becomes "what keeps re-creating this thing I keep deleting". The answer is the controller, and kubectl get <res> -o jsonpath='{.metadata.ownerReferences}' or the app.kubernetes.io/instance label proves it. With self-heal off, argocd app diff is the tool that shows exactly what diverged.

The descent, one level at a time

resist skipping levels
Application / Kustomization   status.conditions, last sync, revision
        │
        ▼
rendered manifests            argocd app manifests · kustomize build · flux build
        │
        ▼
the apply                     sync operation message, API server rejection text
        │
        ▼
controller objects            Deployment → ReplicaSet → Pod   ← quota and PSS errors live on the RS
        │
        ▼
pod                           describe → events → logs → logs --previous → exec/debug
        │
        ▼
prove recovery with the signal that showed the failure

Two habits that pay for themselves. First, compare rendered manifests with live objects rather than reading source YAML: argocd app manifests <app> shows exactly what would be applied, which catches overlay mistakes that source review misses. Second, check the middle layer: a Deployment that is created but produces no pods has its error on the ReplicaSet, and that indirection is easy to miss.

The status vocabulary, engine by engine

every string you will see in a status field, and the bucket it points at

Bucketing is faster when you read the exact word rather than the color. Both engines write stable strings; a scenario quotes them.

Argo CD

FieldValueBucket
status.sync.statusOutOfSync1 if auto-sync is off and git moved; 4 if auto-sync should have acted; check syncPolicy and sync windows
status.sync.statusUnknown3: comparison failed; read conditions
status.health.statusDegraded · Missing1: the workload cannot run, or a desired resource is absent (often pruned by another app, or refused: bucket 5)
status.health.statusProgressing (for long)1, or a CRD with no health check that never settles
status.health.statusSuspendeda paused Deployment, suspended CronJob or paused Rollout: intentional, not a fault
status.operationState.phaseFailed2 or 5: the API server rejected an apply; the text is in syncResult.resources[].message
status.operationState.phaseError3: Argo could not run the operation (render, repo, hook creation)
conditions[].typeComparisonError3: repo unreachable, bad credentials, render failure, cluster unreachable
conditions[].typeInvalidSpecError3: the AppProject forbids the repo, destination or kind; or the spec references a missing project
conditions[].typeSyncError2 or 5
conditions[].typeSharedResourceWarning · RepeatedResourceWarning · OrphanedResourceWarning · ExcludedResourceWarningwarnings: two apps own one resource; one source yields a resource twice; untracked resources in the namespace; a kind excluded by resource.exclusions
conditions[].typeDeletionErrora PreDelete or PostDelete hook failed; the Application stays Terminating

Flux

ObjectReady reasonBucket
GitRepository / OCIRepositoryAuthenticationFailed · GitOperationFailed · OCIArtifactPullFailed · VerificationError3: credentials, URL, ref, signature
KustomizationArtifactFailed3, upstream: the source has no artifact; fix the source first
KustomizationBuildFailed1: bad path, invalid kustomization, missing substituteFrom ConfigMap
KustomizationDependencyNotReadylook at the named dependency, same ladder
KustomizationReconciliationFailed2 or 5: the apply was rejected; the message quotes the API server or webhook
KustomizationHealthCheckFailed1: applied, never became healthy within timeout
KustomizationPruneFailed2: garbage collection blocked (finalizers, protected resource)
HelmReleaseInstallFailed · UpgradeFailed1 or 2: Helm's own error; run the descend with kubectl describe hr and flux logs --kind HelmRelease
HelmReleaseTestFailed1: the chart's tests failed; remediation follows unless ignored
HelmReleaseRollbackSucceeded · UninstallSucceeded · RetriesExceededremediated: last good revision is running; a new generation or flux reconcile hr --reset tries again
anyReconciling=True, reason ProgressingWithRetrythe last attempt failed and a retry is scheduled; the Ready reason names why
anyspec.suspend: true (Suspended column)4
Reflex

Argo: argocd app get X then kubectl -n argocd get app X -o jsonpath='{.status.conditions[*].type}'. Flux: flux get all -A --status-selector ready=false then kubectl get kustomization X -n flux-system -o jsonpath='{.status.conditions[?(@.type=="Ready")].reason}'. The reason string is the bucket; only then read the message.

Bucket 5: admission refused it

webhooks and policy engines between the controller and etcd

The four buckets assume the API server is the last word. On a platform with Kyverno, Gatekeeper, ValidatingAdmissionPolicy or an operator's webhook, there is a fifth place a manifest dies, and it produces its own vocabulary.

Message fragmentWhoFix lives in
admission webhook "validate.kyverno.svc-fail" denied the request: ... policy X rule YKyverno validate, Enforcethe manifest (satisfy the rule) or the policy (exclude, Audit); never the controller
admission webhook "validation.gatekeeper.sh" denied the request: [constraint-name] ...Gatekeeper constraintsame choice; enforcementAction: dryrun is the audit mode
ValidatingAdmissionPolicy 'x' with binding 'y' denied request: ...VAP (GA 1.30), MutatingAdmissionPolicy (GA 1.36) for mutationsthe policy's CEL or the binding's match; no webhook pod involved
failed calling webhook "x": ... connection refused / context deadline exceeded / no endpoints availablea webhook whose backing pod is down, with failurePolicy: Failthe webhook's Deployment or Service; every matching write is blocked until it is back, including the controller's own
dry-run failed ... admission webhook "x" does not support dry runFlux (server-side dry-run before apply) meets a webhook without sideEffects: None or NoneOnDryRunthe webhook configuration
the server could not find the requested resourcea CR whose CRD is not installed yet (Argo CD dry-run)SkipDryRunOnMissingResource=true or a wave; Flux: dependsOn the CRD Kustomization

The tell that separates bucket 5 from bucket 2: the rejection names a webhook or policy rather than a field. The controller-side symptom is identical (Argo SyncFailed on the resource, Flux ReconciliationFailed), and for Deployments the refusal can again land one level down on the ReplicaSet when a pod-level policy is violated. A mutating policy that changes what git said (adds labels, injects a sidecar, sets requests) shows up as permanent OutOfSync in Argo CD and as "configured" events every interval in Flux: the answer there is ignoreDifferences or kustomize.toolkit.fluxcd.io/ssa: Merge, not fighting the policy.

The webhook that blocks its own fix

A dead webhook with failurePolicy: Fail and a broad match (all pods, all namespaces) stops the GitOps controller from applying the very manifest that would repair it. Diagnose from kubectl get validatingwebhookconfiguration,mutatingwebhookconfiguration and the webhook Service's endpoints; the sanctioned repair is fixing the pod, and the emergency one is switching that webhook to Ignore or narrowing its namespaceSelector, which is itself a change to record in git.

Exercises

each one stages a failure; diagnose before fixing

Do the diagnosis before the fix.

Push newTag: 9.9.9-nope to the staging overlay in the platform repo. Watch demo-staging stay Synced while health goes Progressing (Degraded arrives only after the Deployment's ten-minute progress deadline; do not wait for it, the pod evidence is immediate). Diagnose down the stack: argocd app get demo-staging → kubectl -n team-a get pods → describe pod shows ImagePullBackOff. Fix in git only.

verify: revert commit pushed, app back to Healthy without any kubectl mutation.

Add a patch to the staging overlay's kustomization setting spec.clusterIP: 10.96.99.99 on the demo Service (the Service manifest itself lives in base; overlays change it through patches, the standard kustomize idiom). Sync and read the error.

verify: you can quote where the error text came from (the API server, relayed by the sync operation) and fix it by removing the patch. Bucket 2, and the error message named it.

Deleting the repo secret outright would not do it here, and knowing why is the exercise's first half: the lab's repos are public, so anonymous cloning still works. Wrong credentials fail where absent ones would not, because Gitea rejects a bad password even on a public repo. So corrupt the secret:

kubectl -n argocd patch secret gitea-repo -p '{"stringData":{"password":"wrong"}}'
argocd app get demo-staging --refresh
outputcaptured 2026-08-26
$ kubectl -n argocd patch secret gitea-repo -p '{"stringData":{"password":"wrong"}}'
secret/gitea-repo patched
$ argocd app get demo-staging --refresh
Name:               argocd/demo-staging
Project:            default
Server:             https://kubernetes.default.svc
Namespace:          staging
URL:                https://argocd.example.com/applications/demo-staging
Source:
- Repo:             http://gitea.lab:3000/lab/platform.git
  Target:           main
  Path:             demo-app/overlays/staging
SyncWindow:         Sync Allowed
Sync Policy:        Automated (Prune)
Sync Status:        Unknown
Health Status:      Healthy

CONDITION        MESSAGE  LAST TRANSITION
ComparisonError  Failed to load target state: failed to generate manifest for source 1 of 1: rpc error: code = Unknown desc = failed to list refs: authentication required: Failed to authenticate user
                 2026-08-26 22:21:09 -0400 EDT


GROUP  KIND        NAMESPACE  NAME          STATUS   HEALTH   HOOK  MESSAGE
apps   Deployment  team-a     staging-demo  Unknown  Healthy        deployment.apps/staging-demo configured
       Service     team-a     staging-demo  Unknown  Healthy        

The error mentions authentication, not manifests. Restore by re-running make gitops (idempotent) or patching the real token back from .gitea-token.

verify: the refresh succeeds again. Notice how different this error was: no pod was ever involved.

flux suspend kustomization demo-flux (from 2.3), push any change to demo-app/base in the platform repo, and observe that demo-flux's targets never move and nothing errors (the Argo apps do pick the base change up; only the suspended consumer goes quietly stale). The absence of failure is the symptom.

verify: flux get kustomizations shows Suspended, resume, and the change lands. Train yourself to check suspension first; it costs three seconds.

FAULT=config make break corrupts something in team-a's delivery path under a 7-minute clock. Bucket it, fix it, then make break-answer to compare.

verify: run it until the bucketing step takes under a minute.

An admission policy that the manifests violate stops both engines, but each reports it in its own vocabulary. Learn both strings, because the task will show you one of them and expect you to name the cause.

# the LoadBalancer IP only exists when cloud-provider-kind runs; the NodePort is always there
ARGO=$(kubectl get node cnpe-control-plane -o jsonpath='{.status.addresses[?(@.type=="InternalIP")].address}'):$(kubectl -n argocd get svc argocd-server -o jsonpath='{.spec.ports[?(@.port==80)].nodePort}')
argocd login "$ARGO" --username admin --password "$(kubectl -n argocd get secret argocd-initial-admin-secret -o jsonpath='{.data.password}' | base64 -d)" --plaintext --grpc-web
kubectl apply -f - <<'EOF'
apiVersion: kyverno.io/v1
kind: ClusterPolicy
metadata: { name: require-owner }
spec:
  validationFailureAction: Enforce
  rules:
    - name: require-owner-label
      match:
        any:
          - resources:
              kinds: [Deployment]
              namespaces: [team-a, flux-demo]
      validate:
        message: "every Deployment needs an owner label"
        pattern:
          metadata:
            labels:
              owner: "?*"
EOF
sleep 10
# an unchanged resource is never sent to admission, so each engine has to create one
kubectl -n team-a delete deploy staging-demo --ignore-not-found
kubectl -n flux-demo delete deploy demo --ignore-not-found
# this app syncs automatically, so wait out the controller's own operation first
argocd app wait demo-staging --operation --timeout 180 || true
argocd app sync demo-staging --timeout 60 || true
argocd app wait demo-staging --operation --timeout 180 || true
kubectl -n argocd get app demo-staging -o jsonpath='{.status.operationState.phase}{"\n"}{.status.operationState.message}{"\n"}'
flux reconcile kustomization demo-flux --with-source || true
kubectl -n flux-system get kustomization demo-flux -o jsonpath='{.status.conditions}' | jq '.[] | {type, status, reason, message}'
kubectl delete clusterpolicy require-owner
# this app syncs automatically, so wait out the controller's own operation first
argocd app wait demo-staging --operation --timeout 180 || true
argocd app sync demo-staging --timeout 120
outputcaptured 2026-09-13
$ # the LoadBalancer IP only exists when cloud-provider-kind runs; the NodePort is always there
$ ARGO=$(kubectl get node cnpe-control-plane -o jsonpath='{.status.addresses[?(@.type=="InternalIP")].address}'):$(kubectl -n argocd get svc argocd-server -o jsonpath='{.spec.ports[?(@.port==80)].nodePort}')
$ argocd login "$ARGO" --username admin --password "$(kubectl -n argocd get secret argocd-initial-admin-secret -o jsonpath='{.data.password}' | base64 -d)" --plaintext --grpc-web
'admin:login' logged in successfully
Context '172.18.0.4:32015' updated
$ kubectl apply -f - <<'EOF'
apiVersion: kyverno.io/v1
kind: ClusterPolicy
metadata: { name: require-owner }
spec:
  validationFailureAction: Enforce
  rules:
    - name: require-owner-label
      match:
        any:
          - resources:
              kinds: [Deployment]
              namespaces: [team-a, flux-demo]
      validate:
        message: "every Deployment needs an owner label"
        pattern:
          metadata:
            labels:
              owner: "?*"
EOF
Warning: kyverno.io/v1 ClusterPolicy is deprecated and will be removed in a future release; migrate to ValidatingPolicy, MutatingPolicy, GeneratingPolicy or ImageValidatingPolicy (policies.kyverno.io), see https://kyverno.io/docs/guides/migration-to-cel/
clusterpolicy.kyverno.io/require-owner created
$ sleep 10
$ # an unchanged resource is never sent to admission, so each engine has to create one
$ kubectl -n team-a delete deploy staging-demo --ignore-not-found
$ kubectl -n flux-demo delete deploy demo --ignore-not-found
deployment.apps "demo" deleted from flux-demo namespace
$ # this app syncs automatically, so wait out the controller's own operation first
$ argocd app wait demo-staging --operation --timeout 180 || true

Name:               argocd/demo-staging
Project:            default
Server:             https://kubernetes.default.svc
Namespace:          staging
URL:                https://argocd.example.com/applications/demo-staging
Source:
- Repo:             http://gitea.lab:3000/lab/platform.git
  Target:           main
  Path:             demo-app/overlays/staging
SyncWindow:         Sync Allowed
Sync Policy:        Automated (Prune)
Sync Status:        OutOfSync from main (83a41e2)
Health Status:      Healthy

Operation:          Sync
Sync Revision:      83a41e21322e15eff2307bc1e9d5c89e33226d0d
Phase:              Failed
Start:              2026-09-13 01:28:53 -0400 EDT
Finished:           2026-09-13 01:34:21 -0400 EDT
Duration:           5m28s
Message:            one or more objects failed to apply, reason: Internal error occurred: failed calling webhook "mpol.validate.kyverno.svc-fail": failed to call webhook: Post "https://kyverno-svc.kyverno.svc:443/mpol/add-cost-centre-label?timeout=10s": dial tcp 10.96.38.245:443: connect: connection refused (retried 5 times).

GROUP  KIND        NAMESPACE  NAME          STATUS     HEALTH   HOOK  MESSAGE
apps   Deployment  team-a     staging-demo  OutOfSync  Missing        Internal error occurred: failed calling webhook "mpol.validate.kyverno.svc-fail": failed to call webhook: Post "https://kyverno-svc.kyverno.svc:443/mpol/add-cost-centre-label?timeout=10s": dial tcp 10.96.38.245:443: connect: connection refused
       Service     team-a     staging-demo  Synced     Healthy        
$ argocd app sync demo-staging --timeout 60 || true
TIMESTAMP                  GROUP        KIND   NAMESPACE                  NAME    STATUS    HEALTH        HOOK  MESSAGE
2026-09-13T01:34:57-04:00            Service      team-a          staging-demo    Synced   Healthy              
2026-09-13T01:34:57-04:00   apps  Deployment      team-a          staging-demo  OutOfSync  Missing              
2026-09-13T01:35:03-04:00            Service      team-a          staging-demo    Synced   Healthy              service/staging-demo unchanged
2026-09-13T01:35:03-04:00   apps  Deployment      team-a          staging-demo  OutOfSync  Missing              admission webhook "validate.kyverno.svc-fail" denied the request: 

resource Deployment/team-a/staging-demo was blocked due to the following policies 

require-owner:
  require-owner-label: 'validation error: every Deployment needs an owner label. rule require-owner-label failed at path /metadata/labels/owner/'

Name:               argocd/demo-staging
Project:            default
Server:             https://kubernetes.default.svc
Namespace:          staging
URL:                https://argocd.example.com/applications/demo-staging
Source:
- Repo:             http://gitea.lab:3000/lab/platform.git
  Target:           main
  Path:             demo-app/overlays/staging
SyncWindow:         Sync Allowed
Sync Policy:        Automated (Prune)
Sync Status:        OutOfSync from main (83a41e2)
Health Status:      Healthy

Operation:          Sync
Sync Revision:      83a41e21322e15eff2307bc1e9d5c89e33226d0d
Phase:              Failed
Start:              2026-09-13 01:34:57 -0400 EDT
Finished:           2026-09-13 01:35:03 -0400 EDT
Duration:           6s
Message:            one or more synchronization tasks completed unsuccessfully, reason: admission webhook "validate.kyverno.svc-fail" denied the request: 

resource Deployment/team-a/staging-demo was blocked due to the following policies 

require-owner:
  require-owner-label: 'validation error: every Deployment needs an owner label. rule require-owner-label failed at path /metadata/labels/owner/'

GROUP  KIND        NAMESPACE  NAME          STATUS     HEALTH   HOOK  MESSAGE
       Service     team-a     staging-demo  Synced     Healthy        service/staging-demo unchanged
... 7 more lines
$ argocd app wait demo-staging --operation --timeout 180 || true

Name:               argocd/demo-staging
Project:            default
Server:             https://kubernetes.default.svc
Namespace:          staging
URL:                https://argocd.example.com/applications/demo-staging
Source:
- Repo:             http://gitea.lab:3000/lab/platform.git
  Target:           main
  Path:             demo-app/overlays/staging
SyncWindow:         Sync Allowed
Sync Policy:        Automated (Prune)
Sync Status:        OutOfSync from main (83a41e2)
Health Status:      Healthy

Operation:          Sync
Sync Revision:      83a41e21322e15eff2307bc1e9d5c89e33226d0d
Phase:              Failed
Start:              2026-09-13 01:34:57 -0400 EDT
Finished:           2026-09-13 01:35:03 -0400 EDT
Duration:           6s
Message:            one or more synchronization tasks completed unsuccessfully, reason: admission webhook "validate.kyverno.svc-fail" denied the request: 

resource Deployment/team-a/staging-demo was blocked due to the following policies 

require-owner:
  require-owner-label: 'validation error: every Deployment needs an owner label. rule require-owner-label failed at path /metadata/labels/owner/'

GROUP  KIND        NAMESPACE  NAME          STATUS     HEALTH   HOOK  MESSAGE
       Service     team-a     staging-demo  Synced     Healthy        service/staging-demo unchanged
apps   Deployment  team-a     staging-demo  OutOfSync  Missing        admission webhook "validate.kyverno.svc-fail" denied the request: 

resource Deployment/team-a/staging-demo was blocked due to the following policies 

require-owner:
  require-owner-label: 'validation error: every Deployment needs an owner label. rule require-owner-label failed at path /metadata/labels/owner/'
$ kubectl -n argocd get app demo-staging -o jsonpath='{.status.operationState.phase}{"\n"}{.status.operationState.message}{"\n"}'
Failed
one or more synchronization tasks completed unsuccessfully, reason: admission webhook "validate.kyverno.svc-fail" denied the request: 

resource Deployment/team-a/staging-demo was blocked due to the following policies 

require-owner:
  require-owner-label: 'validation error: every Deployment needs an owner label. rule require-owner-label failed at path /metadata/labels/owner/'
$ flux reconcile kustomization demo-flux --with-source || true
► annotating GitRepository platform in flux-system namespace
✔ GitRepository annotated
◎ waiting for GitRepository reconciliation
✔ fetched revision main@sha1:83a41e21322e15eff2307bc1e9d5c89e33226d0d
► annotating Kustomization demo-flux in flux-system namespace
✔ Kustomization annotated
◎ waiting for Kustomization reconciliation
✗ context deadline exceeded
$ kubectl -n flux-system get kustomization demo-flux -o jsonpath='{.status.conditions}' | jq '.[] | {type, status, reason, message}'
{
  "type": "Reconciling",
  "status": "True",
  "reason": "ProgressingWithRetry",
  "message": "Detecting drift for revision main@sha1:83a41e21322e15eff2307bc1e9d5c89e33226d0d with a timeout of 30s"
}
{
  "type": "Ready",
  "status": "False",
  "reason": "ReconciliationFailed",
  "message": "Deployment/flux-demo/demo dry-run failed: admission webhook \"validate.kyverno.svc-fail\" denied the request: \n\nresource Deployment/flux-demo/demo was blocked due to the following policies \n\nrequire-owner:\n  require-owner-label: 'validation error: every Deployment needs an owner label. rule require-owner-label failed at path /metadata/labels/owner/'\n\n"
}
$ kubectl delete clusterpolicy require-owner
Warning: kyverno.io/v1 ClusterPolicy is deprecated and will be removed in a future release; migrate to ValidatingPolicy, MutatingPolicy, GeneratingPolicy or ImageValidatingPolicy (policies.kyverno.io), see https://kyverno.io/docs/guides/migration-to-cel/
clusterpolicy.kyverno.io "require-owner" deleted
$ # this app syncs automatically, so wait out the controller's own operation first
$ argocd app wait demo-staging --operation --timeout 180 || true

Name:               argocd/demo-staging
Project:            default
Server:             https://kubernetes.default.svc
Namespace:          staging
URL:                https://argocd.example.com/applications/demo-staging
Source:
- Repo:             http://gitea.lab:3000/lab/platform.git
  Target:           main
  Path:             demo-app/overlays/staging
SyncWindow:         Sync Allowed
Sync Policy:        Automated (Prune)
Sync Status:        OutOfSync from main (83a41e2)
Health Status:      Healthy

Operation:          Sync
Sync Revision:      83a41e21322e15eff2307bc1e9d5c89e33226d0d
Phase:              Failed
Start:              2026-09-13 01:34:57 -0400 EDT
Finished:           2026-09-13 01:35:03 -0400 EDT
Duration:           6s
Message:            one or more synchronization tasks completed unsuccessfully, reason: admission webhook "validate.kyverno.svc-fail" denied the request: 

resource Deployment/team-a/staging-demo was blocked due to the following policies 

require-owner:
  require-owner-label: 'validation error: every Deployment needs an owner label. rule require-owner-label failed at path /metadata/labels/owner/'

GROUP  KIND        NAMESPACE  NAME          STATUS     HEALTH   HOOK  MESSAGE
       Service     team-a     staging-demo  Synced     Healthy        service/staging-demo unchanged
apps   Deployment  team-a     staging-demo  OutOfSync  Missing        admission webhook "validate.kyverno.svc-fail" denied the request: 

resource Deployment/team-a/staging-demo was blocked due to the following policies 

require-owner:
  require-owner-label: 'validation error: every Deployment needs an owner label. rule require-owner-label failed at path /metadata/labels/owner/'
$ argocd app sync demo-staging --timeout 120
TIMESTAMP                  GROUP        KIND   NAMESPACE                  NAME    STATUS    HEALTH        HOOK  MESSAGE
2026-09-13T01:40:10-04:00            Service      team-a          staging-demo    Synced   Healthy              
2026-09-13T01:40:10-04:00   apps  Deployment      team-a          staging-demo  OutOfSync  Missing              
2026-09-13T01:40:11-04:00   apps  Deployment      team-a          staging-demo  OutOfSync  Progressing              deployment.apps/staging-demo created
2026-09-13T01:40:11-04:00            Service      team-a          staging-demo    Synced   Healthy                  service/staging-demo unchanged
2026-09-13T01:40:12-04:00   apps  Deployment      team-a          staging-demo    Synced  Progressing              deployment.apps/staging-demo created

Name:               argocd/demo-staging
Project:            default
Server:             https://kubernetes.default.svc
Namespace:          staging
URL:                https://argocd.example.com/applications/demo-staging
Source:
- Repo:             http://gitea.lab:3000/lab/platform.git
  Target:           main
  Path:             demo-app/overlays/staging
SyncWindow:         Sync Allowed
Sync Policy:        Automated (Prune)
Sync Status:        Synced to main (83a41e2)
Health Status:      Progressing

Operation:          Sync
Sync Revision:      83a41e21322e15eff2307bc1e9d5c89e33226d0d
Phase:              Succeeded
Start:              2026-09-13 01:40:10 -0400 EDT
Finished:           2026-09-13 01:40:11 -0400 EDT
Duration:           1s
Message:            successfully synced (all tasks run)

GROUP  KIND        NAMESPACE  NAME          STATUS  HEALTH       HOOK  MESSAGE
       Service     team-a     staging-demo  Synced  Healthy            service/staging-demo unchanged
apps   Deployment  team-a     staging-demo  Synced  Progressing        deployment.apps/staging-demo created
verify: Argo CD reports the operation Failed quoting the admission webhook by name, and Flux reports Ready False with a reconciliation failure quoting the same webhook. Same cause, two vocabularies; be able to translate either into "a policy refused the manifest".

A webhook with failurePolicy: Fail and no running backend does not fail open, it fails everything. This is the outage that looks like the API server broke, and it is one scale command away in either direction.

kubectl get validatingwebhookconfiguration -o custom-columns=NAME:.metadata.name,POLICY:.webhooks[*].failurePolicy | grep -i kyverno
# scaling Kyverno to zero does not reproduce this on this install: its webhooks still admit.
# own the webhook instead, pointed at a Service with no backend, scoped to team-a.
kubectl apply -f - <<'EOF'
apiVersion: admissionregistration.k8s.io/v1
kind: ValidatingWebhookConfiguration
metadata: { name: no-endpoints }
webhooks:
  - name: no-endpoints.lab.local
    admissionReviewVersions: ["v1"]
    sideEffects: None
    failurePolicy: Fail
    timeoutSeconds: 5
    namespaceSelector:
      matchLabels: { kubernetes.io/metadata.name: team-a }
    clientConfig:
      service: { name: absent, namespace: team-a, path: /validate, port: 443 }
    rules:
      - { apiGroups: ["apps"], apiVersions: ["v1"], operations: ["CREATE","UPDATE"], resources: ["deployments"], scope: Namespaced }
EOF
sleep 5
kubectl -n team-a create deployment probe --image=nginx:1.27-alpine
# an unchanged resource is never sent to admission, so make the sync create one
kubectl -n team-a delete deploy staging-demo --ignore-not-found
# this app syncs automatically, so wait out the controller's own operation first
argocd app wait demo-staging --operation --timeout 180 || true
argocd app sync demo-staging --timeout 60 || true
kubectl -n argocd get app demo-staging -o jsonpath='{.status.operationState.phase}{"\n"}{.status.operationState.message}{"\n"}'
kubectl delete validatingwebhookconfiguration no-endpoints
kubectl -n team-a delete deployment probe --ignore-not-found
# this app syncs automatically, so wait out the controller's own operation first
argocd app wait demo-staging --operation --timeout 180 || true
argocd app sync demo-staging --timeout 120
outputcaptured 2026-09-12
$ kubectl get validatingwebhookconfiguration -o custom-columns=NAME:.metadata.name,POLICY:.webhooks[*].failurePolicy | grep -i kyverno
kyverno-cel-exception-validating-webhook-cfg    Fail
kyverno-cleanup-validating-webhook-cfg          Fail
kyverno-exception-validating-webhook-cfg        Fail
kyverno-global-context-validating-webhook-cfg   Fail
kyverno-policy-validating-webhook-cfg           Fail
kyverno-resource-validating-webhook-cfg         Fail
kyverno-ttl-validating-webhook-cfg              Ignore
$ # scaling Kyverno to zero does not reproduce this on this install: its webhooks still admit.
$ # own the webhook instead, pointed at a Service with no backend, scoped to team-a.
$ kubectl apply -f - <<'EOF'
apiVersion: admissionregistration.k8s.io/v1
kind: ValidatingWebhookConfiguration
metadata: { name: no-endpoints }
webhooks:
  - name: no-endpoints.lab.local
    admissionReviewVersions: ["v1"]
    sideEffects: None
    failurePolicy: Fail
    timeoutSeconds: 5
    namespaceSelector:
      matchLabels: { kubernetes.io/metadata.name: team-a }
    clientConfig:
      service: { name: absent, namespace: team-a, path: /validate, port: 443 }
    rules:
      - { apiGroups: ["apps"], apiVersions: ["v1"], operations: ["CREATE","UPDATE"], resources: ["deployments"], scope: Namespaced }
EOF
validatingwebhookconfiguration.admissionregistration.k8s.io/no-endpoints created
$ sleep 5
$ kubectl -n team-a create deployment probe --image=nginx:1.27-alpine
Warning: would violate PodSecurity "restricted:latest": allowPrivilegeEscalation != false (container "nginx" must set securityContext.allowPrivilegeEscalation=false), unrestricted capabilities (container "nginx" must set securityContext.capabilities.drop=["ALL"]), runAsNonRoot != true (pod or container "nginx" must set securityContext.runAsNonRoot=true), seccompProfile (pod or container "nginx" must set securityContext.seccompProfile.type to "RuntimeDefault" or "Localhost")
error: failed to create deployment: Internal error occurred: failed calling webhook "no-endpoints.lab.local": failed to call webhook: Post "https://absent.team-a.svc:443/validate?timeout=5s": service "absent" not found
$ # an unchanged resource is never sent to admission, so make the sync create one
$ kubectl -n team-a delete deploy staging-demo --ignore-not-found
deployment.apps "staging-demo" deleted from team-a namespace
$ # this app syncs automatically, so wait out the controller's own operation first
$ argocd app wait demo-staging --operation --timeout 180 || true

Name:               argocd/demo-staging
Project:            default
Server:             https://kubernetes.default.svc
Namespace:          staging
URL:                https://argocd.example.com/applications/demo-staging
Source:
- Repo:             http://gitea.lab:3000/lab/platform.git
  Target:           main
  Path:             demo-app/overlays/staging
SyncWindow:         Sync Allowed
Sync Policy:        Automated (Prune)
Sync Status:        Synced to main (b793ef2)
Health Status:      Healthy

Operation:          Sync
Sync Revision:      b793ef22cee9cf7bc5eb3d41e236dda5c43670fb
Phase:              Succeeded
Start:              2026-09-13 12:44:52 -0400 EDT
Finished:           2026-09-13 12:45:02 -0400 EDT
Duration:           10s
Message:            successfully synced (all tasks run)

GROUP  KIND        NAMESPACE  NAME          STATUS  HEALTH   HOOK  MESSAGE
apps   Deployment  team-a     staging-demo  Synced  Missing        deployment.apps/staging-demo created
       Service     team-a     staging-demo  Synced  Healthy        
$ argocd app sync demo-staging --timeout 60 || true
{"level":"fatal","msg":"rpc error: code = FailedPrecondition desc = another operation is already in progress","time":"2026-09-13T12:50:07-04:00"}
$ kubectl -n argocd get app demo-staging -o jsonpath='{.status.operationState.phase}{"\n"}{.status.operationState.message}{"\n"}'
Running
$ kubectl delete validatingwebhookconfiguration no-endpoints
validatingwebhookconfiguration.admissionregistration.k8s.io "no-endpoints" deleted
$ kubectl -n team-a delete deployment probe --ignore-not-found
$ # this app syncs automatically, so wait out the controller's own operation first
$ argocd app wait demo-staging --operation --timeout 180 || true
TIMESTAMP                  GROUP        KIND   NAMESPACE                  NAME    STATUS    HEALTH        HOOK  MESSAGE
2026-09-13T12:50:07-04:00   apps  Deployment      team-a          staging-demo  OutOfSync  Missing              Internal error occurred: failed calling webhook "no-endpoints.lab.local": failed to call webhook: Post "https://absent.team-a.svc:443/validate?timeout=5s": service "absent" not found
2026-09-13T12:50:07-04:00            Service      team-a          staging-demo    Synced   Healthy              

Name:               argocd/demo-staging
Project:            default
Server:             https://kubernetes.default.svc
Namespace:          staging
URL:                https://argocd.example.com/applications/demo-staging
Source:
- Repo:             http://gitea.lab:3000/lab/platform.git
  Target:           main
  Path:             demo-app/overlays/staging
SyncWindow:         Sync Allowed
Sync Policy:        Automated (Prune)
Sync Status:        Synced to main (b793ef2)
Health Status:      Progressing

Operation:          Sync
Sync Revision:      b793ef22cee9cf7bc5eb3d41e236dda5c43670fb
Phase:              Succeeded
Start:              2026-09-13 12:50:07 -0400 EDT
Finished:           2026-09-13 12:50:17 -0400 EDT
Duration:           10s
Message:            successfully synced (all tasks run)

GROUP  KIND        NAMESPACE  NAME          STATUS  HEALTH       HOOK  MESSAGE
apps   Deployment  team-a     staging-demo  Synced  Progressing        deployment.apps/staging-demo created
       Service     team-a     staging-demo  Synced  Healthy            
$ argocd app sync demo-staging --timeout 120
TIMESTAMP                  GROUP        KIND   NAMESPACE                  NAME    STATUS   HEALTH            HOOK  MESSAGE
2026-09-13T12:50:18-04:00            Service      team-a          staging-demo    Synced  Healthy                  
2026-09-13T12:50:18-04:00   apps  Deployment      team-a          staging-demo    Synced  Progressing              
2026-09-13T12:50:19-04:00            Service      team-a          staging-demo    Synced  Healthy                  service/staging-demo unchanged
2026-09-13T12:50:19-04:00   apps  Deployment      team-a          staging-demo    Synced  Progressing              deployment.apps/staging-demo unchanged

Name:               argocd/demo-staging
Project:            default
Server:             https://kubernetes.default.svc
Namespace:          staging
URL:                https://argocd.example.com/applications/demo-staging
Source:
- Repo:             http://gitea.lab:3000/lab/platform.git
  Target:           main
  Path:             demo-app/overlays/staging
SyncWindow:         Sync Allowed
Sync Policy:        Automated (Prune)
Sync Status:        Synced to main (b793ef2)
Health Status:      Progressing

Operation:          Sync
Sync Revision:      b793ef22cee9cf7bc5eb3d41e236dda5c43670fb
Phase:              Succeeded
Start:              2026-09-13 12:50:18 -0400 EDT
Finished:           2026-09-13 12:50:19 -0400 EDT
Duration:           1s
Message:            successfully synced (all tasks run)

GROUP  KIND        NAMESPACE  NAME          STATUS  HEALTH       HOOK  MESSAGE
       Service     team-a     staging-demo  Synced  Healthy            service/staging-demo unchanged
apps   Deployment  team-a     staging-demo  Synced  Progressing        deployment.apps/staging-demo unchanged
verify: the plain kubectl create deployment fails with Internal error occurred: failed calling webhook "no-endpoints.lab.local" ... service "absent" not found, and Argo CD hits the same wall: the wait shows staging-demo OutOfSync Missing carrying that message, because this app self-heals and the controller tried the write before you did. That is also why the manual sync comes back another operation is already in progress rather than with the webhook error itself. Two things have to be true for this outage: failurePolicy: Fail, and a rule matching what is being written, which is why the Deployment is deleted first. Scaling the Kyverno admission controller to zero does not reproduce it on this install.

Every failure on this page shows up as an Application condition, and the type is a faster triage signal than the message. Run through the three buckets and write the type next to each one. Expect the list to be empty here: conditions describe what is wrong now, and by this point in the page everything has been put back.

kubectl -n argocd get app demo-staging -o jsonpath='{range .status.conditions[*]}{.type}: {.message}{"\n"}{end}'
kubectl -n argocd get app -o custom-columns=NAME:.metadata.name,SYNC:.status.sync.status,HEALTH:.status.health.status,CONDITIONS:.status.conditions[*].type
kubectl explain application.status.conditions.type
outputcaptured 2026-09-12
$ kubectl -n argocd get app demo-staging -o jsonpath='{range .status.conditions[*]}{.type}: {.message}{"\n"}{end}'
$ kubectl -n argocd get app -o custom-columns=NAME:.metadata.name,SYNC:.status.sync.status,HEALTH:.status.health.status,CONDITIONS:.status.conditions[*].type
NAME           SYNC     HEALTH    CONDITIONS
demo-prod      Synced   Healthy   <none>
demo-staging   Synced   Healthy   <none>
payments-dev   Synced   Healthy   <none>
$ kubectl explain application.status.conditions.type
GROUP:      argoproj.io
KIND:       Application
VERSION:    v1alpha1

FIELD: type <string>


DESCRIPTION:
    Type is an application condition type
    
verify: the conditions list is empty on all three Applications, because nothing is failing at this point in the page. That is the exercise: you should be able to name the condition type each earlier failure produced without one in front of you, for a repo the project forbids, for a manifest admission refuses, and for a controller whose credentials are wrong. Three buckets, three types, and only one of them is about the cluster.

Self-check

bucket each of these in one sentence
"The app is Synced and Healthy but the feature we merged an hour ago is missing."

Bucket 4, stale. Check for suspension, then the last sync timestamp and the resolved revision against your commit SHA, then whether the app tracks the branch you pushed to. Refresh/reconcile before anything else.

"Sync fails: Service "demo" is invalid: spec.clusterIP: Invalid value: field is immutable."

Bucket 2: the cluster refused the manifest. The text came from the API server, relayed by the sync operation. Fix the manifest, or force replace semantics if the field genuinely must change (accepting the churn that implies).

"argocd-repo-server logs show authentication required; the workload is untouched and running fine."

Bucket 3: access. No pod of yours is implicated; the controller's identity is. Repo credentials, token scope, or a rotated secret. The workload keeps running because reconciliation is what broke, not the deployment.

"Deployment exists, no pods, no error on the Deployment."

Look one level down at the ReplicaSet: quota, PSS, or an admission webhook rejected the pod template. Same indirection as sections 1.4 and 5.3. Not strictly a delivery bug: the delivery worked, admission refused.

"I keep deleting this ConfigMap and it keeps coming back."

Something reconciles it: self-heal, a Flux Kustomization at interval, or an operator that owns it. Prove which with ownerReferences, the app.kubernetes.io/instance label, or flux trace. Then change the source of truth instead of the live object, or suspend first if you need a temporary window.

"Sync failed: admission webhook "validate.kyverno.svc-fail" denied the request". Which bucket, and where does the fix go?

Bucket 5, admission. The manifest reached the API server and a policy engine refused it. Fix either the manifest so it satisfies the policy (the usual grading intent) or, if the task says the policy is wrong, the policy itself. Re-syncing changes nothing; Argo CD marks the resource SyncFailed, Flux reports ReconciliationFailed with the same text.

A Flux Kustomization reports ArtifactFailed. What is broken and where do you look first?

The source it references has no artifact: the GitRepository or OCIRepository is not Ready (AuthenticationFailed, GitOperationFailed, bad ref) or has not fetched yet. Bucket 3, upstream of the Kustomization. flux get sources all -A and kubectl describe gitrepository first; the Kustomization will recover on its own once the source publishes.

Argo CD shows sync status Unknown with a ComparisonError condition. Bucket?

Bucket 3: comparison itself failed, so Argo does not even know whether live matches git. Causes are repo access (credentials, URL, revision), a render error in repo-server, or an unreachable destination cluster. The condition message names it; the workload is untouched, which is the tell.

Docs to know your way around

study time, not exam time
  • argo-cd.readthedocs.io: sync options (Replace, ServerSideApply), resource health checks, the troubleshooting section.
  • fluxcd.io: the troubleshooting cheatsheet (flux logs, flux events, tracing a resource to its Kustomization).
  • Offline: argocd app manifests, argocd app diff, flux build kustomization <name> --path ./…, and kubectl get events -A --sort-by=.lastTimestamp.
  • argo-cd.readthedocs.io: Application specification (status fields) and the gitops-engine health package: the exact sync, health, operation phase and condition strings.
  • fluxcd.io troubleshooting cheatsheet ("webhook does not support dry run") and each component's Conditions section: the reason strings and the Flux-specific admission failure.