Broken delivery is the most likely domain 2 exam task. Failures sort into four buckets, and the first diagnostic move is deciding which bucket you are in. Everything here is practice for that decision.
make coreOrientation
Under time pressure, the expensive mistake is not a wrong fix; it is fixing in the wrong layer. Bucket first, then descend. Say the bucket out loud before you touch anything; it costs three seconds and it stops you editing live state when the problem is a commit.
argocd app get <app> # sync status + health + last sync time + conditions
flux get kustomizations -A # Ready / Suspended / revision
kubectl -n <ns> get pods # is anything even running
kubectl -n <ns> get events --sort-by=.lastTimestamp | tail -20outputcaptured 2026-08-26
$ argocd app get demo-staging # sync status + health + last sync time + conditions
Name: argocd/demo-staging
Project: default
Server: https://kubernetes.default.svc
Namespace: staging
URL: https://argocd.example.com/applications/demo-staging
Source:
- Repo: http://gitea.lab:3000/lab/platform.git
Target: main
Path: demo-app/overlays/staging
SyncWindow: Sync Allowed
Sync Policy: Automated (Prune)
Sync Status: Synced to main (c7cb9a1)
Health Status: Healthy
GROUP KIND NAMESPACE NAME STATUS HEALTH HOOK MESSAGE
apps Deployment team-a staging-demo Synced Healthy deployment.apps/staging-demo configured
Service team-a staging-demo Synced Healthy
$ flux get kustomizations -A # Ready / Suspended / revision
NAMESPACE NAME REVISION SUSPENDED READY MESSAGE
flux-system demo-flux main@sha1:c7cb9a1e False True Applied revision: main@sha1:c7cb9a1e
$ kubectl -n team-a get pods # is anything even running
NAME READY STATUS RESTARTS AGE
staging-demo-66955f8974-2xxd2 1/1 Running 0 4m21s
staging-demo-66955f8974-pzxdh 1/1 Running 0 4m21s
staging-demo-66955f8974-tzpqx 1/1 Running 0 4m22s
$ kubectl -n team-a get events --sort-by=.lastTimestamp | tail -20
4m16s Normal Created pod/staging-demo-66955f8974-2xxd2 Container created
4m15s Normal Created pod/staging-demo-66955f8974-tzpqx Container created
4m15s Normal Started pod/staging-demo-66955f8974-pzxdh Container started
4m15s Normal Started pod/staging-demo-66955f8974-tzpqx Container started
4m1s Normal SuccessfulCreate replicaset/staging-demo-66955f8974 Created pod: staging-demo-66955f8974-7xnfb
4m1s Normal ScalingReplicaSet deployment/staging-demo Scaled up replica set staging-demo-66955f8974 from 3 to 5
4m1s Normal Scheduled pod/staging-demo-66955f8974-7xnfb Successfully assigned team-a/staging-demo-66955f8974-7xnfb to cnpe-worker
4m1s Normal SuccessfulCreate replicaset/staging-demo-66955f8974 Created pod: staging-demo-66955f8974-jh7rr
4m Normal Scheduled pod/staging-demo-66955f8974-jh7rr Successfully assigned team-a/staging-demo-66955f8974-jh7rr to cnpe-worker2
3m59s Normal SuccessfulDelete replicaset/staging-demo-66955f8974 Deleted pod: staging-demo-66955f8974-7xnfb
3m59s Normal ScalingReplicaSet deployment/staging-demo Scaled down replica set staging-demo-66955f8974 from 5 to 3
3m59s Normal SuccessfulDelete replicaset/staging-demo-66955f8974 Deleted pod: staging-demo-66955f8974-jh7rr
3m58s Normal Pulled pod/staging-demo-66955f8974-7xnfb Container image "ghcr.io/nginxinc/nginx-unprivileged:1.27-alpine" already present on machine and can be accessed by the pod
3m58s Normal Pulled pod/staging-demo-66955f8974-jh7rr Container image "ghcr.io/nginxinc/nginx-unprivileged:1.27-alpine" already present on machine and can be accessed by the pod
3m57s Normal Killing pod/staging-demo-66955f8974-jh7rr Stopping container web
3m57s Normal Started pod/staging-demo-66955f8974-jh7rr Container started
3m57s Normal Created pod/staging-demo-66955f8974-jh7rr Container created
3m57s Normal Killing pod/staging-demo-66955f8974-7xnfb Stopping container web
3m57s Normal Started pod/staging-demo-66955f8974-7xnfb Container started
3m57s Normal Created pod/staging-demo-66955f8974-7xnfb Container createdThose four commands place you in one of the four buckets below almost every time.
The four buckets
1 · Git is wrong, the cluster is faithful
Synced + Degraded. Bad image tag, impossible resource request, missing ConfigMap key, a probe that can never pass. The controller did its job; fix the commit. Evidence: app health Degraded (or Progressing on its way there) while sync status is green, and pod events carry the real error: ImagePullBackOff, CreateContainerConfigError, CrashLoopBackOff.
2 · The cluster refuses what git says
Sync fails outright. An admission policy denies the manifest, the API version does not exist on this cluster, a field is immutable (Service clusterIP, label selectors, PVC shrink), the CRD is not installed yet. Evidence: the sync operation's error message, which quotes the API server's rejection verbatim. Immutable-field errors mean delete-and-recreate or Replace=true.
3 · The controller lacks permission or access
Repo unreachable, credentials rotated, an Application targeting a namespace the controller's RBAC cannot touch, an ApplicationSet token missing a scope. Evidence: errors mention the controller's own identity, appear in controller logs (kubectl -n argocd logs deploy/argocd-repo-server, flux logs), and say nothing about your workload. No pod is ever involved; that absence is the tell.
4 · Nobody is wrong, the state is stale
Webhook lost, refresh interval not elapsed, reconciliation suspended and forgotten, a branch that moved while the app points at a tag. Evidence: everything green but old. flux get kustomizations shows a Suspended row; argocd app get shows a last-sync timestamp from an hour ago. Fix with argocd app get --refresh / flux reconcile … --with-source, and check for suspension first.
| Bucket | Sync | Health | Where the text comes from | Fix lives in |
|---|---|---|---|---|
| 1 git wrong | Synced | Degraded | pod events | a commit |
| 2 cluster refuses | Failed / OutOfSync | n/a | API server, quoted by the sync op | a commit + a sync option |
| 3 access | Unknown / error | n/a | controller logs | a Secret, RBAC, or a token |
| 4 stale | Synced | Healthy | timestamps, Suspended flag | a reconcile or a resume |
With self-heal on, drift self-corrects and the interesting question becomes "what keeps re-creating this thing I keep deleting". The answer is the controller, and kubectl get <res> -o jsonpath='{.metadata.ownerReferences}' or the app.kubernetes.io/instance label proves it. With self-heal off, argocd app diff is the tool that shows exactly what diverged.
The descent, one level at a time
Application / Kustomization status.conditions, last sync, revision
│
▼
rendered manifests argocd app manifests · kustomize build · flux build
│
▼
the apply sync operation message, API server rejection text
│
▼
controller objects Deployment → ReplicaSet → Pod ← quota and PSS errors live on the RS
│
▼
pod describe → events → logs → logs --previous → exec/debug
│
▼
prove recovery with the signal that showed the failure
Two habits that pay for themselves. First, compare rendered manifests with live objects rather than reading source YAML: argocd app manifests <app> shows exactly what would be applied, which catches overlay mistakes that source review misses. Second, check the middle layer: a Deployment that is created but produces no pods has its error on the ReplicaSet, and that indirection is easy to miss.
The status vocabulary, engine by engine
Bucketing is faster when you read the exact word rather than the color. Both engines write stable strings; a scenario quotes them.
Argo CD
| Field | Value | Bucket |
|---|---|---|
| status.sync.status | OutOfSync | 1 if auto-sync is off and git moved; 4 if auto-sync should have acted; check syncPolicy and sync windows |
| status.sync.status | Unknown | 3: comparison failed; read conditions |
| status.health.status | Degraded · Missing | 1: the workload cannot run, or a desired resource is absent (often pruned by another app, or refused: bucket 5) |
| status.health.status | Progressing (for long) | 1, or a CRD with no health check that never settles |
| status.health.status | Suspended | a paused Deployment, suspended CronJob or paused Rollout: intentional, not a fault |
| status.operationState.phase | Failed | 2 or 5: the API server rejected an apply; the text is in syncResult.resources[].message |
| status.operationState.phase | Error | 3: Argo could not run the operation (render, repo, hook creation) |
| conditions[].type | ComparisonError | 3: repo unreachable, bad credentials, render failure, cluster unreachable |
| conditions[].type | InvalidSpecError | 3: the AppProject forbids the repo, destination or kind; or the spec references a missing project |
| conditions[].type | SyncError | 2 or 5 |
| conditions[].type | SharedResourceWarning · RepeatedResourceWarning · OrphanedResourceWarning · ExcludedResourceWarning | warnings: two apps own one resource; one source yields a resource twice; untracked resources in the namespace; a kind excluded by resource.exclusions |
| conditions[].type | DeletionError | a PreDelete or PostDelete hook failed; the Application stays Terminating |
Flux
| Object | Ready reason | Bucket |
|---|---|---|
| GitRepository / OCIRepository | AuthenticationFailed · GitOperationFailed · OCIArtifactPullFailed · VerificationError | 3: credentials, URL, ref, signature |
| Kustomization | ArtifactFailed | 3, upstream: the source has no artifact; fix the source first |
| Kustomization | BuildFailed | 1: bad path, invalid kustomization, missing substituteFrom ConfigMap |
| Kustomization | DependencyNotReady | look at the named dependency, same ladder |
| Kustomization | ReconciliationFailed | 2 or 5: the apply was rejected; the message quotes the API server or webhook |
| Kustomization | HealthCheckFailed | 1: applied, never became healthy within timeout |
| Kustomization | PruneFailed | 2: garbage collection blocked (finalizers, protected resource) |
| HelmRelease | InstallFailed · UpgradeFailed | 1 or 2: Helm's own error; run the descend with kubectl describe hr and flux logs --kind HelmRelease |
| HelmRelease | TestFailed | 1: the chart's tests failed; remediation follows unless ignored |
| HelmRelease | RollbackSucceeded · UninstallSucceeded · RetriesExceeded | remediated: last good revision is running; a new generation or flux reconcile hr --reset tries again |
| any | Reconciling=True, reason ProgressingWithRetry | the last attempt failed and a retry is scheduled; the Ready reason names why |
| any | spec.suspend: true (Suspended column) | 4 |
Argo: argocd app get X then kubectl -n argocd get app X -o jsonpath='{.status.conditions[*].type}'. Flux: flux get all -A --status-selector ready=false then kubectl get kustomization X -n flux-system -o jsonpath='{.status.conditions[?(@.type=="Ready")].reason}'. The reason string is the bucket; only then read the message.
Bucket 5: admission refused it
The four buckets assume the API server is the last word. On a platform with Kyverno, Gatekeeper, ValidatingAdmissionPolicy or an operator's webhook, there is a fifth place a manifest dies, and it produces its own vocabulary.
| Message fragment | Who | Fix lives in |
|---|---|---|
| admission webhook "validate.kyverno.svc-fail" denied the request: ... policy X rule Y | Kyverno validate, Enforce | the manifest (satisfy the rule) or the policy (exclude, Audit); never the controller |
| admission webhook "validation.gatekeeper.sh" denied the request: [constraint-name] ... | Gatekeeper constraint | same choice; enforcementAction: dryrun is the audit mode |
| ValidatingAdmissionPolicy 'x' with binding 'y' denied request: ... | VAP (GA 1.30), MutatingAdmissionPolicy (GA 1.36) for mutations | the policy's CEL or the binding's match; no webhook pod involved |
| failed calling webhook "x": ... connection refused / context deadline exceeded / no endpoints available | a webhook whose backing pod is down, with failurePolicy: Fail | the webhook's Deployment or Service; every matching write is blocked until it is back, including the controller's own |
| dry-run failed ... admission webhook "x" does not support dry run | Flux (server-side dry-run before apply) meets a webhook without sideEffects: None or NoneOnDryRun | the webhook configuration |
| the server could not find the requested resource | a CR whose CRD is not installed yet (Argo CD dry-run) | SkipDryRunOnMissingResource=true or a wave; Flux: dependsOn the CRD Kustomization |
The tell that separates bucket 5 from bucket 2: the rejection names a webhook or policy rather than a field. The controller-side symptom is identical (Argo SyncFailed on the resource, Flux ReconciliationFailed), and for Deployments the refusal can again land one level down on the ReplicaSet when a pod-level policy is violated. A mutating policy that changes what git said (adds labels, injects a sidecar, sets requests) shows up as permanent OutOfSync in Argo CD and as "configured" events every interval in Flux: the answer there is ignoreDifferences or kustomize.toolkit.fluxcd.io/ssa: Merge, not fighting the policy.
A dead webhook with failurePolicy: Fail and a broad match (all pods, all namespaces) stops the GitOps controller from applying the very manifest that would repair it. Diagnose from kubectl get validatingwebhookconfiguration,mutatingwebhookconfiguration and the webhook Service's endpoints; the sanctioned repair is fixing the pod, and the emergency one is switching that webhook to Ignore or narrowing its namespaceSelector, which is itself a change to record in git.
Exercises
Do the diagnosis before the fix.
Push newTag: 9.9.9-nope to the staging overlay in the platform repo. Watch demo-staging stay Synced while health goes Progressing (Degraded arrives only after the Deployment's ten-minute progress deadline; do not wait for it, the pod evidence is immediate). Diagnose down the stack: argocd app get demo-staging → kubectl -n team-a get pods → describe pod shows ImagePullBackOff. Fix in git only.
Add a patch to the staging overlay's kustomization setting spec.clusterIP: 10.96.99.99 on the demo Service (the Service manifest itself lives in base; overlays change it through patches, the standard kustomize idiom). Sync and read the error.
Deleting the repo secret outright would not do it here, and knowing why is the exercise's first half: the lab's repos are public, so anonymous cloning still works. Wrong credentials fail where absent ones would not, because Gitea rejects a bad password even on a public repo. So corrupt the secret:
kubectl -n argocd patch secret gitea-repo -p '{"stringData":{"password":"wrong"}}'
argocd app get demo-staging --refreshoutputcaptured 2026-08-26
$ kubectl -n argocd patch secret gitea-repo -p '{"stringData":{"password":"wrong"}}'
secret/gitea-repo patched
$ argocd app get demo-staging --refresh
Name: argocd/demo-staging
Project: default
Server: https://kubernetes.default.svc
Namespace: staging
URL: https://argocd.example.com/applications/demo-staging
Source:
- Repo: http://gitea.lab:3000/lab/platform.git
Target: main
Path: demo-app/overlays/staging
SyncWindow: Sync Allowed
Sync Policy: Automated (Prune)
Sync Status: Unknown
Health Status: Healthy
CONDITION MESSAGE LAST TRANSITION
ComparisonError Failed to load target state: failed to generate manifest for source 1 of 1: rpc error: code = Unknown desc = failed to list refs: authentication required: Failed to authenticate user
2026-08-26 22:21:09 -0400 EDT
GROUP KIND NAMESPACE NAME STATUS HEALTH HOOK MESSAGE
apps Deployment team-a staging-demo Unknown Healthy deployment.apps/staging-demo configured
Service team-a staging-demo Unknown Healthy The error mentions authentication, not manifests. Restore by re-running make gitops (idempotent) or patching the real token back from .gitea-token.
flux suspend kustomization demo-flux (from 2.3), push any change to demo-app/base in the platform repo, and observe that demo-flux's targets never move and nothing errors (the Argo apps do pick the base change up; only the suspended consumer goes quietly stale). The absence of failure is the symptom.
flux get kustomizations shows Suspended, resume, and the change lands. Train yourself to check suspension first; it costs three seconds.FAULT=config make break corrupts something in team-a's delivery path under a 7-minute clock. Bucket it, fix it, then make break-answer to compare.
An admission policy that the manifests violate stops both engines, but each reports it in its own vocabulary. Learn both strings, because the task will show you one of them and expect you to name the cause.
# the LoadBalancer IP only exists when cloud-provider-kind runs; the NodePort is always there
ARGO=$(kubectl get node cnpe-control-plane -o jsonpath='{.status.addresses[?(@.type=="InternalIP")].address}'):$(kubectl -n argocd get svc argocd-server -o jsonpath='{.spec.ports[?(@.port==80)].nodePort}')
argocd login "$ARGO" --username admin --password "$(kubectl -n argocd get secret argocd-initial-admin-secret -o jsonpath='{.data.password}' | base64 -d)" --plaintext --grpc-web
kubectl apply -f - <<'EOF'
apiVersion: kyverno.io/v1
kind: ClusterPolicy
metadata: { name: require-owner }
spec:
validationFailureAction: Enforce
rules:
- name: require-owner-label
match:
any:
- resources:
kinds: [Deployment]
namespaces: [team-a, flux-demo]
validate:
message: "every Deployment needs an owner label"
pattern:
metadata:
labels:
owner: "?*"
EOF
sleep 10
# an unchanged resource is never sent to admission, so each engine has to create one
kubectl -n team-a delete deploy staging-demo --ignore-not-found
kubectl -n flux-demo delete deploy demo --ignore-not-found
# this app syncs automatically, so wait out the controller's own operation first
argocd app wait demo-staging --operation --timeout 180 || true
argocd app sync demo-staging --timeout 60 || true
argocd app wait demo-staging --operation --timeout 180 || true
kubectl -n argocd get app demo-staging -o jsonpath='{.status.operationState.phase}{"\n"}{.status.operationState.message}{"\n"}'
flux reconcile kustomization demo-flux --with-source || true
kubectl -n flux-system get kustomization demo-flux -o jsonpath='{.status.conditions}' | jq '.[] | {type, status, reason, message}'
kubectl delete clusterpolicy require-owner
# this app syncs automatically, so wait out the controller's own operation first
argocd app wait demo-staging --operation --timeout 180 || true
argocd app sync demo-staging --timeout 120outputcaptured 2026-09-13
$ # the LoadBalancer IP only exists when cloud-provider-kind runs; the NodePort is always there
$ ARGO=$(kubectl get node cnpe-control-plane -o jsonpath='{.status.addresses[?(@.type=="InternalIP")].address}'):$(kubectl -n argocd get svc argocd-server -o jsonpath='{.spec.ports[?(@.port==80)].nodePort}')
$ argocd login "$ARGO" --username admin --password "$(kubectl -n argocd get secret argocd-initial-admin-secret -o jsonpath='{.data.password}' | base64 -d)" --plaintext --grpc-web
'admin:login' logged in successfully
Context '172.18.0.4:32015' updated
$ kubectl apply -f - <<'EOF'
apiVersion: kyverno.io/v1
kind: ClusterPolicy
metadata: { name: require-owner }
spec:
validationFailureAction: Enforce
rules:
- name: require-owner-label
match:
any:
- resources:
kinds: [Deployment]
namespaces: [team-a, flux-demo]
validate:
message: "every Deployment needs an owner label"
pattern:
metadata:
labels:
owner: "?*"
EOF
Warning: kyverno.io/v1 ClusterPolicy is deprecated and will be removed in a future release; migrate to ValidatingPolicy, MutatingPolicy, GeneratingPolicy or ImageValidatingPolicy (policies.kyverno.io), see https://kyverno.io/docs/guides/migration-to-cel/
clusterpolicy.kyverno.io/require-owner created
$ sleep 10
$ # an unchanged resource is never sent to admission, so each engine has to create one
$ kubectl -n team-a delete deploy staging-demo --ignore-not-found
$ kubectl -n flux-demo delete deploy demo --ignore-not-found
deployment.apps "demo" deleted from flux-demo namespace
$ # this app syncs automatically, so wait out the controller's own operation first
$ argocd app wait demo-staging --operation --timeout 180 || true
Name: argocd/demo-staging
Project: default
Server: https://kubernetes.default.svc
Namespace: staging
URL: https://argocd.example.com/applications/demo-staging
Source:
- Repo: http://gitea.lab:3000/lab/platform.git
Target: main
Path: demo-app/overlays/staging
SyncWindow: Sync Allowed
Sync Policy: Automated (Prune)
Sync Status: OutOfSync from main (83a41e2)
Health Status: Healthy
Operation: Sync
Sync Revision: 83a41e21322e15eff2307bc1e9d5c89e33226d0d
Phase: Failed
Start: 2026-09-13 01:28:53 -0400 EDT
Finished: 2026-09-13 01:34:21 -0400 EDT
Duration: 5m28s
Message: one or more objects failed to apply, reason: Internal error occurred: failed calling webhook "mpol.validate.kyverno.svc-fail": failed to call webhook: Post "https://kyverno-svc.kyverno.svc:443/mpol/add-cost-centre-label?timeout=10s": dial tcp 10.96.38.245:443: connect: connection refused (retried 5 times).
GROUP KIND NAMESPACE NAME STATUS HEALTH HOOK MESSAGE
apps Deployment team-a staging-demo OutOfSync Missing Internal error occurred: failed calling webhook "mpol.validate.kyverno.svc-fail": failed to call webhook: Post "https://kyverno-svc.kyverno.svc:443/mpol/add-cost-centre-label?timeout=10s": dial tcp 10.96.38.245:443: connect: connection refused
Service team-a staging-demo Synced Healthy
$ argocd app sync demo-staging --timeout 60 || true
TIMESTAMP GROUP KIND NAMESPACE NAME STATUS HEALTH HOOK MESSAGE
2026-09-13T01:34:57-04:00 Service team-a staging-demo Synced Healthy
2026-09-13T01:34:57-04:00 apps Deployment team-a staging-demo OutOfSync Missing
2026-09-13T01:35:03-04:00 Service team-a staging-demo Synced Healthy service/staging-demo unchanged
2026-09-13T01:35:03-04:00 apps Deployment team-a staging-demo OutOfSync Missing admission webhook "validate.kyverno.svc-fail" denied the request:
resource Deployment/team-a/staging-demo was blocked due to the following policies
require-owner:
require-owner-label: 'validation error: every Deployment needs an owner label. rule require-owner-label failed at path /metadata/labels/owner/'
Name: argocd/demo-staging
Project: default
Server: https://kubernetes.default.svc
Namespace: staging
URL: https://argocd.example.com/applications/demo-staging
Source:
- Repo: http://gitea.lab:3000/lab/platform.git
Target: main
Path: demo-app/overlays/staging
SyncWindow: Sync Allowed
Sync Policy: Automated (Prune)
Sync Status: OutOfSync from main (83a41e2)
Health Status: Healthy
Operation: Sync
Sync Revision: 83a41e21322e15eff2307bc1e9d5c89e33226d0d
Phase: Failed
Start: 2026-09-13 01:34:57 -0400 EDT
Finished: 2026-09-13 01:35:03 -0400 EDT
Duration: 6s
Message: one or more synchronization tasks completed unsuccessfully, reason: admission webhook "validate.kyverno.svc-fail" denied the request:
resource Deployment/team-a/staging-demo was blocked due to the following policies
require-owner:
require-owner-label: 'validation error: every Deployment needs an owner label. rule require-owner-label failed at path /metadata/labels/owner/'
GROUP KIND NAMESPACE NAME STATUS HEALTH HOOK MESSAGE
Service team-a staging-demo Synced Healthy service/staging-demo unchanged
... 7 more lines
$ argocd app wait demo-staging --operation --timeout 180 || true
Name: argocd/demo-staging
Project: default
Server: https://kubernetes.default.svc
Namespace: staging
URL: https://argocd.example.com/applications/demo-staging
Source:
- Repo: http://gitea.lab:3000/lab/platform.git
Target: main
Path: demo-app/overlays/staging
SyncWindow: Sync Allowed
Sync Policy: Automated (Prune)
Sync Status: OutOfSync from main (83a41e2)
Health Status: Healthy
Operation: Sync
Sync Revision: 83a41e21322e15eff2307bc1e9d5c89e33226d0d
Phase: Failed
Start: 2026-09-13 01:34:57 -0400 EDT
Finished: 2026-09-13 01:35:03 -0400 EDT
Duration: 6s
Message: one or more synchronization tasks completed unsuccessfully, reason: admission webhook "validate.kyverno.svc-fail" denied the request:
resource Deployment/team-a/staging-demo was blocked due to the following policies
require-owner:
require-owner-label: 'validation error: every Deployment needs an owner label. rule require-owner-label failed at path /metadata/labels/owner/'
GROUP KIND NAMESPACE NAME STATUS HEALTH HOOK MESSAGE
Service team-a staging-demo Synced Healthy service/staging-demo unchanged
apps Deployment team-a staging-demo OutOfSync Missing admission webhook "validate.kyverno.svc-fail" denied the request:
resource Deployment/team-a/staging-demo was blocked due to the following policies
require-owner:
require-owner-label: 'validation error: every Deployment needs an owner label. rule require-owner-label failed at path /metadata/labels/owner/'
$ kubectl -n argocd get app demo-staging -o jsonpath='{.status.operationState.phase}{"\n"}{.status.operationState.message}{"\n"}'
Failed
one or more synchronization tasks completed unsuccessfully, reason: admission webhook "validate.kyverno.svc-fail" denied the request:
resource Deployment/team-a/staging-demo was blocked due to the following policies
require-owner:
require-owner-label: 'validation error: every Deployment needs an owner label. rule require-owner-label failed at path /metadata/labels/owner/'
$ flux reconcile kustomization demo-flux --with-source || true
► annotating GitRepository platform in flux-system namespace
✔ GitRepository annotated
◎ waiting for GitRepository reconciliation
✔ fetched revision main@sha1:83a41e21322e15eff2307bc1e9d5c89e33226d0d
► annotating Kustomization demo-flux in flux-system namespace
✔ Kustomization annotated
◎ waiting for Kustomization reconciliation
✗ context deadline exceeded
$ kubectl -n flux-system get kustomization demo-flux -o jsonpath='{.status.conditions}' | jq '.[] | {type, status, reason, message}'
{
"type": "Reconciling",
"status": "True",
"reason": "ProgressingWithRetry",
"message": "Detecting drift for revision main@sha1:83a41e21322e15eff2307bc1e9d5c89e33226d0d with a timeout of 30s"
}
{
"type": "Ready",
"status": "False",
"reason": "ReconciliationFailed",
"message": "Deployment/flux-demo/demo dry-run failed: admission webhook \"validate.kyverno.svc-fail\" denied the request: \n\nresource Deployment/flux-demo/demo was blocked due to the following policies \n\nrequire-owner:\n require-owner-label: 'validation error: every Deployment needs an owner label. rule require-owner-label failed at path /metadata/labels/owner/'\n\n"
}
$ kubectl delete clusterpolicy require-owner
Warning: kyverno.io/v1 ClusterPolicy is deprecated and will be removed in a future release; migrate to ValidatingPolicy, MutatingPolicy, GeneratingPolicy or ImageValidatingPolicy (policies.kyverno.io), see https://kyverno.io/docs/guides/migration-to-cel/
clusterpolicy.kyverno.io "require-owner" deleted
$ # this app syncs automatically, so wait out the controller's own operation first
$ argocd app wait demo-staging --operation --timeout 180 || true
Name: argocd/demo-staging
Project: default
Server: https://kubernetes.default.svc
Namespace: staging
URL: https://argocd.example.com/applications/demo-staging
Source:
- Repo: http://gitea.lab:3000/lab/platform.git
Target: main
Path: demo-app/overlays/staging
SyncWindow: Sync Allowed
Sync Policy: Automated (Prune)
Sync Status: OutOfSync from main (83a41e2)
Health Status: Healthy
Operation: Sync
Sync Revision: 83a41e21322e15eff2307bc1e9d5c89e33226d0d
Phase: Failed
Start: 2026-09-13 01:34:57 -0400 EDT
Finished: 2026-09-13 01:35:03 -0400 EDT
Duration: 6s
Message: one or more synchronization tasks completed unsuccessfully, reason: admission webhook "validate.kyverno.svc-fail" denied the request:
resource Deployment/team-a/staging-demo was blocked due to the following policies
require-owner:
require-owner-label: 'validation error: every Deployment needs an owner label. rule require-owner-label failed at path /metadata/labels/owner/'
GROUP KIND NAMESPACE NAME STATUS HEALTH HOOK MESSAGE
Service team-a staging-demo Synced Healthy service/staging-demo unchanged
apps Deployment team-a staging-demo OutOfSync Missing admission webhook "validate.kyverno.svc-fail" denied the request:
resource Deployment/team-a/staging-demo was blocked due to the following policies
require-owner:
require-owner-label: 'validation error: every Deployment needs an owner label. rule require-owner-label failed at path /metadata/labels/owner/'
$ argocd app sync demo-staging --timeout 120
TIMESTAMP GROUP KIND NAMESPACE NAME STATUS HEALTH HOOK MESSAGE
2026-09-13T01:40:10-04:00 Service team-a staging-demo Synced Healthy
2026-09-13T01:40:10-04:00 apps Deployment team-a staging-demo OutOfSync Missing
2026-09-13T01:40:11-04:00 apps Deployment team-a staging-demo OutOfSync Progressing deployment.apps/staging-demo created
2026-09-13T01:40:11-04:00 Service team-a staging-demo Synced Healthy service/staging-demo unchanged
2026-09-13T01:40:12-04:00 apps Deployment team-a staging-demo Synced Progressing deployment.apps/staging-demo created
Name: argocd/demo-staging
Project: default
Server: https://kubernetes.default.svc
Namespace: staging
URL: https://argocd.example.com/applications/demo-staging
Source:
- Repo: http://gitea.lab:3000/lab/platform.git
Target: main
Path: demo-app/overlays/staging
SyncWindow: Sync Allowed
Sync Policy: Automated (Prune)
Sync Status: Synced to main (83a41e2)
Health Status: Progressing
Operation: Sync
Sync Revision: 83a41e21322e15eff2307bc1e9d5c89e33226d0d
Phase: Succeeded
Start: 2026-09-13 01:40:10 -0400 EDT
Finished: 2026-09-13 01:40:11 -0400 EDT
Duration: 1s
Message: successfully synced (all tasks run)
GROUP KIND NAMESPACE NAME STATUS HEALTH HOOK MESSAGE
Service team-a staging-demo Synced Healthy service/staging-demo unchanged
apps Deployment team-a staging-demo Synced Progressing deployment.apps/staging-demo createdFailed quoting the admission webhook by name, and Flux reports Ready False with a reconciliation failure quoting the same webhook. Same cause, two vocabularies; be able to translate either into "a policy refused the manifest".A webhook with failurePolicy: Fail and no running backend does not fail open, it fails everything. This is the outage that looks like the API server broke, and it is one scale command away in either direction.
kubectl get validatingwebhookconfiguration -o custom-columns=NAME:.metadata.name,POLICY:.webhooks[*].failurePolicy | grep -i kyverno
# scaling Kyverno to zero does not reproduce this on this install: its webhooks still admit.
# own the webhook instead, pointed at a Service with no backend, scoped to team-a.
kubectl apply -f - <<'EOF'
apiVersion: admissionregistration.k8s.io/v1
kind: ValidatingWebhookConfiguration
metadata: { name: no-endpoints }
webhooks:
- name: no-endpoints.lab.local
admissionReviewVersions: ["v1"]
sideEffects: None
failurePolicy: Fail
timeoutSeconds: 5
namespaceSelector:
matchLabels: { kubernetes.io/metadata.name: team-a }
clientConfig:
service: { name: absent, namespace: team-a, path: /validate, port: 443 }
rules:
- { apiGroups: ["apps"], apiVersions: ["v1"], operations: ["CREATE","UPDATE"], resources: ["deployments"], scope: Namespaced }
EOF
sleep 5
kubectl -n team-a create deployment probe --image=nginx:1.27-alpine
# an unchanged resource is never sent to admission, so make the sync create one
kubectl -n team-a delete deploy staging-demo --ignore-not-found
# this app syncs automatically, so wait out the controller's own operation first
argocd app wait demo-staging --operation --timeout 180 || true
argocd app sync demo-staging --timeout 60 || true
kubectl -n argocd get app demo-staging -o jsonpath='{.status.operationState.phase}{"\n"}{.status.operationState.message}{"\n"}'
kubectl delete validatingwebhookconfiguration no-endpoints
kubectl -n team-a delete deployment probe --ignore-not-found
# this app syncs automatically, so wait out the controller's own operation first
argocd app wait demo-staging --operation --timeout 180 || true
argocd app sync demo-staging --timeout 120outputcaptured 2026-09-12
$ kubectl get validatingwebhookconfiguration -o custom-columns=NAME:.metadata.name,POLICY:.webhooks[*].failurePolicy | grep -i kyverno
kyverno-cel-exception-validating-webhook-cfg Fail
kyverno-cleanup-validating-webhook-cfg Fail
kyverno-exception-validating-webhook-cfg Fail
kyverno-global-context-validating-webhook-cfg Fail
kyverno-policy-validating-webhook-cfg Fail
kyverno-resource-validating-webhook-cfg Fail
kyverno-ttl-validating-webhook-cfg Ignore
$ # scaling Kyverno to zero does not reproduce this on this install: its webhooks still admit.
$ # own the webhook instead, pointed at a Service with no backend, scoped to team-a.
$ kubectl apply -f - <<'EOF'
apiVersion: admissionregistration.k8s.io/v1
kind: ValidatingWebhookConfiguration
metadata: { name: no-endpoints }
webhooks:
- name: no-endpoints.lab.local
admissionReviewVersions: ["v1"]
sideEffects: None
failurePolicy: Fail
timeoutSeconds: 5
namespaceSelector:
matchLabels: { kubernetes.io/metadata.name: team-a }
clientConfig:
service: { name: absent, namespace: team-a, path: /validate, port: 443 }
rules:
- { apiGroups: ["apps"], apiVersions: ["v1"], operations: ["CREATE","UPDATE"], resources: ["deployments"], scope: Namespaced }
EOF
validatingwebhookconfiguration.admissionregistration.k8s.io/no-endpoints created
$ sleep 5
$ kubectl -n team-a create deployment probe --image=nginx:1.27-alpine
Warning: would violate PodSecurity "restricted:latest": allowPrivilegeEscalation != false (container "nginx" must set securityContext.allowPrivilegeEscalation=false), unrestricted capabilities (container "nginx" must set securityContext.capabilities.drop=["ALL"]), runAsNonRoot != true (pod or container "nginx" must set securityContext.runAsNonRoot=true), seccompProfile (pod or container "nginx" must set securityContext.seccompProfile.type to "RuntimeDefault" or "Localhost")
error: failed to create deployment: Internal error occurred: failed calling webhook "no-endpoints.lab.local": failed to call webhook: Post "https://absent.team-a.svc:443/validate?timeout=5s": service "absent" not found
$ # an unchanged resource is never sent to admission, so make the sync create one
$ kubectl -n team-a delete deploy staging-demo --ignore-not-found
deployment.apps "staging-demo" deleted from team-a namespace
$ # this app syncs automatically, so wait out the controller's own operation first
$ argocd app wait demo-staging --operation --timeout 180 || true
Name: argocd/demo-staging
Project: default
Server: https://kubernetes.default.svc
Namespace: staging
URL: https://argocd.example.com/applications/demo-staging
Source:
- Repo: http://gitea.lab:3000/lab/platform.git
Target: main
Path: demo-app/overlays/staging
SyncWindow: Sync Allowed
Sync Policy: Automated (Prune)
Sync Status: Synced to main (b793ef2)
Health Status: Healthy
Operation: Sync
Sync Revision: b793ef22cee9cf7bc5eb3d41e236dda5c43670fb
Phase: Succeeded
Start: 2026-09-13 12:44:52 -0400 EDT
Finished: 2026-09-13 12:45:02 -0400 EDT
Duration: 10s
Message: successfully synced (all tasks run)
GROUP KIND NAMESPACE NAME STATUS HEALTH HOOK MESSAGE
apps Deployment team-a staging-demo Synced Missing deployment.apps/staging-demo created
Service team-a staging-demo Synced Healthy
$ argocd app sync demo-staging --timeout 60 || true
{"level":"fatal","msg":"rpc error: code = FailedPrecondition desc = another operation is already in progress","time":"2026-09-13T12:50:07-04:00"}
$ kubectl -n argocd get app demo-staging -o jsonpath='{.status.operationState.phase}{"\n"}{.status.operationState.message}{"\n"}'
Running
$ kubectl delete validatingwebhookconfiguration no-endpoints
validatingwebhookconfiguration.admissionregistration.k8s.io "no-endpoints" deleted
$ kubectl -n team-a delete deployment probe --ignore-not-found
$ # this app syncs automatically, so wait out the controller's own operation first
$ argocd app wait demo-staging --operation --timeout 180 || true
TIMESTAMP GROUP KIND NAMESPACE NAME STATUS HEALTH HOOK MESSAGE
2026-09-13T12:50:07-04:00 apps Deployment team-a staging-demo OutOfSync Missing Internal error occurred: failed calling webhook "no-endpoints.lab.local": failed to call webhook: Post "https://absent.team-a.svc:443/validate?timeout=5s": service "absent" not found
2026-09-13T12:50:07-04:00 Service team-a staging-demo Synced Healthy
Name: argocd/demo-staging
Project: default
Server: https://kubernetes.default.svc
Namespace: staging
URL: https://argocd.example.com/applications/demo-staging
Source:
- Repo: http://gitea.lab:3000/lab/platform.git
Target: main
Path: demo-app/overlays/staging
SyncWindow: Sync Allowed
Sync Policy: Automated (Prune)
Sync Status: Synced to main (b793ef2)
Health Status: Progressing
Operation: Sync
Sync Revision: b793ef22cee9cf7bc5eb3d41e236dda5c43670fb
Phase: Succeeded
Start: 2026-09-13 12:50:07 -0400 EDT
Finished: 2026-09-13 12:50:17 -0400 EDT
Duration: 10s
Message: successfully synced (all tasks run)
GROUP KIND NAMESPACE NAME STATUS HEALTH HOOK MESSAGE
apps Deployment team-a staging-demo Synced Progressing deployment.apps/staging-demo created
Service team-a staging-demo Synced Healthy
$ argocd app sync demo-staging --timeout 120
TIMESTAMP GROUP KIND NAMESPACE NAME STATUS HEALTH HOOK MESSAGE
2026-09-13T12:50:18-04:00 Service team-a staging-demo Synced Healthy
2026-09-13T12:50:18-04:00 apps Deployment team-a staging-demo Synced Progressing
2026-09-13T12:50:19-04:00 Service team-a staging-demo Synced Healthy service/staging-demo unchanged
2026-09-13T12:50:19-04:00 apps Deployment team-a staging-demo Synced Progressing deployment.apps/staging-demo unchanged
Name: argocd/demo-staging
Project: default
Server: https://kubernetes.default.svc
Namespace: staging
URL: https://argocd.example.com/applications/demo-staging
Source:
- Repo: http://gitea.lab:3000/lab/platform.git
Target: main
Path: demo-app/overlays/staging
SyncWindow: Sync Allowed
Sync Policy: Automated (Prune)
Sync Status: Synced to main (b793ef2)
Health Status: Progressing
Operation: Sync
Sync Revision: b793ef22cee9cf7bc5eb3d41e236dda5c43670fb
Phase: Succeeded
Start: 2026-09-13 12:50:18 -0400 EDT
Finished: 2026-09-13 12:50:19 -0400 EDT
Duration: 1s
Message: successfully synced (all tasks run)
GROUP KIND NAMESPACE NAME STATUS HEALTH HOOK MESSAGE
Service team-a staging-demo Synced Healthy service/staging-demo unchanged
apps Deployment team-a staging-demo Synced Progressing deployment.apps/staging-demo unchangedkubectl create deployment fails with Internal error occurred: failed calling webhook "no-endpoints.lab.local" ... service "absent" not found, and Argo CD hits the same wall: the wait shows staging-demo OutOfSync Missing carrying that message, because this app self-heals and the controller tried the write before you did. That is also why the manual sync comes back another operation is already in progress rather than with the webhook error itself. Two things have to be true for this outage: failurePolicy: Fail, and a rule matching what is being written, which is why the Deployment is deleted first. Scaling the Kyverno admission controller to zero does not reproduce it on this install.Every failure on this page shows up as an Application condition, and the type is a faster triage signal than the message. Run through the three buckets and write the type next to each one. Expect the list to be empty here: conditions describe what is wrong now, and by this point in the page everything has been put back.
kubectl -n argocd get app demo-staging -o jsonpath='{range .status.conditions[*]}{.type}: {.message}{"\n"}{end}'
kubectl -n argocd get app -o custom-columns=NAME:.metadata.name,SYNC:.status.sync.status,HEALTH:.status.health.status,CONDITIONS:.status.conditions[*].type
kubectl explain application.status.conditions.typeoutputcaptured 2026-09-12
$ kubectl -n argocd get app demo-staging -o jsonpath='{range .status.conditions[*]}{.type}: {.message}{"\n"}{end}'
$ kubectl -n argocd get app -o custom-columns=NAME:.metadata.name,SYNC:.status.sync.status,HEALTH:.status.health.status,CONDITIONS:.status.conditions[*].type
NAME SYNC HEALTH CONDITIONS
demo-prod Synced Healthy <none>
demo-staging Synced Healthy <none>
payments-dev Synced Healthy <none>
$ kubectl explain application.status.conditions.type
GROUP: argoproj.io
KIND: Application
VERSION: v1alpha1
FIELD: type <string>
DESCRIPTION:
Type is an application condition type
Self-check
"The app is Synced and Healthy but the feature we merged an hour ago is missing."
Bucket 4, stale. Check for suspension, then the last sync timestamp and the resolved revision against your commit SHA, then whether the app tracks the branch you pushed to. Refresh/reconcile before anything else.
"Sync fails: Service "demo" is invalid: spec.clusterIP: Invalid value: field is immutable."
Bucket 2: the cluster refused the manifest. The text came from the API server, relayed by the sync operation. Fix the manifest, or force replace semantics if the field genuinely must change (accepting the churn that implies).
"argocd-repo-server logs show authentication required; the workload is untouched and running fine."
Bucket 3: access. No pod of yours is implicated; the controller's identity is. Repo credentials, token scope, or a rotated secret. The workload keeps running because reconciliation is what broke, not the deployment.
"Deployment exists, no pods, no error on the Deployment."
Look one level down at the ReplicaSet: quota, PSS, or an admission webhook rejected the pod template. Same indirection as sections 1.4 and 5.3. Not strictly a delivery bug: the delivery worked, admission refused.
"I keep deleting this ConfigMap and it keeps coming back."
Something reconciles it: self-heal, a Flux Kustomization at interval, or an operator that owns it. Prove which with ownerReferences, the app.kubernetes.io/instance label, or flux trace. Then change the source of truth instead of the live object, or suspend first if you need a temporary window.
"Sync failed: admission webhook "validate.kyverno.svc-fail" denied the request". Which bucket, and where does the fix go?
Bucket 5, admission. The manifest reached the API server and a policy engine refused it. Fix either the manifest so it satisfies the policy (the usual grading intent) or, if the task says the policy is wrong, the policy itself. Re-syncing changes nothing; Argo CD marks the resource SyncFailed, Flux reports ReconciliationFailed with the same text.
A Flux Kustomization reports ArtifactFailed. What is broken and where do you look first?
The source it references has no artifact: the GitRepository or OCIRepository is not Ready (AuthenticationFailed, GitOperationFailed, bad ref) or has not fetched yet. Bucket 3, upstream of the Kustomization. flux get sources all -A and kubectl describe gitrepository first; the Kustomization will recover on its own once the source publishes.
Argo CD shows sync status Unknown with a ComparisonError condition. Bucket?
Bucket 3: comparison itself failed, so Argo does not even know whether live matches git. Causes are repo access (credentials, URL, revision), a render error in repo-server, or an unreachable destination cluster. The condition message names it; the workload is untouched, which is the tell.
Docs to know your way around
- argo-cd.readthedocs.io: sync options (Replace, ServerSideApply), resource health checks, the troubleshooting section.
- fluxcd.io: the troubleshooting cheatsheet (flux logs, flux events, tracing a resource to its Kustomization).
- Offline:
argocd app manifests,argocd app diff,flux build kustomization <name> --path ./…, andkubectl get events -A --sort-by=.lastTimestamp. - argo-cd.readthedocs.io: Application specification (status fields) and the gitops-engine health package: the exact sync, health, operation phase and condition strings.
- fluxcd.io troubleshooting cheatsheet ("webhook does not support dry run") and each component's Conditions section: the reason strings and the Flux-specific admission failure.