The exam names both Argo and Flux, and the lab installs both against the same Gitea for a reason: you should be able to express the same delivery in either. Flux has no Application object. It decomposes GitOps into small CRDs that reference each other, and reading those references is the skill.
make coreOrientation
Argo CD is one controller with one big CRD. Flux is a toolkit of small controllers with small CRDs, composed by reference. Neither is better; they fail differently, and the exam may hand you either. What transfers is the mental translation, which is the last exercise below.
| Controller | Owns | Does |
|---|---|---|
| source-controller | GitRepository, OCIRepository, HelmRepository, HelmChart, Bucket | fetches and publishes an artifact (a tarball + revision) |
| kustomize-controller | Kustomization | an artifact + a path → builds and applies |
| helm-controller | HelmRelease | a chart source → installs/upgrades a release |
| notification-controller | Alert, Provider, Receiver | events out (Slack, webhooks) and events in (git push → instant reconcile) |
| image-*-controller | ImageRepository, ImagePolicy, ImageUpdateAutomation | registry tags → commits back to git |
Know that the last row exists: Flux can close the CI→CD loop by writing the new image tag into git itself, which is the pull-based answer to "how does a new build get deployed without CI touching the cluster".
Working model: sources and appliers
GitRepository/platform (interval 1m, ref: main)
│ produces artifact @ revision main/9f2c1ab
├────────────▶ Kustomization/apps path ./apps prune ✓ interval 5m
├────────────▶ Kustomization/infra path ./infra prune ✓ dependsOn: []
└────────────▶ Kustomization/tenants path ./tenants dependsOn: [infra]
HelmRepository/podinfo ──▶ HelmRelease/podinfo ──▶ release in target namespace
The two-step split is the design: one GitRepository can feed many Kustomizations, each watching a different path with its own interval, health checks and prune setting. dependsOn then orders them, which is Flux's equivalent of Argo's sync waves, except it works between top-level objects rather than inside one sync.
Flux's Kustomization (kustomize.toolkit.fluxcd.io) is not kustomize's Kustomization (kustomize.config.k8s.io). The Flux one points at a directory; that directory may contain the other one. Exam tasks love this ambiguity, and kubectl get kustomizations.kustomize.toolkit.fluxcd.io versus reading a file resolves it.
Kustomization fields that decide behavior
| Field | Effect |
|---|---|
| interval | how often it re-applies; drift correction is a side effect of re-applying, there is no selfHeal toggle |
| prune | delete resources removed from git; same blast radius as Argo's |
| wait / healthChecks / timeout | block Ready until the listed objects are healthy; this is what makes dependsOn meaningful |
| targetNamespace | override the namespace for everything the path applies |
| postBuild.substitute / substituteFrom | variable substitution from ConfigMaps/Secrets after kustomize build; Flux's answer to "one manifest, per-cluster values" |
| serviceAccountName | apply as that SA, the multi-tenancy control: a tenant's Kustomization cannot exceed its own RBAC |
| decryption | SOPS: decrypt secrets in the repo at apply time |
HelmRelease
It references a chart (from a HelmRepository, a GitRepository, or an OCIRepository), sets values inline and/or valuesFrom ConfigMaps and Secrets, and controls failure behavior with install.remediation and upgrade.remediation (retries, and whether to roll back). Newer versions also detect and correct drift on the release's rendered manifests. When one fails, the useful discrimination is which layer: source resolution (the chart could not be fetched), install/upgrade (Helm itself errored), or health (the release installed but the workload never became ready). Each surfaces in a different condition, and reading the right one is the whole diagnosis.
Daily verbs
flux get sources git # and: helm, oci, bucket, all
flux get kustomizations # Ready, revision, last applied
flux reconcile kustomization demo-flux --with-source # force the loop now, refetch first
flux suspend kustomization demo-flux # the sanctioned pause for surgery
flux resume kustomization demo-flux
flux tree kustomization demo-flux # what did this thing create
flux events --for Kustomization/demo-flux
flux logs --level=error --all-namespaces
flux trace deploy/demo -n flux-demo # which Flux object owns this resourceoutputcaptured 2026-08-26
$ flux get sources git # and: helm, oci, bucket, all
NAME REVISION SUSPENDED READY MESSAGE
platform main@sha1:c7cb9a1e False True stored artifact for revision 'main@sha1:c7cb9a1e'
$ flux get kustomizations # Ready, revision, last applied
NAME REVISION SUSPENDED READY MESSAGE
demo-flux main@sha1:c7cb9a1e False True Applied revision: main@sha1:c7cb9a1e
$ flux reconcile kustomization demo-flux --with-source # force the loop now, refetch first
► annotating GitRepository platform in flux-system namespace
✔ GitRepository annotated
◎ waiting for GitRepository reconciliation
✔ fetched revision main@sha1:c7cb9a1e37949219ff64e2c5dd17654109deea6d
► annotating Kustomization demo-flux in flux-system namespace
✔ Kustomization annotated
◎ waiting for Kustomization reconciliation
✔ applied revision main@sha1:c7cb9a1e37949219ff64e2c5dd17654109deea6d
$ flux suspend kustomization demo-flux # the sanctioned pause for surgery
► suspending kustomization demo-flux in flux-system namespace
✔ kustomization suspended
$ flux resume kustomization demo-flux
► resuming kustomization demo-flux in flux-system namespace
✔ kustomization resumed
◎ waiting for Kustomization reconciliation
✔ Kustomization demo-flux reconciliation completed
✔ applied revision main@sha1:c7cb9a1e37949219ff64e2c5dd17654109deea6d
$ flux tree kustomization demo-flux # what did this thing create
Kustomization/flux-system/demo-flux
├── Service/flux-demo/demo
└── Deployment/flux-demo/demo
$ flux events --for Kustomization/demo-flux
LAST SEEN TYPE REASON OBJECT MESSAGE
31m Normal NewArtifact GitRepository/platform stored artifact for commit 'seed lab manifests'
5m34s (x26 over 30m) Normal GitOperationSucceeded GitRepository/platform no changes since last reconciliation: observed revision 'main@sha1:62d527f00a730f9a9b30cbd70a20b8e138f70bd6'
4m35s Normal NewArtifact GitRepository/platform stored artifact for commit 'staging: 3 replicas'
3m33s Normal GarbageCollectionSucceeded GitRepository/platform garbage collected 1 artifacts
2m10s Normal Progressing Kustomization/demo-flux Service/flux-demo/demo created
Deployment/flux-demo/demo created
2m10s Normal ReconciliationSucceeded Kustomization/demo-flux Reconciliation finished in 653.93345ms, next run in 1m0s
92s (x3 over 3m33s) Normal GitOperationSucceeded GitRepository/platform no changes since last reconciliation: observed revision 'main@sha1:c7cb9a1e37949219ff64e2c5dd17654109deea6d'
68s Normal Progressing Kustomization/demo-flux Deployment/flux-demo/demo configured
68s Normal ReconciliationSucceeded Kustomization/demo-flux Reconciliation finished in 553.601679ms, next run in 1m0s
6s Normal ReconciliationSucceeded Kustomization/demo-flux Reconciliation finished in 481.338357ms, next run in 1m0s
4s Normal ReconciliationSucceeded Kustomization/demo-flux Reconciliation finished in 227.246168ms, next run in 1m0s
1s Normal ReconciliationSucceeded Kustomization/demo-flux Reconciliation finished in 343.998075ms, next run in 1m0s
$ flux logs --level=error --all-namespaces
$ flux trace deploy/demo -n flux-demo # which Flux object owns this resource
Object: Deployment/demo
Namespace: flux-demo
Status: Managed by Flux
---
Kustomization: demo-flux
Namespace: flux-system
Target: flux-demo
Path: ./demo-app/base
Revision: main@sha1:c7cb9a1e37949219ff64e2c5dd17654109deea6d
Status: Last reconciled at 2026-08-26 22:20:37 -0400 EDT
Message: Applied revision: main@sha1:c7cb9a1e37949219ff64e2c5dd17654109deea6d
---
GitRepository: platform
Namespace: flux-system
URL: http://gitea.lab:3000/lab/platform.git
Branch: main
Revision: main@sha1:c7cb9a1e37949219ff64e2c5dd17654109deea6d
Status: Last reconciled at 2026-08-26 22:16:03 -0400 EDT
Message: stored artifact for revision 'main@sha1:c7cb9a1e37949219ff64e2c5dd17654109deea6d'flux trace is the reverse lookup: hand it any live resource and it tells you which Kustomization or HelmRelease put it there. On an unfamiliar cluster that is the fastest orientation command Flux has.
"Stop the controller overwriting my hotfix while I debug" has a named, auditable answer in both engines: flux suspend, or disabling auto-sync in Argo. Doing it with kubectl scale deploy/kustomize-controller --replicas=0 works and will cost you the mark, because it stops everything and leaves no record on the object.
More Kustomization fields and the per-resource annotations
retryInterval: how soon to retry after a failure (defaults tointerval);timeoutbounds build, apply and health checks together.deletionPolicy:MirrorPrune(default: delete managed objects on Kustomization deletion only ifprune: true),Delete,WaitForTermination,Orphan.Orphanis the answer when a namespace deletion would otherwise race the controller.healthCheckExprs: CELcurrent,failedandinProgressexpressions perapiVersion/kind, for custom resources whose status kstatus cannot read;dependsOn[].readyExprdoes the same for a dependency.patches,images,components,namePrefix/nameSuffix,commonMetadata: kustomize edits applied by the controller without touching the repo.kubeConfig.secretRef: apply to another cluster with a kubeconfig Secret; Flux's multi-cluster story is one Kustomization per remote cluster.
Annotations you put on the manifests in the source, not on the Kustomization: kustomize.toolkit.fluxcd.io/prune: disabled (never garbage-collect this object, for PVCs and namespaces), kustomize.toolkit.fluxcd.io/ssa: Merge (keep fields other tools add), IfNotPresent (create once, never overwrite: cert-manager-managed Secrets, webhook configurations) or Ignore (skip entirely), and kustomize.toolkit.fluxcd.io/force: enabled (recreate on immutable-field changes, for Jobs; dangerous on StatefulSets). Fields you edit with kubectl survive only if you apply them with --field-manager=flux-client-side-apply. A reconcile outside the interval is the annotation reconcile.fluxcd.io/requestedAt, which is what flux reconcile sets.
The translation table
| Intent | Argo CD | Flux |
|---|---|---|
| where the manifests are | Application.spec.source | GitRepository + Kustomization.spec.path |
| auto-apply on change | syncPolicy.automated | implicit: every interval |
| revert live drift | automated.selfHeal | implicit: re-apply at interval |
| delete removed resources | automated.prune | Kustomization.spec.prune |
| ordering | sync waves + hooks | dependsOn + healthChecks + wait |
| pause | disable auto-sync | flux suspend |
| per-env values | kustomize overlays / Helm values | same, plus postBuild.substitute |
| tenant guardrails | AppProject | serviceAccountName + namespace isolation |
| instant trigger | webhook to argocd-server | Receiver (notification-controller) |
Practice moving across this table in both directions; the exam can hand you either tool.
Condition vocabulary and failure remediation
Flux objects follow kstatus: a Ready condition summarizes, a Reconciling condition (negative polarity, present only while true) says work is in progress with reason Progressing or ProgressingWithRetry, and Stalled means the controller gave up until the spec changes. The reason strings are stable API and a task can quote them.
| Object | Ready=True reasons | Ready=False reasons | Other condition types |
|---|---|---|---|
| GitRepository / OCIRepository / Bucket | Succeeded | AuthenticationFailed · GitOperationFailed · OCIArtifactPullFailed · OCIArtifactLayerOperationFailed · VerificationError | FetchFailed · IncludeUnavailable · StorageOperationFailed · SourceVerified · ArtifactInStorage |
| Kustomization | ReconciliationSucceeded | ArtifactFailed · BuildFailed · HealthCheckFailed · PruneFailed · DependencyNotReady · ReconciliationFailed | Healthy |
| HelmRelease | InstallSucceeded · UpgradeSucceeded · TestSucceeded | InstallFailed · UpgradeFailed · TestFailed · RollbackSucceeded · RollbackFailed · UninstallSucceeded · UninstallFailed · RetriesExceeded | Released · TestSuccess · Remediated · Drifted (DriftDetected / NoDriftDetected) |
Reading the table as a ladder: ArtifactFailed means the source is not Ready or has no artifact yet (fix the source); BuildFailed is kustomize build (path, invalid YAML, a missing substituteFrom ConfigMap unless optional: true); DependencyNotReady points at another Kustomization; ReconciliationFailed carries the API server's rejection ("field is immutable", "admission webhook denied"); HealthCheckFailed means it applied but never became healthy within timeout. For a HelmRelease, Ready=False with reason RollbackSucceeded is not a contradiction: the upgrade failed, remediation rolled back, and the object stays not-ready until you fix the chart or values.
HelmRelease failure handling, field by field
install.remediation.retries(default 0): between attempts the release is uninstalled;-1retries forever.upgrade.remediation.retrieswithstrategy: rollback(default) oruninstall;remediateLastFailuredefaults to true for upgrades once retries is above 0, so the release lands on the last good revision rather than a failed one.install.strategy.name/upgrade.strategy.name:RemediateOnFailure(the behavior above) orRetryOnFailurewithretryInterval(Flux 2.7+), which simply retries without rolling back.test.enable: trueruns Helm tests after install/upgrade; failures count as release failures unlesstest.ignoreFailures: true.install.crdsdefaults toCreate, butupgrade.crdsdefaults toSkip: Helm does not upgrade CRDs, so a chart bump with new CRD fields needsupgrade.crds: CreateReplaceor the CRDs delivered by a separate Kustomization.driftDetection.mode:disabled(default),warn(emit an event and theDriftedcondition),enabled(correct it with a server-side apply);driftDetection.ignoretakes JSON 6902 paths and a target, the equivalent of Argo'signoreDifferences, for HPA-managed replicas and webhook-injected fields.- Since Flux 2.8 the helm-controller applies with server-side apply by default (
install.serverSideApply,upgrade.serverSideApply: auto) and records astatus.inventoryof everything the release owns; the wait strategy defaults to kstatus polling rather than Helm's own wait. valuesFrom[]:kindConfigMap or Secret,name,valuesKey(defaultvalues.yaml),targetPathto graft one value,optional. Inlinevalueswin overvaluesFrom.
The three annotations behind flux reconcile
flux reconcile helmrelease X sets reconcile.fluxcd.io/requestedAt; --force adds reconcile.fluxcd.io/forceAt (run the install or upgrade again even though nothing changed); --reset adds reconcile.fluxcd.io/resetAt (zero the failure counters after RetriesExceeded). Without --reset, a release that exhausted its retries stays failed until the chart, values or generation changes. Kustomizations know requestedAt only; a stuck one is fixed by fixing the cause.
A chart whose first install fails is left in a failed state and not retried until the spec changes; only the Ready reason tells you (InstallFailed). Set install.remediation.retries: 3 and upgrade.remediation.retries: 3 on anything you cannot babysit, and know that flux reconcile hr X --reset is the manual restart.
Sources, image automation, notifications and bootstrap
Source kinds
| Kind | Pointer fields | Extras a task may set |
|---|---|---|
| GitRepository | url · ref.branch | ref.tag | ref.semver | ref.name (refs/pull/1/head) | ref.commit | secretRef (basic auth, token, SSH key, TLS CA), ignore (or a .sourceignore file), include other GitRepositories, verify.mode for signed commits, sparseCheckout, recurseSubmodules, provider: github|azure|aws for app or workload identity |
| OCIRepository | url: oci://... · ref.tag | ref.semver | ref.semverFilter | ref.digest | layerSelector, verify.provider: cosign|notation (keyless via OIDC identities or public keys in a Secret), provider: aws|azure|gcp, insecure for plain HTTP registries |
| HelmRepository | url (https:// index, or oci:// with type: oci) | secretRef, passCredentials; a HelmChart is generated per HelmRelease and is what flux get sources chart lists |
| Bucket | bucketName · endpoint · provider: generic|aws|gcp|azure | for S3-style config stores |
Every source publishes an Artifact (a tarball plus a revision such as main@sha1:9f2c1ab) and consumers see only the artifact, which is why a failing source starves every Kustomization behind it with ArtifactFailed.
Image automation, the loop that closes CI to CD
ImageRepositoryscans a registry path on aninterval(credentials viasecretRefor cloudprovider).ImagePolicypicks one tag from the scan:policy.semver.range: 5.0.x,alphabetical, ornumerical, optionally afterfilterTags.patternwith anextractgroup for tags likemain-abc123-1699999999.status.latestRefholds image, tag and digest.ImageUpdateAutomationwrites the chosen tag into files underupdate.pathin thesourceRefGitRepository (which needs write credentials) wherever a marker comment appears:image: ghcr.io/org/app:1.0 # {"$imagepolicy": "flux-system:app"}, with:tagor:namesuffixes to update only that part. It commits withgit.commit.authorand pushes togit.push.branch, which may be a different branch so a PR gates production.
Notification controller
Provider(notification.toolkit.fluxcd.io/v1beta3):typeslack, msteams, discord, github, gitlab, generic webhook and others;address,channel,secretRef. Thegithub/gitlabtypes post commit statuses, which is how a red X appears on the commit that broke reconciliation.Alert(v1beta3):providerRef,eventSeverity: info|error,eventSources(kind, name or*, namespace),inclusionList/exclusionListregexes,eventMetadataadded to every message.Receiver(v1):typegithub, gitea, gitlab, harbor, dockerhub, generic, generic-hmac, cdevents;eventsto accept;secretRefwith the webhook token;resourcesto reconcile (a GitRepository, or ImageRepositories for registry pushes).status.webhookPathis the path to give the git server, served by thewebhook-receiverService that you must expose yourself. A Receiver turns a one-minute poll into an instant pull.
Bootstrap and lockdown recap
flux bootstrap <provider> commits gotk-components.yaml and gotk-sync.yaml to --path, installs the controllers, and creates the flux-system GitRepository and Kustomization that manage Flux from then on; --components-extra adds the image controllers; --token-auth uses HTTPS with a token instead of a deploy key. Multi-tenant lockdown is three controller flags (--no-cross-namespace-refs=true, --no-remote-bases=true, --default-service-account=default) plus spec.serviceAccountName: kustomize-controller on the flux-system Kustomization, applied as kustomize patches in the bootstrap directory. After lockdown a tenant Kustomization without a serviceAccountName impersonates the namespace's default account, which has no rights, and fails with a Forbidden in ReconciliationFailed.
flux get all -A --status-selector ready=false lists only what is broken; kubectl get fluxcd -A (the CRD category) lists every Flux object; flux events --for Kustomization/x, flux logs --level=error -A, flux trace and flux tree descend from there. flux build kustomization x --path ./dir renders locally and flux diff kustomization x --path ./dir shows what a commit would change.
Exercises
Deploy the demo base (not the overlays, which Argo owns in 2.2; two controllers must never manage the same resource) into a Flux-owned namespace:
kubectl create ns flux-demo
flux create kustomization demo-flux \
--source=GitRepository/platform \
--path=./demo-app/base \
--target-namespace=flux-demo \
--prune=true --interval=1m \
--health-check-timeout=2m
flux get kustomizationsoutputcaptured 2026-08-26
$ kubectl create ns flux-demo
namespace/flux-demo created
$ flux create kustomization demo-flux \
--source=GitRepository/platform \
--path=./demo-app/base \
--target-namespace=flux-demo \
--prune=true --interval=1m \
--health-check-timeout=2m
✚ generating Kustomization
► applying Kustomization
✔ Kustomization created
◎ waiting for Kustomization reconciliation
✔ Kustomization demo-flux is ready
✔ applied revision main@sha1:c7cb9a1e37949219ff64e2c5dd17654109deea6d
$ flux get kustomizations
NAME REVISION SUSPENDED READY MESSAGE
demo-flux main@sha1:c7cb9a1e False True Applied revision: main@sha1:c7cb9a1e kubectl -n flux-demo get deploy demo shows 2/2. Then flux tree kustomization demo-flux and confirm it lists exactly the deployment, service, and nothing else.Scale the deployment by hand, then wait out one interval:
kubectl -n flux-demo scale deploy demo --replicas=5
sleep 70 && kubectl -n flux-demo get deploy demo -o jsonpath='{.spec.replicas}{"\n"}'outputcaptured 2026-08-26
$ kubectl -n flux-demo scale deploy demo --replicas=5
deployment.apps/demo scaled
$ sleep 70 && kubectl -n flux-demo get deploy demo -o jsonpath='{.spec.replicas}{"\n"}'
2flux suspend kustomization demo-flux, scale to 5 again, confirm it stays at 5 for two intervals, then flux resume kustomization demo-flux and confirm it snaps back.
flux get kustomizations as Suspended True. This is the answer to "make the controller stop overwriting my hotfix while I debug".Sources need not be git:
flux create source helm podinfo --url=https://stefanprodan.github.io/podinfo --interval=10m
flux create helmrelease podinfo --source=HelmRepository/podinfo \
--chart=podinfo --release-name=podinfo --target-namespace=flux-demo --interval=5m
kubectl -n flux-demo get deploy podinfooutputcaptured 2026-08-26
$ flux create source helm podinfo --url=https://stefanprodan.github.io/podinfo --interval=10m
✚ generating HelmRepository source
► applying HelmRepository source
✔ source created
◎ waiting for HelmRepository source reconciliation
✔ HelmRepository source reconciliation completed
✔ fetched revision: sha256:616e9b4128d1df741234ee73d2411fd56d17c47046c506a7f47b82416c946a9b
$ flux create helmrelease podinfo --source=HelmRepository/podinfo \
--chart=podinfo --release-name=podinfo --target-namespace=flux-demo --interval=5m
✚ generating HelmRelease
► applying HelmRelease
✔ HelmRelease created
◎ waiting for HelmRelease reconciliation
✔ HelmRelease podinfo is ready
✔ applied revision 6.14.1
$ kubectl -n flux-demo get deploy podinfo
NAME READY UP-TO-DATE AVAILABLE AGE
podinfo 1/1 1 1 48sBreak it once: set --chart-version='>99.0.0' on a new helmrelease, read the failure in flux get helmreleases and kubectl describe helmrelease, then delete it. Failed source resolution vs failed install vs failed health check appear in different conditions, and knowing which layer failed is the troubleshooting pattern.
flux get helmreleases Ready True, release revision 1, and helm list -n flux-system shows the release: Helm storage lives with the HelmRelease, not the workload. (Without --release-name, helm-controller would have named the release flux-demo-podinfo and the Deployment with it.)Write down, from memory, the Flux equivalent of the Argo CD Application demo-staging from 2.2 (GitRepository exists; you need one Kustomization spec: path, targetNamespace, prune, interval). Verify against flux create kustomization --export.
Remediation is the part of a HelmRelease you only meet when something breaks, which is the worst time to read the field names. Break one deliberately and watch the retry count, the reason strings, and the recovery.
# a leftover release makes this an upgrade, and the install remediation never runs
flux -n flux-demo delete hr broken -s --timeout=90s 2>/dev/null || true
kubectl -n flux-demo patch hr broken --type merge -p '{"metadata":{"finalizers":null}}' 2>/dev/null || true
kubectl -n flux-demo delete hr broken --ignore-not-found --timeout=60s
kubectl -n flux-demo wait --for=delete hr/broken --timeout=60s 2>/dev/null || true
flux create source helm podinfo --url=https://stefanprodan.github.io/podinfo --interval=10m
kubectl create ns flux-demo --dry-run=client -o yaml | kubectl apply -f -
kubectl apply -f - <<'EOF'
apiVersion: helm.toolkit.fluxcd.io/v2
kind: HelmRelease
metadata: { name: broken, namespace: flux-demo }
spec:
interval: 1m
chart:
spec:
chart: podinfo
sourceRef: { kind: HelmRepository, name: podinfo, namespace: flux-system }
install:
remediation: { retries: 2 }
timeout: 1m
values:
image:
repository: kind-registry:5000/does-not-exist
tag: nope
EOF
sleep 210
kubectl -n flux-demo get hr broken -o jsonpath='{.status.conditions}' | jq '.[] | {type, status, reason, message}'
kubectl -n flux-demo get hr broken -o jsonpath='{.status.installFailures}{"\n"}'
flux -n flux-demo events --for HelmRelease/broken | tail -15
# the install remediation already uninstalled, so the recovery runs as an upgrade:
# give it a timeout it can meet and a retry budget, or --reset clears a counter and the retry stalls anyway
kubectl -n flux-demo patch hr broken --type merge -p '{"spec":{"timeout":"5m","upgrade":{"remediation":{"retries":2}},"values":{"image":{"repository":"ghcr.io/stefanprodan/podinfo","tag":"6.7.1"}}}}'
flux -n flux-demo reconcile hr broken --reset --timeout=60s || true
kubectl -n flux-demo get hr broken -o jsonpath='{.status.conditions}' | jq '.[] | {type, status, reason}'
flux -n flux-demo delete hr broken -s --timeout=90s
flux -n flux-system delete source helm podinfo -soutputcaptured 2026-09-12
$ # a leftover release makes this an upgrade, and the install remediation never runs
$ flux -n flux-demo delete hr broken -s --timeout=90s 2>/dev/null || true
$ kubectl -n flux-demo patch hr broken --type merge -p '{"metadata":{"finalizers":null}}' 2>/dev/null || true
$ kubectl -n flux-demo delete hr broken --ignore-not-found --timeout=60s
$ kubectl -n flux-demo wait --for=delete hr/broken --timeout=60s 2>/dev/null || true
$ flux create source helm podinfo --url=https://stefanprodan.github.io/podinfo --interval=10m
✚ generating HelmRepository source
► applying HelmRepository source
✔ source created
◎ waiting for HelmRepository source reconciliation
✔ HelmRepository source reconciliation completed
✔ fetched revision: sha256:e7dc68a4dec90a35c2c6d8cdfedb7eaaee17fde45dced5898289df85069ec089
$ kubectl create ns flux-demo --dry-run=client -o yaml | kubectl apply -f -
namespace/flux-demo unchanged
$ kubectl apply -f - <<'EOF'
apiVersion: helm.toolkit.fluxcd.io/v2
kind: HelmRelease
metadata: { name: broken, namespace: flux-demo }
spec:
interval: 1m
chart:
spec:
chart: podinfo
sourceRef: { kind: HelmRepository, name: podinfo, namespace: flux-system }
install:
remediation: { retries: 2 }
timeout: 1m
values:
image:
repository: kind-registry:5000/does-not-exist
tag: nope
EOF
helmrelease.helm.toolkit.fluxcd.io/broken created
$ sleep 210
$ kubectl -n flux-demo get hr broken -o jsonpath='{.status.conditions}' | jq '.[] | {type, status, reason, message}'
{
"type": "Stalled",
"status": "True",
"reason": "RetriesExceeded",
"message": "Failed to install after 3 attempt(s)"
}
{
"type": "Ready",
"status": "False",
"reason": "InstallFailed",
"message": "Helm install failed for release flux-demo/broken with chart podinfo@6.15.0: timeout waiting for: [Deployment/flux-demo/broken-podinfo status: 'InProgress']"
}
{
"type": "Released",
"status": "False",
"reason": "InstallFailed",
"message": "Helm install failed for release flux-demo/broken with chart podinfo@6.15.0: timeout waiting for: [Deployment/flux-demo/broken-podinfo status: 'InProgress']"
}
$ kubectl -n flux-demo get hr broken -o jsonpath='{.status.installFailures}{"\n"}'
3
$ flux -n flux-demo events --for HelmRelease/broken | tail -15
2026-09-13T16:35:17.741388114Z: creating resource(s): {"resources":2}
2026-09-13T16:35:17.741398875Z: using server-side apply for resource creation: {"dryRun":false,"fieldValidationDirective":"Strict","forceConflicts":false}
2026-09-13T16:35:17.925407257Z: Created resource via patch: {"gvk":"/v1, Kind=Service","name":"broken-podinfo","namespace":"flux-demo"}
2026-09-13T16:35:18.0179015Z: Created resource via patch: {"gvk":"apps/v1, Kind=Deployment","name":"broken-podinfo","namespace":"flux-demo"}
flux-demo 74s (x2 over 2m26s) Normal UninstallSucceeded HelmRelease/broken Helm uninstall remediation for release flux-demo/broken.v1 with chart podinfo@6.15.0 succeeded
flux-system 29s (x3 over 2m28s) Normal ArtifactUpToDate HelmChart/flux-demo-broken artifact up-to-date with remote revision: '6.15.0'
flux-demo 1s Warning InstallFailed HelmRelease/broken Helm install failed for release flux-demo/broken with chart podinfo@6.15.0: timeout waiting for: [Deployment/flux-demo/broken-podinfo status: 'InProgress']
Last Helm logs:
2026-09-13T16:36:31.153330997Z: creating resource(s): {"resources":2}
2026-09-13T16:36:31.153335997Z: using server-side apply for resource creation: {"dryRun":false,"fieldValidationDirective":"Strict","forceConflicts":false}
2026-09-13T16:36:31.198687584Z: Created resource via patch: {"gvk":"/v1, Kind=Service","name":"broken-podinfo","namespace":"flux-demo"}
2026-09-13T16:36:31.405086633Z: Created resource via patch: {"gvk":"apps/v1, Kind=Deployment","name":"broken-podinfo","namespace":"flux-demo"}
$ # the install remediation already uninstalled, so the recovery runs as an upgrade:
$ # give it a timeout it can meet and a retry budget, or --reset clears a counter and the retry stalls anyway
$ kubectl -n flux-demo patch hr broken --type merge -p '{"spec":{"timeout":"5m","upgrade":{"remediation":{"retries":2}},"values":{"image":{"repository":"ghcr.io/stefanprodan/podinfo","tag":"6.7.1"}}}}'
helmrelease.helm.toolkit.fluxcd.io/broken patched
$ flux -n flux-demo reconcile hr broken --reset --timeout=60s || true
► annotating HelmRelease broken in flux-demo namespace
✔ HelmRelease annotated
◎ waiting for HelmRelease reconciliation
✗ context deadline exceeded
$ kubectl -n flux-demo get hr broken -o jsonpath='{.status.conditions}' | jq '.[] | {type, status, reason}'
{
"type": "Reconciling",
"status": "True",
"reason": "Progressing"
}
{
"type": "Ready",
"status": "Unknown",
"reason": "Progressing"
}
{
"type": "Released",
"status": "False",
"reason": "InstallFailed"
}
$ flux -n flux-demo delete hr broken -s --timeout=90s
► deleting helmrelease broken in flux-demo namespace
✔ helmrelease deleted
$ flux -n flux-system delete source helm podinfo -s
► deleting source helm podinfo in flux-system namespace
✔ source helm deletedStalled with reason RetriesExceeded, the message Failed to install after 3 attempt(s), and status.installFailures at 3. --reset then clears that counter and the release goes back to Reconciling / Progressing, which is why the 60 second wait returns context deadline exceeded rather than a verdict: resetting starts the attempts again, it does not finish them. Delete any earlier release first, and clear its finalizer: with one already in history this is an upgrade rather than an install, and you get UpgradeFailed and MissingRollbackTarget from an install.remediation block that never ran.A Kustomization corrects drift because it applies every interval. A HelmRelease does not, unless you turn drift detection on: Helm only acts when the release changes. This is the single most surprising difference between the two.
kubectl -n flux-demo get hr podinfo -o jsonpath='{.spec.driftDetection}{"\n"}'
kubectl -n flux-demo patch hr podinfo --type merge -p '{"spec":{"driftDetection":{"mode":"enabled"}}}'
kubectl -n flux-demo scale deploy podinfo --replicas=5
kubectl -n flux-demo get deploy podinfo -o jsonpath='{.spec.replicas}{"\n"}'
sleep 120
kubectl -n flux-demo get deploy podinfo -o jsonpath='{.spec.replicas}{"\n"}'
flux -n flux-demo events --for HelmRelease/podinfo | tail -10
kubectl -n flux-demo get hr podinfo -o jsonpath='{.status.conditions}' | jq '.[] | {type, status, reason}'outputcaptured 2026-09-12
$ kubectl -n flux-demo get hr podinfo -o jsonpath='{.spec.driftDetection}{"\n"}'
$ kubectl -n flux-demo patch hr podinfo --type merge -p '{"spec":{"driftDetection":{"mode":"enabled"}}}'
helmrelease.helm.toolkit.fluxcd.io/podinfo patched
$ kubectl -n flux-demo scale deploy podinfo --replicas=5
deployment.apps/podinfo scaled
$ kubectl -n flux-demo get deploy podinfo -o jsonpath='{.spec.replicas}{"\n"}'
5
$ sleep 120
$ kubectl -n flux-demo get deploy podinfo -o jsonpath='{.spec.replicas}{"\n"}'
1
$ flux -n flux-demo events --for HelmRelease/podinfo | tail -10
flux-system 13m Normal NewArtifact HelmRepository/podinfo stored fetched index of size 81.8kB from 'https://stefanprodan.github.io/podinfo'
flux-demo 13m Normal HelmChartCreated HelmRelease/podinfo Created HelmChart/flux-system/flux-demo-podinfo with SourceRef 'HelmRepository/flux-system/podinfo'
flux-system 13m Normal ChartPullSucceeded HelmChart/flux-demo-podinfo pulled 'podinfo' chart with version '6.15.0'
flux-demo 12m Normal InstallSucceeded HelmRelease/podinfo Helm install succeeded for release flux-demo/podinfo.v1 with chart podinfo@6.15.0
flux-system 2m57s Normal ArtifactUpToDate HelmRepository/podinfo artifact up-to-date with remote revision: 'sha256:e7dc68a4dec90a35c2c6d8cdfedb7eaaee17fde45dced5898289df85069ec089'
flux-demo 60s Warning DriftDetected HelmRelease/podinfo Cluster state of release flux-demo/podinfo.v1 has drifted from the desired state:
Deployment/flux-demo/podinfo changed (0 additions, 1 changes, 0 removals)
flux-demo 60s Normal DriftCorrected HelmRelease/podinfo Cluster state of release flux-demo/podinfo.v1 has been corrected:
Deployment/flux-demo/podinfo configured
flux-system 14s (x13 over 12m) Normal ArtifactUpToDate HelmChart/flux-demo-podinfo artifact up-to-date with remote revision: '6.15.0'
$ kubectl -n flux-demo get hr podinfo -o jsonpath='{.status.conditions}' | jq '.[] | {type, status, reason}'
{
"type": "Ready",
"status": "True",
"reason": "InstallSucceeded"
}
{
"type": "Drifted",
"status": "False",
"reason": "NoDriftDetected"
}
{
"type": "Released",
"status": "True",
"reason": "InstallSucceeded"
}podinfo is not installed, do that one first.Pruning is what makes a Kustomization authoritative, and sometimes one object must survive a delete in git anyway. The exemption is an annotation on the object, which means the escape hatch lives in the manifests, not in the controller.
START=$PWD # the cd below would otherwise follow you into every later command
rm -rf /tmp/platform-prune-flux && git clone "http://lab:${GITEA_PASS}@gitea.lab:3000/lab/platform.git" /tmp/platform-prune-flux && cd /tmp/platform-prune-flux
kubectl -n flux-demo annotate deploy demo kustomize.toolkit.fluxcd.io/prune=disabled --overwrite
git rm demo-app/base/deployment.yaml
sed -i '/deployment.yaml/d' demo-app/base/kustomization.yaml
git commit -am 'remove the deployment from git' && git push
flux reconcile kustomization demo-flux --with-source
sleep 20
kubectl -n flux-demo get deploy demo
flux events --for Kustomization/demo-flux | tail -10
cd /tmp/platform-prune-flux && git revert --no-edit HEAD && git push
flux reconcile kustomization demo-flux --with-source
kubectl -n flux-demo annotate deploy demo kustomize.toolkit.fluxcd.io/prune-
cd "$START"outputcaptured 2026-09-12
$ START=$PWD # the cd below would otherwise follow you into every later command
$ rm -rf /tmp/platform-prune-flux && git clone "http://lab:${GITEA_PASS}@gitea.lab:3000/lab/platform.git" /tmp/platform-prune-flux && cd /tmp/platform-prune-flux
Cloning into '/tmp/platform-prune-flux'...
$ kubectl -n flux-demo annotate deploy demo kustomize.toolkit.fluxcd.io/prune=disabled --overwrite
deployment.apps/demo annotated
$ git rm demo-app/base/deployment.yaml
rm 'demo-app/base/deployment.yaml'
$ sed -i '/deployment.yaml/d' demo-app/base/kustomization.yaml
$ git commit -am 'remove the deployment from git' && git push
[main 793d5b3] remove the deployment from git
2 files changed, 40 deletions(-)
delete mode 100644 demo-app/base/deployment.yaml
To http://gitea.lab:3000/lab/platform.git
83a41e2..793d5b3 main -> main
$ flux reconcile kustomization demo-flux --with-source
► annotating GitRepository platform in flux-system namespace
✔ GitRepository annotated
◎ waiting for GitRepository reconciliation
✔ fetched revision main@sha1:793d5b378e2e11c0eb62f8f322a7f6acb0846077
► annotating Kustomization demo-flux in flux-system namespace
✔ Kustomization annotated
◎ waiting for Kustomization reconciliation
✔ applied revision main@sha1:793d5b378e2e11c0eb62f8f322a7f6acb0846077
$ sleep 20
$ kubectl -n flux-demo get deploy demo
NAME READY UP-TO-DATE AVAILABLE AGE
demo 2/2 2 2 5h34m
$ flux events --for Kustomization/demo-flux | tail -10
5h34m Normal ReconciliationSucceeded Kustomization/demo-flux Reconciliation finished in 1.695245538s, next run in 1m0s
5h33m Normal ReconciliationSucceeded Kustomization/demo-flux Reconciliation finished in 403.342017ms, next run in 1m0s
5h32m Normal ReconciliationSucceeded Kustomization/demo-flux Reconciliation finished in 522.080036ms, next run in 1m0s
5h31m Normal ReconciliationSucceeded Kustomization/demo-flux Reconciliation finished in 1.613648954s, next run in 1m0s
5h30m Normal ReconciliationSucceeded Kustomization/demo-flux Reconciliation finished in 1.983666745s, next run in 1m0s
5h29m Normal ReconciliationSucceeded Kustomization/demo-flux Reconciliation finished in 442.161038ms, next run in 1m0s
5h28m Normal ReconciliationSucceeded Kustomization/demo-flux Reconciliation finished in 269.029392ms, next run in 1m0s
5h27m Normal ReconciliationSucceeded Kustomization/demo-flux Reconciliation finished in 669.242058ms, next run in 1m0s
3m47s (x85 over 6h20m) Normal GitOperationSucceeded GitRepository/platform no changes since last reconciliation: observed revision 'main@sha1:83a41e21322e15eff2307bc1e9d5c89e33226d0d'
2m24s (x32 over 5h26m) Normal ReconciliationSucceeded Kustomization/demo-flux (combined from similar events): Reconciliation finished in 435.434931ms, next run in 1m0s
$ cd /tmp/platform-prune-flux && git revert --no-edit HEAD && git push
[main dfe88a6] Revert "remove the deployment from git"
Date: Sun Sep 13 07:15:13 2026 -0400
2 files changed, 40 insertions(+)
create mode 100644 demo-app/base/deployment.yaml
To http://gitea.lab:3000/lab/platform.git
793d5b3..dfe88a6 main -> main
$ flux reconcile kustomization demo-flux --with-source
► annotating GitRepository platform in flux-system namespace
✔ GitRepository annotated
◎ waiting for GitRepository reconciliation
✔ fetched revision main@sha1:dfe88a660f9591488e000cc24ba120ac1186b7b1
► annotating Kustomization demo-flux in flux-system namespace
✔ Kustomization annotated
◎ waiting for Kustomization reconciliation
✔ applied revision main@sha1:dfe88a660f9591488e000cc24ba120ac1186b7b1
$ kubectl -n flux-demo annotate deploy demo kustomize.toolkit.fluxcd.io/prune-
deployment.apps/demo annotated
$ cd "$START"Half the Flux tickets you will be handed are a path that does not exist in the repo. The condition is not "cannot connect" and not "cannot apply"; learn the exact reason so you can skip straight to the right half of the problem.
kubectl apply -f - <<'EOF'
apiVersion: kustomize.toolkit.fluxcd.io/v1
kind: Kustomization
metadata: { name: nowhere, namespace: flux-system }
spec:
interval: 1m
prune: true
sourceRef: { kind: GitRepository, name: platform }
path: ./does-not-exist
targetNamespace: flux-demo
EOF
sleep 30
kubectl -n flux-system get kustomization nowhere -o jsonpath='{.status.conditions}' | jq '.[] | {type, status, reason, message}'
flux events --for Kustomization/nowhere | tail -5
flux delete kustomization nowhere -soutputcaptured 2026-09-13
$ kubectl apply -f - <<'EOF'
apiVersion: kustomize.toolkit.fluxcd.io/v1
kind: Kustomization
metadata: { name: nowhere, namespace: flux-system }
spec:
interval: 1m
prune: true
sourceRef: { kind: GitRepository, name: platform }
path: ./does-not-exist
targetNamespace: flux-demo
EOF
kustomization.kustomize.toolkit.fluxcd.io/nowhere created
$ sleep 30
$ kubectl -n flux-system get kustomization nowhere -o jsonpath='{.status.conditions}' | jq '.[] | {type, status, reason, message}'
{
"type": "Reconciling",
"status": "True",
"reason": "ProgressingWithRetry",
"message": "Fetching manifests for revision main@sha1:2cb925cb149a0393023a477f919d901cfec7e371 with a timeout of 30s"
}
{
"type": "Ready",
"status": "False",
"reason": "ArtifactFailed",
"message": "kustomization path not found: stat /tmp/kustomization-552514629/does-not-exist: no such file or directory"
}
$ flux events --for Kustomization/nowhere | tail -5
LAST SEEN TYPE REASON OBJECT MESSAGE
27m (x85 over 6h44m) Normal GitOperationSucceeded GitRepository/platform no changes since last reconciliation: observed revision 'main@sha1:83a41e21322e15eff2307bc1e9d5c89e33226d0d'
22m (x11 over 13h) Normal GarbageCollectionSucceeded GitRepository/platform garbage collected 1 artifacts
2m4s (x20 over 21m) Normal GitOperationSucceeded GitRepository/platform no changes since last reconciliation: observed revision 'main@sha1:2cb925cb149a0393023a477f919d901cfec7e371'
29s Warning ArtifactFailed Kustomization/nowhere kustomization path not found: stat /tmp/kustomization-552514629/does-not-exist: no such file or directory
$ flux delete kustomization nowhere -s
► deleting kustomization nowhere in flux-system namespace
✔ kustomization deletedAn interval is a promise that you will be no more than a minute behind. A Receiver is how you stop waiting: the git server calls Flux, and the reconcile starts on the push.
TOKEN=$(head -c 24 /dev/urandom | base64 | tr -d '=+/')
kubectl -n flux-system create secret generic receiver-token --from-literal=token="$TOKEN"
kubectl apply -f - <<'EOF'
apiVersion: notification.toolkit.fluxcd.io/v1
kind: Receiver
metadata: { name: gitea-receiver, namespace: flux-system }
spec:
# this Flux build has no "gitea" type; Gitea sends GitHub-format webhooks
type: github
events: [push]
secretRef: { name: receiver-token }
resources:
- { kind: GitRepository, name: platform }
EOF
sleep 10
kubectl -n flux-system get receiver gitea-receiver -o jsonpath='{.status.webhookPath}{"\n"}'
kubectl -n flux-system patch svc webhook-receiver -p '{"spec":{"type":"NodePort"}}'
# a LoadBalancer IP needs cloud-provider-kind; the NodePort is always reachable from the host
NODE_IP=$(kubectl get node cnpe-control-plane -o jsonpath='{.status.addresses[?(@.type=="InternalIP")].address}')
NODE_PORT=$(kubectl -n flux-system get svc webhook-receiver -o jsonpath='{.spec.ports[0].nodePort}')
HOOK=$(kubectl -n flux-system get receiver gitea-receiver -o jsonpath='{.status.webhookPath}')
echo "webhook URL for Gitea: http://$NODE_IP:$NODE_PORT$HOOK"
flux events --for GitRepository/platform | tail -5
# put the lab back: the Receiver and its token are the only things this exercise needed
kubectl -n flux-system delete receiver gitea-receiver
kubectl -n flux-system delete secret receiver-token
kubectl -n flux-system patch svc webhook-receiver -p '{"spec":{"type":"ClusterIP"}}'outputcaptured 2026-09-13
$ TOKEN=$(head -c 24 /dev/urandom | base64 | tr -d '=+/')
$ kubectl -n flux-system create secret generic receiver-token --from-literal=token="$TOKEN"
secret/receiver-token created
$ kubectl apply -f - <<'EOF'
apiVersion: notification.toolkit.fluxcd.io/v1
kind: Receiver
metadata: { name: gitea-receiver, namespace: flux-system }
spec:
# this Flux build has no "gitea" type; Gitea sends GitHub-format webhooks
type: github
events: [push]
secretRef: { name: receiver-token }
resources:
- { kind: GitRepository, name: platform }
EOF
receiver.notification.toolkit.fluxcd.io/gitea-receiver created
$ sleep 10
$ kubectl -n flux-system get receiver gitea-receiver -o jsonpath='{.status.webhookPath}{"\n"}'
/hook/a2b76ad1fe000998707d8a990b50c26a825e15ea38b8b30872ae0f7062646270
$ kubectl -n flux-system patch svc webhook-receiver -p '{"spec":{"type":"NodePort"}}'
service/webhook-receiver patched
$ # a LoadBalancer IP needs cloud-provider-kind; the NodePort is always reachable from the host
$ NODE_IP=$(kubectl get node cnpe-control-plane -o jsonpath='{.status.addresses[?(@.type=="InternalIP")].address}')
$ NODE_PORT=$(kubectl -n flux-system get svc webhook-receiver -o jsonpath='{.spec.ports[0].nodePort}')
$ HOOK=$(kubectl -n flux-system get receiver gitea-receiver -o jsonpath='{.status.webhookPath}')
$ echo "webhook URL for Gitea: http://$NODE_IP:$NODE_PORT$HOOK"
webhook URL for Gitea: http://172.18.0.4:31536/hook/a2b76ad1fe000998707d8a990b50c26a825e15ea38b8b30872ae0f7062646270
$ flux events --for GitRepository/platform | tail -5
LAST SEEN TYPE REASON OBJECT MESSAGE
27m (x85 over 6h44m) Normal GitOperationSucceeded GitRepository/platform no changes since last reconciliation: observed revision 'main@sha1:83a41e21322e15eff2307bc1e9d5c89e33226d0d'
22m (x11 over 14h) Normal GarbageCollectionSucceeded GitRepository/platform garbage collected 1 artifacts
2m15s (x20 over 21m) Normal GitOperationSucceeded GitRepository/platform no changes since last reconciliation: observed revision 'main@sha1:2cb925cb149a0393023a477f919d901cfec7e371'
$ # put the lab back: the Receiver and its token are the only things this exercise needed
$ kubectl -n flux-system delete receiver gitea-receiver
receiver.notification.toolkit.fluxcd.io "gitea-receiver" deleted from flux-system namespace
$ kubectl -n flux-system delete secret receiver-token
secret "receiver-token" deleted from flux-system namespace
$ kubectl -n flux-system patch svc webhook-receiver -p '{"spec":{"type":"ClusterIP"}}'
service/webhook-receiver patchedstatus.webhookPath of the form /hook/<64 hex characters>, and the last line prints the full URL to paste into Gitea, for example http://172.18.0.4:31536/hook/a2b76ad1.... Note the type: this Flux build has no gitea receiver type, and Gitea sends GitHub-format webhooks, so github is the right answer. Point the repository's webhook at that URL with the same token and a push reconciles within seconds instead of at the next interval.Image automation is Flux writing to git on your behalf: a policy picks a tag, and a controller commits the change back to the overlay. The marker comment in the manifest is what tells it where to write.
flux create image repository demo --image=kind-registry:5000/demo --interval=1m
kubectl -n flux-system patch imagerepository demo --type merge -p '{"spec":{"insecure":true}}'
flux create image policy demo --image-ref=demo --select-semver='>=1.0.0'
skopeo copy --dest-tls-verify=false docker://ghcr.io/stefanprodan/podinfo:6.7.1 docker://localhost:5001/demo:1.0.1
sleep 90
kubectl -n flux-system get imagerepository demo -o jsonpath='{.status.lastScanResult}' | jq
kubectl -n flux-system get imagepolicy demo -o jsonpath='{.status.latestRef}' | jq
# the automation rewrites the line carrying the marker comment, and there is none in the repo yet.
# keep it out of demo-app/overlays: Argo CD syncs that tree and would deploy the rewritten tag into team-a.
rm -rf /tmp/platform-img && git clone "http://lab:${GITEA_PASS}@gitea.lab:3000/lab/platform.git" /tmp/platform-img
mkdir -p /tmp/platform-img/image-automation
cat > /tmp/platform-img/image-automation/deployment.yaml <<'EOF'
apiVersion: apps/v1
kind: Deployment
metadata: { name: image-automation-demo, namespace: flux-demo }
spec:
replicas: 1
selector: { matchLabels: { app: image-automation-demo } }
template:
metadata: { labels: { app: image-automation-demo } }
spec:
containers:
- name: app
image: kind-registry:5000/demo:1.0.0 # {"$imagepolicy": "flux-system:demo"}
EOF
git -C /tmp/platform-img add image-automation/deployment.yaml
git -C /tmp/platform-img -c user.name=lab -c user.email=lab@lab.local commit -m 'a manifest for the automation to rewrite'
git -C /tmp/platform-img push
flux create image update demo-auto \
--interval=1m \
--git-repo-ref=platform \
--git-repo-path=./image-automation \
--checkout-branch=main \
--push-branch=main \
--author-name=fluxcdbot \
--author-email=fluxcdbot@lab.local
flux -n flux-system reconcile image update demo-auto
sleep 30
flux events --for ImageUpdateAutomation/demo-auto | tail -10
git -C /tmp/platform-img fetch && git -C /tmp/platform-img log --format='%an %s' -1 origin/main
# three objects that would otherwise keep pushing to main forever
flux -n flux-system delete image update demo-auto -s
flux -n flux-system delete image policy demo -s
flux -n flux-system delete image repository demo -s
git -C /tmp/platform-img pull --rebase && git -C /tmp/platform-img rm -r image-automation && git -C /tmp/platform-img -c user.name=lab -c user.email=lab@lab.local commit -m 'remove the automation target' && git -C /tmp/platform-img push
rm -rf /tmp/platform-imgoutputcaptured 2026-09-13
$ flux create image repository demo --image=kind-registry:5000/demo --interval=1m
✚ generating ImageRepository
► applying ImageRepository
✔ ImageRepository updated
◎ waiting for ImageRepository reconciliation
✗ context deadline exceeded
$ kubectl -n flux-system patch imagerepository demo --type merge -p '{"spec":{"insecure":true}}'
imagerepository.image.toolkit.fluxcd.io/demo patched
$ flux create image policy demo --image-ref=demo --select-semver='>=1.0.0'
✚ generating ImagePolicy
► applying ImagePolicy
✔ ImagePolicy created
◎ waiting for ImagePolicy reconciliation
✔ ImagePolicy reconciliation completed
$ skopeo copy --dest-tls-verify=false docker://ghcr.io/stefanprodan/podinfo:6.7.1 docker://localhost:5001/demo:1.0.1
Getting image source signatures
Copying blob sha256:7e517c53cd4ca936c3bc3d3dc66419278a680214f37444ad0ad1b8bd65968940
Copying blob sha256:43c4264eed91be63b206e17d93e75256a6097070ce643c5e8f0379998b44f170
Copying blob sha256:4f4fb700ef54461cfa02571ae0db9a0dc1e0cdb5577484a6d75e68dc38e8acc1
Copying blob sha256:35d14af0e611499048effc4d39a2341345f96dd9103c8c186b7cc239a687de5c
Copying blob sha256:2bf70b2badb9fa00aa36d50d61c6b9cc094d7a18fb4a3c7a91ce950dd7202900
Copying blob sha256:31eaa2883f08a0ce854d061714a111f052444a3ba1b2323719cb12b872a0dbe3
Copying blob sha256:47f87f374862e0c7555781b116554936999ad2e8e4c6a4e7417688ffe57d8b88
Copying config sha256:c875de4397634b0dc138954cbe2c11b9d86d948b98b17bd9fc3227b3ad992eae
Writing manifest to image destination
$ sleep 90
$ kubectl -n flux-system get imagerepository demo -o jsonpath='{.status.lastScanResult}' | jq
{
"latestTags": [
"v1",
"unsigned",
"sha256-f451e838d60a425de0f81bd3126670198c3f7f02b89e5e010430abafe66466fa",
"sha256-ed73e9871dfba28276b28daab3463cc2f9b3c15337dfe0445e8b0f79694e0cd6",
"sha256-33213719f2b030a8a1260d1c5358f55524f8f690a790e35b14b00c1f50b46ddf",
"fake",
"1.0.1"
],
"revision": "2060731329",
"scanTime": "2026-09-13T12:31:14Z",
"tagCount": 7
}
$ kubectl -n flux-system get imagepolicy demo -o jsonpath='{.status.latestRef}' | jq
{
"name": "kind-registry:5000/demo",
"tag": "1.0.1"
}
$ # the automation rewrites the line carrying the marker comment, and there is none in the repo yet.
$ # keep it out of demo-app/overlays: Argo CD syncs that tree and would deploy the rewritten tag into team-a.
$ rm -rf /tmp/platform-img && git clone "http://lab:${GITEA_PASS}@gitea.lab:3000/lab/platform.git" /tmp/platform-img
Cloning into '/tmp/platform-img'...
$ mkdir -p /tmp/platform-img/image-automation
$ cat > /tmp/platform-img/image-automation/deployment.yaml <<'EOF'
apiVersion: apps/v1
kind: Deployment
metadata: { name: image-automation-demo, namespace: flux-demo }
spec:
replicas: 1
selector: { matchLabels: { app: image-automation-demo } }
template:
metadata: { labels: { app: image-automation-demo } }
spec:
containers:
- name: app
image: kind-registry:5000/demo:1.0.0 # {"$imagepolicy": "flux-system:demo"}
EOF
$ git -C /tmp/platform-img add image-automation/deployment.yaml
$ git -C /tmp/platform-img -c user.name=lab -c user.email=lab@lab.local commit -m 'a manifest for the automation to rewrite'
[main 226a068] a manifest for the automation to rewrite
1 file changed, 12 insertions(+)
create mode 100644 image-automation/deployment.yaml
$ git -C /tmp/platform-img push
To http://gitea.lab:3000/lab/platform.git
1d38064..226a068 main -> main
$ flux create image update demo-auto \
--interval=1m \
--git-repo-ref=platform \
--git-repo-path=./image-automation \
--checkout-branch=main \
--push-branch=main \
--author-name=fluxcdbot \
--author-email=fluxcdbot@lab.local
✚ generating ImageUpdateAutomation
► applying ImageUpdateAutomation
✔ ImageUpdateAutomation created
◎ waiting for ImageUpdateAutomation reconciliation
✔ ImageUpdateAutomation reconciliation completed
$ flux -n flux-system reconcile image update demo-auto
► annotating ImageUpdateAutomation demo-auto in flux-system namespace
✔ ImageUpdateAutomation annotated
◎ waiting for ImageUpdateAutomation reconciliation
✔ repository up-to-date
$ sleep 30
$ flux events --for ImageUpdateAutomation/demo-auto | tail -10
LAST SEEN TYPE REASON OBJECT MESSAGE
34s Normal Succeeded ImageUpdateAutomation/demo-auto pushed commit 'ae19a2b' to branch 'main'
Update from image update automation
31s Normal Succeeded ImageUpdateAutomation/demo-auto no change since last reconciliation
$ git -C /tmp/platform-img fetch && git -C /tmp/platform-img log --format='%an %s' -1 origin/main
From http://gitea.lab:3000/lab/platform
226a068..ae19a2b main -> origin/main
fluxcdbot Update from image update automation
$ # three objects that would otherwise keep pushing to main forever
$ flux -n flux-system delete image update demo-auto -s
► deleting image update automation demo-auto in flux-system namespace
✔ image update automation deleted
$ flux -n flux-system delete image policy demo -s
► deleting image policy demo in flux-system namespace
✔ image policy deleted
$ flux -n flux-system delete image repository demo -s
► deleting image repository demo in flux-system namespace
✔ image repository deleted
$ git -C /tmp/platform-img pull --rebase && git -C /tmp/platform-img rm -r image-automation && git -C /tmp/platform-img -c user.name=lab -c user.email=lab@lab.local commit -m 'remove the automation target' && git -C /tmp/platform-img push
Updating 226a068..ae19a2b
Fast-forward
image-automation/deployment.yaml | 8 ++++----
1 file changed, 4 insertions(+), 4 deletions(-)
rm 'image-automation/deployment.yaml'
[main b793ef2] remove the automation target
1 file changed, 12 deletions(-)
delete mode 100644 image-automation/deployment.yaml
To http://gitea.lab:3000/lab/platform.git
ae19a2b..b793ef2 main -> main
$ rm -rf /tmp/platform-img1.0.1, and the newest commit on main is authored by fluxcdbot and rewrites the image line to kind-registry:5000/demo:1.0.1. Without the marker comment on that line the automation has nothing to rewrite and reports repository up-to-date, which is the failure you will be handed. Keep the marked manifest out of demo-app/overlays: Argo CD syncs that tree and would deploy the rewritten tag straight into team-a.Self-check
Flux has no selfHeal setting. Why is drift still corrected?
Because the Kustomization re-applies its rendered manifests every interval; correcting drift is a side effect of applying, not a separate feature. The corollary: your drift window is bounded by the interval, and shortening it costs API server load.
A Kustomization is Ready False. Name the three layers where it could have failed.
Source (the GitRepository is not Ready: auth, ref, or network), build (kustomize build failed, bad path or invalid YAML), apply/health (the API server rejected something, or health checks timed out). The conditions and flux logs name the layer; do not skip straight to the manifests.
How do you order "install CRDs, then the operator, then the tenant CRs" in Flux?
Three Kustomizations with dependsOn, each with wait: true (or explicit healthChecks) so a dependency counts as satisfied only when its objects are actually healthy. Without wait/healthChecks, dependsOn only orders the apply, not the readiness.
What is postBuild.substitute for, and what is its Argo CD analog?
Injecting per-cluster or per-environment values into manifests after kustomize build, sourced inline or from ConfigMaps/Secrets. Argo's analogs are kustomize overlays, Helm values per Application, or ApplicationSet template parameters; Flux just lets you do it without another overlay directory.
How does Flux stop one tenant's Kustomization from deploying cluster-admin-level resources?
spec.serviceAccountName: the controller impersonates that ServiceAccount when applying, so the tenant's manifests can never exceed the tenant's RBAC. Combined with disallowing cross-namespace source references, that is Flux's multi-tenancy story: the structural equivalent of Argo's AppProject.
A HelmRelease shows Ready False with reason RollbackSucceeded. Is the release broken or healthy right now?
Healthy on the last good revision: the upgrade failed and remediation rolled back. The object stays not-ready because the desired state (new chart or values) is not what is running. Fix the chart or values in git, or after fixing something out of band run flux reconcile hr X --reset to zero the retry counters if RetriesExceeded stopped further attempts.
A chart upgrade adds a CRD field but the CR that uses it is rejected as unknown. Why, and which field fixes it?
spec.upgrade.crds defaults to Skip: Helm does not upgrade CRDs on upgrade, so the old schema is still served and prunes or rejects the new field. Set upgrade.crds: CreateReplace (and install.crds: CreateReplace), or deliver CRDs through their own Kustomization that the HelmRelease depends on.
Make a git push reconcile Flux immediately instead of at the next interval. Objects and the field you hand to the git server?
A Receiver of the matching type (github, gitea, gitlab, generic) with events: [push], a secretRef holding a token, and resources listing the GitRepository. Expose the notification-controller's webhook-receiver Service, then configure the webhook URL as that host plus status.webhookPath with the same token as the secret.
After multi-tenant lockdown, a tenant's Kustomization fails with Forbidden. What did they leave out?
spec.serviceAccountName. With --default-service-account=default, a Kustomization or HelmRelease without an explicit account impersonates the namespace's default ServiceAccount, which has no RBAC; the API server refuses the apply and the Ready condition reads ReconciliationFailed with the Forbidden text. The platform must have created a ServiceAccount and RoleBinding for the tenant to name.
Docs to know your way around
- fluxcd.io: GitRepository, Kustomization, HelmRelease API references; the "flux CLI" cheat sheet; the multi-tenancy guide for the serviceAccountName pattern.
- Offline:
flux --helpandflux create <kind> --help --exportgenerate correct YAML without docs, which is faster than searching during the exam. - fluxcd.io/flux/components/kustomize/kustomizations (Kustomization Status, Conditions) and components/helm/helmreleases (Conditions, Install/Upgrade remediation, Drift detection): the reason strings quoted above and the remediation fields.
- fluxcd.io/flux/components/image and guides/image-update: ImageRepository, ImagePolicy, ImageUpdateAutomation and the marker comment syntax.
- fluxcd.io/flux/components/notification (Alerts, Providers, Receivers) and installation/configuration/multitenancy: the notification objects and the lockdown patches.