Progressive delivery is a Deployment with feedback: the new version goes to a slice of traffic, something measures whether it is worse, and the rollout continues or reverses based on evidence.
make core obsmake meshOrientation
Two strategies to know cold, and one distinction to keep straight: what actually shifts the traffic.
| Strategy | How it works | Rollback | Costs |
|---|---|---|---|
| Canary | shift traffic in steps, measure between steps | shift back; only the canary slice was exposed | slow; needs metrics worth trusting |
| Blue-green | run both versions full size, cut over at once | instant: flip the selector back | double capacity for the window |
| Rolling (baseline) | replace pods gradually, no analysis | roll forward or undo, no traffic control | what you already had |
| A/B or shadow | route by header/cookie, or mirror traffic | n/a: no user impact for shadow | needs L7 routing (mesh or gateway) |
Argo Rollouts without a traffic provider approximates a 20% canary by running 20% of the replicas: the split is statistical, granularity is limited by replica count, and every client is load-balanced by the Service. With a traffic provider (Istio, Gateway API, NGINX, ALB) or with Flagger driving a mesh, the split is a route weight: exact, independent of replica count, and able to key on headers.
Argo Rollouts
A Rollout is a drop-in replacement for a Deployment (same pod template, same selector semantics) whose spec.strategy is richer. It owns ReplicaSets exactly like a Deployment does (the stable one and the canary one), which is why kubectl get rs during a rollout is so legible.
strategy:
canary:
steps:
- setWeight: 20
- pause: { duration: 60s } # omit duration → wait for a human promote
- analysis: # run an AnalysisTemplate, act on the result
templates: [{ templateName: success-rate }]
- setWeight: 50
- pause: {}
# optional: canaryService / stableService / trafficRouting for real weightsThe step list is the whole grammar: setWeight, pause, analysis, plus setCanaryScale and experiment steps. Blue-green swaps that for two Services and a cutover:
strategy:
blueGreen:
activeService: demo-active
previewService: demo-preview
autoPromotionEnabled: false # wait for `promote`
scaleDownDelaySeconds: 30 # keep the old RS warm for instant rollbackAnalysisTemplate is the measurement half: a list of metrics, each with a provider (Prometheus, Datadog, a Job, a web request), an interval, a successCondition or failureCondition, plus failureLimit, count and inconclusiveLimit. An AnalysisRun is one execution of it. Its status carries every measurement, which is where you look when a rollout aborts and you want to know which number failed it. Analyzes can run as a rollout step, as a background analysis for the whole rollout, or as a pre-promotion/post-promotion gate in blue-green.
If the metric provider returns an error (bad query, missing series, unreachable Prometheus), the measurement is Error, and errors have their own budget: blow consecutiveErrorLimit (four in a row, by default) and the analysis kills the rollout just as dead as failures blowing failureLimit. So a canary can abort because your PromQL was wrong, not because the new version was. Read the AnalysisRun's measurements before blaming the release. In production, a broken metrics pipeline blocking a rollout is a feature, not a bug.
Step through a rollout and watch the difference the traffic provider makes:
The verbs, via the kubectl plugin
kubectl argo rollouts get rollout demo -n team-a --watch # the best mental-model builder in domain 2
kubectl argo rollouts set image demo web=<image> -n team-a
kubectl argo rollouts promote demo -n team-a # advance past a pause; --full skips remaining steps
kubectl argo rollouts abort demo -n team-a # stop and shift back to stable
kubectl argo rollouts undo demo -n team-a # roll the spec back to the previous revision
kubectl argo rollouts status demo -n team-a --timeout 60s # scriptable, exits non-zero on failureoutputcaptured 2026-08-26
$ kubectl argo rollouts get rollout demo -n team-a --watch # the best mental-model builder in domain 2
Name: demo
Namespace: team-a
Status: ✖ Degraded
Message: RolloutAborted: Rollout aborted update to revision 2: Background analysis phase error/failed: Metric "success-rate" assessed Error due to consecutiveErrors (5) > consecutiveErrorLimit (4): "Error Message: reflect: slice index out of range"
Strategy: Canary
Step: 0/5
SetWeight: 0
ActualWeight: 0
Images: ghcr.io/nginxinc/nginx-unprivileged:1.27-alpine (stable)
Replicas:
Desired: 4
Current: 4
Updated: 0
Ready: 4
Available: 4
NAME KIND STATUS AGE INFO
⟳ demo Rollout ✖ Degraded 4m16s
├──# revision:2
│ ├──⧉ demo-dff9796f6 ReplicaSet • ScaledDown 4m4s canary
│ └──α demo-dff9796f6-2 AnalysisRun ⚠ Error 3m30s ⚠ 5
└──# revision:1
└──⧉ demo-5cfc49f565 ReplicaSet ✔ Healthy 4m16s stable
├──□ demo-5cfc49f565-nx4r8 Pod ✔ Running 4m16s ready:1/1
├──□ demo-5cfc49f565-trpgb Pod ✔ Running 4m16s ready:1/1
├──□ demo-5cfc49f565-8rhc9 Pod ✔ Running 2m50s ready:1/1
└──□ demo-5cfc49f565-sqb6q Pod ✔ Running 2m50s ready:1/1
^C
$ kubectl argo rollouts set image demo web=ghcr.io/nginxinc/nginx-unprivileged:1.27-alpine -n team-a
rollout "demo" image updated
$ kubectl argo rollouts promote demo -n team-a # advance past a pause; --full skips remaining steps
rollout 'demo' promoted
$ kubectl argo rollouts abort demo -n team-a # stop and shift back to stable
rollout 'demo' aborted
$ kubectl argo rollouts undo demo -n team-a # roll the spec back to the previous revision
rollout 'demo' undo
$ kubectl argo rollouts status demo -n team-a --timeout 60s # scriptable, exits non-zero on failure
Paused - CanaryPauseStep
Progressing - more replicas need to be updated
Paused - CanaryPauseStep
Error: Rollout status watch exceeded timeoutabort stops the rollout and sends all traffic back to the stable ReplicaSet, but spec still asks for the new image, so status shows Degraded and it will try again if you touch it. undo changes the spec back. "Safely roll back" on an exam means both: abort to stop the bleeding, undo to make the desired state honest. And in a GitOps world, undo's real form is a revert commit, or Argo CD will just re-apply the bad image.
Flagger, for contrast
Flagger inverts the authoring model. You keep authoring a plain Deployment; Flagger's Canary custom resource generates everything else: a <name>-primary Deployment that actually serves, the <name>, <name>-primary and <name>-canary Services, and the mesh routing objects, then shifts real traffic while running its analysis. You never write a Rollout.
kubectl get deploy carefully during a Flagger task
Between rollouts, Flagger scales your original Deployment to zero and serves from the primary copy; your Deployment is the template, not the workload. Seeing podinfo 0/0 next to podinfo-primary 2/2 is the system working, and mistaking it for a broken deployment is the standard first-time reaction.
| Aspect | Argo Rollouts | Flagger |
|---|---|---|
| What you author | a Rollout (replaces the Deployment) | a Canary next to an untouched Deployment |
| Traffic control | replica ratio, or a traffic provider | mesh/gateway route weights, always |
| Analysis | AnalysisTemplate + AnalysisRun | analysis.metrics + webhooks (load test, acceptance, confirm-rollout) |
| Promotion | steps, manual promote | automatic once thresholds hold for N intervals |
| Rollback | abort / undo | automatic on threshold failures; the Canary reports Failed |
Flagger's webhooks are how it does things metrics cannot: confirm-rollout (a gate before starting), pre-rollout (acceptance test against the canary), rollout (during each step, typically a load generator), confirm-promotion and post-rollout. A canary with no traffic produces no metrics, so the load-test webhook is what makes the analysis meaningful in a lab.
The mesh cluster runs no Prometheus, and make mesh wires Flagger at a metrics address that does not exist, so the built-in request-success-rate check can never pass there. Build the analysis from webhooks only (a load-test webhook against flagger-loadtester.test, plus a pre-rollout acceptance check if you want a gate). The mechanics you are practicing (Canary spec, generated objects, weight progression, events) are entirely real.
How to think about it on exam day
- Pick the strategy from the constraint. Cannot afford double capacity → canary. Need instant rollback and a clean cutover → blue-green. Need to compare behavior per user segment → A/B with header routing, which requires L7.
- Stateful and schema changes break both. Two versions run at once, so the database must tolerate both. Expand-and-contract migrations are the standard answer.
- Choose metrics the user feels. Success rate and latency percentiles over the canary's own series; not CPU, not pod restarts. And make sure the query selects only the canary: an analysis that accidentally measures the stable version will happily promote a broken release.
- Enough traffic to be significant. With three requests a minute, a 20% canary measures noise. Either drive load (Flagger's loadtester, or your own) or lengthen the interval.
- Where GitOps meets this. The rollout object lives in git like everything else; the promotion decision does not. Argo CD syncing a Rollout is fine, but if you
abortwithout reverting git, the next sync re-applies the bad image.
Fields that decide the outcome
Argo Rollouts: the analysis budget
failureLimitdefaults to 0: a single failed measurement fails the analysis and aborts the rollout. Set it deliberately.consecutiveSuccessLimit(1.8+) requires N successes in a row; withfailureLimit: -1it turns the analysis into "wait until the metric is good".inconclusiveLimitcounts measurements that met neithersuccessConditionnorfailureCondition; a run endingInconclusivepauses the rollout for a human, it does not abort. A metric with no conditions at all is always inconclusive by design.countandintervalbound how many measurements; nocountmeans run until the rollout finishes (background analysis).initialDelayper metric andanalysis.startingStepon the rollout delay the start so the canary has traffic before it is judged.- Where analysis runs:
canary.steps[].analysis(inline, blocks the step),canary.analysis(background, whole rollout),blueGreen.prePromotionAnalysis(blocks the selector switch),blueGreen.postPromotionAnalysis(aborts and switches back on failure).argsflow from the rollout;ClusterAnalysisTemplateis the cluster-scoped form. - Prometheus provider:
address,query, optionalrangeQuery; results are a vector, so conditions readresult[0], and for range queriesall(result, # < 1000). Other providers:web(jsonPath over an HTTP response),job(a Kubernetes Job's exit),datadog,newrelic,wavefront,graphite,influxdb,skywalking, and metric plugins. dryRunper metric records results without affecting the rollout;measurementRetentionkeeps more than the default 10 measurements.
Rollout spec fields not yet named
| Field | Default | Effect |
|---|---|---|
| progressDeadlineSeconds · progressDeadlineAbort | 600 · false | stalled for this long marks Degraded; with abort true it also rolls back |
| workloadRef + scaleDown | never | adopt an existing Deployment as the template; onsuccess or progressively scale the Deployment down as the Rollout takes over (the migration path) |
| rollbackWindow.revisions | unset | a rollback to one of the last N revisions skips analysis and pauses |
| canary.dynamicStableScale · abortScaleDownDelaySeconds | false · 30 | with traffic routing, shrink the stable set as canary weight grows; delay shrinking the canary after abort |
| canary.maxSurge · maxUnavailable | 25% · 25% | only meaningful without traffic routing; with it the stable set stays at 100% |
| blueGreen.autoPromotionSeconds · previewReplicaCount · scaleDownDelayRevisionLimit | unset · full · unset | timed auto-promote; a smaller preview stack until promotion; how many old ReplicaSets to keep warm |
| paused: true · restartAt | manual pause outside the step list; a timestamp that triggers a restart of all pods without a spec change |
Traffic routers and header routing
canary.trafficRouting takes one provider block: istio (a VirtualService and route names, or DestinationRule subsets), nginx (stableIngress), alb, smi, traefik, apisix, or plugins such as argoproj-labs/gatewayAPI with httpRoute and namespace, which drives HTTPRoute backendRefs weights on any conformant gateway (Cilium, Envoy Gateway, kgateway, Traefik). With a router present, canaryService and stableService are required. Extra steps become available: setHeaderRoute (match headerName with exact, prefix or regex, so testers reach the canary at weight 0), setMirrorRoute (Istio: shadow a percentage), and setCanaryScale (replicas independent of weight). Routes Rollouts creates must be listed in trafficRouting.managedRoutes and are removed on completion or abort. Plugin-style traffic routers and canary step plugins are how Rollouts 1.9 and 1.10 add integrations; the core no longer accepts new built-in providers.
Flagger: analysis fields, phases, strategy selection
Flagger has one CR and picks the strategy from which analysis fields you fill:
| You set | You get | Timing |
|---|---|---|
| stepWeight + maxWeight (or stepWeights list) | canary with progressive traffic shift; stepWeightPromotion makes the promotion gradual too | minimum interval × (maxWeight / stepWeight); rollback after interval × threshold |
| match (headers or cookie regex) + iterations | A/B testing: matched users go to the canary for the whole analysis; maxWeight and stepWeight are ignored | interval × iterations |
| iterations only | blue/green: canary is tested with conformance and load webhooks, then traffic is switched at once; works with provider: kubernetes (L4, no mesh) | interval × iterations |
| mirror: true (+ mirrorWeight) | blue/green with traffic shadowing (Istio, Gateway API) | as above |
| skipAnalysis: true | promote as soon as the canary is healthy; the emergency lever | immediate |
thresholdis the number of failed checks (metrics or webhooks) before rollback, default 1;intervaldefault 60s;progressDeadlineSeconds(default 600) rolls back a canary that never becomes ready.- Built-in metrics:
request-success-ratewiththresholdRange.minin percent andrequest-durationwiththresholdRange.maxin milliseconds (P99). Custom metrics are aMetricTemplate(provider.typeprometheus, datadog, cloudwatch, newrelic, graphite, influxdb, dynatrace, keptn, splunk, external metrics;querywith{{ namespace }},{{ target }},{{ interval }}variables) referenced bytemplateRefand judged bythresholdRange. service.port,targetPort,portDiscovery, and for Gateway APIservice.gatewayRefsandhosts; the target Deployment must have a single-label selector (app,nameorapp.kubernetes.io/name). ConfigMaps and Secrets mounted by the target are tracked and copied to-primaryunless annotatedflagger.app/config-tracking: disabled.- Phases in
status.phase:Initializing→Initialized→ (Waitingif a confirm-rollout webhook gates) →Progressing→ (WaitingPromotion) →Promoting→Finalising→SucceededorFailed.kubectl wait canary/x --for=condition=promotedis the scriptable gate.suspend: truefreezes the Canary entirely;revertOnDeletion: truerestores the original Deployment and Service when the Canary is deleted. - Webhook types beyond the ones above:
confirm-traffic-increase(gate each weight step),rollback(a 2xx from it fails the canary on demand),event(receive every Flagger event as JSON). Non-2xx from a rollout webhook counts towardthreshold; the sum of rollout webhook timeouts must fit insideinterval.
Rollouts: failureLimit 0 means one bad sample aborts. Flagger: threshold 1 means one failed check rolls back, and a canary with maxWeight unset runs to 50 in steps of stepWeight. Both tools promote a first-ever deployment straight through with no analysis, because there is no stable version to compare against; do not read that as "analysis is broken".
The baseline and the decision table
What a plain Deployment already gives you
| Field | Default | Behavior |
|---|---|---|
| strategy.type | RollingUpdate | Recreate kills everything first; the only option when two versions cannot coexist |
| rollingUpdate.maxSurge · maxUnavailable | 25% · 25% | extra pods allowed above replicas · pods allowed missing; maxSurge: 0, maxUnavailable: 1 for capacity-bound clusters, maxUnavailable: 0 for zero-dip |
| minReadySeconds | 0 | a pod must be Ready this long before it counts, the poor man's soak |
| progressDeadlineSeconds | 600 | no progress for this long sets condition Progressing=False, reason ProgressDeadlineExceeded; nothing rolls back automatically, and Argo CD flips health to Degraded at the same moment |
| revisionHistoryLimit | 10 | old ReplicaSets kept for kubectl rollout undo --to-revision |
| paused | false | kubectl rollout pause / resume; a paused Deployment accepts spec changes without rolling |
A rolling update has no traffic control and no analysis: readiness is the only gate, and the Service sends new pods their share the moment they are Ready. That is the baseline every progressive strategy is measured against, and under GitOps an "undo" is a revert commit, since kubectl rollout undo is drift.
Choosing
| Constraint in the scenario | Strategy | What it needs | How rollback works |
|---|---|---|---|
| "cannot afford double capacity", "measure before exposing everyone" | canary | metrics that reflect users (success rate, latency) and a router for exact weights; enough traffic to be significant | shift weight back; only the slice was exposed |
| "instant cutover", "instant rollback", "test the new version privately first" | blue-green | 2× capacity for the window; a preview Service or hostname; schema compatible both ways | flip the selector back; old stack kept for scaleDownDelaySeconds |
| "specific users or a header must see the new version", "same user must stay on one version" | A/B (header or cookie routing) | L7 routing: mesh or Gateway API; sticky sessions for cookies | remove the match; nobody else was affected |
| "validate against real traffic with zero user risk" | shadow / mirror | Istio or Gateway API mirroring; side-effect-free requests | none needed; responses were discarded |
| "two versions can never run at once" | Recreate | a maintenance window | redeploy the old version |
Two things decide between the tools. First, what the app team authors: a Rollout (Argo Rollouts replaces the Deployment kind and works without a mesh by replica ratio) or a plain Deployment with a Canary beside it (Flagger always needs a mesh or gateway, except for L4 blue/green). Second, who promotes: a step list with manual promote points, or an automatic threshold loop with optional confirm webhooks. Both integrate with Argo CD and Flux the same way: the Rollout or Canary lives in git and the promotion decision stays outside it.
A task that says "route requests with header X-Beta: true to the new version and nobody else" is A/B rather than canary: in Rollouts a setHeaderRoute step with a router, in Flagger analysis.match with iterations. A task that says "10% of traffic, then 50%, then all, only if error rate stays under 1%" is a canary with analysis, and the grader will look for the weight steps and the metric threshold in the object, not for a working dashboard.
Exercises
kubectl apply -f examples/rollouts/canary.yaml
kubectl argo rollouts get rollout demo -n team-a --watchoutputcaptured 2026-08-26
$ kubectl apply -f examples/rollouts/canary.yaml
analysistemplate.argoproj.io/success-rate created
rollout.argoproj.io/demo created
$ kubectl argo rollouts get rollout demo -n team-a --watch
Name: demo
Namespace: team-a
Status: ✔ Healthy
Strategy: Canary
Step: 5/5
SetWeight: 100
ActualWeight: 100
Images: ghcr.io/nginxinc/nginx-unprivileged:1.27-alpine (stable)
Replicas:
Desired: 4
Current: 4
Updated: 4
Ready: 4
Available: 4
NAME KIND STATUS AGE INFO
⟳ demo Rollout ✔ Healthy 6s
└──# revision:1
└──⧉ demo-5cfc49f565 ReplicaSet ✔ Healthy 6s stable
├──□ demo-5cfc49f565-nx4r8 Pod ✔ Running 6s ready:1/1
├──□ demo-5cfc49f565-trpgb Pod ✔ Running 6s ready:1/1
├──□ demo-5cfc49f565-xj4xl Pod ✔ Running 6s ready:1/1
└──□ demo-5cfc49f565-znz6p Pod ✔ Running 5s ready:1/1
… (frame repeats elided; next state change)
Name: demo
Namespace: team-a
Status: ◌ Progressing
Message: waiting for rollout spec update to be observed
Strategy: Canary
Step: 5/5
SetWeight: 100
ActualWeight: 100
Images: ghcr.io/nginxinc/nginx-unprivileged:1.27-alpine (stable)
Replicas:
Desired: 4
Current: 4
Updated: 4
Ready: 4
Available: 4
NAME KIND STATUS AGE INFO
⟳ demo Rollout ◌ Progressing 12s
└──# revision:1
└──⧉ demo-5cfc49f565 ReplicaSet ✔ Healthy 12s stable
├──□ demo-5cfc49f565-nx4r8 Pod ✔ Running 12s ready:1/1
├──□ demo-5cfc49f565-trpgb Pod ✔ Running 12s ready:1/1
├──□ demo-5cfc49f565-xj4xl Pod ✔ Running 12s ready:1/1
└──□ demo-5cfc49f565-znz6p Pod ✔ Running 11s ready:1/1
… (frame repeats elided; next state change)
Name: demo
Namespace: team-a
Status: ◌ Progressing
Message: more replicas need to be updated
Strategy: Canary
Step: 0/5
SetWeight: 25
ActualWeight: 0
Images: ghcr.io/nginxinc/nginx-unprivileged:1.27-alpine (stable)
Replicas:
Desired: 4
Current: 4
Updated: 0
Ready: 4
Available: 4
NAME KIND STATUS AGE INFO
⟳ demo Rollout ◌ Progressing 12s
├──# revision:2
│ └──⧉ demo-dff9796f6 ReplicaSet • ScaledDown 0s canary
└──# revision:1
└──⧉ demo-5cfc49f565 ReplicaSet ✔ Healthy 12s stable
├──□ demo-5cfc49f565-nx4r8 Pod ✔ Running 12s ready:1/1
├──□ demo-5cfc49f565-trpgb Pod ✔ Running 12s ready:1/1
├──□ demo-5cfc49f565-xj4xl Pod ✔ Running 12s ready:1/1
└──□ demo-5cfc49f565-znz6p Pod ✔ Running 11s ready:1/1
… (frame repeats elided; next state change)
Name: demo
Namespace: team-a
Status: ◌ Progressing
Message: more replicas need to be updated
Strategy: Canary
Step: 0/5
SetWeight: 25
ActualWeight: 0
Images: ghcr.io/nginxinc/nginx-unprivileged:1.26-alpine (canary)
ghcr.io/nginxinc/nginx-unprivileged:1.27-alpine (stable)
Replicas:
Desired: 4
Current: 3
Updated: 0
Ready: 3
Available: 3
NAME KIND STATUS AGE INFO
⟳ demo Rollout ◌ Progressing 13s
├──# revision:2
│ └──⧉ demo-dff9796f6 ReplicaSet ◌ Progressing 1s canary
│ └──□ demo-dff9796f6-x5d6g Pod ◌ ContainerCreating 1s ready:0/1
└──# revision:1
└──⧉ demo-5cfc49f565 ReplicaSet ✔ Healthy 13s stable
├──□ demo-5cfc49f565-nx4r8 Pod ✔ Running 13s ready:1/1
├──□ demo-5cfc49f565-trpgb Pod ✔ Running 13s ready:1/1
├──□ demo-5cfc49f565-xj4xl Pod ◌ Terminating 13s ready:0/1
└──□ demo-5cfc49f565-znz6p Pod ✔ Running 12s ready:1/1
… (frame repeats elided; next state change)
Name: demo
… (77 lines omitted)
├──# revision:2
│ ├──⧉ demo-dff9796f6 ReplicaSet ✔ Healthy 74s canary
│ │ ├──□ demo-dff9796f6-x5d6g Pod ◌ Terminating 74s ready:1/1
│ │ └──□ demo-dff9796f6-fxzp8 Pod ✔ Running 39s ready:1/1
│ └──α demo-dff9796f6-2 AnalysisRun ⚠ Error 40s ⚠ 5
└──# revision:1
└──⧉ demo-5cfc49f565 ReplicaSet ◌ Progressing 86s stable
├──□ demo-5cfc49f565-nx4r8 Pod ✔ Running 86s ready:1/1
├──□ demo-5cfc49f565-trpgb Pod ✔ Running 86s ready:1/1
└──□ demo-5cfc49f565-8rhc9 Pod ◌ ContainerCreating 0s ready:0/1
… (frame repeats elided; next state change)
Name: demo
Namespace: team-a
Status: ✖ Degraded
Message: RolloutAborted: Rollout aborted update to revision 2: Background analysis phase error/failed: Metric "success-rate" assessed Error due to consecutiveErrors (5) > consecutiveErrorLimit (4): "Error Message: reflect: slice index out of range"
Strategy: Canary
Step: 0/5
SetWeight: 0
ActualWeight: 0
Images: ghcr.io/nginxinc/nginx-unprivileged:1.27-alpine (stable)
Replicas:
Desired: 4
Current: 5
Updated: 1
Ready: 4
Available: 4
NAME KIND STATUS AGE INFO
⟳ demo Rollout ✖ Degraded 89s
├──# revision:2
│ ├──⧉ demo-dff9796f6 ReplicaSet • ScaledDown 77s canary
│ │ └──□ demo-dff9796f6-fxzp8 Pod ◌ Terminating 42s ready:1/1
│ └──α demo-dff9796f6-2 AnalysisRun ⚠ Error 43s ⚠ 5
└──# revision:1
└──⧉ demo-5cfc49f565 ReplicaSet ◌ Progressing 89s stable
├──□ demo-5cfc49f565-nx4r8 Pod ✔ Running 89s ready:1/1
├──□ demo-5cfc49f565-trpgb Pod ✔ Running 89s ready:1/1
├──□ demo-5cfc49f565-8rhc9 Pod ✔ Running 3s ready:1/1
└──□ demo-5cfc49f565-sqb6q Pod ◌ ContainerCreating 3s ready:0/1
^CFirst rollout of a new Rollout goes straight to healthy (nothing to compare against). Now change the image to trigger a real canary, in a second terminal:
kubectl argo rollouts set image demo web=ghcr.io/nginxinc/nginx-unprivileged:1.26-alpine -n team-aoutputcaptured 2026-08-26
$ kubectl argo rollouts set image demo web=ghcr.io/nginxinc/nginx-unprivileged:1.26-alpine -n team-a
rollout "demo" image updated(1.26-alpine because it exists on ghcr.io, whose nginx-unprivileged mirror currently stops at 1.27; any tag that differs from the running one triggers the canary, direction irrelevant.)
Check the analysis itself: kubectl -n team-a get analysisrun and read one with -o yaml; the measured value and the success condition are both in status. If the analysis errors because the demo app exposes no http_requests_total, that is faithful to production life; read the AnalysisRun error, then either drive traffic that produces the metric or loosen the query, and understand that an erroring analysis aborts the rollout too, on its own budget (consecutiveErrorLimit).
Trigger another image change, and while it pauses: kubectl argo rollouts abort demo -n team-a. Then kubectl argo rollouts undo demo -n team-a and confirm Healthy.
Rewrite the Rollout: strategy blueGreen, two Services (demo-active, demo-preview, both selecting app: demo), autoPromotionEnabled: false. Push a new image and inspect both Services' spec.selector before and after kubectl argo rollouts promote demo -n team-a.
rollouts-pod-template-hash while active still points at the old one; after promotion both point at the new. That selector flip is blue-green.Flagger is on the exam's tool list, so this one is not optional. On kind-mesh (kubectx kind-mesh): deploy podinfo and Flagger's loadtester (both from Flagger's podinfo tutorial manifests), then write a Canary CR targeting the podinfo Deployment with provider: istio. Build the analysis from webhooks only: a load-test webhook against the loadtester (http://flagger-loadtester.test/), and if you want a gate, a pre-rollout acceptance webhook. Then bump the podinfo image and watch kubectl describe canary podinfo events walk the weights up.
kubectl get canary -A), and while it runs, kubectl get virtualservice podinfo -o yaml shows Flagger moving real route weights between primary and canary. If the Canary sticks in Progressing, its events name the failing check.A metric with neither a success nor a failure condition cannot decide anything, and Argo Rollouts records that honestly rather than guessing: the run ends Inconclusive, which is neither Successful nor Failed. Know the third outcome, because it is the one people forget when they write the alert on the other two.
kubectl apply -f - <<'EOF'
apiVersion: argoproj.io/v1alpha1
kind: AnalysisTemplate
metadata: { name: undecided, namespace: team-a }
spec:
metrics:
- name: undecided
interval: 15s
count: 1
# no successCondition and no failureCondition: nothing can decide, so the
# measurement is Inconclusive. A job provider could not show this: its
# verdict is the job's exit code, which is always a pass or a fail.
provider:
prometheus:
address: http://prometheus-kube-prometheus-prometheus.monitoring.svc:9090
query: vector(1)
EOF
kubectl -n team-a patch rollout demo --type merge -p '{"spec":{"strategy":{"canary":{"analysis":{"templates":[{"templateName":"undecided"}],"startingStep":0,"args":null}}}}}'
kubectl argo rollouts set image demo web=ghcr.io/nginxinc/nginx-unprivileged:1.26-alpine -n team-a
sleep 90
kubectl -n team-a get analysisrun -o jsonpath='{range .items[*]}{.metadata.name} {.status.phase}{"\n"}{end}'
kubectl argo rollouts status demo -n team-a --timeout 10s || true
kubectl argo rollouts get rollout demo -n team-a --no-color | head -20
kubectl argo rollouts promote demo -n team-a
kubectl argo rollouts status demo -n team-a --timeout 120s
kubectl -n team-a delete analysisrun --all
kubectl -n team-a delete analysistemplate undecided
kubectl -n team-a apply -f examples/rollouts/canary.yamloutputcaptured 2026-09-13
$ kubectl apply -f - <<'EOF'
apiVersion: argoproj.io/v1alpha1
kind: AnalysisTemplate
metadata: { name: undecided, namespace: team-a }
spec:
metrics:
- name: undecided
interval: 15s
count: 1
# no successCondition and no failureCondition: nothing can decide, so the
# measurement is Inconclusive. A job provider could not show this: its
# verdict is the job's exit code, which is always a pass or a fail.
provider:
prometheus:
address: http://prometheus-kube-prometheus-prometheus.monitoring.svc:9090
query: vector(1)
EOF
analysistemplate.argoproj.io/undecided created
$ kubectl -n team-a patch rollout demo --type merge -p '{"spec":{"strategy":{"canary":{"analysis":{"templates":[{"templateName":"undecided"}],"startingStep":0,"args":null}}}}}'
rollout.argoproj.io/demo patched
$ kubectl argo rollouts set image demo web=ghcr.io/nginxinc/nginx-unprivileged:1.26-alpine -n team-a
rollout "demo" image updated
$ sleep 90
$ kubectl -n team-a get analysisrun -o jsonpath='{range .items[*]}{.metadata.name} {.status.phase}{"\n"}{end}'
demo-5cfc49f565-13 Inconclusive
demo-5cfc49f565-13.1 Error
$ kubectl argo rollouts status demo -n team-a --timeout 10s || true
Healthy
$ kubectl argo rollouts get rollout demo -n team-a --no-color | head -20
Name: demo
Namespace: team-a
Status: ✔ Healthy
Strategy: Canary
Step: 5/5
SetWeight: 100
ActualWeight: 100
Images: ghcr.io/nginxinc/nginx-unprivileged:1.26-alpine (stable)
Replicas:
Desired: 4
Current: 4
Updated: 4
Ready: 4
Available: 4
NAME KIND STATUS AGE INFO
⟳ demo Rollout ✔ Healthy 3h17m
├──# revision:14
│ └──⧉ demo-dff9796f6 ReplicaSet ✔ Healthy 3h1m stable
│ ├──□ demo-dff9796f6-x8b2r Pod ✔ Running 3h1m ready:1/1
$ kubectl argo rollouts promote demo -n team-a
rollout 'demo' promoted
$ kubectl argo rollouts status demo -n team-a --timeout 120s
Healthy
$ kubectl -n team-a delete analysisrun --all
analysisrun.argoproj.io "demo-5cfc49f565-13" deleted from team-a namespace
analysisrun.argoproj.io "demo-5cfc49f565-13.1" deleted from team-a namespace
$ kubectl -n team-a delete analysistemplate undecided
analysistemplate.argoproj.io "undecided" deleted from team-a namespace
$ kubectl -n team-a apply -f examples/rollouts/canary.yaml
analysistemplate.argoproj.io/success-rate unchanged
rollout.argoproj.io/demo configuredInconclusive, which is neither of the two outcomes people expect: not Successful, not Failed. That is the whole point of the exercise. Note the second run in the list ending Error: that is the ordinary success-rate template coming back when the original analysis is restored, and it errors because nothing is serving the metric it queries. Restore the original analysis from examples/rollouts/canary.yaml before the next exercise.failureLimit is how much bad news the analysis tolerates before it aborts the rollout. The default is zero, which means the first failed measurement is fatal, and that surprises people who set count: 5 expecting a vote.
# name order would hand you a stale AnalysisRun from an earlier exercise
kubectl -n team-a delete analysisrun --all
kubectl apply -f - <<'EOF'
apiVersion: argoproj.io/v1alpha1
kind: AnalysisTemplate
metadata: { name: always-fails, namespace: team-a }
spec:
metrics:
- name: always-fails
interval: 15s
count: 5
failureLimit: 0
successCondition: result == 'never'
provider:
job:
spec:
template:
spec:
restartPolicy: Never
containers:
- name: check
image: busybox:1.36
command: ["sh", "-c", "exit 1"]
backoffLimit: 0
EOF
kubectl -n team-a patch rollout demo --type merge -p '{"spec":{"strategy":{"canary":{"analysis":{"templates":[{"templateName":"always-fails"}],"startingStep":0,"args":null}}}}}'
kubectl argo rollouts set image demo web=ghcr.io/nginxinc/nginx-unprivileged:1.25-alpine -n team-a
sleep 75
kubectl -n team-a get analysisrun --sort-by=.metadata.creationTimestamp -o jsonpath='{.items[-1:].status.metricResults[0]}' | jq
kubectl argo rollouts get rollout demo -n team-a --no-color | head -6
kubectl argo rollouts undo demo -n team-a
kubectl patch analysistemplate always-fails -n team-a --type merge -p '{"spec":{"metrics":[{"name":"always-fails","interval":"15s","count":5,"failureLimit":2,"successCondition":"result == \"never\"","provider":{"job":{"spec":{"backoffLimit":0,"template":{"spec":{"restartPolicy":"Never","containers":[{"name":"check","image":"busybox:1.36","command":["sh","-c","exit 1"]}]}}}}}}]}}'
kubectl argo rollouts set image demo web=ghcr.io/nginxinc/nginx-unprivileged:1.24-alpine -n team-a
sleep 120
kubectl -n team-a get analysisrun --sort-by=.metadata.creationTimestamp -o jsonpath='{.items[-1:].status.metricResults[0]}' | jq
kubectl argo rollouts get rollout demo -n team-a --no-color | head -6
kubectl argo rollouts undo demo -n team-a
kubectl -n team-a delete analysisrun --all
kubectl -n team-a delete analysistemplate always-fails
kubectl -n team-a apply -f examples/rollouts/canary.yamloutputcaptured 2026-09-12
$ # name order would hand you a stale AnalysisRun from an earlier exercise
$ kubectl -n team-a delete analysisrun --all
No resources found
$ kubectl apply -f - <<'EOF'
apiVersion: argoproj.io/v1alpha1
kind: AnalysisTemplate
metadata: { name: always-fails, namespace: team-a }
spec:
metrics:
- name: always-fails
interval: 15s
count: 5
failureLimit: 0
successCondition: result == 'never'
provider:
job:
spec:
template:
spec:
restartPolicy: Never
containers:
- name: check
image: busybox:1.36
command: ["sh", "-c", "exit 1"]
backoffLimit: 0
EOF
analysistemplate.argoproj.io/always-fails created
$ kubectl -n team-a patch rollout demo --type merge -p '{"spec":{"strategy":{"canary":{"analysis":{"templates":[{"templateName":"always-fails"}],"startingStep":0,"args":null}}}}}'
rollout.argoproj.io/demo patched
$ kubectl argo rollouts set image demo web=ghcr.io/nginxinc/nginx-unprivileged:1.25-alpine -n team-a
rollout "demo" image updated
$ sleep 75
$ kubectl -n team-a get analysisrun --sort-by=.metadata.creationTimestamp -o jsonpath='{.items[-1:].status.metricResults[0]}' | jq
{
"count": 1,
"failed": 1,
"measurements": [
{
"finishedAt": "2026-09-13T16:56:49Z",
"metadata": {
"job-name": "c59c8d09-fe03-4eb2-b6d7-269647fa9645.always-fails.1",
"job-namespace": "team-a"
},
"phase": "Failed",
"resumeAt": "2026-09-13T16:56:49Z",
"startedAt": "2026-09-13T16:56:40Z"
}
],
"name": "always-fails",
"phase": "Failed"
}
$ kubectl argo rollouts get rollout demo -n team-a --no-color | head -6
Name: demo
Namespace: team-a
Status: ✖ Degraded
Message: RolloutAborted: Rollout aborted update to revision 23: Background analysis phase error/failed: Metric "always-fails" assessed Failed due to failed (1) > failureLimit (0)
Strategy: Canary
Step: 0/5
$ kubectl argo rollouts undo demo -n team-a
rollout 'demo' undo
$ kubectl patch analysistemplate always-fails -n team-a --type merge -p '{"spec":{"metrics":[{"name":"always-fails","interval":"15s","count":5,"failureLimit":2,"successCondition":"result == \"never\"","provider":{"job":{"spec":{"backoffLimit":0,"template":{"spec":{"restartPolicy":"Never","containers":[{"name":"check","image":"busybox:1.36","command":["sh","-c","exit 1"]}]}}}}}}]}}'
analysistemplate.argoproj.io/always-fails patched
$ kubectl argo rollouts set image demo web=ghcr.io/nginxinc/nginx-unprivileged:1.24-alpine -n team-a
rollout "demo" image updated
$ sleep 120
$ kubectl -n team-a get analysisrun --sort-by=.metadata.creationTimestamp -o jsonpath='{.items[-1:].status.metricResults[0]}' | jq
{
"count": 3,
"failed": 3,
"measurements": [
{
"finishedAt": "2026-09-13T16:58:05Z",
"metadata": {
"job-name": "cbfcfe8a-a830-4350-8fc1-08f7c2d63670.always-fails.1",
"job-namespace": "team-a"
},
"phase": "Failed",
"resumeAt": "2026-09-13T16:58:05Z",
"startedAt": "2026-09-13T16:57:56Z"
},
{
"finishedAt": "2026-09-13T16:58:26Z",
"metadata": {
"job-name": "cbfcfe8a-a830-4350-8fc1-08f7c2d63670.always-fails.2",
"job-namespace": "team-a"
},
"phase": "Failed",
"resumeAt": "2026-09-13T16:58:26Z",
"startedAt": "2026-09-13T16:58:20Z"
},
{
"finishedAt": "2026-09-13T16:58:50Z",
"metadata": {
"job-name": "cbfcfe8a-a830-4350-8fc1-08f7c2d63670.always-fails.3",
"job-namespace": "team-a"
},
"phase": "Failed",
"resumeAt": "2026-09-13T16:58:50Z",
"startedAt": "2026-09-13T16:58:41Z"
}
],
"name": "always-fails",
"phase": "Failed"
}
$ kubectl argo rollouts get rollout demo -n team-a --no-color | head -6
Name: demo
Namespace: team-a
Status: ✖ Degraded
Message: RolloutAborted: Rollout aborted update to revision 25: Background analysis phase error/failed: Metric "always-fails" assessed Failed due to failed (3) > failureLimit (2)
Strategy: Canary
Step: 0/5
$ kubectl argo rollouts undo demo -n team-a
rollout 'demo' undo
$ kubectl -n team-a delete analysisrun --all
analysisrun.argoproj.io "demo-584f94bcff-23" deleted from team-a namespace
analysisrun.argoproj.io "demo-5cfc49f565-24" deleted from team-a namespace
analysisrun.argoproj.io "demo-6bcd89b78b-25" deleted from team-a namespace
$ kubectl -n team-a delete analysistemplate always-fails
analysistemplate.argoproj.io "always-fails" deleted from team-a namespace
$ kubectl -n team-a apply -f examples/rollouts/canary.yaml
analysistemplate.argoproj.io/success-rate unchanged
rollout.argoproj.io/demo configuredfailureLimit: 0 the run records count 1, failed 1, phase Failed and the rollout aborts on that single measurement; with failureLimit: 2 it records count 3, failed 3, phase Failed and aborts on the third. Both abort, at different prices. count is how many measurements you asked for, failureLimit is how many you will tolerate before the run is Failed. Sort by creationTimestamp: name order hands you a stale AnalysisRun from an earlier exercise, and a terminated one reads as Inconclusive rather than Failed.Blue-green keeps a whole second stack alive, which is the expensive part. previewReplicaCount makes the preview cheap until promotion, and autoPromotionSeconds decides how long you have to look at it.
kubectl apply -f - <<'EOF'
apiVersion: v1
kind: Service
metadata: { name: bg-active, namespace: team-a }
spec: { selector: { app: bg }, ports: [{ port: 80, targetPort: 8080 }] }
---
apiVersion: v1
kind: Service
metadata: { name: bg-preview, namespace: team-a }
spec: { selector: { app: bg }, ports: [{ port: 80, targetPort: 8080 }] }
EOF
# the Services first: created together, the controller sees previewService missing and parks the Rollout InvalidSpec
kubectl apply -f - <<'EOF'
apiVersion: argoproj.io/v1alpha1
kind: Rollout
metadata: { name: bg, namespace: team-a }
spec:
replicas: 4
selector: { matchLabels: { app: bg } }
template:
metadata: { labels: { app: bg } }
spec:
containers:
- name: web
image: ghcr.io/nginxinc/nginx-unprivileged:1.27-alpine
ports: [{ containerPort: 8080 }]
resources:
requests: { cpu: 10m, memory: 32Mi }
strategy:
blueGreen:
activeService: bg-active
previewService: bg-preview
previewReplicaCount: 1
autoPromotionSeconds: 30
EOF
kubectl argo rollouts status bg -n team-a --timeout 180s
kubectl argo rollouts set image bg web=ghcr.io/nginxinc/nginx-unprivileged:1.26-alpine -n team-a
sleep 15
kubectl -n team-a get rs -l app=bg -o custom-columns=NAME:.metadata.name,DESIRED:.spec.replicas,READY:.status.readyReplicas
sleep 45
kubectl -n team-a get rs -l app=bg -o custom-columns=NAME:.metadata.name,DESIRED:.spec.replicas,READY:.status.readyReplicas
kubectl -n team-a delete rollout bg
kubectl -n team-a delete svc bg-active bg-previewoutputcaptured 2026-09-13
$ kubectl apply -f - <<'EOF'
apiVersion: v1
kind: Service
metadata: { name: bg-active, namespace: team-a }
spec: { selector: { app: bg }, ports: [{ port: 80, targetPort: 8080 }] }
---
apiVersion: v1
kind: Service
metadata: { name: bg-preview, namespace: team-a }
spec: { selector: { app: bg }, ports: [{ port: 80, targetPort: 8080 }] }
EOF
service/bg-active created
service/bg-preview created
$ # the Services first: created together, the controller sees previewService missing and parks the Rollout InvalidSpec
$ kubectl apply -f - <<'EOF'
apiVersion: argoproj.io/v1alpha1
kind: Rollout
metadata: { name: bg, namespace: team-a }
spec:
replicas: 4
selector: { matchLabels: { app: bg } }
template:
metadata: { labels: { app: bg } }
spec:
containers:
- name: web
image: ghcr.io/nginxinc/nginx-unprivileged:1.27-alpine
ports: [{ containerPort: 8080 }]
resources:
requests: { cpu: 10m, memory: 32Mi }
strategy:
blueGreen:
activeService: bg-active
previewService: bg-preview
previewReplicaCount: 1
autoPromotionSeconds: 30
EOF
rollout.argoproj.io/bg created
$ kubectl argo rollouts status bg -n team-a --timeout 180s
Progressing - more replicas need to be updated
Progressing - updated replicas are still becoming available
Healthy
$ kubectl argo rollouts set image bg web=ghcr.io/nginxinc/nginx-unprivileged:1.26-alpine -n team-a
rollout "bg" image updated
$ sleep 15
$ kubectl -n team-a get rs -l app=bg -o custom-columns=NAME:.metadata.name,DESIRED:.spec.replicas,READY:.status.readyReplicas
NAME DESIRED READY
bg-5cb944994f 4 4
bg-845689b9dd 1 1
$ sleep 45
$ kubectl -n team-a get rs -l app=bg -o custom-columns=NAME:.metadata.name,DESIRED:.spec.replicas,READY:.status.readyReplicas
NAME DESIRED READY
bg-5cb944994f 4 4
bg-845689b9dd 4 4
$ kubectl -n team-a delete rollout bg
rollout.argoproj.io "bg" deleted from team-a namespace
$ kubectl -n team-a delete svc bg-active bg-preview
service "bg-active" deleted from team-a namespace
service "bg-preview" deleted from team-a namespacepreviewService bg-preview missing and parks the Rollout Degraded with InvalidSpec, and a set image issued before revision 1 is promoted skips previewReplicaCount entirely because there is no active selector to preview against.Weight sends a slice of everyone's traffic to the new version. A header route sends all of one person's, which is how you let a tester see the canary before any customer does. It needs a mesh, so this one is on kind-mesh, which has no Rollouts controller until the first line installs one. Note the --server-side: a plain apply of these CRDs is refused because the last-applied annotation would exceed the 256KB limit on annotations.
kubectl --context kind-mesh create ns argo-rollouts --dry-run=client -o yaml | kubectl --context kind-mesh apply -f -
kubectl --context kind-mesh -n argo-rollouts apply --server-side -f https://github.com/argoproj/argo-rollouts/releases/latest/download/install.yaml
kubectl --context kind-mesh -n argo-rollouts rollout status deploy/argo-rollouts --timeout=180s
kubectl --context kind-mesh apply -f - <<'EOF'
apiVersion: argoproj.io/v1alpha1
kind: Rollout
metadata: { name: hdr, namespace: default }
spec:
replicas: 2
selector: { matchLabels: { app: hdr } }
template:
metadata: { labels: { app: hdr } }
spec:
containers:
- name: web
image: ghcr.io/stefanprodan/podinfo:6.7.1
ports: [{ containerPort: 9898 }]
strategy:
canary:
canaryService: hdr-canary
stableService: hdr-stable
trafficRouting:
managedRoutes: [{ name: header-route }]
istio:
virtualService:
name: hdr
routes: [primary]
steps:
- setCanaryScale: { weight: 25 }
- setHeaderRoute:
name: header-route
match:
- headerName: x-canary
headerValue: { exact: "true" }
- pause: {}
EOF
kubectl --context kind-mesh apply -f - <<'EOF'
apiVersion: v1
kind: Service
metadata: { name: hdr-canary, namespace: default }
spec: { selector: { app: hdr }, ports: [{ port: 80, targetPort: 9898 }] }
---
apiVersion: v1
kind: Service
metadata: { name: hdr-stable, namespace: default }
spec: { selector: { app: hdr }, ports: [{ port: 80, targetPort: 9898 }] }
---
apiVersion: networking.istio.io/v1
kind: VirtualService
metadata: { name: hdr, namespace: default }
spec:
hosts: [hdr-stable]
http:
- name: primary
route:
- destination: { host: hdr-stable }
weight: 100
- destination: { host: hdr-canary }
weight: 0
EOF
kubectl argo rollouts status hdr --context kind-mesh --timeout 90s || true
kubectl argo rollouts set image hdr web=ghcr.io/stefanprodan/podinfo:6.7.0 --context kind-mesh
sleep 45
kubectl --context kind-mesh get virtualservice hdr -o yaml | sed -n '/http:/,$p'
kubectl argo rollouts abort hdr --context kind-mesh
kubectl --context kind-mesh delete rollout hdr
kubectl --context kind-mesh delete svc hdr-canary hdr-stable
kubectl --context kind-mesh delete virtualservice hdroutputcaptured 2026-09-12
$ kubectl --context kind-mesh create ns argo-rollouts --dry-run=client -o yaml | kubectl --context kind-mesh apply -f -
namespace/argo-rollouts unchanged
$ kubectl --context kind-mesh -n argo-rollouts apply --server-side -f https://github.com/argoproj/argo-rollouts/releases/latest/download/install.yaml
customresourcedefinition.apiextensions.k8s.io/analysisruns.argoproj.io serverside-applied
customresourcedefinition.apiextensions.k8s.io/analysistemplates.argoproj.io serverside-applied
customresourcedefinition.apiextensions.k8s.io/clusteranalysistemplates.argoproj.io serverside-applied
customresourcedefinition.apiextensions.k8s.io/experiments.argoproj.io serverside-applied
customresourcedefinition.apiextensions.k8s.io/rollouts.argoproj.io serverside-applied
serviceaccount/argo-rollouts serverside-applied
clusterrole.rbac.authorization.k8s.io/argo-rollouts serverside-applied
clusterrole.rbac.authorization.k8s.io/argo-rollouts-aggregate-to-admin serverside-applied
clusterrole.rbac.authorization.k8s.io/argo-rollouts-aggregate-to-edit serverside-applied
clusterrole.rbac.authorization.k8s.io/argo-rollouts-aggregate-to-view serverside-applied
clusterrolebinding.rbac.authorization.k8s.io/argo-rollouts serverside-applied
configmap/argo-rollouts-config serverside-applied
secret/argo-rollouts-notification-secret serverside-applied
service/argo-rollouts-metrics serverside-applied
deployment.apps/argo-rollouts serverside-applied
$ kubectl --context kind-mesh -n argo-rollouts rollout status deploy/argo-rollouts --timeout=180s
deployment "argo-rollouts" successfully rolled out
$ kubectl --context kind-mesh apply -f - <<'EOF'
apiVersion: argoproj.io/v1alpha1
kind: Rollout
metadata: { name: hdr, namespace: default }
spec:
replicas: 2
selector: { matchLabels: { app: hdr } }
template:
metadata: { labels: { app: hdr } }
spec:
containers:
- name: web
image: ghcr.io/stefanprodan/podinfo:6.7.1
ports: [{ containerPort: 9898 }]
strategy:
canary:
canaryService: hdr-canary
stableService: hdr-stable
trafficRouting:
managedRoutes: [{ name: header-route }]
istio:
virtualService:
name: hdr
routes: [primary]
steps:
- setCanaryScale: { weight: 25 }
- setHeaderRoute:
name: header-route
match:
- headerName: x-canary
headerValue: { exact: "true" }
- pause: {}
EOF
rollout.argoproj.io/hdr created
$ kubectl --context kind-mesh apply -f - <<'EOF'
apiVersion: v1
kind: Service
metadata: { name: hdr-canary, namespace: default }
spec: { selector: { app: hdr }, ports: [{ port: 80, targetPort: 9898 }] }
---
apiVersion: v1
kind: Service
metadata: { name: hdr-stable, namespace: default }
spec: { selector: { app: hdr }, ports: [{ port: 80, targetPort: 9898 }] }
---
apiVersion: networking.istio.io/v1
kind: VirtualService
metadata: { name: hdr, namespace: default }
spec:
hosts: [hdr-stable]
http:
- name: primary
route:
- destination: { host: hdr-stable }
weight: 100
- destination: { host: hdr-canary }
weight: 0
EOF
service/hdr-canary unchanged
service/hdr-stable unchanged
virtualservice.networking.istio.io/hdr unchanged
$ kubectl argo rollouts status hdr --context kind-mesh --timeout 90s || true
Progressing - more replicas need to be updated
Progressing - updated replicas are still becoming available
Healthy
$ kubectl argo rollouts set image hdr web=ghcr.io/stefanprodan/podinfo:6.7.0 --context kind-mesh
rollout "hdr" image updated
$ sleep 45
$ kubectl --context kind-mesh get virtualservice hdr -o yaml | sed -n '/http:/,$p'
http:
- match:
- headers:
x-canary:
exact: "true"
name: header-route
route:
- destination:
host: hdr-canary
weight: 100
- name: primary
route:
- destination:
host: hdr-stable
weight: 100
- destination:
host: hdr-canary
weight: 0
$ kubectl argo rollouts abort hdr --context kind-mesh
rollout 'hdr' aborted
$ kubectl --context kind-mesh delete rollout hdr
rollout.argoproj.io "hdr" deleted from default namespace
$ kubectl --context kind-mesh delete svc hdr-canary hdr-stable
service "hdr-canary" deleted from default namespace
service "hdr-stable" deleted from default namespace
$ kubectl --context kind-mesh delete virtualservice hdr
virtualservice.networking.istio.io "hdr" deleted from default namespaceFlagger can run the same analysis without shifting any weight: match a header, run a fixed number of iterations, then promote. The events are the difference, and with provider: kubernetes it needs no mesh at all.
kubectl --context kind-mesh -n test get canary podinfo -o jsonpath='{.spec.analysis}' | jq
# a merge patch leaves maxWeight and stepWeight in place beside iterations, and the schema's oneOf
# then matches two alternatives and rejects the object: replace the whole analysis instead
kubectl --context kind-mesh -n test patch canary podinfo --type json -p '[{"op":"replace","path":"/spec/analysis","value":{"interval":"15s","iterations":3,"threshold":2,"match":[{"headers":{"x-canary":{"regex":"^insider$"}}}]}}]'
kubectl --context kind-mesh -n test get canary podinfo -o jsonpath='{.spec.analysis}' | jq
kubectl --context kind-mesh -n test set image deploy/podinfo podinfod=ghcr.io/stefanprodan/podinfo:6.7.0
sleep 180
kubectl --context kind-mesh -n test describe canary podinfo | sed -n '/Events/,$p' | tail -12
kubectl --context kind-mesh -n test get canary podinfo -o jsonpath='{.status.phase} {.status.failedChecks}{"\n"}'
kubectl --context kind-mesh -n test patch canary podinfo --type json -p '[{"op":"replace","path":"/spec/analysis","value":{"interval":"15s","maxWeight":50,"stepWeight":10,"threshold":5,"metrics":[{"name":"request-success-rate","interval":"1m","thresholdRange":{"min":99}}]}}]'outputcaptured 2026-09-13
$ kubectl --context kind-mesh -n test get canary podinfo -o jsonpath='{.spec.analysis}' | jq
{
"interval": "15s",
"maxWeight": 50,
"metrics": [
{
"interval": "1m",
"name": "request-success-rate",
"thresholdRange": {
"min": 99
}
}
],
"stepWeight": 10,
"threshold": 5
}
$ # a merge patch leaves maxWeight and stepWeight in place beside iterations, and the schema's oneOf
$ # then matches two alternatives and rejects the object: replace the whole analysis instead
$ kubectl --context kind-mesh -n test patch canary podinfo --type json -p '[{"op":"replace","path":"/spec/analysis","value":{"interval":"15s","iterations":3,"threshold":2,"match":[{"headers":{"x-canary":{"regex":"^insider$"}}}]}}]'
canary.flagger.app/podinfo patched
$ kubectl --context kind-mesh -n test get canary podinfo -o jsonpath='{.spec.analysis}' | jq
{
"interval": "15s",
"iterations": 3,
"match": [
{
"headers": {
"x-canary": {
"regex": "^insider$"
}
}
}
],
"threshold": 2
}
$ kubectl --context kind-mesh -n test set image deploy/podinfo podinfod=ghcr.io/stefanprodan/podinfo:6.7.0
deployment.apps/podinfo image updated
$ sleep 180
$ kubectl --context kind-mesh -n test describe canary podinfo | sed -n '/Events/,$p' | tail -12
Warning Synced 3m15s (x3 over 3m45s) flagger Error checking metric providers: prometheus not avaiable: running query failed: request failed: Get "http://prometheus:9090/api/v1/query?query=vector%281%29": dial tcp: lookup prometheus on 10.96.0.10:53: server misbehaving
Normal Synced 3m14s flagger Initialization done! podinfo.test
Normal Synced 2m45s flagger New revision detected! Scaling up podinfo.test
Warning Synced 2m30s flagger canary deployment podinfo.test not ready: waiting for rollout to finish: 0 of 2 (readyThreshold 100%) updated replicas are available
Normal Synced 2m15s flagger Starting canary analysis for podinfo.test
Normal Synced 2m15s flagger Advance podinfo.test canary iteration 1/3
Normal Synced 2m flagger Advance podinfo.test canary iteration 2/3
Normal Synced 105s flagger Advance podinfo.test canary iteration 3/3
Normal Synced 90s flagger Copying podinfo.test template spec to podinfo-primary.test
Warning Synced 75s flagger podinfo-primary.test not ready: waiting for rollout to finish: 1 old replicas are pending termination
Normal Synced 60s flagger Routing all traffic to primary
Normal Synced 45s flagger Promotion completed! Scaling down podinfo.test
$ kubectl --context kind-mesh -n test get canary podinfo -o jsonpath='{.status.phase} {.status.failedChecks}{"\n"}'
Succeeded 0
$ kubectl --context kind-mesh -n test patch canary podinfo --type json -p '[{"op":"replace","path":"/spec/analysis","value":{"interval":"15s","maxWeight":50,"stepWeight":10,"threshold":5,"metrics":[{"name":"request-success-rate","interval":"1m","thresholdRange":{"min":99}}]}}]'
canary.flagger.app/podinfo patchedAdvance podinfo.test canary iteration 1/3 rather than a weight ladder, then promote.Before any of this tooling, a plain Deployment already has a clock: progressDeadlineSeconds. When it expires the rollout stops advancing and says so in a condition, and nothing rolls back on its own.
kubectl create deployment stuck --image=nginx:1.27-alpine
kubectl patch deployment stuck -p '{"spec":{"progressDeadlineSeconds":60}}'
kubectl set image deployment/stuck nginx=kind-registry:5000/does-not-exist:nope
kubectl rollout status deployment/stuck || true
kubectl get deployment stuck -o jsonpath='{.status.conditions}' | jq '.[] | {type, status, reason, message}'
kubectl rollout undo deployment/stuck
kubectl rollout status deployment/stuck
kubectl delete deployment stuckoutputcaptured 2026-09-12
$ kubectl create deployment stuck --image=nginx:1.27-alpine
deployment.apps/stuck created
$ kubectl patch deployment stuck -p '{"spec":{"progressDeadlineSeconds":60}}'
deployment.apps/stuck patched
$ kubectl set image deployment/stuck nginx=kind-registry:5000/does-not-exist:nope
deployment.apps/stuck image updated
$ kubectl rollout status deployment/stuck || true
Waiting for deployment spec update to be observed...
Waiting for deployment "stuck" rollout to finish: 0 out of 1 new replicas have been updated...
Waiting for deployment "stuck" rollout to finish: 1 old replicas are pending termination...
Waiting for deployment "stuck" rollout to finish: 1 old replicas are pending termination...
error: deployment "stuck" exceeded its progress deadline
$ kubectl get deployment stuck -o jsonpath='{.status.conditions}' | jq '.[] | {type, status, reason, message}'
{
"type": "Available",
"status": "True",
"reason": "MinimumReplicasAvailable",
"message": "Deployment has minimum availability."
}
{
"type": "Progressing",
"status": "False",
"reason": "ProgressDeadlineExceeded",
"message": "ReplicaSet \"stuck-57cb855fcc\" has timed out progressing."
}
$ kubectl rollout undo deployment/stuck
deployment.apps/stuck rolled back
$ kubectl rollout status deployment/stuck
Waiting for deployment spec update to be observed...
error: deployment "stuck" exceeded its progress deadline
$ kubectl delete deployment stuck
deployment.apps "stuck" deleted from default namespacekubectl rollout status exits non-zero quoting the progress deadline, and the Progressing condition is False with reason ProgressDeadlineExceeded. The old pods are still serving the whole time; say why that is the correct default.Self-check
You set setWeight: 10 with 3 replicas and no traffic provider. What actually happens?
Rollouts rounds to whole pods: it runs one canary pod alongside the stable set, so the real split is nearer 25–33%, not 10%. Exact weights require a traffic provider (Istio, Gateway API, NGINX) or more replicas.
A canary aborted and the AnalysisRun shows measurements in phase Error. What do you check?
The metric query and its provider: wrong series name, no data yet, unreachable Prometheus, or a selector that matches nothing. Errors count toward consecutiveErrorLimit (default 4), so a broken query aborts releases exactly like a bad release does; check the measurement message before touching the application.
Blue-green with autoPromotionEnabled: false: which Service points where, before and after promote?
Before: active selects the old ReplicaSet's rollouts-pod-template-hash, preview selects the new one. After: both select the new hash, and the old ReplicaSet lingers for scaleDownDelaySeconds so rollback is a selector flip away.
Argo CD manages your Rollout. You abort a bad canary. What happens next, and what should you do?
The Rollout spec in git still names the bad image, so the next sync re-applies it and the rollout starts again. Revert the commit (or pin the tag back): abort is an operational stop, git is the desired state. Same lesson as any manual fix under GitOps.
Name two things that make progressive delivery unsafe regardless of tooling.
Database or API changes that both versions cannot tolerate simultaneously (fix with expand-and-contract migrations and backwards-compatible contracts), and metrics too sparse or too coarse to detect harm within the canary window. A third: analysis that measures the stable version by accident.
An AnalysisRun ended Inconclusive and the rollout stopped. Did it abort?
No: inconclusive pauses the rollout at its current step and waits for a human to promote or abort. It happens when measurements satisfy neither successCondition nor failureCondition (or when a metric defines neither) more than inconclusiveLimit times. Failure aborts; error beyond consecutiveErrorLimit aborts; inconclusive asks.
Flagger: you set match on a header and also stepWeight/maxWeight. What runs?
A/B testing: when Flagger finds an HTTP match condition it ignores maxWeight and stepWeight and runs iterations rounds with matched traffic routed to the canary. Remove match for a weighted canary; keep only iterations for blue/green.
What does progressDeadlineSeconds do on a plain Deployment, and what does it not do?
After 600 s (default) without progress it sets the Deployment condition Progressing=False with reason ProgressDeadlineExceeded, which is what Argo CD reads as Degraded. It does not roll back or pause anything; the old ReplicaSet keeps whatever pods maxUnavailable left it. Rollouts' progressDeadlineAbort: true is the version that also rolls back.
Why must canaryService and stableService exist when you add trafficRouting to a Rollout?
The router splits traffic between two Services, so Rollouts needs one Service whose selector it points at the canary ReplicaSet and one for the stable set (it injects rollouts-pod-template-hash into their selectors). Without a router there is one Service and the split is by replica count, so neither field is needed.
Docs to know your way around
- argo-rollouts.readthedocs.io: canary and blueGreen strategy references, AnalysisTemplate spec, the kubectl plugin page, traffic-router support matrix.
- flagger.app: the Istio canary tutorial and the webhook reference.
- Offline:
kubectl argo rollouts --help,kubectl explain rollout.spec.strategy.canary --recursive,kubectl explain canary.spec.analysis. - argo-rollouts.readthedocs.io: Analysis ("Failure Conditions and Failure Limit", "Inconclusive Runs"), Traffic Management ("Managed Routes", "Header Values"), Specification: every default quoted above is on those pages.
- docs.flagger.app: Deployment Strategies, Webhooks, Metrics: the strategy-selection rules (match, iterations, mirror), the eight webhook types and the MetricTemplate variables.
make down-obsmake down-mesh