Argo Workflows runs DAGs of containers. It is not CI (Tekton's job here) and not reconciliation (operators' job); it is imperative orchestration with dependencies, retries and parameters, which makes it the right engine for provisioning sequences: validate the request, create the resources, register them elsewhere, notify.

needsmake core

Orientation

competency 3.2 · workflows for self-service provisioning

Run to completion, not converge forever. That distinction is section 3.6's whole decision table, and it is the reason a workflow is the right tool for "stamp this out once when asked" and the wrong tool for "keep this true for the next three years".

Lab note

Workflows installs with the gitops layer and the controller watches the argo and default namespaces. Workflows created elsewhere will sit untouched forever, looking exactly like a broken controller.

The model

entrypoint · templates · parameters

A Workflow is a spec with an entrypoint and a list of templates. Templates come in flavors, and picking the right one is most of the authoring skill:

Template typeDoesUse for
containerruns an imageanything with a CLI
scriptinline code with an interpreter; its stdout becomes a resultvalidation, small glue, computed values
resourcecreate/apply/patch/delete a Kubernetes object, optionally waiting on a success conditionthe provisioning workhorse
dagtasks with dependenciesthe one to default to
stepssequential groups (- parallel within a group, - - serial between groups)simple linear flows
suspendpauses until resumed (manually or after a duration)approval gates

Parameters flow via {{workflow.parameters.x}} and between tasks via outputs ({{tasks.a.outputs.result}}); when gates conditional branches; withItems/withParam fan a template out over a list. Artifacts (S3/MinIO-backed files) move large data between steps; not configured in this lab, but know the word, because "results are small strings, artifacts are files" is the same distinction Tekton draws.

Reusable pieces: WorkflowTemplate (namespaced library, invoked with workflowTemplateRef or submitted from the UI), ClusterWorkflowTemplate (cluster-scoped version), and CronWorkflow (scheduled). Promoting a Workflow to a WorkflowTemplate is what turns a script into a self-service endpoint: it becomes a thing with a name, parameters and a submit button.

WorkflowTemplate "provision-tenant"   ← the product: named, parameterized, submittable
        │ submit(team=team-d)
        ▼
Workflow (run-to-completion)
   dag:
     check ──▶ namespace ──▶ quota ──▶ register
     (script)    (resource)      (resource)   (container)
        │
        └─ each node = one pod, running as spec.serviceAccountName

RBAC is the part that actually fails

and the Argo-specific wrinkle

Workflow pods run as a ServiceAccount, and a resource template creating namespaces needs cluster-scoped rights that default does not have. The skill the exam wants is wiring SA → Role/ClusterRole → binding → spec.serviceAccountName, and then recognizing a Forbidden error inside a workflow node as an RBAC problem rather than a workflow problem.

The one nobody guesses

Argo's executor reports each step's outcome through its own CRD, workflowtaskresults.argoproj.io. Every workflow ServiceAccount therefore needs create,patch on that resource; forget it and even a perfectly correct workflow fails, usually before your actual permission problem surfaces.

Two related permissions worth knowing exist: the SA also needs get,list,watch on pods to report progress in some configurations, and if you use artifacts, access to the artifact repository Secret. The general rule is the same as section 5.1's: least privilege, then prove it with kubectl auth can-i --as=system:serviceaccount:<ns>:<sa> before you blame the tool.

Diagnosing a failed workflow

kubectl get workflow <name> -o jsonpath='{.status.nodes}' | jq '.[] | {displayName, phase, message}'
kubectl get workflow <name> -o jsonpath='{.status.phase}{"\n"}'
kubectl logs -l workflows.argoproj.io/workflow=<name> --all-containers --tail=50
outputcaptured 2026-08-26
$ kubectl get workflow provision-tenant-29926 -o jsonpath='{.status.nodes}' | jq '.[] | {displayName, phase, message}'
{
  "displayName": "provision-tenant-29926",
  "phase": "Failed",
  "message": null
}
{
  "displayName": "namespace",
  "phase": "Failed",
  "message": "main: Error (exit code 64): no more retries Error from server (Forbidden): error when creating \"/tmp/manifest.yaml\": namespaces is forbidden: User \"system:serviceaccount:default:default\" cannot create resource \"namespaces\" in API group \"\" at the cluster scope"
}
$ kubectl get workflow provision-tenant-29926 -o jsonpath='{.status.phase}{"\n"}'
Failed
$ kubectl logs -l workflows.argoproj.io/workflow=provision-tenant-29926 --all-containers --tail=50
time=2026-08-27T03:31:16.589Z level=INFO msg="Starting Workflow Executor" gitTag=v4.1.2 gitTreeState=clean goVersion=go1.26.5 argo=true version=v4.1.2 buildDate=2026-08-21T11:10:01Z gitCommit=16a52d67daf2f4a8a76fa8bec02a76a46aa46257
time=2026-08-27T03:31:16.592Z level=INFO msg="Using executor retry strategy" argo=true Jitter=0.5 Steps=5 Duration=1s Factor=1.6
time=2026-08-27T03:31:16.592Z level=INFO msg="Executor initialized" goVersion=go1.26.5 deadline=0001-01-01T00:00:00.000Z gitTag=v4.1.2 podName=provision-tenant-29926-make-ns-596359171 buildDate=2026-08-21T11:10:01Z gitTreeState=clean templateName=make-ns includeScriptOutput=false version=v4.1.2 namespace=default argo=true gitCommit=16a52d67daf2f4a8a76fa8bec02a76a46aa46257
time=2026-08-27T03:31:16.718Z level=INFO msg="Loading manifest" argo=true path=/tmp/manifest.yaml
time=2026-08-27T03:31:16.718Z level=INFO msg="Start loading input artifacts..." argo=true pluginName=""
time="2026-08-27T03:31:16.718Z" level=info msg="Alloc=11003 TotalAlloc=16525 Sys=28866 NumGC=4 Goroutines=3"
time=2026-08-27T03:31:18.326Z level=INFO msg="waiting for signals" argo=true signalPath=/var/run/argo/ctr/main/signal
time=2026-08-27T03:31:18.368Z level=INFO msg="Starting Workflow Executor" gitTreeState=clean goVersion=go1.26.5 argo=true version=v4.1.2 buildDate=2026-08-21T11:10:01Z gitCommit=16a52d67daf2f4a8a76fa8bec02a76a46aa46257 gitTag=v4.1.2
time=2026-08-27T03:31:18.371Z level=INFO msg="Using executor retry strategy" argo=true Jitter=0.5 Steps=5 Duration=1s Factor=1.6
time=2026-08-27T03:31:18.371Z level=INFO msg="Executor initialized" gitTreeState=clean namespace=default podName=provision-tenant-29926-make-ns-596359171 gitTag=v4.1.2 gitCommit=16a52d67daf2f4a8a76fa8bec02a76a46aa46257 deadline=0001-01-01T00:00:00.000Z goVersion=go1.26.5 buildDate=2026-08-21T11:10:01Z templateName=make-ns version=v4.1.2 argo=true includeScriptOutput=false
time=2026-08-27T03:31:18.375Z level=WARN msg="failed to patch task result, see https://argo-workflows.readthedocs.io/en/latest/workflow-rbac/" error="workflowtaskresults.argoproj.io is forbidden: User \"system:serviceaccount:default:default\" cannot create resource \"workflowtaskresults\" in API group \"argoproj.io\" in the namespace \"default\"" argo=true attempt=0
time=2026-08-27T03:31:18.375Z level=WARN msg="Non-transient error" argo=true error="workflowtaskresults.argoproj.io is forbidden: User \"system:serviceaccount:default:default\" cannot create resource \"workflowtaskresults\" in API group \"argoproj.io\" in the namespace \"default\""
… (rest of the executor log omitted)

The node map is the whole diagnosis: each node's phase and message, in one JSON blob. A workflow's failure diagnostics are pod diagnostics plus that one layer.

Workflows as a self-service interface

what makes it a product rather than a script
  • Parameters at submit time: a WorkflowTemplate with typed, defaulted, documented parameters is an API. A Workflow with hard-coded values is a script you happened to run in a pod.
  • Validation before side effects: a first DAG node that rejects bad input (a name that is not a DNS label, a quota over policy) is admission control one layer earlier, and it is much cheaper than half-provisioning and rolling back.
  • Idempotency: a re-run should not double-create. resource templates with action: apply rather than create, or a when guard on an existence check, are the two usual answers. Workflows do not reconcile, so idempotency is your job.
  • Observability: the workflow's own status is the audit trail; ship it somewhere. And verify the outputs, not the runner: a green workflow that produced nothing is still a failure.
  • Triggering: CronWorkflow for schedules, the API/CLI for portals, or Argo Events (not installed here) for event-driven runs. Backstage templates commonly call one of these underneath, which is how a form becomes infrastructure.
Verify the output, not the runner

Any provisioning task should be graded twice: the workflow reached Succeeded, and the objects it was supposed to create exist with the right content. Exam graders check the second; get in the habit of checking both.

The fields that decide behavior

retries, deadlines, exit handlers, locks, cleanup

The six template types above cover the lab. The exam's cluster may hand you a workflow using one of the other four, so recognize them on sight:

Template typeDoesWrinkle
httpissues an HTTP request; the response body becomes outputs.result; successCondition evaluates response.statusCode or response.bodyruns in the per-workflow agent pod, which talks to the controller through the WorkflowTaskSet CRD, so the SA needs an agent role too
containerSetseveral containers in one pod, ordered with dependencies, sharing an emptyDir workspacethe container named main must finish last if you want base-layer output artifacts
datalists artifact paths from a repository (source.artifactPaths) and filters them with transformation[].expressionbare-bones by design; the only source today is the artifact repository
pluginhands the step to an executor plugin (an ExecutorPlugin CR) running in the agent podan admin-installed extension; treat "plugin" in a task as "check the ExecutorPlugin exists"

Retries, deadlines, garbage collection

  • retryStrategy. limit (a string, "3"), retryPolicy and optional backoff (duration, factor, maxDuration). The policy chooses what counts: OnFailure retries a main container that exited non-zero (the default), OnError retries controller errors and init/wait container failures, OnTransientError retries only errors matching Argo's transient list, Always retries everything. An expression over lastRetry.exitCode, lastRetry.status, lastRetry.duration and lastRetry.message narrows it further; from v3.5, if you give an expression and no policy the expression alone decides.
  • activeDeadlineSeconds. On the workflow spec it bounds the whole run; on a template it bounds one step. The node message when it trips is Pod was active on the node longer than the specified deadline. A workflow that must never run forever gets this field, not a hope.
  • ttlStrategy (secondsAfterCompletion, secondsAfterSuccess, secondsAfterFailure) deletes the Workflow object; podGC.strategy (OnPodCompletion, OnPodSuccess, OnWorkflowCompletion, OnWorkflowSuccess) deletes the pods. Set cluster-wide defaults for both under workflowDefaults in the controller ConfigMap; a Workflow's own value wins over the default. Upstream that ConfigMap is workflow-controller-configmap, but the Helm chart prefixes it with the release name, so in this lab it is argo-workflows-workflow-controller-configmap.
  • parallelism caps concurrent pods per workflow (or globally in the config map).

Exit handlers, hooks and richer dependencies

spec.onExit: <template> names a template that always runs after the entrypoint finishes; inside it, {{workflow.status}} is Succeeded, Failed or Error, and {{workflow.failures}} lists what broke. That is how notification, cleanup and "register the tenant only on success" get expressed. hooks generalize this: a workflow-level or template-level hooks.exit plus arbitrary named hooks with an expression such as workflow.status == "Running".

In a DAG, depends replaces dependencies with a boolean over task results: task-a.Succeeded, .Failed, .Errored, .Skipped, .Omitted, .Daemoned, combined with &&, || and !. A bare task name means "succeeded, skipped or daemoned". You cannot mix dependencies and depends in one DAG, and continueOn disappears when you use depends (write .Failed instead).

spec:
  entrypoint: main
  onExit: notify                      # always runs; sees {{workflow.status}}
  activeDeadlineSeconds: 900
  ttlStrategy: { secondsAfterSuccess: 3600 }
  podGC: { strategy: OnWorkflowSuccess }
  synchronization:
    mutexes: [{ name: provision-lock }]   # one run at a time per namespace
  templates:
    - name: flaky-step
      retryStrategy:
        limit: "3"
        retryPolicy: OnFailure
        backoff: { duration: "10s", factor: "2", maxDuration: "2m" }
      container: { image: alpine:3.20, command: [sh, -c, "exit 1"] }

Locks and resource templates

  • synchronization. mutexes allow one holder; semaphores take a size from a ConfigMap key (configMapKeyRef) and allow N. Both work at workflow level or template level, and both are namespace-scoped unless you configure the database-backed multi-controller variant. A workflow waiting on a lock sits in Pending with a message naming the lock, which is easy to misread as a scheduling problem. Note the plural field names (mutexes, semaphores): v3.6 made them lists.
  • resource templates. action is any kubectl verb (create, apply, patch, delete, get); successCondition and failureCondition poll the created object (for example status.succeeded > 0 / status.failed > 3 on a Job) so a step waits for a real outcome; patch takes mergeStrategy: strategic|merge|json, and custom resources cannot use strategic; setOwnerReference: true makes the created object die with the workflow, which is the opposite of what a provisioning workflow usually wants.
  • Outputs. outputs.parameters[].valueFrom.path reads a file the container wrote; outputs.result captures up to 256 kB of stdout (or the HTTP body). A reference to an output the producer never wrote fails the consuming node terminally unless you give valueFrom.default or an expression fallback (??).
  • Artifacts without plumbing. A key-only artifact (s3: { key: my-file }, nothing else) inherits bucket and credentials from the configured artifact repository (artifactRepositoryRef or the controller default), which is how templates stay portable between clusters.
Trap

The executor has been emissary only since v3.4; an Error (exit code 64) on a node is the emissary itself failing, not your command. The lab's workflowtaskresults RBAC lesson is the usual cause of the wait container failing shortly after.

The submission surface and what fails on it

templates, cron, events, server, archive, messages

Three ways to reference a template

  • workflowTemplateRef at the top of a Workflow spec runs a whole WorkflowTemplate; the Workflow's own arguments and entrypoint override the template's. argo submit --from workflowtemplate/<name> -p key=value is the CLI form; the UI's submit button is the same call.
  • templateRef inside a step or task borrows one template from a WorkflowTemplate (name + template); add clusterScope: true to reach a ClusterWorkflowTemplate. Forgetting clusterScope produces a "not found" error against a namespaced lookup.
  • Restrictions. A controller with workflowRestrictions.templateReferencing: Strict accepts only Workflows that use workflowTemplateRef; Secure additionally fails a running Workflow if its template changes underneath it. Under either, a submitted Workflow may only set an allow-listed set of fields (arguments, entrypoint, ttlStrategy, podGC, and a few more); a serviceAccountName or volumes on the submission is rejected. If your submit is refused for a field you are sure is valid, that is the reason.

CronWorkflow, the fields that bite

FieldDefaultMeans
schedulesrequireda list of cron expressions (v3.6+; older manifests use singular schedule)
concurrencyPolicyAllowForbid skips a tick while the previous run is alive; Replace kills the old run first
startingDeadlineSeconds0grace period after a missed tick (controller was down) during which one catch-up run still starts
timezonemachine tzIANA name; DST is honored, so 02:30 can be skipped or run twice
suspendfalsethe GitOps-friendly pause; argo cron suspend flips the same field
successfulJobsHistoryLimit / failedJobsHistoryLimit3 / 1how many finished Workflows to keep

Submitting from outside the shell

The Argo Server exposes POST /api/v1/workflows/{namespace} and, for webhooks, POST /api/v1/events/{namespace}/{discriminator}. A WorkflowEventBinding turns an event into a submission: its event.selector is an expression over payload, metadata (header names lower-cased) and discriminator, and its submit block names the WorkflowTemplate and maps payload fields into parameters. Clients authenticate with a ServiceAccount bearer token (Authorization: Bearer ...), and the server's --auth-mode decides what it accepts: client (default since v3.0, use the caller's token), server (act as the server's own SA), sso. Under SSO, RBAC maps OIDC groups to ServiceAccounts through the workflows.argoproj.io/rbac-rule and rbac-rule-precedence annotations on the SA. Argo Events (Sensor plus EventSource) is the separate project that turns queues and webhooks into Workflows at scale; the built-in events endpoint is the lightweight version.

The archive

Finished Workflows are Kubernetes objects and vanish with ttlStrategy. Setting persistence.archive: true with a PostgreSQL or MySQL connection in the controller config map copies every completed Workflow (not pod logs) into argo_archived_workflows, queryable from the UI and CLI (argo archive list). "Keep an auditable record of every provisioning run" is the exam sentence that means the archive plus a log shipper for pod output.

Reading failures

Node messageMeansFix lives in
Error (exit code 1)the main container failedyour command; read the pod log
OOMKilled (exit code 137)memory limitthe template's resources
Pod was active on the node longer than the specified deadlineactiveDeadlineSeconds trippedthe deadline or the step's speed
No more retries leftretryStrategy.limit exhaustedthe underlying child's message
... is forbidden: User "system:serviceaccount:..." cannot create resource "workflowtaskresults"executor RBACRole/RoleBinding for the workflow SA
child 'name' faileda DAG or steps parent summarizing a childdescend to the child node
Error (exit code 64)emissary executor failureusually RBAC or image command resolution
node stays Pending, message names a lock or a ConfigMap keywaiting on synchronizationthe other holder, or the semaphore size

Two other silent stalls: a Workflow in a namespace the controller does not watch (this lab: only argo and default) shows no status at all, and a Workflow whose SA lacks pods permissions in server auth mode fails at the server, not in a node. argo retry resumes a failed Workflow from the failed nodes; argo resubmit starts a fresh one; argo stop runs exit handlers, argo terminate does not.

How this gets tested

Expect tasks like "make this workflow retry up to three times with backoff", "notify on completion regardless of outcome", "run this nightly and never overlap", or "expose this Workflow so a user can submit it with a parameter". Each maps to one field above: retryStrategy, onExit, CronWorkflow with concurrencyPolicy: Forbid, and promotion to a WorkflowTemplate. Verify by reading status.nodes, not the UI's color.

Exercises

tick the dot when its check passes

The self-service story from 3.1, made concrete. First the identity:

kubectl -n default create sa provisioner
kubectl create clusterrole tenant-provisioner --verb=create,get --resource=namespaces,resourcequotas
kubectl create clusterrolebinding tenant-provisioner --clusterrole=tenant-provisioner --serviceaccount=default:provisioner
# Argo's executor reports each step's outcome through its own CRD, so every
# workflow SA needs this too; forget it and even a correct workflow fails:
kubectl -n default create role wf-taskresults --verb=create,patch --resource=workflowtaskresults.argoproj.io
kubectl -n default create rolebinding wf-taskresults --role=wf-taskresults --serviceaccount=default:provisioner
outputcaptured 2026-08-26
$ kubectl -n default create sa provisioner
serviceaccount/provisioner created
$ kubectl create clusterrole tenant-provisioner --verb=create,get --resource=namespaces,resourcequotas
clusterrole.rbac.authorization.k8s.io/tenant-provisioner created
$ kubectl create clusterrolebinding tenant-provisioner --clusterrole=tenant-provisioner --serviceaccount=default:provisioner
clusterrolebinding.rbac.authorization.k8s.io/tenant-provisioner created
# Argo's executor reports each step's outcome through its own CRD, so every
# workflow SA needs this too; forget it and even a correct workflow fails:
$ kubectl -n default create role wf-taskresults --verb=create,patch --resource=workflowtaskresults.argoproj.io
role.rbac.authorization.k8s.io/wf-taskresults created
$ kubectl -n default create rolebinding wf-taskresults --role=wf-taskresults --serviceaccount=default:provisioner
rolebinding.rbac.authorization.k8s.io/wf-taskresults created

Then the workflow, a two-node DAG using resource templates:

apiVersion: argoproj.io/v1alpha1
kind: Workflow
metadata: { generateName: provision-tenant-, namespace: default }
spec:
  entrypoint: provision
  serviceAccountName: provisioner
  arguments: { parameters: [{ name: team, value: team-d }] }
  templates:
    - name: provision
      dag:
        tasks:
          - name: namespace
            template: make-ns
          - name: quota
            template: make-quota
            dependencies: [namespace]
    - name: make-ns
      resource:
        action: create
        manifest: |
          apiVersion: v1
          kind: Namespace
          metadata:
            name: "{{workflow.parameters.team}}"
            labels: { tenant: "{{workflow.parameters.team}}" }
    - name: make-quota
      resource:
        action: create
        manifest: |
          apiVersion: v1
          kind: ResourceQuota
          metadata:
            name: default-quota
            namespace: "{{workflow.parameters.team}}"
          spec:
            hard: { requests.cpu: "1", requests.memory: 2Gi, pods: "10" }

kubectl create -f it (generateName forbids apply), then watch: kubectl get workflow -w.

verify: the workflow reaches Succeeded, and kubectl get ns team-d --show-labels plus kubectl -n team-d get resourcequota show the outputs.

Rerun with serviceAccountName: default and a new team name. Which rule it names first is instructive: with the bare default SA the executor usually trips over workflowtaskresults (its own reporting channel) before your namespace rule even gets a chance. That is an Argo-specific wrinkle worth seeing once.

verify: the workflow fails, and kubectl get workflow <name> -o jsonpath='{.status.nodes}' | jq '.[] | {phase, message}' contains a forbidden message naming a missing verb and resource. The message structure is identical for every RBAC failure you will ever debug.

Submit the same workflow for team-e without editing the file. The UI's submit flow lists WorkflowTemplates, not bare Workflows, so either promote yours to a WorkflowTemplate first (change the kind, drop generateName for a name) and submit it from the UI with the parameter overridden, or stay in the shell and sed 's/team-d/team-e/' workflow.yaml | kubectl create -f -.

verify: two tenants exist, one workflow spec. Parameters-at-submit is what makes a workflow a self-service endpoint rather than a script, and the WorkflowTemplate promotion is exactly how you would productize it.

Add a first DAG node check using a script template that exits non-zero when {{workflow.parameters.team}} doesn't match ^team-[a-z]+$, and make the other nodes depend on it.

verify: a run with team=Team_X fails at check and creates nothing. Clean up the test tenants when done: kubectl delete ns team-d team-e.

A bare Workflow is a script someone ran. A WorkflowTemplate is a thing with a name that other people can submit with their own parameters, which is the difference between automation and a platform API.

kubectl delete ns team-g --ignore-not-found
kubectl create role wf-executor --verb=create,patch --resource=workflowtaskresults.argoproj.io --dry-run=client -o yaml | kubectl apply -f -
kubectl create rolebinding default-wf-executor --role=wf-executor --serviceaccount=default:default --dry-run=client -o yaml | kubectl apply -f -
kubectl apply -f - <<'EOF'
apiVersion: argoproj.io/v1alpha1
kind: WorkflowTemplate
metadata: { name: provision-tenant, namespace: default }
spec:
  entrypoint: provision
  serviceAccountName: provisioner
  arguments: { parameters: [{ name: team, value: team-d }] }
  templates:
    - name: provision
      dag:
        tasks:
          - name: namespace
            template: make-ns
          - name: quota
            template: make-quota
            dependencies: [namespace]
    - name: make-ns
      resource:
        action: create
        manifest: |
          apiVersion: v1
          kind: Namespace
          metadata:
            name: "{{workflow.parameters.team}}"
            labels: { tenant: "{{workflow.parameters.team}}" }
    - name: make-quota
      resource:
        action: create
        manifest: |
          apiVersion: v1
          kind: ResourceQuota
          metadata:
            name: default-quota
            namespace: "{{workflow.parameters.team}}"
          spec:
            hard: { requests.cpu: "1", requests.memory: 2Gi, pods: "10" }
EOF
argo submit --from workflowtemplate/provision-tenant -p team=team-g --wait
argo get @latest
kubectl get ns team-g
kubectl -n team-g get resourcequota default-quota
kubectl delete ns team-g --ignore-not-found --wait=false
kubectl delete workflowtemplate provision-tenant
outputcaptured 2026-09-12
$ kubectl delete ns team-g --ignore-not-found
namespace "team-g" deleted
$ kubectl create role wf-executor --verb=create,patch --resource=workflowtaskresults.argoproj.io --dry-run=client -o yaml | kubectl apply -f -
role.rbac.authorization.k8s.io/wf-executor unchanged
$ kubectl create rolebinding default-wf-executor --role=wf-executor --serviceaccount=default:default --dry-run=client -o yaml | kubectl apply -f -
rolebinding.rbac.authorization.k8s.io/default-wf-executor unchanged
$ kubectl apply -f - <<'EOF'
apiVersion: argoproj.io/v1alpha1
kind: WorkflowTemplate
metadata: { name: provision-tenant, namespace: default }
spec:
  entrypoint: provision
  serviceAccountName: provisioner
  arguments: { parameters: [{ name: team, value: team-d }] }
  templates:
    - name: provision
      dag:
        tasks:
          - name: namespace
            template: make-ns
          - name: quota
            template: make-quota
            dependencies: [namespace]
    - name: make-ns
      resource:
        action: create
        manifest: |
          apiVersion: v1
          kind: Namespace
          metadata:
            name: "{{workflow.parameters.team}}"
            labels: { tenant: "{{workflow.parameters.team}}" }
    - name: make-quota
      resource:
        action: create
        manifest: |
          apiVersion: v1
          kind: ResourceQuota
          metadata:
            name: default-quota
            namespace: "{{workflow.parameters.team}}"
          spec:
            hard: { requests.cpu: "1", requests.memory: 2Gi, pods: "10" }
EOF
workflowtemplate.argoproj.io/provision-tenant unchanged
$ argo submit --from workflowtemplate/provision-tenant -p team=team-g --wait
Name:                provision-tenant-524qt
Namespace:           default
ServiceAccount:      unset
Status:              Pending
Created:             Sat Sep 12 23:12:05 -0400 (now)
Progress:            
Parameters:          
  team:              team-g
provision-tenant-524qt Succeeded at 2026-09-12 23:12:25 -0400 EDT
$ argo get @latest
Name:                provision-tenant-524qt
Namespace:           default
ServiceAccount:      provisioner
Status:              Succeeded
Conditions:          
 PodRunning          False
 Completed           True
Created:             Sat Sep 12 23:12:05 -0400 (21 seconds ago)
Started:             Sat Sep 12 23:12:05 -0400 (21 seconds ago)
Finished:            Sat Sep 12 23:12:25 -0400 (1 second ago)
Duration:            20 seconds
Progress:            2/2
ResourcesDuration:   2s*(100Mi memory),0s*(1 cpu)
Parameters:          
  team:              team-g

STEP                       TEMPLATE    PODNAME                                       DURATION  MESSAGE
 ✔ provision-tenant-524qt  provision                                                             
 ├─✔ namespace             make-ns     provision-tenant-524qt-make-ns-3648215291     5s          
 └─✔ quota                 make-quota  provision-tenant-524qt-make-quota-3757520604  5s          
$ kubectl get ns team-g
NAME     STATUS   AGE
team-g   Active   17s
$ kubectl -n team-g get resourcequota default-quota
NAME            REQUEST                                                 LIMIT   AGE
default-quota   pods: 0/10, requests.cpu: 0/1, requests.memory: 0/2Gi           6s
$ kubectl delete ns team-g --ignore-not-found --wait=false
namespace "team-g" deleted
$ kubectl delete workflowtemplate provision-tenant
workflowtemplate.argoproj.io "provision-tenant" deleted from default namespace
verify: the submit prints a live node tree that ends Succeeded with both resource nodes green, and the namespace and quota exist.

A retry strategy turns one node into a parent with children, and the node tree shows every attempt. The backoff fields are the ones people guess wrong, so read them off a real run.

kubectl create role wf-executor --verb=create,patch --resource=workflowtaskresults.argoproj.io --dry-run=client -o yaml | kubectl apply -f -
kubectl create rolebinding default-wf-executor --role=wf-executor --serviceaccount=default:default --dry-run=client -o yaml | kubectl apply -f -
kubectl create -f - <<'EOF'
apiVersion: argoproj.io/v1alpha1
kind: Workflow
metadata: { generateName: retryer-, namespace: default }
spec:
  entrypoint: flaky
  templates:
    - name: flaky
      retryStrategy:
        limit: "2"
        retryPolicy: OnFailure
        backoff: { duration: "5s", factor: "2" }
      container:
        image: busybox:1.36
        command: [sh, -c]
        args: ["exit 1"]
EOF
sleep 90
WF=$(kubectl get wf --sort-by=.metadata.creationTimestamp -o jsonpath='{.items[-1:].metadata.name}')
kubectl get wf "$WF" -o jsonpath='{.status.nodes}' | jq '.[] | {displayName, type, phase, message}'
kubectl get wf "$WF" -o jsonpath='{.status.phase} {.status.message}{"\n"}'
outputcaptured 2026-09-12
$ kubectl create role wf-executor --verb=create,patch --resource=workflowtaskresults.argoproj.io --dry-run=client -o yaml | kubectl apply -f -
role.rbac.authorization.k8s.io/wf-executor unchanged
$ kubectl create rolebinding default-wf-executor --role=wf-executor --serviceaccount=default:default --dry-run=client -o yaml | kubectl apply -f -
rolebinding.rbac.authorization.k8s.io/default-wf-executor unchanged
$ kubectl create -f - <<'EOF'
apiVersion: argoproj.io/v1alpha1
kind: Workflow
metadata: { generateName: retryer-, namespace: default }
spec:
  entrypoint: flaky
  templates:
    - name: flaky
      retryStrategy:
        limit: "2"
        retryPolicy: OnFailure
        backoff: { duration: "5s", factor: "2" }
      container:
        image: busybox:1.36
        command: [sh, -c]
        args: ["exit 1"]
EOF
workflow.argoproj.io/retryer-mhqnn created
$ sleep 90
$ WF=$(kubectl get wf --sort-by=.metadata.creationTimestamp -o jsonpath='{.items[-1:].metadata.name}')
$ kubectl get wf "$WF" -o jsonpath='{.status.nodes}' | jq '.[] | {displayName, type, phase, message}'
{
  "displayName": "retryer-mhqnn",
  "type": "Retry",
  "phase": "Failed",
  "message": "No more retries left"
}
{
  "displayName": "retryer-mhqnn(1)",
  "type": "Pod",
  "phase": "Failed",
  "message": "main: Error (exit code 1)"
}
{
  "displayName": "retryer-mhqnn(0)",
  "type": "Pod",
  "phase": "Failed",
  "message": "main: Error (exit code 1)"
}
{
  "displayName": "retryer-mhqnn(2)",
  "type": "Pod",
  "phase": "Failed",
  "message": "main: Error (exit code 1)"
}
$ kubectl get wf "$WF" -o jsonpath='{.status.phase} {.status.message}{"\n"}'
Failed No more retries left
verify: there are three child nodes under one retry node, and the final message says there are no more retries left. limit: 2 means two retries after the first attempt, not two attempts in total; that off-by-one is the thing to remember.

activeDeadlineSeconds is enforced by the pod, not by the controller, so the message you get back is Kubernetes' own. Recognize it, because it looks nothing like an Argo error.

kubectl create role wf-executor --verb=create,patch --resource=workflowtaskresults.argoproj.io --dry-run=client -o yaml | kubectl apply -f -
kubectl create rolebinding default-wf-executor --role=wf-executor --serviceaccount=default:default --dry-run=client -o yaml | kubectl apply -f -
kubectl create -f - <<'EOF'
apiVersion: argoproj.io/v1alpha1
kind: Workflow
metadata: { generateName: deadline-, namespace: default }
spec:
  entrypoint: slow
  templates:
    - name: slow
      activeDeadlineSeconds: 10
      container:
        image: busybox:1.36
        command: [sh, -c]
        args: ["sleep 60"]
EOF
sleep 60
WF=$(kubectl get wf --sort-by=.metadata.creationTimestamp -o jsonpath='{.items[-1:].metadata.name}')
kubectl get wf "$WF" -o jsonpath='{.status.nodes}' | jq '.[] | {displayName, phase, message}'
kubectl get wf -o name | grep 'deadline-' | xargs -r kubectl delete
outputcaptured 2026-09-12
$ kubectl create role wf-executor --verb=create,patch --resource=workflowtaskresults.argoproj.io --dry-run=client -o yaml | kubectl apply -f -
role.rbac.authorization.k8s.io/wf-executor unchanged
$ kubectl create rolebinding default-wf-executor --role=wf-executor --serviceaccount=default:default --dry-run=client -o yaml | kubectl apply -f -
rolebinding.rbac.authorization.k8s.io/default-wf-executor unchanged
$ kubectl create -f - <<'EOF'
apiVersion: argoproj.io/v1alpha1
kind: Workflow
metadata: { generateName: deadline-, namespace: default }
spec:
  entrypoint: slow
  templates:
    - name: slow
      activeDeadlineSeconds: 10
      container:
        image: busybox:1.36
        command: [sh, -c]
        args: ["sleep 60"]
EOF
workflow.argoproj.io/deadline-pwgzt created
$ sleep 60
$ WF=$(kubectl get wf --sort-by=.metadata.creationTimestamp -o jsonpath='{.items[-1:].metadata.name}')
$ kubectl get wf "$WF" -o jsonpath='{.status.nodes}' | jq '.[] | {displayName, phase, message}'
{
  "displayName": "deadline-pwgzt",
  "phase": "Failed",
  "message": "Pod was active on the node longer than the specified deadline"
}
$ kubectl get wf -o name | grep 'deadline-' | xargs -r kubectl delete
workflow.argoproj.io "deadline-pwgzt" deleted from default namespace
verify: the node message says the pod was active on the node longer than the specified deadline.

Self-service means two people submit at once. A mutex serializes them without either one failing, and the waiting workflow records what it is waiting for in status.message, which is where you look when someone asks why their run has not started.

kubectl create role wf-executor --verb=create,patch --resource=workflowtaskresults.argoproj.io --dry-run=client -o yaml | kubectl apply -f -
kubectl create rolebinding default-wf-executor --role=wf-executor --serviceaccount=default:default --dry-run=client -o yaml | kubectl apply -f -
cat > /tmp/locked.yaml <<'EOF'
apiVersion: argoproj.io/v1alpha1
kind: Workflow
metadata: { generateName: locked-, namespace: default }
spec:
  entrypoint: work
  synchronization:
    mutexes: [{ name: lab-lock }]
  templates:
    - name: work
      container:
        image: busybox:1.36
        command: [sh, -c]
        args: ["sleep 60"]
EOF
kubectl create -f /tmp/locked.yaml
kubectl create -f /tmp/locked.yaml
sleep 15
kubectl get wf -l workflows.argoproj.io/phase --sort-by=.metadata.creationTimestamp | tail -3
kubectl get wf --sort-by=.metadata.creationTimestamp -o jsonpath='{range .items[-2:]}{.metadata.name} {.status.phase}{"\n"}{end}'
kubectl get wf --sort-by=.metadata.creationTimestamp -o jsonpath='{.items[-1:].status.message}{"\n"}'
kubectl -n argo get cm argo-workflows-workflow-controller-configmap -o jsonpath='{.data}' | jq
kubectl delete wf -l workflows.argoproj.io/workflow --ignore-not-found 2>/dev/null; kubectl get wf -o name | grep locked- | xargs -r kubectl delete
outputcaptured 2026-09-12
$ kubectl create role wf-executor --verb=create,patch --resource=workflowtaskresults.argoproj.io --dry-run=client -o yaml | kubectl apply -f -
role.rbac.authorization.k8s.io/wf-executor unchanged
$ kubectl create rolebinding default-wf-executor --role=wf-executor --serviceaccount=default:default --dry-run=client -o yaml | kubectl apply -f -
rolebinding.rbac.authorization.k8s.io/default-wf-executor unchanged
$ cat > /tmp/locked.yaml <<'EOF'
apiVersion: argoproj.io/v1alpha1
kind: Workflow
metadata: { generateName: locked-, namespace: default }
spec:
  entrypoint: work
  synchronization:
    mutexes: [{ name: lab-lock }]
  templates:
    - name: work
      container:
        image: busybox:1.36
        command: [sh, -c]
        args: ["sleep 60"]
EOF
$ kubectl create -f /tmp/locked.yaml
workflow.argoproj.io/locked-ccpvq created
$ kubectl create -f /tmp/locked.yaml
workflow.argoproj.io/locked-2nmfw created
$ sleep 15
$ kubectl get wf -l workflows.argoproj.io/phase --sort-by=.metadata.creationTimestamp | tail -3
rightscope-ccj6r         Succeeded   5m2s    
locked-ccpvq             Running     16s     
locked-2nmfw             Pending             Waiting for default/Mutex/lab-lock lock. Lock status: 0/1
$ kubectl get wf --sort-by=.metadata.creationTimestamp -o jsonpath='{range .items[-2:]}{.metadata.name} {.status.phase}{"\n"}{end}'
locked-ccpvq Running
locked-2nmfw Pending
$ kubectl get wf --sort-by=.metadata.creationTimestamp -o jsonpath='{.items[-1:].status.message}{"\n"}'
Waiting for default/Mutex/lab-lock lock. Lock status: 0/1
$ kubectl -n argo get cm argo-workflows-workflow-controller-configmap -o jsonpath='{.data}' | jq
{
  "config": "nodeEvents:\n  enabled: true\nworkflowEvents:\n  enabled: true\nfailedPodRestart:\n  enabled: false\n  maxRestarts: 3\n"
}
$ kubectl delete wf -l workflows.argoproj.io/workflow --ignore-not-found 2>/dev/null; kubectl get wf -o name | grep locked- | xargs -r kubectl delete
No resources found
workflow.argoproj.io "locked-2nmfw" deleted from default namespace
workflow.argoproj.io "locked-ccpvq" deleted from default namespace
verify: one workflow runs and the other is Pending with a message naming the lock.

onExit is the finally block: it sees the workflow's final status and runs whether the entrypoint succeeded or not. The interesting part is that one kind of cancellation skips it.

kubectl create role wf-executor --verb=create,patch --resource=workflowtaskresults.argoproj.io --dry-run=client -o yaml | kubectl apply -f -
kubectl create rolebinding default-wf-executor --role=wf-executor --serviceaccount=default:default --dry-run=client -o yaml | kubectl apply -f -
kubectl create -f - <<'EOF'
apiVersion: argoproj.io/v1alpha1
kind: Workflow
metadata: { generateName: exiter-, namespace: default }
spec:
  entrypoint: work
  onExit: report
  templates:
    - name: work
      container:
        image: busybox:1.36
        command: [sh, -c]
        args: ["exit 1"]
    - name: report
      container:
        image: busybox:1.36
        command: [sh, -c]
        args: ["echo workflow finished as {{workflow.status}}"]
EOF
sleep 75
WF=$(kubectl get wf --sort-by=.metadata.creationTimestamp -o jsonpath='{.items[-1:].metadata.name}')
argo logs "$WF" | tail -5
kubectl create -f - <<'EOF'
apiVersion: argoproj.io/v1alpha1
kind: Workflow
metadata: { name: stoppable, namespace: default }
spec:
  entrypoint: work
  onExit: report
  templates:
    - name: work
      container:
        image: busybox:1.36
        command: [sh, -c]
        args: ["sleep 600"]
    - name: report
      container:
        image: busybox:1.36
        command: [sh, -c]
        args: ["echo exit handler ran after {{workflow.status}}"]
EOF
sleep 20
argo stop stoppable
sleep 45
kubectl get wf stoppable -o jsonpath='{.status.nodes}' | jq '.[] | {displayName, phase}'
kubectl delete wf stoppable
kubectl get wf -o name | grep 'exiter-' | xargs -r kubectl delete
outputcaptured 2026-09-12
$ kubectl create role wf-executor --verb=create,patch --resource=workflowtaskresults.argoproj.io --dry-run=client -o yaml | kubectl apply -f -
role.rbac.authorization.k8s.io/wf-executor unchanged
$ kubectl create rolebinding default-wf-executor --role=wf-executor --serviceaccount=default:default --dry-run=client -o yaml | kubectl apply -f -
rolebinding.rbac.authorization.k8s.io/default-wf-executor unchanged
$ kubectl create -f - <<'EOF'
apiVersion: argoproj.io/v1alpha1
kind: Workflow
metadata: { generateName: exiter-, namespace: default }
spec:
  entrypoint: work
  onExit: report
  templates:
    - name: work
      container:
        image: busybox:1.36
        command: [sh, -c]
        args: ["exit 1"]
    - name: report
      container:
        image: busybox:1.36
        command: [sh, -c]
        args: ["echo workflow finished as {{workflow.status}}"]
EOF
workflow.argoproj.io/exiter-dl5qg created
$ sleep 75
$ WF=$(kubectl get wf --sort-by=.metadata.creationTimestamp -o jsonpath='{.items[-1:].metadata.name}')
$ argo logs "$WF" | tail -5
exiter-dl5qg: Error: exit status 1
exiter-dl5qg-report-4006390619: time=2026-09-13T12:59:01.702Z level=INFO msg="waiting for signals" argo=true signalPath=/var/run/argo/ctr/main/signal
exiter-dl5qg-report-4006390619: workflow finished as Failed
exiter-dl5qg-report-4006390619: time=2026-09-13T12:59:02.702Z level=INFO msg="sub-process exited" error=<nil> argo=true
exiter-dl5qg-report-4006390619: time=2026-09-13T12:59:02.703Z level=INFO msg="file signal handler exiting due to context cancellation" argo=true
$ kubectl create -f - <<'EOF'
apiVersion: argoproj.io/v1alpha1
kind: Workflow
metadata: { name: stoppable, namespace: default }
spec:
  entrypoint: work
  onExit: report
  templates:
    - name: work
      container:
        image: busybox:1.36
        command: [sh, -c]
        args: ["sleep 600"]
    - name: report
      container:
        image: busybox:1.36
        command: [sh, -c]
        args: ["echo exit handler ran after {{workflow.status}}"]
EOF
workflow.argoproj.io/stoppable created
$ sleep 20
$ argo stop stoppable
workflow stoppable stopped
$ sleep 45
$ kubectl get wf stoppable -o jsonpath='{.status.nodes}' | jq '.[] | {displayName, phase}'
{
  "displayName": "stoppable",
  "phase": "Failed"
}
{
  "displayName": "stoppable.onExit",
  "phase": "Succeeded"
}
$ kubectl delete wf stoppable
workflow.argoproj.io "stoppable" deleted from default namespace
$ kubectl get wf -o name | grep 'exiter-' | xargs -r kubectl delete
workflow.argoproj.io "exiter-dl5qg" deleted from default namespace
verify: the failing workflow's exit handler logs the status Failed, and argo stop leaves an exit-handler node behind while argo terminate would not.

A CronWorkflow that fires faster than its work finishes is a self-service outage waiting to happen. concurrencyPolicy: Forbid makes the controller skip the tick instead, and it records that it did.

kubectl create role wf-executor --verb=create,patch --resource=workflowtaskresults.argoproj.io --dry-run=client -o yaml | kubectl apply -f -
kubectl create rolebinding default-wf-executor --role=wf-executor --serviceaccount=default:default --dry-run=client -o yaml | kubectl apply -f -
kubectl create -f - <<'EOF'
apiVersion: argoproj.io/v1alpha1
kind: CronWorkflow
metadata: { name: every-minute, namespace: default }
spec:
  schedules: ["* * * * *"]
  concurrencyPolicy: Forbid
  workflowSpec:
    entrypoint: work
    templates:
      - name: work
        container:
          image: busybox:1.36
          command: [sh, -c]
          args: ["sleep 90"]
EOF
sleep 210
argo cron get every-minute
kubectl get wf -l workflows.argoproj.io/cron-workflow=every-minute
kubectl get cronwf every-minute -o jsonpath='{.status}' | jq
kubectl delete cronwf every-minute
kubectl delete wf -l workflows.argoproj.io/cron-workflow=every-minute
outputcaptured 2026-09-12
$ kubectl create role wf-executor --verb=create,patch --resource=workflowtaskresults.argoproj.io --dry-run=client -o yaml | kubectl apply -f -
role.rbac.authorization.k8s.io/wf-executor unchanged
$ kubectl create rolebinding default-wf-executor --role=wf-executor --serviceaccount=default:default --dry-run=client -o yaml | kubectl apply -f -
rolebinding.rbac.authorization.k8s.io/default-wf-executor unchanged
$ kubectl create -f - <<'EOF'
apiVersion: argoproj.io/v1alpha1
kind: CronWorkflow
metadata: { name: every-minute, namespace: default }
spec:
  schedules: ["* * * * *"]
  concurrencyPolicy: Forbid
  workflowSpec:
    entrypoint: work
    templates:
      - name: work
        container:
          image: busybox:1.36
          command: [sh, -c]
          args: ["sleep 90"]
EOF
cronworkflow.argoproj.io/every-minute created
$ sleep 210
$ argo cron get every-minute
Name:                          every-minute
Namespace:                     default
Created:                       Sat Sep 12 23:16:10 -0400 (3 minutes ago)
Schedules:                     * * * * *
Suspended:                     false
ConcurrencyPolicy:             Forbid
LastScheduledTime:             Sat Sep 12 23:19:00 -0400 (40 seconds ago)
NextScheduledTime:             Sat Sep 12 23:20:00 -0400 (19 seconds from now) (assumes workflow-controller is in UTC)
Active Workflows:              every-minute-1789269540
$ kubectl get wf -l workflows.argoproj.io/cron-workflow=every-minute
NAME                      STATUS      AGE     MESSAGE
every-minute-1789269420   Succeeded   2m41s   
every-minute-1789269540   Running     41s     
$ kubectl get cronwf every-minute -o jsonpath='{.status}' | jq
{
  "active": [
    {
      "apiVersion": "argoproj.io/v1alpha1",
      "kind": "Workflow",
      "name": "every-minute-1789269540",
      "namespace": "default",
      "resourceVersion": "338837",
      "uid": "d13a072b-faf2-4516-ae37-69b988fb3d1b"
    }
  ],
  "failed": 0,
  "lastScheduledTime": "2026-09-13T03:19:00Z",
  "phase": "Active",
  "succeeded": 1
}
$ kubectl delete cronwf every-minute
cronworkflow.argoproj.io "every-minute" deleted from default namespace
$ kubectl delete wf -l workflows.argoproj.io/cron-workflow=every-minute
No resources found
verify: two Workflows exist for three and a half minutes of a one-minute schedule, and argo cron get shows a single Active Workflow. The status keeps no record of the skipped ticks: it carries active, lastScheduledTime, succeeded and failed and nothing else, so the skips only show as the gap between minutes elapsed and workflows created. Replace would have deleted the running Workflow and started the new one; Allow would have let them pile up.

A ClusterWorkflowTemplate and a WorkflowTemplate are different kinds with identical fields. Referencing one as the other produces a not-found error that sends people looking for a typo that is not there.

kubectl create role wf-executor --verb=create,patch --resource=workflowtaskresults.argoproj.io --dry-run=client -o yaml | kubectl apply -f -
kubectl create rolebinding default-wf-executor --role=wf-executor --serviceaccount=default:default --dry-run=client -o yaml | kubectl apply -f -
kubectl apply -f - <<'EOF'
apiVersion: argoproj.io/v1alpha1
kind: ClusterWorkflowTemplate
metadata: { name: shared-say }
spec:
  templates:
    - name: say
      container:
        image: busybox:1.36
        command: [sh, -c]
        args: ["echo from the cluster-scoped template"]
EOF
kubectl create -f - <<'EOF'
apiVersion: argoproj.io/v1alpha1
kind: Workflow
metadata: { generateName: wrongscope-, namespace: default }
spec:
  entrypoint: main
  templates:
    - name: main
      steps:
        - - name: call
            templateRef: { name: shared-say, template: say }
EOF
sleep 30
kubectl get wf --sort-by=.metadata.creationTimestamp -o jsonpath='{.items[-1:].status.message}{"\n"}'
kubectl create -f - <<'EOF'
apiVersion: argoproj.io/v1alpha1
kind: Workflow
metadata: { generateName: rightscope-, namespace: default }
spec:
  entrypoint: main
  templates:
    - name: main
      steps:
        - - name: call
            templateRef: { name: shared-say, template: say, clusterScope: true }
EOF
sleep 45
argo logs @latest | tail -3
kubectl delete clusterworkflowtemplate shared-say
outputcaptured 2026-09-12
$ kubectl create role wf-executor --verb=create,patch --resource=workflowtaskresults.argoproj.io --dry-run=client -o yaml | kubectl apply -f -
role.rbac.authorization.k8s.io/wf-executor unchanged
$ kubectl create rolebinding default-wf-executor --role=wf-executor --serviceaccount=default:default --dry-run=client -o yaml | kubectl apply -f -
rolebinding.rbac.authorization.k8s.io/default-wf-executor unchanged
$ kubectl apply -f - <<'EOF'
apiVersion: argoproj.io/v1alpha1
kind: ClusterWorkflowTemplate
metadata: { name: shared-say }
spec:
  templates:
    - name: say
      container:
        image: busybox:1.36
        command: [sh, -c]
        args: ["echo from the cluster-scoped template"]
EOF
clusterworkflowtemplate.argoproj.io/shared-say created
$ kubectl create -f - <<'EOF'
apiVersion: argoproj.io/v1alpha1
kind: Workflow
metadata: { generateName: wrongscope-, namespace: default }
spec:
  entrypoint: main
  templates:
    - name: main
      steps:
        - - name: call
            templateRef: { name: shared-say, template: say }
EOF
workflow.argoproj.io/wrongscope-pzf62 created
$ sleep 30
$ kubectl get wf --sort-by=.metadata.creationTimestamp -o jsonpath='{.items[-1:].status.message}{"\n"}'
invalid spec: templates.main.steps[0].call template reference shared-say.say not found
$ kubectl create -f - <<'EOF'
apiVersion: argoproj.io/v1alpha1
kind: Workflow
metadata: { generateName: rightscope-, namespace: default }
spec:
  entrypoint: main
  templates:
    - name: main
      steps:
        - - name: call
            templateRef: { name: shared-say, template: say, clusterScope: true }
EOF
workflow.argoproj.io/rightscope-ccj6r created
$ sleep 45
$ argo logs @latest | tail -3
rightscope-ccj6r-say-4102760526: from the cluster-scoped template
rightscope-ccj6r-say-4102760526: time=2026-09-13T03:20:45.688Z level=INFO msg="sub-process exited" argo=true error=<nil>
rightscope-ccj6r-say-4102760526: time=2026-09-13T03:20:45.688Z level=INFO msg="file signal handler exiting due to context cancellation" argo=true
$ kubectl delete clusterworkflowtemplate shared-say
clusterworkflowtemplate.argoproj.io "shared-say" deleted
verify: the first run fails with invalid spec: templates.main.steps[0].call template reference shared-say.say not found, and one boolean fixes it. Note what the message does not say: it names neither the kind it looked for nor the scope it looked in, so a cluster-scoped object sitting right there looks like a typo. clusterScope: true is the difference.

A platform that lets anyone submit arbitrary workflow specs has no guardrails at all. templateReferencing: Strict means every submission must go through a reviewed template, and a ServiceAccount override is refused too. Note where the enforcement happens: there is no admission webhook, so the API accepts all three of these Workflows. The controller is what refuses them, which means you read the verdict out of status.message rather than out of the kubectl create exit code.

kubectl create role wf-executor --verb=create,patch --resource=workflowtaskresults.argoproj.io --dry-run=client -o yaml | kubectl apply -f -
kubectl create rolebinding default-wf-executor --role=wf-executor --serviceaccount=default:default --dry-run=client -o yaml | kubectl apply -f -
kubectl apply -f - <<'EOF'
apiVersion: argoproj.io/v1alpha1
kind: WorkflowTemplate
metadata: { name: sanctioned, namespace: default }
spec:
  entrypoint: work
  serviceAccountName: default
  templates:
    - name: work
      container: { image: busybox:1.36, command: [sh, -c], args: ["echo sanctioned"] }
EOF
# every controller setting lives in the single config key, so append to it rather than adding a sibling
kubectl -n argo get cm argo-workflows-workflow-controller-configmap -o jsonpath='{.data.config}' > /tmp/argocfg.yaml
cp /tmp/argocfg.yaml /tmp/argocfg.orig
printf 'workflowRestrictions:\n  templateReferencing: Strict\n' >> /tmp/argocfg.yaml
kubectl -n argo create cm argo-workflows-workflow-controller-configmap --from-file=config=/tmp/argocfg.yaml --dry-run=client -o yaml | kubectl -n argo apply -f -
kubectl -n argo rollout restart deploy argo-workflows-workflow-controller
kubectl -n argo rollout status deploy argo-workflows-workflow-controller --timeout=180s
kubectl create -f - <<'EOF'
apiVersion: argoproj.io/v1alpha1
kind: Workflow
metadata: { generateName: bare-, namespace: default }
spec:
  entrypoint: work
  templates:
    - name: work
      container: { image: busybox:1.36, command: [sh, -c], args: ["echo hi"] }
EOF
kubectl create -f - <<'EOF'
apiVersion: argoproj.io/v1alpha1
kind: Workflow
metadata: { generateName: override-, namespace: default }
spec:
  serviceAccountName: default
  workflowTemplateRef: { name: sanctioned }
EOF
kubectl create -f - <<'EOF'
apiVersion: argoproj.io/v1alpha1
kind: Workflow
metadata: { generateName: fromtmpl-, namespace: default }
spec:
  workflowTemplateRef: { name: sanctioned }
EOF
sleep 40
kubectl get wf --sort-by=.metadata.creationTimestamp -o jsonpath='{range .items[-3:]}{.metadata.name} {.status.phase} {.status.message}{"\n"}{end}'
kubectl -n argo create cm argo-workflows-workflow-controller-configmap --from-file=config=/tmp/argocfg.orig --dry-run=client -o yaml | kubectl -n argo apply -f -
kubectl -n argo rollout restart deploy argo-workflows-workflow-controller
kubectl -n argo rollout status deploy argo-workflows-workflow-controller --timeout=180s
kubectl delete workflowtemplate sanctioned
kubectl get wf -o name | grep -E 'bare-|override-|fromtmpl-' | xargs -r kubectl delete
outputcaptured 2026-09-12
$ kubectl create role wf-executor --verb=create,patch --resource=workflowtaskresults.argoproj.io --dry-run=client -o yaml | kubectl apply -f -
role.rbac.authorization.k8s.io/wf-executor unchanged
$ kubectl create rolebinding default-wf-executor --role=wf-executor --serviceaccount=default:default --dry-run=client -o yaml | kubectl apply -f -
rolebinding.rbac.authorization.k8s.io/default-wf-executor unchanged
$ kubectl apply -f - <<'EOF'
apiVersion: argoproj.io/v1alpha1
kind: WorkflowTemplate
metadata: { name: sanctioned, namespace: default }
spec:
  entrypoint: work
  serviceAccountName: default
  templates:
    - name: work
      container: { image: busybox:1.36, command: [sh, -c], args: ["echo sanctioned"] }
EOF
workflowtemplate.argoproj.io/sanctioned unchanged
$ # every controller setting lives in the single config key, so append to it rather than adding a sibling
$ kubectl -n argo get cm argo-workflows-workflow-controller-configmap -o jsonpath='{.data.config}' > /tmp/argocfg.yaml
$ cp /tmp/argocfg.yaml /tmp/argocfg.orig
$ printf 'workflowRestrictions:\n  templateReferencing: Strict\n' >> /tmp/argocfg.yaml
$ kubectl -n argo create cm argo-workflows-workflow-controller-configmap --from-file=config=/tmp/argocfg.yaml --dry-run=client -o yaml | kubectl -n argo apply -f -
configmap/argo-workflows-workflow-controller-configmap configured
$ kubectl -n argo rollout restart deploy argo-workflows-workflow-controller
deployment.apps/argo-workflows-workflow-controller restarted
$ kubectl -n argo rollout status deploy argo-workflows-workflow-controller --timeout=180s
Waiting for deployment "argo-workflows-workflow-controller" rollout to finish: 0 out of 1 new replicas have been updated...
Waiting for deployment "argo-workflows-workflow-controller" rollout to finish: 0 out of 1 new replicas have been updated...
Waiting for deployment "argo-workflows-workflow-controller" rollout to finish: 0 out of 1 new replicas have been updated...
Waiting for deployment "argo-workflows-workflow-controller" rollout to finish: 0 out of 1 new replicas have been updated...
Waiting for deployment "argo-workflows-workflow-controller" rollout to finish: 0 of 1 updated replicas are available...
deployment "argo-workflows-workflow-controller" successfully rolled out
$ kubectl create -f - <<'EOF'
apiVersion: argoproj.io/v1alpha1
kind: Workflow
metadata: { generateName: bare-, namespace: default }
spec:
  entrypoint: work
  templates:
    - name: work
      container: { image: busybox:1.36, command: [sh, -c], args: ["echo hi"] }
EOF
workflow.argoproj.io/bare-nt5d8 created
$ kubectl create -f - <<'EOF'
apiVersion: argoproj.io/v1alpha1
kind: Workflow
metadata: { generateName: override-, namespace: default }
spec:
  serviceAccountName: default
  workflowTemplateRef: { name: sanctioned }
EOF
workflow.argoproj.io/override-j74f6 created
$ kubectl create -f - <<'EOF'
apiVersion: argoproj.io/v1alpha1
kind: Workflow
metadata: { generateName: fromtmpl-, namespace: default }
spec:
  workflowTemplateRef: { name: sanctioned }
EOF
workflow.argoproj.io/fromtmpl-tz52d created
$ sleep 40
$ kubectl get wf --sort-by=.metadata.creationTimestamp -o jsonpath='{range .items[-3:]}{.metadata.name} {.status.phase} {.status.message}{"\n"}{end}'
bare-nt5d8 Error workflows must use workflowTemplateRef to be executed when the controller is in reference mode
fromtmpl-tz52d Succeeded 
override-j74f6 Error fields [ServiceAccountName] are not permitted when using workflowTemplateRef with templateReferencing restriction
$ kubectl -n argo create cm argo-workflows-workflow-controller-configmap --from-file=config=/tmp/argocfg.orig --dry-run=client -o yaml | kubectl -n argo apply -f -
configmap/argo-workflows-workflow-controller-configmap configured
$ kubectl -n argo rollout restart deploy argo-workflows-workflow-controller
deployment.apps/argo-workflows-workflow-controller restarted
$ kubectl -n argo rollout status deploy argo-workflows-workflow-controller --timeout=180s
Waiting for deployment "argo-workflows-workflow-controller" rollout to finish: 0 out of 1 new replicas have been updated...
Waiting for deployment "argo-workflows-workflow-controller" rollout to finish: 0 out of 1 new replicas have been updated...
Waiting for deployment "argo-workflows-workflow-controller" rollout to finish: 0 out of 1 new replicas have been updated...
Waiting for deployment "argo-workflows-workflow-controller" rollout to finish: 0 out of 1 new replicas have been updated...
Waiting for deployment "argo-workflows-workflow-controller" rollout to finish: 0 of 1 updated replicas are available...
deployment "argo-workflows-workflow-controller" successfully rolled out
$ kubectl delete workflowtemplate sanctioned
workflowtemplate.argoproj.io "sanctioned" deleted from default namespace
$ kubectl get wf -o name | grep -E 'bare-|override-|fromtmpl-' | xargs -r kubectl delete
workflow.argoproj.io "bare-nt5d8" deleted from default namespace
workflow.argoproj.io "fromtmpl-tz52d" deleted from default namespace
workflow.argoproj.io "override-j74f6" deleted from default namespace
verify: all three Workflows are created, then the controller sorts them out: the bare one ends Error with workflows must use workflowTemplateRef to be executed when the controller is in reference mode, the overriding one ends Error with fields [ServiceAccountName] are not permitted when using workflowTemplateRef with templateReferencing restriction, and the plain reference ends Succeeded.

The UI is a client. A portal, a CI job or a chatbot is another, and all of them submit the same JSON with a ServiceAccount token. Do it once by hand so the integration is not a black box.

kubectl create role wf-executor --verb=create,patch --resource=workflowtaskresults.argoproj.io --dry-run=client -o yaml | kubectl apply -f -
kubectl create rolebinding default-wf-executor --role=wf-executor --serviceaccount=default:default --dry-run=client -o yaml | kubectl apply -f -
kubectl -n argo port-forward svc/argo-workflows-server 2746:2746 & PF1=$!
sleep 5
kubectl -n argo create sa submitter --dry-run=client -o yaml | kubectl apply -f -
kubectl create clusterrolebinding submitter --clusterrole=admin --serviceaccount=argo:submitter --dry-run=client -o yaml | kubectl apply -f -
TOKEN=$(kubectl -n argo create token submitter --duration=10m)
cat > /tmp/submit.json <<'EOF'
{"namespace":"default","serverDryRun":false,"workflow":{"metadata":{"generateName":"via-api-","namespace":"default"},"spec":{"entrypoint":"work","templates":[{"name":"work","container":{"image":"busybox:1.36","command":["sh","-c"],"args":["echo submitted over http"]}}]}}}
EOF
curl -s -o /tmp/submit-response.json -w '%{http_code}\n' -H "Authorization: Bearer $TOKEN" -H 'Content-Type: application/json' -d @/tmp/submit.json http://localhost:2746/api/v1/workflows/default
jq '{name: .metadata.name, phase: .status.phase}' /tmp/submit-response.json
kubectl get wf | grep via-api
kill $PF1
kubectl delete clusterrolebinding submitter
kubectl -n argo delete sa submitter
kubectl get wf -o name | grep via-api- | xargs -r kubectl delete
outputcaptured 2026-09-12
$ kubectl create role wf-executor --verb=create,patch --resource=workflowtaskresults.argoproj.io --dry-run=client -o yaml | kubectl apply -f -
role.rbac.authorization.k8s.io/wf-executor unchanged
$ kubectl create rolebinding default-wf-executor --role=wf-executor --serviceaccount=default:default --dry-run=client -o yaml | kubectl apply -f -
rolebinding.rbac.authorization.k8s.io/default-wf-executor unchanged
$ kubectl -n argo port-forward svc/argo-workflows-server 2746:2746 & PF1=$!
$ sleep 5
Forwarding from 127.0.0.1:2746 -> 2746
Forwarding from [::1]:2746 -> 2746
$ kubectl -n argo create sa submitter --dry-run=client -o yaml | kubectl apply -f -
serviceaccount/submitter created
$ kubectl create clusterrolebinding submitter --clusterrole=admin --serviceaccount=argo:submitter --dry-run=client -o yaml | kubectl apply -f -
clusterrolebinding.rbac.authorization.k8s.io/submitter created
$ TOKEN=$(kubectl -n argo create token submitter --duration=10m)
$ cat > /tmp/submit.json <<'EOF'
{"namespace":"default","serverDryRun":false,"workflow":{"metadata":{"generateName":"via-api-","namespace":"default"},"spec":{"entrypoint":"work","templates":[{"name":"work","container":{"image":"busybox:1.36","command":["sh","-c"],"args":["echo submitted over http"]}}]}}}
EOF
$ curl -s -o /tmp/submit-response.json -w '%{http_code}\n' -H "Authorization: Bearer $TOKEN" -H 'Content-Type: application/json' -d @/tmp/submit.json http://localhost:2746/api/v1/workflows/default
Handling connection for 2746
200
$ jq '{name: .metadata.name, phase: .status.phase}' /tmp/submit-response.json
{
  "name": "via-api-8tft9",
  "phase": null
}
$ kubectl get wf | grep via-api
via-api-8tft9                                
$ kill $PF1
$ kubectl delete clusterrolebinding submitter
clusterrolebinding.rbac.authorization.k8s.io "submitter" deleted
$ kubectl -n argo delete sa submitter
serviceaccount "submitter" deleted from argo namespace
$ kubectl get wf -o name | grep via-api- | xargs -r kubectl delete
workflow.argoproj.io "via-api-8tft9" deleted from default namespace
verify: the call returns 200 with the created workflow's name, and the workflow appears in kubectl get wf.

Self-check

answer before opening
Which template type creates Kubernetes objects, and what is its idempotency story?

resource. With action: create a re-run fails on AlreadyExists; with action: apply it converges. Workflows do not reconcile, so if the same request may be submitted twice, you design for it: apply semantics, or a guard step.

A workflow node says forbidden. Where do you look, and what is the classic first offender?

The workflow's ServiceAccount and its bindings, not the workflow spec. The classic first offender in Argo is missing create,patch on workflowtaskresults.argoproj.io, the executor's own reporting channel, which fails before your intended resource permission is even exercised.

When is a workflow the wrong tool, and what replaces it?

When the desired state must stay true indefinitely; a workflow runs once and forgets. Use a controller/operator, or a declarative composition engine (Crossplane, kro) whose CR keeps reconciling. Workflows are for sequences with an end.

Two DAG tasks with no dependencies: what happens, and how do you gate one on a condition?

They run in parallel. Gate with dependencies for ordering and when for conditions (evaluated against parameters or a previous task's result). A validation node that everything depends on is the standard "fail before side effects" pattern.

What turns a Workflow into a self-service product?

Promotion to a WorkflowTemplate with typed, documented, defaulted parameters, a submit surface (UI, CLI, portal, or a Backstage action), input validation before side effects, and an auditable record of who requested what. Same content, packaged as an interface.

A step must run after build whether build succeeded or failed, but only if lint succeeded. How do you express that in a DAG?

Use depends rather than dependencies: depends: "(build.Succeeded || build.Failed) && lint.Succeeded". The result operands (.Succeeded, .Failed, .Errored, .Skipped, .Omitted, .Daemoned) are what make outcome-based ordering possible; dependencies can only say "after it finished successfully".

Two provisioning runs for the same tenant must never execute concurrently, and a nightly job must skip its tick if the previous one is still running. Fields?

A workflow-level synchronization.mutexes entry named per tenant (or a ConfigMap-backed semaphores entry of size 1) serializes submissions in a namespace. For the schedule, CronWorkflow.spec.concurrencyPolicy: Forbid; Replace would kill the running one instead. A waiting Workflow shows Pending with a message naming the lock, not a scheduling error.

You submit a Workflow with serviceAccountName set and the controller rejects it, although the SA exists. Likely cause?

The controller runs with workflowRestrictions.templateReferencing: Strict or Secure: submissions must use workflowTemplateRef and may only override an allow-listed set of fields (arguments, entrypoint, GC settings and a few more). Security-sensitive fields such as serviceAccountName, volumes or podSpecPatch must live in the WorkflowTemplate. Move the field there and resubmit.

What does an exit handler see that a normal template does not, and what is the difference between argo stop and argo terminate?

An onExit template (or hooks.exit) runs after the entrypoint and can read {{workflow.status}} (Succeeded, Failed, Error) and {{workflow.failures}}, so it can branch with when. argo stop stops the running nodes and still runs exit handlers; argo terminate kills everything and skips them, which is why cleanup logic belongs in a hook you know will run.

Docs to know your way around

study time, not exam time
  • argo-workflows.readthedocs.io: core concepts, the fields reference for resource templates, DAG examples, and the WorkflowTemplate page.
  • The Workflows UI's built-in examples (submit → examples) are a legitimate crib sheet during practice.
  • Offline: kubectl explain workflow.spec.templates --recursive and argo submit --help.
  • argo-workflows.readthedocs.io/en/latest/retries and /synchronization: the two pages whose field names appear verbatim in tasks; /cron-workflows has the option table with defaults.
  • argo-workflows.readthedocs.io/en/latest/workflow-rbac: the minimal executor Role, plus /argo-server-auth-mode and /access-token for submitting through the API with a ServiceAccount token.