Domain 3 hands you four ways to turn a request into resources: CRD+operator, Crossplane, workflows, and portal templates. This section adds the newest option, walks the lab's end-to-end golden path, and then forces the comparison, because "evaluate when to use operators, workflows, or pipelines" is a named outcome and a cheap scenario question.

needsmake core apimake portal

Orientation

competency 3.4 + the judgment call

Everything in domain 3 has been one engine at a time. This section is about how they join up, which is where platform engineering actually happens: a Backstage form that commits to git, that an ApplicationSet notices, that deploys an XR, that a controller reconciles. Being able to narrate that chain is worth more than any single command in it.

kro in one sitting

Crossplane's niche with less machinery

kro (Kube Resource Orchestrator) occupies Crossplane's niche with far less apparatus: one ResourceGraphDefinition declares a schema and the resources it expands to; kro generates the CRD and runs the reconciliation. No providers, no functions, no packages. CEL expressions wire fields together, and references between resources (${ns.metadata.name}) build the dependency graph automatically; kro works out the ordering from the references you wrote.

apiVersion: kro.run/v1alpha1
kind: ResourceGraphDefinition
metadata: { name: tenantspace }
spec:
  schema:
    apiVersion: v1alpha1
    kind: TenantSpace
    spec:
      team: string
      cpu: string | default="1"
  resources:
    - id: ns
      template:
        apiVersion: v1
        kind: Namespace
        metadata:
          name: ${schema.spec.team}
          labels:
            tenant: ${schema.spec.team}
    - id: quota
      template:
        apiVersion: v1
        kind: ResourceQuota
        metadata:
          name: tenant-quota
          namespace: ${ns.metadata.name}
        spec:
          hard:
            requests.cpu: ${schema.spec.cpu}
YAML flow-style trap

Keep kro's ${...} substitutions in block style. Inside flow-style braces (labels: { tenant: ${schema.spec.team} }) the expression's own } closes the map early and the whole manifest fails to parse with a message that points nowhere useful.

kro is young and its API moves. If this manifest disagrees with your installed version, kubectl explain resourcegraphdefinition.spec is the arbiter, and practicing that recovery is worth more than the manifest.

AspectCrossplanekro
Files to writeXRD + Composition (+ providers, functions)one ResourceGraphDefinition
Wiringpatches / function pipelineCEL references, ordering inferred
External cloudsrich provider ecosystemwhatever CRDs already exist in-cluster
Connection secretsfirst-classnot really
Maturityproduction, widely deployedyoung, API in motion

Both sentences are true: "same result, one file, no providers" and "no provider ecosystem, no connection secrets, alpha-grade stability". An exam answer may need either.

The RGD fields beyond the first manifest

  • Schema markers. Simple Schema types (string, integer, boolean, objects, arrays) take validation markers after a pipe: required=true, default=..., enum="a,b", pattern="regex", minimum=1 maximum=100. They become OpenAPI validation on the generated CRD, so a bad instance is rejected by the API server, not by kro.
  • Status. schema.status fields are CEL expressions over resources (url: ${service.status.loadBalancer.ingress[0].hostname}); kro infers their types and writes them onto the instance. That is your ToCompositeFieldPath equivalent.
  • readyWhen. A list of CEL conditions on the resource itself; dependents wait until all are true. Without it a resource counts as ready once it exists and its expressions resolve, which is why a Deployment that never becomes Available can still leave the instance looking fine unless you write readyWhen: [${deployment.status.availableReplicas > 0}].
  • includeWhen. Conditions that decide whether a resource is created at all (${schema.spec.ingress.enabled}), evaluated every reconcile; flipping the field later removes the resource.
  • forEach turns one resource entry into a collection (one object per list element); externalRef reads an existing object without managing it.
  • Ordering. The dependency graph is inferred from references; kubectl get rgd <name> -o jsonpath='{.status.topologicalOrder}' prints the computed order, creation follows it and deletion runs it backwards. A cycle (A reads B, B reads A) is rejected at RGD creation; break it by reading schema.spec instead.
  • Scope. schema.scope: Namespaced (default) or Cluster. Namespaced instances can still create cluster-scoped objects (that is what the lab's Namespace example does), because kro tracks ownership with labels following the ApplySet convention rather than ownerReferences, which cannot cross scopes.

Conditions to read

An RGD carries GraphRevisionsResolved, GraphAccepted, KindReady (the CRD exists), ControllerReady (the dynamic controller registered) and an aggregate Ready; status.state is Active only when the graph, kind and controller are all good, otherwise Inactive with the failing condition's message (typically a CEL type error or an unknown field, caught by kro's static analysis before anything is created). An instance carries InstanceManaged, GraphResolved, ResourcesReady and Ready, and status.state ACTIVE; RGD authors can replace the wire conditions with their own via schema.status.conditions. To freeze an instance for debugging, annotate it kro.run/reconcile: suspended.

Trap

kro's controller needs RBAC for every kind an RGD creates. The Helm chart's rbac.mode: unrestricted grants everything; aggregation mode grants only what ClusterRoles labeled rbac.kro.run/aggregate-to-controller: "true" allow. In aggregation mode an RGD that templates a kind nobody granted goes Active and then every instance fails with forbidden from the kro ServiceAccount. Same lesson as the Crossplane and Argo Workflows RBAC ones, one namespace over.

The golden path, hop by hop

Backstage → Gitea → ApplicationSet → workload

Backstage is the portal half of the white paper's "web portals" capability: a software catalog (components, systems, APIs, owners, described by catalog-info.yaml in each repo), software templates (the scaffolder: parameters form + actions like fetch:template, publish:gitea, catalog:register), TechDocs, and plugins that surface Kubernetes, CI and other systems in one place.

Backstage template  (parameters form)
      │ scaffolder actions: fetch skeleton → publish → register
      ▼
new repo in Gitea org services   (contains k8s/ manifests + catalog-info.yaml)
      │ noticed by the SCM-provider generator
      ▼
ApplicationSet golden-path  ──generates──▶ Argo CD Application
      │
      ▼
running workload  ── and the component appears in the catalog with its owner

Each hop is a thing that can break and a thing you may be asked about: template action failures (credentials to the git host), the generator's filter (which repos match), the generated Application's destination, and the workload itself. Narrating this chain is the "golden path" competency; the commands below just prove each hop.

Why a portal is not a platform

A common scenario trap: the portal is an interface over capabilities, not a capability itself. If the underlying provisioning is a ticket queue, adding Backstage produces a nicer ticket queue. Say that when a question describes a portal rollout with no automation behind it.

Backstage as an exam object

catalog-info.yaml, templates, the two plugins that matter

Backstage is on the exam as an interface, not as a Node.js project: read a catalog-info.yaml, read or fix a software template, and explain what the portal does and does not do. The vocabulary is small and exact.

Catalog kinds and their required fields

kindModelsRequired in spec
Componenta service, website or librarytype, lifecycle (experimental/production/deprecated by convention), owner
APIan interface a component provides or consumestype (openapi, asyncapi, graphql, grpc), lifecycle, owner, definition
Resourceinfrastructure a component depends on (database, bucket, cluster)type, owner
Systema group of components and resources delivering a capabilityowner
Domaina business area grouping systemsowner
Group / Userthe organization; owner fields point hereGroup: type, children; User: memberOf
Locationa pointer to other descriptor files (targets)targets or target
Templatea scaffolder template (apiVersion scaffolder.backstage.io/v1beta3)type, parameters, steps
apiVersion: backstage.io/v1alpha1
kind: Component
metadata:
  name: payments-api
  annotations:
    backstage.io/techdocs-ref: dir:.            # TechDocs builds docs/ from this repo
    backstage.io/kubernetes-id: payments-api     # Kubernetes plugin matches workloads by label
    argocd/app-name: payments-api                # Argo CD plugin lookup
spec:
  type: service
  lifecycle: production
  owner: group:default/team-payments
  system: payments
  providesApis: [payments-v1]

Entity references are kind:namespace/name (group:default/team-payments); the default namespace is assumed when omitted. Relations (ownedBy, partOf, dependsOn, providesApi) are derived from these spec fields; you never write them by hand.

Software templates, the parts that break

  • parameters is JSON Schema rendered as form steps. ui:field picks special widgets: OwnerPicker, EntityPicker, RepoUrlPicker (with ui:options.allowedHosts, which must list your git host or the form refuses it), Secret (read back as ${{ secrets.x }}, never parameters.x).
  • steps run in order in the scaffolder backend; each has id, action, input, optional if. Built-in actions: fetch:template (copy a skeleton, render ${{ values.x }}), fetch:plain, publish:github / publish:gitlab / publish:gitea / publish:bitbucket (create the repo and push), catalog:register (register the new catalog-info.yaml by repoContentsUrl + catalogInfoPath), catalog:write, debug:log, fs:rename, fs:delete. The installed list is at /create/actions in the UI.
  • Templating syntax. ${{ parameters.name }} for inputs, ${{ steps['publish'].output.repoContentsUrl }} for prior outputs, output.links for what the user sees at the end. Backstage's ${{ }} is Nunjucks and is not the same as kro's ${} or Helm's {{ }}; a template that fails to render is very often one of those confused.
  • Failure points. Publish actions need an integration token in app-config.yaml (integrations.github/gitlab/gitea) with rights to create repos; catalog:register fails if the repo is created but the URL is not readable by the catalog; a template that succeeds but produces no Application is the ApplicationSet generator's filter, not Backstage.

Two plugins the exam vocabulary assumes

TechDocs builds MkDocs sites from a repo's docs/ and mkdocs.yml, located by the backstage.io/techdocs-ref annotation (dir:. means "this repo"). The Kubernetes plugin shows pods, deployments and errors for a component by matching the backstage.io/kubernetes-id label on workloads (or a backstage.io/kubernetes-label-selector annotation), reading clusters listed under kubernetes.clusterLocatorMethods with a ServiceAccount token. Both are read-only views over things the platform already produces, which is the point of the next callout.

Mental model

The maturity model's Interfaces aspect runs Custom processes → Standard tooling → Self-service solutions → Integrated services. A portal that opens tickets is level 1 with a nicer font; a portal whose template commits to git and lets GitOps and controllers do the rest is level 3. The exam's "portal is an interface, not the platform" scenario is asking you to place a proposal on that ladder.

Choosing the engine

the differentiators to reason from
EngineTriggerLifetimeBest at
CRD + operatoran API objectforeverdomain behavior: failover, backup, upgrade
Crossplanean API objectforevercomposing many resources, incl. cloud APIs, behind one abstraction
kroan API objectforeverthe same, in-cluster only, with one file
Argo Workflowssubmit, schedule, or eventrun to completionimperative sequences, batch, scheduled maintenance
Tektongit eventrun to completionbuild, test, publish artifacts
Backstage templatea human filling a formone-shot scaffoldingthe front door: create repo + register + hand off to GitOps

Three questions decide almost every case:

  1. Does the state need reconciling forever, or producing once? Forever → a controller (operator, Crossplane, kro). Once → workflow or pipeline.
  2. What is the trigger? An API object, a schedule, a git event, or a human.
  3. Who is the audience? A developer writing YAML, a developer clicking a form, or another system.

And when two engines both work, say so and pick the thinnest one. That answer style scores, because it is the TVP instinct from section 3.1 applied to your own tooling.

Engines the table leaves out

EngineWhat it isPick it whenNot when
Helm charttemplating at install time, no controllerpackaging one application's manifests with valuesanything must keep converging or compose external APIs
tofu-controller (Flux)a Terraform CR reconciled by Flux; approvePlan gates applies, drift detection is a modeexisting Terraform/OpenTofu modules and state that must become GitOps-drivenyou would rather model the API in Kubernetes types (then Crossplane or kro)
KubeVela (OAM)an Application of components, traits and policies with a delivery workflowan application-centric abstraction across clusters with CUE-defined workflow stepsinfrastructure-first APIs; Crossplane is the better fit
Cloud-vendor controllers (ACK, Config Connector, ASO)one CRD per cloud resource typeone cloud, native resource fidelity, wrapped by kro or Crossplane for the abstractionyou need a single abstraction over several clouds

Two more distinctions the drill should surface. Crossplane vs a cloud controller: the controller gives you the raw resource kinds; Crossplane (or kro) gives you the abstraction over them, so they compose rather than compete. Argo Workflows vs Tekton: both are on the official tool list and a task that names one wants that one; the general rule is CI pipelines on git events go to Tekton, parameterized orchestration on request or schedule goes to Argo Workflows.

Exercises

tick the dot when its check passes

Apply the RGD above, then:

kubectl get rgd tenantspace -o jsonpath='{.status.state}{"\n"}'   # wants: Active
kubectl apply -f - <<'EOF'
apiVersion: kro.run/v1alpha1
kind: TenantSpace
metadata: { name: team-h, namespace: default }
spec: { team: team-h, cpu: "2" }
EOF
kubectl get ns team-h && kubectl -n team-h get resourcequota tenant-quota
outputcaptured 2026-08-26
$ kubectl get rgd tenantspace -o jsonpath='{.status.state}{"\n"}'   # wants: Active
Active
$ kubectl apply -f - <<'EOF'
apiVersion: kro.run/v1alpha1
kind: TenantSpace
metadata: { name: team-h, namespace: default }
spec: { team: team-h, cpu: "2" }
EOF
tenantspace.kro.run/team-h created
$ kubectl get ns team-h && kubectl -n team-h get resourcequota tenant-quota
NAME     STATUS   AGE
team-h   Active   25s
NAME           REQUEST             LIMIT   AGE
tenant-quota   requests.cpu: 0/2           22s
verify: the RGD is Active (meaning kro generated and established the TenantSpace CRD), and the namespace and quota exist with your values. Then compare effort honestly against 3.5: same result, one file, no providers; also no provider ecosystem, no connection secrets, alpha-grade stability.

With make portal done and Backstage running (cd portal && yarn start), create a component from the "Golden path service" template. Then trace every hop with commands: new repo in the Gitea services org (http://gitea.lab:3000/services), the SCM-generator ApplicationSet notices it (kubectl -n argocd get applicationset golden-path -o yaml, read the generator block), a generated Application appears (kubectl -n argocd get applications), workload deploys.

verify: make validate flips the portal checks to green, including "Applications auto-generated from git".

For each request, pick the engine and one sentence of why: (a) every team needs a Postgres with backups that self-heals for years; (b) on request, stamp a namespace with quotas and labels from three parameters; (c) nightly, rebuild and rescan all base images, notify on failure; (d) on every push, build, test, and deploy a service; (e) offer non-technical users a form that creates a new service from a skeleton.

verify: defensible answers: (a) operator, long-lived lifecycle with domain logic; (b) Crossplane or kro, declarative expansion, no imperative steps; (c) CronWorkflow, scheduled run-to-completion; (d) Tekton, event-driven CI; (e) Backstage template feeding GitOps. If two engines both work, say so and pick the thinnest one.

You never write the ordering; kro infers it from the references between resources and publishes the result. Read it, then feed it a cycle and watch the rejection, because a cycle is the one authoring mistake the engine cannot work around.

# the RGD is applied by hand earlier on this page; make the block stand on its own
kubectl apply -f - <<'EOF'
apiVersion: kro.run/v1alpha1
kind: ResourceGraphDefinition
metadata: { name: tenantspace }
spec:
  schema:
    apiVersion: v1alpha1
    kind: TenantSpace
    spec:
      team: string
      cpu: string | default="1"
  resources:
    - id: ns
      template:
        apiVersion: v1
        kind: Namespace
        metadata:
          name: ${schema.spec.team}
          labels:
            tenant: ${schema.spec.team}
    - id: quota
      template:
        apiVersion: v1
        kind: ResourceQuota
        metadata:
          name: tenant-quota
          namespace: ${ns.metadata.name}
        spec:
          hard:
            requests.cpu: ${schema.spec.cpu}
EOF
sleep 20
kubectl get rgd tenantspace -o jsonpath='{.status.conditions}' | jq '.[] | {type, status, reason, message}'
kubectl get rgd tenantspace -o jsonpath='{.status.topologicalOrder}{"\n"}'
kubectl apply -f - <<'EOF'
apiVersion: kro.run/v1alpha1
kind: ResourceGraphDefinition
metadata: { name: cyclic }
spec:
  schema:
    apiVersion: v1alpha1
    kind: Cyclic
    spec:
      team: string
  resources:
    - id: first
      template:
        apiVersion: v1
        kind: ConfigMap
        metadata:
          name: first
          namespace: default
        data:
          other: ${second.metadata.name}
    - id: second
      template:
        apiVersion: v1
        kind: ConfigMap
        metadata:
          name: second
          namespace: default
        data:
          other: ${first.metadata.name}
EOF
kubectl get rgd cyclic -o jsonpath='{.status.state}{"\n"}'; kubectl get rgd cyclic -o jsonpath='{.status.conditions}' 2>/dev/null | jq '.[] | {type, status, message}'
kubectl delete rgd cyclic --ignore-not-found
outputcaptured 2026-09-13
$ # the RGD is applied by hand earlier on this page; make the block stand on its own
$ kubectl apply -f - <<'EOF'
apiVersion: kro.run/v1alpha1
kind: ResourceGraphDefinition
metadata: { name: tenantspace }
spec:
  schema:
    apiVersion: v1alpha1
    kind: TenantSpace
    spec:
      team: string
      cpu: string | default="1"
  resources:
    - id: ns
      template:
        apiVersion: v1
        kind: Namespace
        metadata:
          name: ${schema.spec.team}
          labels:
            tenant: ${schema.spec.team}
    - id: quota
      template:
        apiVersion: v1
        kind: ResourceQuota
        metadata:
          name: tenant-quota
          namespace: ${ns.metadata.name}
        spec:
          hard:
            requests.cpu: ${schema.spec.cpu}
EOF
resourcegraphdefinition.kro.run/tenantspace unchanged
$ sleep 20
$ kubectl get rgd tenantspace -o jsonpath='{.status.conditions}' | jq '.[] | {type, status, reason, message}'
{
  "type": "ResourceGraphAccepted",
  "status": "True",
  "reason": "Valid",
  "message": "resource graph and schema are valid"
}
{
  "type": "KindReady",
  "status": "True",
  "reason": "Ready",
  "message": "kind TenantSpace has been accepted and ready"
}
{
  "type": "ControllerReady",
  "status": "True",
  "reason": "Running",
  "message": "controller is running"
}
{
  "type": "Ready",
  "status": "True",
  "reason": "Ready",
  "message": ""
}
$ kubectl get rgd tenantspace -o jsonpath='{.status.topologicalOrder}{"\n"}'
["ns","quota"]
$ kubectl apply -f - <<'EOF'
apiVersion: kro.run/v1alpha1
kind: ResourceGraphDefinition
metadata: { name: cyclic }
spec:
  schema:
    apiVersion: v1alpha1
    kind: Cyclic
    spec:
      team: string
  resources:
    - id: first
      template:
        apiVersion: v1
        kind: ConfigMap
        metadata:
          name: first
          namespace: default
        data:
          other: ${second.metadata.name}
    - id: second
      template:
        apiVersion: v1
        kind: ConfigMap
        metadata:
          name: second
          namespace: default
        data:
          other: ${first.metadata.name}
EOF
resourcegraphdefinition.kro.run/cyclic created
$ kubectl get rgd cyclic -o jsonpath='{.status.state}{"\n"}'; kubectl get rgd cyclic -o jsonpath='{.status.conditions}' 2>/dev/null | jq '.[] | {type, status, message}'
Inactive
{
  "type": "KindReady",
  "status": "Unknown",
  "message": "condition \"KindReady\" is awaiting reconciliation"
}
{
  "type": "ControllerReady",
  "status": "Unknown",
  "message": "condition \"ControllerReady\" is awaiting reconciliation"
}
{
  "type": "ResourceGraphAccepted",
  "status": "False",
  "message": "failed to build dependency graph: graph contains a cycle: first -> second -> first"
}
{
  "type": "Ready",
  "status": "False",
  "message": "failed to build dependency graph: graph contains a cycle: first -> second -> first"
}
$ kubectl delete rgd cyclic --ignore-not-found
resourcegraphdefinition.kro.run "cyclic" deleted
verify: tenantspace is Active with a topologicalOrder of ["ns","quota"], the namespace before the quota that lives in it. The cyclic one is not refused at admission: kro accepts the object and then leaves it out of Active with a condition naming the cycle, so read status.state and the conditions after a reconcile rather than expecting the apply to fail.

readyWhen makes a dependent wait for something real rather than for the object to exist, and includeWhen leaves a resource out of the graph entirely. Together they are how one API serves a tier that gets a database and a tier that does not.

kubectl apply -f - <<'EOF'
apiVersion: kro.run/v1alpha1
kind: ResourceGraphDefinition
metadata: { name: tieredspace }
spec:
  schema:
    apiVersion: v1alpha1
    kind: TieredSpace
    spec:
      team: string
      cpu: string | default="1"
  resources:
    - id: ns
      template:
        apiVersion: v1
        kind: Namespace
        metadata:
          name: ${schema.spec.team}
    - id: quota
      readyWhen:
        - ${quota.status.hard != null}
      template:
        apiVersion: v1
        kind: ResourceQuota
        metadata:
          name: tenant-quota
          namespace: ${ns.metadata.name}
        spec:
          hard:
            requests.cpu: ${schema.spec.cpu}
    - id: extra
      includeWhen:
        - ${schema.spec.cpu != "0"}
      template:
        apiVersion: v1
        kind: ConfigMap
        metadata:
          name: tier-note
          namespace: ${ns.metadata.name}
        data:
          note: included because cpu was not zero
EOF
sleep 20
kubectl get rgd tieredspace -o jsonpath='{.status.state} {.status.topologicalOrder}{"\n"}'
kubectl apply -f - <<'EOF'
apiVersion: kro.run/v1alpha1
kind: TieredSpace
metadata: { name: team-j, namespace: default }
spec: { team: team-j, cpu: "0" }
EOF
sleep 30
kubectl -n team-j get configmap tier-note || echo 'omitted by includeWhen'
kubectl get tieredspace team-j -o jsonpath='{.status.conditions}' | jq '.[] | {type, status, reason}'
kubectl delete tieredspace team-j
kubectl delete rgd tieredspace
kubectl delete ns team-j --ignore-not-found
outputcaptured 2026-09-12
$ kubectl apply -f - <<'EOF'
apiVersion: kro.run/v1alpha1
kind: ResourceGraphDefinition
metadata: { name: tieredspace }
spec:
  schema:
    apiVersion: v1alpha1
    kind: TieredSpace
    spec:
      team: string
      cpu: string | default="1"
  resources:
    - id: ns
      template:
        apiVersion: v1
        kind: Namespace
        metadata:
          name: ${schema.spec.team}
    - id: quota
      readyWhen:
        - ${quota.status.hard != null}
      template:
        apiVersion: v1
        kind: ResourceQuota
        metadata:
          name: tenant-quota
          namespace: ${ns.metadata.name}
        spec:
          hard:
            requests.cpu: ${schema.spec.cpu}
    - id: extra
      includeWhen:
        - ${schema.spec.cpu != "0"}
      template:
        apiVersion: v1
        kind: ConfigMap
        metadata:
          name: tier-note
          namespace: ${ns.metadata.name}
        data:
          note: included because cpu was not zero
EOF
resourcegraphdefinition.kro.run/tieredspace created
$ sleep 20
$ kubectl get rgd tieredspace -o jsonpath='{.status.state} {.status.topologicalOrder}{"\n"}'
Active ["ns","quota","extra"]
$ kubectl apply -f - <<'EOF'
apiVersion: kro.run/v1alpha1
kind: TieredSpace
metadata: { name: team-j, namespace: default }
spec: { team: team-j, cpu: "0" }
EOF
tieredspace.kro.run/team-j created
$ sleep 30
$ kubectl -n team-j get configmap tier-note || echo 'omitted by includeWhen'
Error from server (NotFound): configmaps "tier-note" not found
omitted by includeWhen
$ kubectl get tieredspace team-j -o jsonpath='{.status.conditions}' | jq '.[] | {type, status, reason}'
{
  "type": "InstanceSynced",
  "status": "True",
  "reason": "ReconciliationSucceeded"
}
$ kubectl delete tieredspace team-j
tieredspace.kro.run "team-j" deleted from default namespace
$ kubectl delete rgd tieredspace
resourcegraphdefinition.kro.run "tieredspace" deleted
$ kubectl delete ns team-j --ignore-not-found
verify: the ConfigMap is absent for the instance whose cpu is zero, and the instance still reports ready because the quota satisfied its own readyWhen.

The suspend annotation stops kro reconciling one instance. Delete something it owns while it is suspended and nothing comes back, which is both the escape hatch and the demonstration that the controller is what keeps the graph true.

kubectl get rgd tenantspace -o jsonpath='{.status.state}{"\n"}'
kubectl apply -f - <<'EOF'
apiVersion: kro.run/v1alpha1
kind: TenantSpace
metadata: { name: team-h, namespace: default }
spec: { team: team-h, cpu: "2" }
EOF
# kro reconciles an instance when the instance changes, so bump the spec to be sure of a fresh one
kubectl patch tenantspace team-h --type merge -p '{"spec":{"cpu":"2"}}'
sleep 45
kubectl -n team-h get resourcequota tenant-quota
kubectl annotate tenantspace team-h kro.run/reconcile=suspended --overwrite
kubectl -n team-h delete resourcequota tenant-quota
sleep 30
kubectl -n team-h get resourcequota || echo 'still gone: the instance is suspended'
kubectl annotate tenantspace team-h kro.run/reconcile-
sleep 45
kubectl -n team-h get resourcequota || echo 'still gone: unsuspending alone does not trigger a reconcile'
# the trigger is a change to the instance, not drift in what it composed
kubectl patch tenantspace team-h --type merge -p '{"spec":{"cpu":"3"}}'
sleep 45
kubectl -n team-h get resourcequota tenant-quota
kubectl get tenantspace team-h -o jsonpath='{.status.conditions}' | jq '.[] | {type, status, reason}'
outputcaptured 2026-09-12
$ kubectl get rgd tenantspace -o jsonpath='{.status.state}{"\n"}'
Active
$ kubectl apply -f - <<'EOF'
apiVersion: kro.run/v1alpha1
kind: TenantSpace
metadata: { name: team-h, namespace: default }
spec: { team: team-h, cpu: "2" }
EOF
tenantspace.kro.run/team-h configured
$ # kro reconciles an instance when the instance changes, so bump the spec to be sure of a fresh one
$ kubectl patch tenantspace team-h --type merge -p '{"spec":{"cpu":"2"}}'
tenantspace.kro.run/team-h patched (no change)
$ sleep 45
$ kubectl -n team-h get resourcequota tenant-quota
NAME           REQUEST             LIMIT   AGE
tenant-quota   requests.cpu: 0/2           45s
$ kubectl annotate tenantspace team-h kro.run/reconcile=suspended --overwrite
tenantspace.kro.run/team-h annotated
$ kubectl -n team-h delete resourcequota tenant-quota
resourcequota "tenant-quota" deleted from team-h namespace
$ sleep 30
$ kubectl -n team-h get resourcequota || echo 'still gone: the instance is suspended'
No resources found in team-h namespace.
$ kubectl annotate tenantspace team-h kro.run/reconcile-
tenantspace.kro.run/team-h annotated
$ sleep 45
$ kubectl -n team-h get resourcequota || echo 'still gone: unsuspending alone does not trigger a reconcile'
No resources found in team-h namespace.
$ # the trigger is a change to the instance, not drift in what it composed
$ kubectl patch tenantspace team-h --type merge -p '{"spec":{"cpu":"3"}}'
tenantspace.kro.run/team-h patched
$ sleep 45
$ kubectl -n team-h get resourcequota tenant-quota
NAME           REQUEST             LIMIT   AGE
tenant-quota   requests.cpu: 0/3           45s
$ kubectl get tenantspace team-h -o jsonpath='{.status.conditions}' | jq '.[] | {type, status, reason}'
{
  "type": "InstanceSynced",
  "status": "True",
  "reason": "ReconciliationSucceeded"
}
verify: the quota exists, then stays deleted while kro.run/reconcile=suspended is on the instance. Removing the annotation does not bring it back: kro reconciles an instance when the instance changes, not when something it composed drifts, so the quota only returns after the cpu patch. Note that InstanceSynced reads True / ReconciliationSucceeded the whole way through, including while the quota is missing, which is exactly the status you must not trust on its own.

kro's controller can only create what its ServiceAccount may create. A resource type outside that set fails at the instance, not at the RGD, which is a confusing place to find a permissions problem. This block does not narrow anything: it asks for a kind outside the usual set and then reads the two places the answer could be.

This block makes the failure deliberately: it narrows the controller's aggregated permissions, applies an RGD that needs the missing verb, then puts the rule back.

kubectl get clusterrole -o name | grep -i kro
kubectl apply -f - <<'EOF'
apiVersion: kro.run/v1alpha1
kind: ResourceGraphDefinition
metadata: { name: needsmore }
spec:
  schema:
    apiVersion: v1alpha1
    kind: NeedsMore
    spec:
      team: string
  resources:
    - id: pdb
      template:
        apiVersion: policy/v1
        kind: PodDisruptionBudget
        metadata:
          name: tenant-pdb
          namespace: default
        spec:
          minAvailable: 1
          selector:
            matchLabels:
              team: ${schema.spec.team}
EOF
sleep 20
kubectl apply -f - <<'EOF'
apiVersion: kro.run/v1alpha1
kind: NeedsMore
metadata: { name: team-k, namespace: default }
spec: { team: team-k }
EOF
sleep 45
kubectl get needsmore team-k -o jsonpath='{.status.conditions}' | jq '.[] | {type, status, reason, message}'
kubectl get pdb tenant-pdb -n default || echo 'not created'
kubectl delete needsmore team-k
kubectl delete rgd needsmore
kubectl delete pdb tenant-pdb -n default --ignore-not-found
outputcaptured 2026-09-13
$ kubectl get clusterrole -o name | grep -i kro
clusterrole.rbac.authorization.k8s.io/kro-cluster-role
$ kubectl apply -f - <<'EOF'
apiVersion: kro.run/v1alpha1
kind: ResourceGraphDefinition
metadata: { name: needsmore }
spec:
  schema:
    apiVersion: v1alpha1
    kind: NeedsMore
    spec:
      team: string
  resources:
    - id: pdb
      template:
        apiVersion: policy/v1
        kind: PodDisruptionBudget
        metadata:
          name: tenant-pdb
          namespace: default
        spec:
          minAvailable: 1
          selector:
            matchLabels:
              team: ${schema.spec.team}
EOF
resourcegraphdefinition.kro.run/needsmore created
$ sleep 20
$ kubectl apply -f - <<'EOF'
apiVersion: kro.run/v1alpha1
kind: NeedsMore
metadata: { name: team-k, namespace: default }
spec: { team: team-k }
EOF
needsmore.kro.run/team-k created
$ sleep 45
$ kubectl get needsmore team-k -o jsonpath='{.status.conditions}' | jq '.[] | {type, status, reason, message}'
{
  "type": "InstanceSynced",
  "status": "True",
  "reason": "ReconciliationSucceeded",
  "message": "Instance reconciled successfully"
}
$ kubectl get pdb tenant-pdb -n default || echo 'not created'
NAME         MIN AVAILABLE   MAX UNAVAILABLE   ALLOWED DISRUPTIONS   AGE
tenant-pdb   1               N/A               0                     46s
$ kubectl delete needsmore team-k
needsmore.kro.run "team-k" deleted from default namespace
$ kubectl delete rgd needsmore
resourcegraphdefinition.kro.run "needsmore" deleted
$ kubectl delete pdb tenant-pdb -n default --ignore-not-found
verify: either the PodDisruptionBudget appears, in which case the controller already had the permission and you can say which ClusterRole granted it, or the instance condition carries a forbidden message naming the verb. Both are the answer; the point is knowing which object to read.

A golden path fails in two places: the form, before anything is created, and the publish step, after the user has already committed to it. This block does not break anything; it reads the template through the Backstage API so you can see both places before you go near them. The form is the allowedHosts list on the RepoUrlPicker, and the publish step is publish:gitea.

Backstage runs on the host from make portal, which is not part of make full; start it with cd portal && yarn start. Its API wants a token even for reads, so the first line trades the guest provider for one. These commands read the API rather than clicking through the UI.

TOKEN=$(curl -s http://localhost:7007/api/auth/guest/refresh | jq -r .backstageIdentity.token)
curl -s -H "Authorization: Bearer $TOKEN" 'http://localhost:7007/api/catalog/entities?filter=kind=template' | jq '.[] | {name: .metadata.name, steps: [.spec.steps[].action]}'
curl -s -H "Authorization: Bearer $TOKEN" 'http://localhost:7007/api/catalog/entities?filter=kind=template' | jq '.[] | .spec.parameters[]? | .properties.repoUrl?."ui:options"?.allowedHosts'
curl -s -H "Authorization: Bearer $TOKEN" http://localhost:7007/api/scaffolder/v2/tasks | jq '{tasks: (.tasks | length), first: (.tasks[0] // null | if . == null then null else {id, status} end)}'
outputcaptured 2026-09-12
$ TOKEN=$(curl -s http://localhost:7007/api/auth/guest/refresh | jq -r .backstageIdentity.token)
$ curl -s -H "Authorization: Bearer $TOKEN" 'http://localhost:7007/api/catalog/entities?filter=kind=template' | jq '.[] | {name: .metadata.name, steps: [.spec.steps[].action]}'
{
  "name": "golden-path-service",
  "steps": [
    "fetch:template",
    "publish:gitea",
    "catalog:register"
  ]
}
{
  "name": "example-nodejs-template",
  "steps": [
    "fetch:template",
    "publish:github",
    "catalog:register",
    "notification:send"
  ]
}
$ curl -s -H "Authorization: Bearer $TOKEN" 'http://localhost:7007/api/catalog/entities?filter=kind=template' | jq '.[] | .spec.parameters[]? | .properties.repoUrl?."ui:options"?.allowedHosts'
null
null
[
  "gitea.lab:3000"
]
null
[
  "github.com"
]
$ curl -s -H "Authorization: Bearer $TOKEN" http://localhost:7007/api/scaffolder/v2/tasks | jq '{tasks: (.tasks | length), first: (.tasks[0] // null | if . == null then null else {id, status} end)}'
{
  "tasks": 0,
  "first": null
}
verify: golden-path-service lists steps fetch:template, publish:gitea, catalog:register, and its RepoUrlPicker allows exactly one host, gitea.lab:3000. The task list is empty because nobody has run the template yet, which is what an untouched scaffolder looks like. To see either failure for yourself, drop that host from allowedHosts and the form refuses before a task is ever created; break the Gitea token instead and a task appears and then dies at publish:gitea.

The catalog is an API, not a web page, and it needs make portal running on the host. Note the last line: an unauthenticated read is refused, which is worth knowing before you wire anything to it. Query it directly and the entities, their owners and the relations between them come back as data you can check in a script, which is how a portal stays true as the platform grows.

TOKEN=$(curl -s http://localhost:7007/api/auth/guest/refresh | jq -r .backstageIdentity.token)
curl -s -H "Authorization: Bearer $TOKEN" 'http://localhost:7007/api/catalog/entities?filter=kind=component' | jq '[.[] | {name: .metadata.name, owner: .spec.owner, type: .spec.type}]'
curl -s -H "Authorization: Bearer $TOKEN" 'http://localhost:7007/api/catalog/entities?filter=kind=component' | jq '[.[] | {name: .metadata.name, relations: [.relations[]? | {type, targetRef}]}]'
curl -s -H "Authorization: Bearer $TOKEN" 'http://localhost:7007/api/catalog/entities?filter=kind=group' | jq '[.[] | .metadata.name]'
curl -s -o /dev/null -w 'without a token: %{http_code}\n' 'http://localhost:7007/api/catalog/entities?filter=kind=component'
outputcaptured 2026-09-12
$ TOKEN=$(curl -s http://localhost:7007/api/auth/guest/refresh | jq -r .backstageIdentity.token)
$ curl -s -H "Authorization: Bearer $TOKEN" 'http://localhost:7007/api/catalog/entities?filter=kind=component' | jq '[.[] | {name: .metadata.name, owner: .spec.owner, type: .spec.type}]'
[
  {
    "name": "example-website",
    "owner": "guests",
    "type": "website"
  }
]
$ curl -s -H "Authorization: Bearer $TOKEN" 'http://localhost:7007/api/catalog/entities?filter=kind=component' | jq '[.[] | {name: .metadata.name, relations: [.relations[]? | {type, targetRef}]}]'
[
  {
    "name": "example-website",
    "relations": [
      {
        "type": "ownedBy",
        "targetRef": "group:default/guests"
      },
      {
        "type": "partOf",
        "targetRef": "system:default/examples"
      },
      {
        "type": "providesApi",
        "targetRef": "api:default/example-grpc-api"
      }
    ]
  }
]
$ curl -s -H "Authorization: Bearer $TOKEN" 'http://localhost:7007/api/catalog/entities?filter=kind=group' | jq '[.[] | .metadata.name]'
[
  "guests"
]
$ curl -s -o /dev/null -w 'without a token: %{http_code}\n' 'http://localhost:7007/api/catalog/entities?filter=kind=component'
without a token: 401
verify: every component names an owner that exists as a group, and the relations show ownedBy and partOf derived from the descriptor files rather than typed in twice. A component whose owner does not resolve is the catalog's version of a dangling pointer; say how you would catch that in CI.

Self-check

answer before opening
How does kro know to create the namespace before the quota?

From the reference: the quota template uses ${ns.metadata.name}, so kro builds a dependency graph from the CEL references and orders accordingly. You never write dependsOn, which also means a refactor that removes a reference can silently reorder your graph.

A team asks for "a portal so developers can self-serve". What do you ask before agreeing?

What happens after the form is submitted. If fulfillment is still manual, the portal is a prettier ticket queue; the automation (templates → git → GitOps → controllers) is the actual work, and the portal is the last 10%. That is the white paper's interface-versus-capability distinction.

Nightly image rebuild-and-rescan: which engine, and why not the others?

A CronWorkflow: scheduled, run-to-completion, DAG with retries and notifications. Not an operator (nothing to reconcile forever), not Tekton (no git event trigger, though it could run the pipeline), not Crossplane (no resource graph to maintain).

Give the three questions that decide the engine.

Does state need reconciling forever or producing once? What is the trigger: API object, schedule, git event, human? Who is the audience: YAML author, form filler, or another system? Answer those three and the table picks itself.

What does the SCM-provider generator actually watch, and how would you debug a repo that never appears?

It queries the git host's API for repos in an org matching filters (name pattern, topic, path existence), then templates one Application per match. Debug by reading the generator block (org, filters, credentials Secret), then checking the ApplicationSet's status/conditions and controller logs for API errors, usually a token scope or a filter that does not match.

An RGD is Active, but every instance sits with ResourcesReady=False and the kro controller logs forbidden. Cause?

kro runs in RBAC aggregation mode and nobody granted the controller rights on the kind the RGD templates. The RGD passes static analysis (so it is Active) but the controller cannot create the child objects. Add a ClusterRole for that kind labeled rbac.kro.run/aggregate-to-controller: "true". Same failure family as a Crossplane provider without RBAC or a workflow SA without workflowtaskresults.

A Backstage template's publish step fails with an authentication error although the form rendered fine. Where do you look, and what is the other common form-time failure?

The scaffolder backend's integration configuration (integrations.github|gitlab|gitea in app-config.yaml) and the token it holds: repo-creation rights on the target org. The form-time failure is RepoUrlPicker with allowedHosts not listing your git host, which blocks submission before any step runs.

How do you make a kro resource wait for a Deployment to actually be available before creating the Service that points at it, and why is the default insufficient?

Add readyWhen: [${deployment.status.availableReplicas > 0}] (or a check on the Available condition) to the Deployment resource. By default a resource is "ready" as soon as it exists and its expressions resolve, so dependents proceed even if the pods never start. includeWhen is the different question of whether to create the resource at all.

Which Backstage kinds must a catalog contain before spec.owner: group:default/team-x resolves, and what happens if they are missing?

A Group named team-x (and usually User entities in memberOf). Without it the Component still ingests, but the ownedBy relation dangles: the UI shows a broken owner link and ownership filters miss the entity. Groups and Users usually come from an org provider (GitHub, LDAP) rather than hand-written YAML.

Docs to know your way around

study time, not exam time
  • kro.run: the ResourceGraphDefinition documentation and examples.
  • backstage.io: Software Templates (actions and parameters) and the catalog model, plus the lab's own template at backstage/template/template.yaml, which is small enough to read whole.
  • Offline: kubectl explain resourcegraphdefinition.spec, kubectl -n argocd get applicationset <name> -o yaml.
  • kro.run/docs/concepts/rgd (schema markers, readiness, conditional creation, dependencies) and kro.run/docs/advanced/access-control for the RBAC modes.
  • backstage.io/docs/features/software-catalog/descriptor-format (every kind's required fields) and backstage.io/docs/features/software-templates/writing-templates; the builtin actions list is the /create/actions page of any running Backstage.
  • tag-app-delivery.cncf.io/whitepapers/platform-eng-maturity-model: the Interfaces and Measurement rows are the ones scenario questions borrow.
free before 4.1make down-gitopsmake down-api