A CRD teaches the API server a new noun. That is all it does: storage, validation, RBAC integration, kubectl support and watch semantics, all inherited for free. Behavior needs a controller (section 3.3). Splitting those two in your head (schema versus behavior) is what makes the whole domain legible.

needsmake upmake api

Orientation

competency 3.1 · designing and creating CRDs

Nothing in this section needs a controller running; the whole point is that a CRD exists without one. Apply one and you immediately have a typed, validated, RBAC-aware, watchable, kubectl explain-documented API, with zero behavior behind it. That gap is where platform engineering lives.

What a task looks like

"Create a CRD for X with required field Y, an enum, and a printer column; create one valid instance; show that an invalid one is rejected." Every clause is graded by an object or an error message. Typing a CRD from memory under time pressure is a genuine skill; do it by hand at least three times before exam day.

Anatomy: the fields that matter

group · names · scope · versions

group + names (plural, singular, kind, shortNames, categories) + scope (Namespaced or Cluster) + versions[]. The metadata name is not free-form: it must be exactly <plural>.<group>, and getting that wrong is the most common first error.

FieldWhy it matters
scopeNamespaced gets you RBAC per namespace and quota via count/<plural>.<group>. Cluster-scoped resources cannot be owned by namespaced ones, a real constraint in Crossplane (3.5).
names.shortNamescluster-global and first-come. Two CRDs claiming tb collide, and kubectl warns you. A platform API designer owns that namespace collision.
names.categoriesputs your kind into kubectl get all-style groupings. Cheap usability win.
versions[].servedwhether the API answers for this version at all
versions[].storageexactly one version is what etcd holds; the rest are converted on the fly
versions[].deprecatedemits a warning to clients using it, the polite way to sunset

Versioning, at exam depth: deprecating a version means served: true with storage moved on; removing it means served: false. That is the whole mechanism unless the schemas actually differ, in which case you need conversion: strategy: None (same fields, different name) or a Webhook converter (real transformation). Know that the choice exists and what each costs.

Schema: where design happens

The schema is OpenAPI v3: types, required, enum, pattern, minimum/maximum, default, format. Two extensions worth knowing by name:

  • x-kubernetes-validations: CEL rules with custom messages, evaluated by the API server. This is how you express cross-field constraints (self.max >= self.min) and immutability (self == oldSelf) without a webhook. The message you write is the message your user sees, so write it like a human.
  • x-kubernetes-preserve-unknown-fields: opts a subtree out of pruning. By default anything not in the schema is silently pruned on write, which looks like data loss and is actually the schema doing its job.
  • Also worth knowing: x-kubernetes-list-type: map with x-kubernetes-list-map-keys (makes server-side apply merge lists sanely instead of replacing them), and x-kubernetes-int-or-string.
Pruning is not an error

Apply a field the schema does not declare and it disappears with no complaint, no event, no warning by default. Users report "my setting is ignored"; the truth is it was never stored. This is why a permissive preserve-unknown-fields subtree is tempting and why a platform API mostly should not have one: you would be trading a clear rejection for a silent failure.

Subresources and presentation

  • subresources.status splits /status into its own endpoint, and that split cuts both ways: writes to the main resource then ignore the status stanza, and writes to /status ignore everything else. So the classic "my controller's status writes vanish" is the subresource working: the controller is calling Update() where it needs UpdateStatus(), or kubectl patch without --subresource=status. Without the subresource, status is just another field and anyone can clobber it, which is why every operator convention in section 3.3 assumes it is on.
  • subresources.scale makes kubectl scale and HPAs work against your kind; three JSONPaths and your CRD is autoscalable.
  • additionalPrinterColumns decides what kubectl get shows. A platform API without a useful READY column is user-hostile, per section 3.1.

Once applied, wait for the Established condition; until then, requests for the new kind may 404. After that the kind behaves like any built-in: RBAC rules can name it, quota can count/ it, admission webhooks and policy engines see it, and kubectl explain documents it from the descriptions you wrote. Descriptions are not decoration; they are the docs your users get.

Where CRDs stop

and what you reach for instead
  • Aggregated API servers: for when you need custom storage, huge object counts, or non-etcd backends. Named in the docs, almost never the answer.
  • Admission webhooks: for validation CEL cannot express (cross-object lookups) or mutation with logic. Every webhook is a new availability dependency for the whole API server; failurePolicy is where you choose which risk you prefer (section 5.2).
  • Operators: for behavior. A CRD with no controller is a typed ConfigMap, which is occasionally exactly what you want (Crossplane's XRs, Argo's Applications and Tekton's Tasks are all "data with a controller elsewhere").
  • Generators: Crossplane's XRD and kro's ResourceGraphDefinition both generate CRDs for you. Sections 3.5 and 3.6 are exactly that: the same thing you hand-write here, produced by a higher-level API.
Command reflex on an unfamiliar cluster

kubectl api-resources | grep <tool> to find the nouns, kubectl explain <kind> --recursive to read the schema, kubectl get crd <name> -o yaml to see printer columns, versions and validation. Three commands and you can operate an operator you have never met, which is precisely the exam scenario.

Schema mechanics the API server enforces

structural rules · CEL in detail · ratcheting · defaulting and nullable · list types

Structural schema: the four rules

apiextensions.k8s.io/v1 requires a structural schema, and a CRD that is not structural is accepted but marked with the NonStructural condition and loses pruning, defaulting and CEL. The rules: every node (root, each property, each array item) has a non-empty type, except nodes with x-kubernetes-int-or-string: true or x-kubernetes-preserve-unknown-fields: true; anything mentioned inside allOf, anyOf, oneOf or not must also be declared outside it; description, type, default, additionalProperties and nullable may not appear inside those junctors; and metadata may only be constrained on name and generateName. The usual failure in hand-written CRDs is a property without type two levels down.

Pruning and its exceptions

Unknown fields are pruned on every write, at every version. x-kubernetes-preserve-unknown-fields: true on a node keeps everything beneath it, but pruning switches back on for any property you do declare under that node. x-kubernetes-embedded-resource: true marks a field that holds a whole Kubernetes object (apiVersion, kind, metadata are then validated). kubectl apply with strict field validation (the default) rejects unknown fields before pruning ever happens, which is why the same manifest can fail from kubectl and silently succeed from a controller using a client without validation.

CEL rules, all the fields

  • rule: scoped to where the x-kubernetes-validations block sits; self is that node. At the root, self.spec and self.status are both visible, which is how you write a rule that compares them; elsewhere only the subtree is. apiVersion, kind, metadata.name and metadata.generateName are the only metadata you can read.
  • message (static) or messageExpression (CEL returning a string; use string(x) to concatenate numbers; falls back to message on error). Without either, the error prints failed rule: <expression>.
  • reason: FieldValueInvalid (default), FieldValueForbidden, FieldValueRequired, FieldValueDuplicate; sets the HTTP status. fieldPath: a relative JSON path so the error points at the offending field rather than the scope.
  • Transition rules: any rule mentioning oldSelf runs only on updates where both values exist, and only on correlatable paths (parents must be objects or x-kubernetes-list-type: map lists, never set or atomic). Idioms: immutability self == oldSelf; append-only self.all(e, e in oldSelf); monotonic self >= oldSelf. optionalOldSelf: true makes oldSelf a CEL optional so the rule also runs on create and on newly set fields.
  • Property names that collide with CEL keywords or contain -, ., / are escaped: self.__namespace__, self.x__dash__prop.
  • Cost budget: rules are cost-estimated when the CRD is written. Iterating a list or string without maxItems, maxProperties or maxLength is assumed worst-case and rejected with "CEL rule exceeded budget by more than 100x (try simplifying the rule, or adding maxItems, maxProperties, and maxLength where arrays, maps, and strings are used)". Bound the collection and the rule is accepted. There is also a per-CRD total and a runtime cap.
  • Compilation errors that fail the CRD write: no_matching_overload (comparing an int to a bool), no_such_field (a typo in a property name), invalid argument to has() macro.
  • Type mapping worth knowing: integer is 64-bit int, number is double (so self.replicas > 1.5 against an integer fails to compile), format: date-time becomes a timestamp, format: duration a duration, a map is additionalProperties, and a listType: set or map compares without order.

Validation ratcheting (GA since 1.33)

An update that leaves an already-invalid field unchanged is accepted even though the field fails the current schema, so you can tighten a schema without first migrating every stored object. Ratcheting does not cover: transition rules (anything using oldSelf), changes to x-kubernetes-list-type, changes to the required list, renamed properties, and metadata validation. Create requests are never ratcheted. Combine it with optionalOldSelf when you want explicit grandfathering logic.

Defaulting and nullable

  • Defaults apply on the request (in the request version), when reading from etcd (storage version), and after a mutating webhook returns a patch; defaults applied on read are not written back until the object is next updated.
  • A default for a non-leaf field must itself validate against the schema and be pruned (no unknown keys inside it).
  • nullable: false (default): a null value is pruned and then defaulted if a default exists. nullable: true: null is stored as null and not defaulted. So foo: null with default: "x" becomes "x" only when the field is not nullable.
  • Defaults are why the status subresource matters: a CRD that defaults into status confuses controllers that compare generations.

List semantics

x-kubernetes-list-type: atomic (default for lists of primitives and structs): server-side apply replaces the whole list and one manager owns it. set: scalar elements, unique, merged as a set. map with x-kubernetes-list-map-keys: [name]: elements are identified by key, so two field managers can each own their entries, and CEL can correlate items in transition rules. Conditions lists use map keyed on type. Argo CD and Flux server-side apply both behave differently on atomic lists (any change reverts the whole list), which is the source of "my sidecar keeps disappearing" reports.

The error you will read most

"The X "y" is invalid: spec.tier: Unsupported value: "platinum": supported values: "bronze", "silver", "gold"" is the enum; "spec.sizeGB: Invalid value: 900: spec.sizeGB in body should be less than or equal to 500" is a bound; "spec: Invalid value: ...: bronze tier is capped at 50GB" is your CEL message; "strict decoding error: unknown field "spec.color"" is kubectl's field validation, not the schema. Four validators, four sentence styles; knowing which produced the text tells you which line of the CRD to edit.

Versions, subresources and the CRD-level errors

conversion · storedVersions · selectableFields · printer columns · scale · what kubectl says when the CRD itself is wrong

Versions in practice

  • Exactly one version has storage: true; any number are served. kubectl defaults to the highest-priority served version by name sorting: GA (v2, v1) before beta before alpha, higher numbers first, so v1 beats v1beta1 beats v2alpha1.
  • deprecated: true plus deprecationWarning makes every request to that version return a warning header (kubectl prints it). served: false removes the endpoint; objects stored at that version stay in etcd.
  • spec.conversion.strategy: None only rewrites apiVersion; use it when the schemas are identical. Webhook calls your service with a ConversionReview containing a batch of objects to convert; it may change anything except metadata other than labels and annotations.
  • Objects are rewritten to the current storage version only when they are next written. status.storedVersions lists every version anything might still be stored at, and you cannot drop a version from spec.versions while it is listed there. The clean path is a StorageVersionMigration (storagemigration.k8s.io) or a no-op update of every object, then patch status.storedVersions.

Subresources, precisely

  • subresources.status: {}: writes to the main endpoint ignore .status, writes to /status ignore everything else and validate only the status stanza, and metadata.generation increments only for changes outside metadata and status. Without it, generation bumps on every status write and observedGeneration means nothing.
  • subresources.scale: specReplicasPath (under .spec, required), statusReplicasPath (under .status, required), labelSelectorPath (a string field holding a serialized selector; required for HPA). Then kubectl scale, HPA and VPA work on your kind.
  • selectableFields (GA since 1.32): spec.versions[].selectableFields[].jsonPath (up to 8 per version, scalar fields only) enables kubectl get shirts --field-selector spec.color=blue and lets controllers watch a subset without labels.

Printer columns and names

additionalPrinterColumns[]: name, type (integer, number, string, boolean, date), optional format, jsonPath, description and priority (0 shows in the default view, greater than 0 only with -o wide). type: date renders as an age. NAME is implicit; AGE is not, so most CRDs add {name: Age, type: date, jsonPath: .metadata.creationTimestamp}. Columns come from the served version you request, so two versions can print differently. names.shortNames are cluster-global and first-come; names.categories: [all] puts the kind into kubectl get all; a category of your own (fluxcd, platform) lets kubectl get platform -A list every kind you ship.

Lifecycle of the CRD object itself

  • After creation wait for Established: True (and NamesAccepted: True); before that, requests for the kind 404. A NamesAccepted: False with "is already in use" means a plural, kind or shortName collides with another CRD.
  • Deleting a CRD deletes every custom object of that kind; the finalizer customresourcecleanup.apiextensions.k8s.io holds the CRD in Terminating until they are gone, and a stuck CR (its own finalizer, a dead controller) therefore sticks the CRD too. Argo CD and Flux prune CRDs like any resource unless you protect them (Prune=false, kustomize.toolkit.fluxcd.io/prune: disabled), which is the single most destructive prune there is.
  • Quota counts custom objects with count/<plural>.<group>; RBAC names them by apiGroups: [group], resources: [plural, plural/status].

Errors from applying the CRD, and their causes

Message containsCause
metadata.name: Invalid value: must be spec.names.plural+"."+spec.groupthe name is not <plural>.<group>
spec.versions: Invalid value: must have exactly one version marked as storage versionzero or two storage: true
spec.versions[0].schema.openAPIV3Schema: Required value: schemas are requireda version without a schema in v1
must be structural / must not be inside allOf / NonStructural conditiona rule-1 to rule-4 violation above
compilation failed: ERROR: ... found no matching overload / undefined fielda CEL type mismatch or misspelt property
CEL rule exceeded budgetunbounded iteration; add maxItems, maxLength
spec.conversion.webhook: Required value when strategy is Webhookstrategy: Webhook without webhook.clientConfig and conversionReviewVersions
status.storedVersions[0]: Invalid value: must appear in spec.versionsyou removed a version that objects are still stored at
additionalPrinterColumns[1].jsonPath: Invalid valuea path that does not start with . or is malformed
Reflex

kubectl explain <kind> --recursive and kubectl explain <kind>.spec.<field> read the served schema and its descriptions; add --api-version group/v1beta1 to read an older version. kubectl get crd <name> -o jsonpath='{.status.conditions}' before anything else when a brand new CRD "does not work". kubectl get crd <name> -o jsonpath='{.status.storedVersions}' before removing a version.

Exercises

tick the dot when its check passes

A platform-flavored example, typed out rather than pasted, because the exam gives you a task description, not a starting file:

apiVersion: apiextensions.k8s.io/v1
kind: CustomResourceDefinition
metadata:
  name: tenantbuckets.platform.lab.local
spec:
  group: platform.lab.local
  scope: Namespaced
  names: { plural: tenantbuckets, singular: tenantbucket, kind: TenantBucket, shortNames: [tb] }
  versions:
    - name: v1alpha1
      served: true
      storage: true
      subresources: { status: {} }
      additionalPrinterColumns:
        - { name: Tier,     type: string,  jsonPath: .spec.tier }
        - { name: SizeGB,   type: integer, jsonPath: .spec.sizeGB }
      schema:
        openAPIV3Schema:
          type: object
          properties:
            spec:
              type: object
              required: [tier]
              properties:
                tier:   { type: string, enum: [bronze, silver, gold], description: "Service tier; sets replication and backup policy." }
                sizeGB: { type: integer, minimum: 1, maximum: 500, default: 10 }
              x-kubernetes-validations:
                - rule: "self.tier != 'bronze' || self.sizeGB <= 50"
                  message: "bronze tier is capped at 50GB"
            status:
              type: object
              properties:
                phase: { type: string }

Apply it, then make the API server prove each design decision:

kubectl wait --for=condition=Established crd/tenantbuckets.platform.lab.local
kubectl apply -f - <<'EOF'
apiVersion: platform.lab.local/v1alpha1
kind: TenantBucket
metadata: { name: good, namespace: default }
spec: { tier: silver, sizeGB: 100 }
EOF
kubectl get tb    # printer columns show Tier and SizeGB; apply one without sizeGB to watch the default appear
outputcaptured 2026-08-26
$ kubectl wait --for=condition=Established crd/tenantbuckets.platform.lab.local
customresourcedefinition.apiextensions.k8s.io/tenantbuckets.platform.lab.local condition met
$ kubectl apply -f - <<'EOF'
apiVersion: platform.lab.local/v1alpha1
kind: TenantBucket
metadata: { name: good, namespace: default }
spec: { tier: silver, sizeGB: 100 }
EOF
tenantbucket.platform.lab.local/good created
$ kubectl get tb    # printer columns show Tier and SizeGB; apply one without sizeGB to watch the default appear
Warning: short name "tb" could also match lower priority resource triggerbindings.triggers.tekton.dev
NAME   TIER     SIZEGB
good   silver   100

With cicd installed, that last command also prints a warning that tb could match Tekton's triggerbindings. Keep it; it demonstrates that short names are cluster-global and first-come.

verify: the rejections, which are the real test of the schema: tier: platinum must fail on the enum, sizeGB: 900 on the maximum, and tier: bronze, sizeGB: 100 on the CEL rule with your message in the error. Three different validators refusing three different ways; know which produced which text.

Apply a TenantBucket with an extra field spec.color: red, read it back, and observe the field is gone with no error. Then kubectl explain tenantbucket.spec --recursive and see your descriptions serving as live documentation.

verify: you can state the flag a schema author would set to keep unknown fields, and why a platform API mostly should not.

kubectl patch tenantbucket good --subresource=status --type=merge -p '{"status":{"phase":"Ready"}}', then confirm kubectl get tb good -o jsonpath='{.status.phase}' says Ready and that a plain spec-edit did not clear it.

verify: status survives spec edits and vice versa. This is the mechanism every operator in section 3.3 relies on.

kubectl get crd appenvironments.platform.lab.local -o yaml (Crossplane generated it from the lab's XRD). Compare its schema and printer columns against yours; note what a machine-generated platform API includes that your hand-rolled one lacks (conditions conventions, connection details).

verify: you can name three things to steal, and add at least one of them to your TenantBucket.

A CRD's metadata.name is not free text: it must be the plural joined to the group. The error is exact and instantly recognizable, which makes it a cheap mark on a task where you typed the name from memory.

kubectl apply -f - <<'EOF'
apiVersion: apiextensions.k8s.io/v1
kind: CustomResourceDefinition
metadata:
  name: buckets.platform.lab.local
spec:
  group: platform.lab.local
  scope: Namespaced
  names: { plural: tenantbuckets, singular: tenantbucket, kind: TenantBucket }
  versions:
    - name: v1alpha1
      served: true
      storage: true
      schema:
        openAPIV3Schema:
          type: object
          properties:
            spec: { type: object }
EOF
outputcaptured 2026-09-13
$ kubectl apply -f - <<'EOF'
apiVersion: apiextensions.k8s.io/v1
kind: CustomResourceDefinition
metadata:
  name: buckets.platform.lab.local
spec:
  group: platform.lab.local
  scope: Namespaced
  names: { plural: tenantbuckets, singular: tenantbucket, kind: TenantBucket }
  versions:
    - name: v1alpha1
      served: true
      storage: true
      schema:
        openAPIV3Schema:
          type: object
          properties:
            spec: { type: object }
EOF
The CustomResourceDefinition "buckets.platform.lab.local" is invalid: metadata.name: Invalid value: "buckets.platform.lab.local": must be spec.names.plural+"."+spec.group
verify: the message names spec.names.plural+"."+spec.group as the required form.

Exactly one version holds the bytes in etcd. Marking two as storage is the mistake behind half of the failed version-migration tasks, and the API server refuses it up front rather than letting you find out later.

kubectl apply -f - <<'EOF'
apiVersion: apiextensions.k8s.io/v1
kind: CustomResourceDefinition
metadata:
  name: twostore.platform.lab.local
spec:
  group: platform.lab.local
  scope: Namespaced
  names: { plural: twostore, singular: twostore, kind: TwoStore }
  versions:
    - name: v1alpha1
      served: true
      storage: true
      schema:
        openAPIV3Schema: { type: object, properties: { spec: { type: object } } }
    - name: v1beta1
      served: true
      storage: true
      schema:
        openAPIV3Schema: { type: object, properties: { spec: { type: object } } }
EOF
outputcaptured 2026-09-13
$ kubectl apply -f - <<'EOF'
apiVersion: apiextensions.k8s.io/v1
kind: CustomResourceDefinition
metadata:
  name: twostore.platform.lab.local
spec:
  group: platform.lab.local
  scope: Namespaced
  names: { plural: twostore, singular: twostore, kind: TwoStore }
  versions:
    - name: v1alpha1
      served: true
      storage: true
      schema:
        openAPIV3Schema: { type: object, properties: { spec: { type: object } } }
    - name: v1beta1
      served: true
      storage: true
      schema:
        openAPIV3Schema: { type: object, properties: { spec: { type: object } } }
EOF
The CustomResourceDefinition "twostore.platform.lab.local" is invalid: 
* spec.versions: Invalid value: [{"Name":"v1alpha1","Served":true,"Storage":true,"Deprecated":false,"DeprecationWarning":null,"Schema":null,"Subresources":null,"AdditionalPrinterColumns":null,"SelectableFields":null},{"Name":"v1beta1","Served":true,"Storage":true,"Deprecated":false,"DeprecationWarning":null,"Schema":null,"Subresources":null,"AdditionalPrinterColumns":null,"SelectableFields":null}]: must have exactly one version marked as storage version
* status.storedVersions: Invalid value: ["v1alpha1"]: must have the storage version v1beta1
verify: the error says exactly one version must be marked as the storage version.

CEL validation runs on every write, so the API server charges it against a budget. An unbounded loop over an unbounded list exceeds that budget at CRD creation time, and bounding the schema is what makes the same rule legal.

kubectl apply -f - <<'EOF'
apiVersion: apiextensions.k8s.io/v1
kind: CustomResourceDefinition
metadata:
  name: budgets.platform.lab.local
spec:
  group: platform.lab.local
  scope: Namespaced
  names: { plural: budgets, singular: budget, kind: Budget }
  versions:
    - name: v1alpha1
      served: true
      storage: true
      schema:
        openAPIV3Schema:
          type: object
          properties:
            spec:
              type: object
              properties:
                names:
                  type: array
                  items: { type: string }
                  x-kubernetes-validations:
                    - rule: "self.all(x, x.contains('a'))"
EOF
kubectl apply -f - <<'EOF'
apiVersion: apiextensions.k8s.io/v1
kind: CustomResourceDefinition
metadata:
  name: budgets.platform.lab.local
spec:
  group: platform.lab.local
  scope: Namespaced
  names: { plural: budgets, singular: budget, kind: Budget }
  versions:
    - name: v1alpha1
      served: true
      storage: true
      schema:
        openAPIV3Schema:
          type: object
          properties:
            spec:
              type: object
              properties:
                names:
                  type: array
                  maxItems: 10
                  items: { type: string, maxLength: 32 }
                  x-kubernetes-validations:
                    - rule: "self.all(x, x.contains('a'))"
EOF
kubectl delete crd budgets.platform.lab.local
outputcaptured 2026-09-13
$ kubectl apply -f - <<'EOF'
apiVersion: apiextensions.k8s.io/v1
kind: CustomResourceDefinition
metadata:
  name: budgets.platform.lab.local
spec:
  group: platform.lab.local
  scope: Namespaced
  names: { plural: budgets, singular: budget, kind: Budget }
  versions:
    - name: v1alpha1
      served: true
      storage: true
      schema:
        openAPIV3Schema:
          type: object
          properties:
            spec:
              type: object
              properties:
                names:
                  type: array
                  items: { type: string }
                  x-kubernetes-validations:
                    - rule: "self.all(x, x.contains('a'))"
EOF
The CustomResourceDefinition "budgets.platform.lab.local" is invalid: 
* spec.validation.openAPIV3Schema.properties[spec].properties[names].x-kubernetes-validations[0].rule: Forbidden: estimated rule cost exceeds budget by factor of more than 100x (try simplifying the rule, or adding maxItems, maxProperties, and maxLength where arrays, maps, and strings are declared)
* spec.validation.openAPIV3Schema.properties[spec].properties[names].x-kubernetes-validations[0].rule: Forbidden: contributed to estimated rule cost total exceeding cost limit for entire OpenAPIv3 schema
* spec.validation.openAPIV3Schema: Forbidden: x-kubernetes-validations estimated rule cost total for entire OpenAPIv3 schema exceeds budget by factor of more than 100x (try simplifying the rule, or adding maxItems, maxProperties, and maxLength where arrays, maps, and strings are declared)
$ kubectl apply -f - <<'EOF'
apiVersion: apiextensions.k8s.io/v1
kind: CustomResourceDefinition
metadata:
  name: budgets.platform.lab.local
spec:
  group: platform.lab.local
  scope: Namespaced
  names: { plural: budgets, singular: budget, kind: Budget }
  versions:
    - name: v1alpha1
      served: true
      storage: true
      schema:
        openAPIV3Schema:
          type: object
          properties:
            spec:
              type: object
              properties:
                names:
                  type: array
                  maxItems: 10
                  items: { type: string, maxLength: 32 }
                  x-kubernetes-validations:
                    - rule: "self.all(x, x.contains('a'))"
EOF
customresourcedefinition.apiextensions.k8s.io/budgets.platform.lab.local created
$ kubectl delete crd budgets.platform.lab.local
customresourcedefinition.apiextensions.k8s.io "budgets.platform.lab.local" deleted
verify: the unbounded version is rejected for exceeding the estimated cost budget, and adding maxItems and maxLength makes the identical rule acceptable. The lesson generalizes: bound your lists and strings before you write CEL over them.

oldSelf turns a validation into a transition rule, and transition rules are skipped on create because there is nothing to compare against. That asymmetry is the whole feature and the whole trap.

kubectl apply -f - <<'EOF'
apiVersion: apiextensions.k8s.io/v1
kind: CustomResourceDefinition
metadata: { name: tenantbuckets.platform.lab.local }
spec:
  group: platform.lab.local
  scope: Namespaced
  names: { plural: tenantbuckets, singular: tenantbucket, kind: TenantBucket, shortNames: [tb] }
  versions:
    - name: v1alpha1
      served: true
      storage: true
      subresources: { status: {} }
      schema:
        openAPIV3Schema:
          type: object
          properties:
            spec:
              type: object
              required: [tier]
              properties:
                tier:   { type: string, enum: [bronze, silver, gold] }
                sizeGB: { type: integer, minimum: 1, maximum: 500, default: 10 }
            status:
              type: object
              properties:
                phase: { type: string }
EOF
kubectl wait --for=condition=Established crd/tenantbuckets.platform.lab.local --timeout=60s
kubectl apply -f - <<'EOF'
apiVersion: platform.lab.local/v1alpha1
kind: TenantBucket
metadata: { name: good, namespace: default }
spec: { tier: silver, sizeGB: 100 }
EOF
kubectl patch crd tenantbuckets.platform.lab.local --type json -p '[{"op":"add","path":"/spec/versions/0/schema/openAPIV3Schema/properties/spec/x-kubernetes-validations","value":[{"rule":"self.tier == oldSelf.tier","message":"tier is immutable once set"}]}]'
sleep 5
kubectl patch tenantbucket good --type merge -p '{"spec":{"tier":"gold"}}'
kubectl apply -f - <<'EOF'
apiVersion: platform.lab.local/v1alpha1
kind: TenantBucket
metadata: { name: fresh, namespace: default }
spec: { tier: gold, sizeGB: 20 }
EOF
kubectl get tb
kubectl delete tb fresh
outputcaptured 2026-09-12
$ kubectl apply -f - <<'EOF'
apiVersion: apiextensions.k8s.io/v1
kind: CustomResourceDefinition
metadata: { name: tenantbuckets.platform.lab.local }
spec:
  group: platform.lab.local
  scope: Namespaced
  names: { plural: tenantbuckets, singular: tenantbucket, kind: TenantBucket, shortNames: [tb] }
  versions:
    - name: v1alpha1
      served: true
      storage: true
      subresources: { status: {} }
      schema:
        openAPIV3Schema:
          type: object
          properties:
            spec:
              type: object
              required: [tier]
              properties:
                tier:   { type: string, enum: [bronze, silver, gold] }
                sizeGB: { type: integer, minimum: 1, maximum: 500, default: 10 }
            status:
              type: object
              properties:
                phase: { type: string }
EOF
customresourcedefinition.apiextensions.k8s.io/tenantbuckets.platform.lab.local created
$ kubectl wait --for=condition=Established crd/tenantbuckets.platform.lab.local --timeout=60s
customresourcedefinition.apiextensions.k8s.io/tenantbuckets.platform.lab.local condition met
$ kubectl apply -f - <<'EOF'
apiVersion: platform.lab.local/v1alpha1
kind: TenantBucket
metadata: { name: good, namespace: default }
spec: { tier: silver, sizeGB: 100 }
EOF
tenantbucket.platform.lab.local/good created
$ kubectl patch crd tenantbuckets.platform.lab.local --type json -p '[{"op":"add","path":"/spec/versions/0/schema/openAPIV3Schema/properties/spec/x-kubernetes-validations","value":[{"rule":"self.tier == oldSelf.tier","message":"tier is immutable once set"}]}]'
customresourcedefinition.apiextensions.k8s.io/tenantbuckets.platform.lab.local patched
$ sleep 5
$ kubectl patch tenantbucket good --type merge -p '{"spec":{"tier":"gold"}}'
The TenantBucket "good" is invalid: spec: Invalid value: tier is immutable once set
$ kubectl apply -f - <<'EOF'
apiVersion: platform.lab.local/v1alpha1
kind: TenantBucket
metadata: { name: fresh, namespace: default }
spec: { tier: gold, sizeGB: 20 }
EOF
tenantbucket.platform.lab.local/fresh created
$ kubectl get tb
Warning: short name "tb" could also match lower priority resource triggerbindings.triggers.tekton.dev
NAME    AGE
fresh   0s
good    6s
$ kubectl delete tb fresh
Warning: short name "tb" could also match lower priority resource triggerbindings.triggers.tekton.dev
tenantbucket.platform.lab.local "fresh" deleted from default namespace
verify: the update is refused with your message and the create of a brand new object with the same tier succeeds.

Tightening a schema used to break every existing object that violated the new rule, including on unrelated edits. Ratcheting validation means an update that does not touch the offending field is allowed through, so you can roll a constraint out before the data is clean.

kubectl apply -f - <<'EOF'
apiVersion: apiextensions.k8s.io/v1
kind: CustomResourceDefinition
metadata: { name: tenantbuckets.platform.lab.local }
spec:
  group: platform.lab.local
  scope: Namespaced
  names: { plural: tenantbuckets, singular: tenantbucket, kind: TenantBucket, shortNames: [tb] }
  versions:
    - name: v1alpha1
      served: true
      storage: true
      subresources: { status: {} }
      schema:
        openAPIV3Schema:
          type: object
          properties:
            spec:
              type: object
              required: [tier]
              properties:
                tier:   { type: string, enum: [bronze, silver, gold] }
                sizeGB: { type: integer, minimum: 1, maximum: 500, default: 10 }
            status:
              type: object
              properties:
                phase: { type: string }
EOF
kubectl wait --for=condition=Established crd/tenantbuckets.platform.lab.local --timeout=60s
kubectl apply -f - <<'EOF'
apiVersion: platform.lab.local/v1alpha1
kind: TenantBucket
metadata: { name: good, namespace: default }
spec: { tier: silver, sizeGB: 100 }
EOF
kubectl apply -f - <<'EOF'
apiVersion: platform.lab.local/v1alpha1
kind: TenantBucket
metadata: { name: legacy, namespace: default }
spec: { tier: silver, sizeGB: 100 }
EOF
kubectl patch crd tenantbuckets.platform.lab.local --type json -p '[{"op":"add","path":"/spec/versions/0/schema/openAPIV3Schema/properties/spec/properties/tier/pattern","value":"^gold$"}]'
sleep 5
kubectl label tenantbucket legacy team=a
kubectl patch tenantbucket legacy --type merge -p '{"spec":{"sizeGB":120}}'
kubectl patch tenantbucket legacy --type merge -p '{"spec":{"tier":"silver"}}'
kubectl patch crd tenantbuckets.platform.lab.local --type json -p '[{"op":"remove","path":"/spec/versions/0/schema/openAPIV3Schema/properties/spec/properties/tier/pattern"}]'
kubectl delete tb legacy
outputcaptured 2026-09-12
$ kubectl apply -f - <<'EOF'
apiVersion: apiextensions.k8s.io/v1
kind: CustomResourceDefinition
metadata: { name: tenantbuckets.platform.lab.local }
spec:
  group: platform.lab.local
  scope: Namespaced
  names: { plural: tenantbuckets, singular: tenantbucket, kind: TenantBucket, shortNames: [tb] }
  versions:
    - name: v1alpha1
      served: true
      storage: true
      subresources: { status: {} }
      schema:
        openAPIV3Schema:
          type: object
          properties:
            spec:
              type: object
              required: [tier]
              properties:
                tier:   { type: string, enum: [bronze, silver, gold] }
                sizeGB: { type: integer, minimum: 1, maximum: 500, default: 10 }
            status:
              type: object
              properties:
                phase: { type: string }
EOF
customresourcedefinition.apiextensions.k8s.io/tenantbuckets.platform.lab.local configured
$ kubectl wait --for=condition=Established crd/tenantbuckets.platform.lab.local --timeout=60s
customresourcedefinition.apiextensions.k8s.io/tenantbuckets.platform.lab.local condition met
$ kubectl apply -f - <<'EOF'
apiVersion: platform.lab.local/v1alpha1
kind: TenantBucket
metadata: { name: good, namespace: default }
spec: { tier: silver, sizeGB: 100 }
EOF
tenantbucket.platform.lab.local/good unchanged
$ kubectl apply -f - <<'EOF'
apiVersion: platform.lab.local/v1alpha1
kind: TenantBucket
metadata: { name: legacy, namespace: default }
spec: { tier: silver, sizeGB: 100 }
EOF
tenantbucket.platform.lab.local/legacy created
$ kubectl patch crd tenantbuckets.platform.lab.local --type json -p '[{"op":"add","path":"/spec/versions/0/schema/openAPIV3Schema/properties/spec/properties/tier/pattern","value":"^gold$"}]'
customresourcedefinition.apiextensions.k8s.io/tenantbuckets.platform.lab.local patched
$ sleep 5
$ kubectl label tenantbucket legacy team=a
tenantbucket.platform.lab.local/legacy labeled
$ kubectl patch tenantbucket legacy --type merge -p '{"spec":{"sizeGB":120}}'
tenantbucket.platform.lab.local/legacy patched
$ kubectl patch tenantbucket legacy --type merge -p '{"spec":{"tier":"silver"}}'
tenantbucket.platform.lab.local/legacy patched (no change)
$ kubectl patch crd tenantbuckets.platform.lab.local --type json -p '[{"op":"remove","path":"/spec/versions/0/schema/openAPIV3Schema/properties/spec/properties/tier/pattern"}]'
customresourcedefinition.apiextensions.k8s.io/tenantbuckets.platform.lab.local patched
$ kubectl delete tb legacy
Warning: short name "tb" could also match lower priority resource triggerbindings.triggers.tekton.dev
tenantbucket.platform.lab.local "legacy" deleted from default namespace
verify: labeling the object succeeds even though it now violates the pattern, and writing the same invalid value back explicitly is refused.

kubectl get --field-selector works on built-in kinds and fails on yours until you declare which fields are selectable. Declaring one is a one-line change and the error before it is unmistakable.

kubectl get tenantbuckets.platform.lab.local --field-selector spec.tier=silver
kubectl patch crd tenantbuckets.platform.lab.local --type json -p '[{"op":"add","path":"/spec/versions/0/selectableFields","value":[{"jsonPath":".spec.tier"}]}]'
sleep 5
kubectl get tenantbuckets.platform.lab.local --field-selector spec.tier=silver
kubectl get tenantbuckets.platform.lab.local --field-selector spec.tier=gold
outputcaptured 2026-09-12
$ kubectl get tenantbuckets.platform.lab.local --field-selector spec.tier=silver
Error from server (BadRequest): Unable to find "platform.lab.local/v1alpha1, Resource=tenantbuckets" that match label selector "", field selector "spec.tier=silver": field label not supported: spec.tier
$ kubectl patch crd tenantbuckets.platform.lab.local --type json -p '[{"op":"add","path":"/spec/versions/0/selectableFields","value":[{"jsonPath":".spec.tier"}]}]'
customresourcedefinition.apiextensions.k8s.io/tenantbuckets.platform.lab.local patched
$ sleep 5
$ kubectl get tenantbuckets.platform.lab.local --field-selector spec.tier=silver
NAME   AGE
good   4h48m
$ kubectl get tenantbuckets.platform.lab.local --field-selector spec.tier=gold
No resources found in default namespace.
verify: the first attempt fails saying the field label is not supported, and after the patch the same query filters. Note that the server does the filtering, which is the reason this matters at ten thousand objects rather than ten.

status.storedVersions is the list of versions that still exist as bytes. The API server will not let you stop serving one that is in that list, and the fix is to rewrite the objects and then correct the status by hand.

kubectl get crd tenantbuckets.platform.lab.local -o jsonpath='{.status.storedVersions}{"\n"}'
kubectl patch crd tenantbuckets.platform.lab.local --type json -p '[{"op":"add","path":"/spec/versions/-","value":{"name":"v1beta1","served":true,"storage":false,"subresources":{"status":{}},"schema":{"openAPIV3Schema":{"type":"object","properties":{"spec":{"type":"object","properties":{"tier":{"type":"string"},"sizeGB":{"type":"integer"}}},"status":{"type":"object","properties":{"phase":{"type":"string"}}}}}}}}]'
kubectl patch crd tenantbuckets.platform.lab.local --type json -p '[{"op":"replace","path":"/spec/versions/0/storage","value":false},{"op":"replace","path":"/spec/versions/1/storage","value":true}]'
kubectl get crd tenantbuckets.platform.lab.local -o jsonpath='{.status.storedVersions}{"\n"}'
kubectl patch crd tenantbuckets.platform.lab.local --type json -p '[{"op":"remove","path":"/spec/versions/0"}]'
kubectl get tb -o name | xargs -r -I{} kubectl patch {} --type merge -p '{"metadata":{"annotations":{"rewrite":"now"}}}'
kubectl patch crd tenantbuckets.platform.lab.local --subresource status --type merge -p '{"status":{"storedVersions":["v1beta1"]}}'
kubectl patch crd tenantbuckets.platform.lab.local --type json -p '[{"op":"remove","path":"/spec/versions/0"}]'
kubectl get crd tenantbuckets.platform.lab.local -o jsonpath='{range .spec.versions[*]}{.name} storage={.storage}{"\n"}{end}'
outputcaptured 2026-09-12
$ kubectl get crd tenantbuckets.platform.lab.local -o jsonpath='{.status.storedVersions}{"\n"}'
["v1alpha1"]
$ kubectl patch crd tenantbuckets.platform.lab.local --type json -p '[{"op":"add","path":"/spec/versions/-","value":{"name":"v1beta1","served":true,"storage":false,"subresources":{"status":{}},"schema":{"openAPIV3Schema":{"type":"object","properties":{"spec":{"type":"object","properties":{"tier":{"type":"string"},"sizeGB":{"type":"integer"}}},"status":{"type":"object","properties":{"phase":{"type":"string"}}}}}}}}]'
customresourcedefinition.apiextensions.k8s.io/tenantbuckets.platform.lab.local patched
$ kubectl patch crd tenantbuckets.platform.lab.local --type json -p '[{"op":"replace","path":"/spec/versions/0/storage","value":false},{"op":"replace","path":"/spec/versions/1/storage","value":true}]'
customresourcedefinition.apiextensions.k8s.io/tenantbuckets.platform.lab.local patched
$ kubectl get crd tenantbuckets.platform.lab.local -o jsonpath='{.status.storedVersions}{"\n"}'
["v1alpha1","v1beta1"]
$ kubectl patch crd tenantbuckets.platform.lab.local --type json -p '[{"op":"remove","path":"/spec/versions/0"}]'
The CustomResourceDefinition "tenantbuckets.platform.lab.local" is invalid: status.storedVersions[0]: Invalid value: "v1alpha1": missing from spec.versions; v1alpha1 was previously a storage version, and must remain in spec.versions until a storage migration ensures no data remains persisted in v1alpha1 and removes v1alpha1 from status.storedVersions
$ kubectl get tb -o name | xargs -r -I{} kubectl patch {} --type merge -p '{"metadata":{"annotations":{"rewrite":"now"}}}'
Warning: short name "tb" could also match lower priority resource triggerbindings.triggers.tekton.dev
tenantbucket.platform.lab.local/good patched
$ kubectl patch crd tenantbuckets.platform.lab.local --subresource status --type merge -p '{"status":{"storedVersions":["v1beta1"]}}'
customresourcedefinition.apiextensions.k8s.io/tenantbuckets.platform.lab.local patched
$ kubectl patch crd tenantbuckets.platform.lab.local --type json -p '[{"op":"remove","path":"/spec/versions/0"}]'
customresourcedefinition.apiextensions.k8s.io/tenantbuckets.platform.lab.local patched
$ kubectl get crd tenantbuckets.platform.lab.local -o jsonpath='{range .spec.versions[*]}{.name} storage={.storage}{"\n"}{end}'
v1beta1 storage=true
verify: the first removal is refused naming status.storedVersions, and it succeeds after every object has been rewritten through the new storage version and the status corrected. That no-op update loop is the real migration step; say why a conversion webhook does not remove the need for it.

Before debugging why a custom resource "does not work", ask whether the API server accepted the type at all. Two conditions answer it, and they are the first thing to read on a CRD you did not write.

kubectl get crd tenantbuckets.platform.lab.local -o jsonpath='{.status.conditions[*].type}{"\n"}'
kubectl get crd tenantbuckets.platform.lab.local -o jsonpath='{.status.conditions}' | jq '.[] | {type, status, reason, message}'
kubectl get crd appenvironments.platform.lab.local -o jsonpath='{.status.conditions[*].type}{"\n"}'
outputcaptured 2026-09-12
$ kubectl get crd tenantbuckets.platform.lab.local -o jsonpath='{.status.conditions[*].type}{"\n"}'
NamesAccepted Established
$ kubectl get crd tenantbuckets.platform.lab.local -o jsonpath='{.status.conditions}' | jq '.[] | {type, status, reason, message}'
{
  "type": "NamesAccepted",
  "status": "True",
  "reason": "NoConflicts",
  "message": "no conflicts found"
}
{
  "type": "Established",
  "status": "True",
  "reason": "InitialNamesAccepted",
  "message": "the initial names have been accepted"
}
$ kubectl get crd appenvironments.platform.lab.local -o jsonpath='{.status.conditions[*].type}{"\n"}'
NamesAccepted Established
verify: both CRDs report NamesAccepted and Established. Say what a missing Established with a present NamesAccepted would tell you, and what a failing NamesAccepted usually means on a cluster that already has other CRDs.

Self-check

answer before opening
A user says their field "does not work". kubectl get -o yaml shows it missing entirely, and no error was printed. What happened?

Pruning: the field is not in the schema, so the API server dropped it on write. Either add it to the schema or, rarely, mark that subtree with x-kubernetes-preserve-unknown-fields. The user experience argument is for adding it properly.

Your controller's status writes vanish. One likely cause?

The status subresource is enabled and the writer is going through the main endpoint, which ignores the status stanza. Use UpdateStatus() (or kubectl patch --subresource=status). The mirror-image bug also exists: with no subresource at all, a spec update from anyone overwrites the status your controller just wrote.

You must forbid changing spec.tier after creation. How, without a webhook?

A CEL validation on the field: rule: "self == oldSelf" with a message like "tier is immutable". x-kubernetes-validations is evaluated by the API server and can compare against the old object, which is exactly the immutability case.

You want to serve v1alpha1 and v1beta1 with different fields. What do you need?

Both versions served, exactly one marked storage, and a conversion strategy: None only if the fields are compatible, otherwise a conversion webhook. Everything read from etcd is stored in the storage version and converted on the way out.

What does a CRD give you for free, and what does it definitively not give you?

Free: persistence, validation, defaulting, RBAC integration, watch/informers, kubectl support and explain docs, quota counting, policy-engine visibility. Not free: any behavior whatsoever. Nothing happens until a controller acts on the object.

Your CEL rule over a list of strings is rejected with "exceeded budget by more than 100x". What do you change?

Bound the inputs the rule iterates: maxItems on the array and maxLength on the string items (or maxProperties on a map). The estimator assumes worst-case sizes for unbounded collections, so the rule itself can stay as it is. The limits also become real validation, which is usually what a platform API wants anyway.

You tightened a pattern on spec.name. Existing objects with old-style names must still be updatable. What makes that work, and what does it not cover?

Validation ratcheting, GA since 1.33: an update is accepted if every part that fails validation is unchanged by that update. It does not ratchet transition rules (oldSelf), list-type changes, changes to required, renamed properties or metadata rules, and it never applies to creates. For finer control write the grandfathering into the rule with optionalOldSelf: true.

You want to drop v1alpha1 from a CRD but the update is refused. Why, and what is the sequence?

status.storedVersions still lists v1alpha1, meaning objects may be stored at it; the API server refuses to remove a version that stored data might need. Make v1 the storage version, rewrite every object (a StorageVersionMigration or a no-op update loop) so they are stored at v1, patch status.storedVersions to [v1], then remove v1alpha1 from spec.versions. Setting served: false first is the polite intermediate step.

Enable kubectl get widgets --field-selector spec.tier=gold. Field, version status, and limits?

spec.versions[].selectableFields: [{jsonPath: .spec.tier}], GA since Kubernetes 1.32. Only scalar fields, at most 8 per version, and the selector must be written against the version that declares them. It replaces the habit of copying spec fields into labels purely so controllers and humans can filter.

Docs to know your way around

study time, not exam time
  • kubernetes.io: "Extend the Kubernetes API with CustomResourceDefinitions" (one long page covering versions, pruning, CEL validation and defaults), and "Versions in CustomResourceDefinitions".
  • Offline: kubectl explain crd.spec.versions --recursive when you forget field placement, which everyone does; kubectl get crd <name> -o yaml to learn from an operator's own schema.
  • kubernetes.io: "Extend the Kubernetes API with CustomResourceDefinitions", sections Specifying a structural schema, Validation rules (messageExpression, reason, fieldPath, transition rules, resource use), Validation ratcheting, Defaulting and Nullable, Field selectors: the versions and error texts above come from there.
  • kubernetes.io: "Versions in CustomResourceDefinitions": version priority sorting, conversion strategies, storedVersions and the storage version migration procedure.
  • Offline: kubectl explain crd.spec.versions.schema.openAPIV3Schema --recursive lists every x-kubernetes-* extension the server understands; kubectl get crd <name> -o jsonpath='{.status}' for Established, NamesAccepted, NonStructural and storedVersions.