The closest thing the CNPE has to a philosophy exam, and it is 13 pages. Read the CNCF Platforms white paper once end to end. What follows is a night-before compression, plus the API-design half that domain 3 then implements five different ways.

needsnothing running

Orientation

the design half of every domain 3 competency

This is the one reading session in the curriculum, and it belongs before the other five sections, because they all implement ideas named here. Keep it short: 45 minutes, then go build something.

How this gets tested

Vocabulary and judgment, not commands. A scenario describes a platform that violates one of the paper's attributes and asks what is wrong; or it hands you a request and asks which capability and which interface should serve it. The words below are the answer key: using the paper's own terms is what makes an answer score.

Definitions that get tested as vocabulary

learn the nouns

A platform is an integrated collection of capabilities, defined and presented according to the needs of its users. The key move in that sentence is users: a platform is a product with internal customers, not an infrastructure inventory.

The attribute list, worth memorizing

AttributeViolated when…
Platform as a productno roadmap, no users consulted, success measured in tickets closed
Consistent user experienceevery capability has its own bespoke interface and conventions
Documentation and onboardingthe platform is knowable only by asking a person in Slack
Self-servicea human approves each request; lead time measured in days
Reduced cognitive loaddevelopers must understand the implementation to use the interface
Optional and composableadoption is mandated, escape hatches are forbidden
Secure by defaultthe safe path is the harder path

Four terms to be able to define cold

  • Thinnest Viable Platform (TVP): the smallest layer that provides consistency and accelerates delivery, deliberately kept small. The paper's example of a minimal platform is a wiki page of provisioning links. The instinct to carry: platform teams build interfaces and experiences, and should not rebuild capabilities that managed providers or upstream projects already offer.
  • Golden path: a templated composition of well-integrated code and capabilities for rapid project development, documentation included. Paved, not mandatory. This lab implements one literally: Backstage template → new Gitea repo → ApplicationSet picks it up → running workload (section 3.6).
  • Self-service: a user requests a capability and receives it automatically, with no human in the loop, through a portal, API or CLI. Every tool in domain 3 is a different way to deliver that property.
  • Capability: a thing the platform offers (build automation, observability, secrets, data services), distinct from the interface through which it is offered. Scenario questions often turn on that distinction.

Capability domains the paper enumerates

Web portals; APIs and CLIs; golden path templates; build and test automation; delivery and verification automation; development environments; observability; infrastructure services; data services; messaging; identity and secrets; security services; artifact storage. A "which capability is this" question is cheap to set, so skim the list until each one has an example attached.

Maturity, in one line each

The paper's model runs provisional (ad hoc, individual effort) → operational (a team owns it, usage is requested) → scalable (self-service, documented, versioned) → optimizing (measured, funded as a product, continuously improved). If a scenario describes tickets and a shared spreadsheet, it is operational at best, and the named next step is self-service with documentation.

The white paper, section by section

the exact lists a vocabulary question draws from

The definitions panel compresses the paper; this one keeps its lists intact, in the paper's order and close to its wording, because a scenario answer scores when it uses the same nouns. The paper is CNCF TAG App Delivery's "Platforms White Paper" v1.0; the maturity model came later and is the next panel.

Why platforms: the five benefits

By investing in platforms, the paper says, enterprises can:

  1. Reduce the cognitive load on product teams and thereby accelerate product development and delivery.
  2. Improve reliability and resiliency of products relying on platform capabilities by dedicating experts to configure and manage them.
  3. Accelerate product development and delivery by reusing and sharing platform tools and knowledge across many teams.
  4. Reduce risk of security, regulatory and functional issues by governing platform capabilities and the users, tools and processes around them.
  5. Enable cost-effective and productive use of services from public clouds and other managed offerings by delegating implementations to those providers while maintaining control over user experience.

The paper's own summary of why these accrue: a few platform teams serve many product teams, they consolidate management of common functionality, and they emphasize user interfaces and experiences above all else.

The definition, verbatim

"A platform for cloud-native computing is an integrated collection of capabilities defined and presented according to the needs of the platform's users. It is a cross-cutting layer that ensures a consistent experience for acquiring and integrating typical capabilities and services for a broad set of applications and use cases." Platform teams provide capabilities but should not always implement them; the platform is "the thinnest reasonable layer that provides consistency across provided implementations". The minimal example is a wiki page linking to provisioning procedures.

Platform maturity: five use cases in order

  1. Product developers can provision capabilities on demand and immediately use them to run systems: compute, storage, databases, identities.
  2. Product developers can provision service spaces on demand and use them to run pipelines and tasks, to store artifacts and configuration, and to collect telemetry.
  3. Administrators of third-party software can provision required dependencies like databases on demand and easily install and run that software.
  4. Product developers can provision complete environments from templates combining run-time and development-time services for specific scenarios, such as web development or MLOps.
  5. Product developers and managers can observe functionality, performance and cost of deployed services through automatic instrumentation and standard dashboards.

Attributes of platforms: the seven headings as printed

Platform as a product; User experience; Documentation and onboarding; Self-service; Reduced cognitive load for users; Optional and composable; Secure by default. Two details from the text that get paraphrased wrongly: "User experience" is about meeting users where they are (GUIs, APIs, CLIs, IDEs, portals for different roles), and "Documentation and onboarding" is where the paper introduces the phrase golden path, as "an initial project template and documentation" bundled with a reusable workflow. "Secure by default" is the shortest attribute: compliance and validation "based on rules and standards defined by the organization".

Attributes of platform teams

Platform teams are responsible for the interfaces to and experiences with platform capabilities. Their three jobs:

  1. Research platform user requirements and plan feature roadmap.
  2. Market, evangelize and advocate for the platform's proposed values.
  3. Manage and develop interfaces for using and observing capabilities and services, including portals, APIs, documentation and templates, and CLI tools.

Ways to learn requirements the paper lists: user interviews, hackathons, issue trackers and surveys, and direct observation through observability tools. A platform team "doesn't necessarily run compute, network, storage or other services"; it should rely on externally-provided capabilities as much as possible and build its own only when nothing else offers them.

Challenges, and enabling platform teams

Three challenges: platform teams must treat their platforms like products and develop them together with users; they must carefully choose their priorities and initial partner application teams; they must seek support of enterprise leadership and show impact on value streams. The paper's advice on the first is to include product managers from the start; on the second, to begin with frequently required, undifferentiated capabilities (pipelines, databases, observability) and a few engaged teams who then champion the platform.

Three ways to reduce the load on the platform team itself: seek to build the thinnest viable platform layer over implementations from managed providers; "leverage open source frameworks and toolkits for creating docs, templates and compositions for application team use" (the paper's own wording); make sure platform teams are staffed appropriately for their domain and number of customers.

Measuring success: three categories and their measures

CategoryMeasures the paper names
User satisfaction and productivityactive users and retention (capabilities provisioned, user growth and churn); Net Promoter Score or similar surveys; developer productivity metrics such as those in the SPACE framework
Organizational efficiencylatency from request to fulfillment of a service or capability (a database, a test environment); latency to build and deploy a brand new service into production; time for a new user to submit their first code changes
Product and feature deliverythe DORA four: deployment frequency, lead time for changes, time to restore services after failure, change failure rate

The paper adds that platform teams should use "the smallest viable effort to gather the feedback they need"; surveys and usage analysis may be most valuable at first.

Capabilities: the thirteen domains and the projects the paper attaches

Capability domainExample CNCF / CDF projects (from the paper's table)
Web portals for provisioning and observing capabilitiesBackstage, Skooner, Ortelius
APIs (and CLIs) for automatically provisioning capabilitiesKubernetes, Crossplane, Operator Framework, Helm, KubeVela
"Golden path" templates and docsArtifactHub
Automation for building and testingTekton, Jenkins, Buildpacks, ko, Carvel
Automation for delivering and verifyingArgo, Flux, Keptn, Flagger, OpenFeature
Development environmentsDevfile, Nocalhost, Telepresence, DevSpace
Observability (functionality, performance and costs)OpenTelemetry, Jaeger, Prometheus, Thanos, Fluentd, Grafana, OpenCost
Infrastructure services (compute runtimes, programmable networks, block and volume storage)Kubernetes, Kubevirt, Knative, WasmEdge; CNI, Istio, Cilium, Envoy, Linkerd, CoreDNS; Rook, Longhorn, Etcd
Data services (databases, caches, object stores)TiKV, Vitess, SchemaHero
Messaging and event servicesStrimzi, NATS, gRPC, Knative, Dapr
Identity and secret managementDex, External Secrets, SPIFFE/SPIRE, Teller, cert-manager
Security services (static and runtime analysis, policy enforcement)Falco, In-toto, KubeArmor, OPA, Kyverno, Cloud Custodian
Artifact storage (images, packages, binaries, source)ArtifactHub, Harbor, Distribution, Porter

Every tool on the CNPE list sits in one of these rows: Argo, Flux, Flagger and Tekton in the two automation rows; Crossplane in APIs; Prometheus, Grafana, Jaeger, OpenTelemetry and OpenCost in observability; Istio and Linkerd in infrastructure; OPA, Gatekeeper and Kyverno in security. A "which capability domain does this tool serve" question is answered from this table.

Glossary roles

  • Platform capability providers develop and maintain the capabilities the platform offers; external organizations or internal teams; infrastructure, runtime or supporting services.
  • Platform engineers develop and maintain the interfaces and tools that let applications provision and integrate platform capabilities, "according to the requirements and instructions provided by platform product managers".
  • Platform product managers understand the experience of platform users, build the roadmap that addresses gaps, requirements and opportunities, and manage platform teams in their daily work.
  • Platform teams own the interfaces and experiences (portals, custom APIs, golden path templates); managed by product managers, staffed by platform engineers, later joined by operators, QA, UX, technical writers and developer advocates.
  • Platform users: app developers and operators, data scientists, COTS software operators, information workers; whoever runs software on the platform or uses its capabilities.
  • Thinnest viable platform (TVP), from Team Topologies (Skelton and Pais): "a careful balance between keeping the platform small and ensuring that the platform is helping to accelerate and simplify software delivery for teams building on the platform."
How this gets tested

A scenario describes a team that runs its own Postgres, writes tickets to get namespaces and has no roadmap. Naming the violated attributes (Self-service, Platform as a product), the missing role (a platform product manager), the measure that would expose the problem (latency from request to fulfillment) and the capability domain that should absorb the database (data services, delegated to a provider or operator under a thin interface) is a complete answer. Paraphrases lose the words the marking key is looking for.

The maturity model grid

five aspects × four levels, one line each

The CNCF Platform Engineering Maturity Model (v1.0) answers "how do we plan to build it" where the white paper answered why and what. Each aspect is scored independently; an organization is not "Level 2" as a whole, and the model says outright that the highest level is not a goal in itself because each level costs more funding and time. Lower levels are tactical, higher levels strategic.

Aspect (its question)1 Provisional2 Operational3 Scalable4 Optimizing
Investment (how are staff and funds allocated to platform capabilities?)Voluntary or temporary: tiger teams, hack days, burnoutDedicated team: a central DevOps or DevEx team funded as a cost center, impact unmeasuredAs product: product management and UX roles, a published roadmap, funding by expected value, possibly chargebackEnabled ecosystem: specialists (security, performance) extend the platform without a central backlog
Adoption (why and how do users discover and use platform capabilities?)Erratic: discovery by rumor, each team its own scriptsExtrinsic push: mandates and incentives; fragmented use; office hours as the support modelIntrinsic pull: users choose it for the value; self-serve portals and golden paths; teams will pay via chargebackParticipatory: users contribute fixes and features; contribution processes and internal ambassadors
Interfaces (how do users interact with and consume capabilities?)Custom processes: person-to-person knowledge, manual requestsStandard tooling: documented paved roads and templates that still need expert help; template driftSelf-service solutions: one-click provisioning with consistent interfaces; thin day-2 experienceIntegrated services: capabilities appear inside existing tools (IDE, pipeline, namespace by default) and compose as building blocks
Operations (how are platforms planned, prioritized, developed and maintained?)By request: reactive, no maintenance plan, bespoke upgradesCentrally tracked: a register of services and owners, burndown of upgrades, still manualCentrally enabled: standard practices for new capabilities, documented upgrade processes, continuous delivery of the platform itselfManaged services: automated lifecycle, no user impact, a shared responsibility model
Measurement (what is the process for gathering and incorporating feedback and learning?)Ad hoc: inconsistent surveys, anecdotesConsistent collection: standard channels and instrumentation, weak link to the roadmapInsights: outcomes chosen first, metrics chosen to track them, a product manager owns the loopQuantitative and qualitative: data democratized, leading indicators, awareness of Goodhart's law

Reading a scenario against the grid: "a central team maintains Terraform modules and CRDs that teams copy and customize" is Interfaces level 2 (standard tooling, template drift). "An API provisions a database and returns a connection string and dashboard" is level 3. "Every new project gets a pipeline space and a namespace automatically, and an OIDC proxy in front" is level 4. The model's own examples use exactly these cases. The model also states that maturing one aspect may require a minimum level in another; a self-service interface (Interfaces 3) with volunteer staffing (Investment 1) is the classic imbalance.

Mental model

White paper for nouns (attributes, capability domains, roles, measures), maturity model for verbs (what to change next, in which aspect). A scenario that asks "what is wrong" wants the paper's vocabulary; one that asks "what should they do next" wants the adjacent level in the aspect that is lagging.

What "API as contract" means when you design one

the part that turns into YAML in 3.2–3.6

"Designing platform APIs" cashes out in domain 3 as five decisions. Hold examples/crossplane/xrd.yaml against this list; it is the lab's concrete instance of every point.

  1. Abstraction level. A developer asks for an AppEnvironment with a team name and a quota, not for a Namespace plus a LimitRange plus two NetworkPolicies. Choose the noun your user already has in their head. Too thin and you have added a layer without removing work; too thick and every request needs an escape hatch you did not build.
  2. Validate at admission. Enums, ranges, patterns, CEL rules with messages a human can act on. Mistakes should fail in the terminal within a second, not in a controller log an hour later.
  3. Report status. Conditions, phases and useful printer columns so consumers can self-diagnose. A platform API without a meaningful READY column is user-hostile, and the support burden lands on you.
  4. Version deliberately. v1alpha1 means you may break it; once someone depends on it, a new served version and a conversion story are the price of changing your mind (section 3.2).
  5. Document in the schema. Field descriptions become kubectl explain output. That is the documentation your users will actually read, because it is available where they already work.
Design test

Write the ten-line YAML your user will type before you write the schema. If those ten lines contain a field whose value the user must look up in your implementation, the abstraction is leaking. examples/crossplane/xr.yaml is exactly this exercise, done: a team name and two quota numbers.

Measuring a platform

CategoryMeasuresWhere the data lives
User satisfaction and productivitysurveys, adoption, time to first contributionoutside the cluster: ask people
Organizational efficiencylatency from request to fulfillment of a capabilitytimestamps on your own CRs (section 4.5)
Product delivery (DORA)deployment frequency, lead time for changes, time to restore, change failure ratethe CD tool's metrics (section 4.5)

Section 4.5 turns the bottom two rows into PromQL against Argo CD's metrics, which is the closest thing to an exam task you can do with this paper.

Anti-patterns worth naming

  • The platform as a gate. Self-service in the docs, a ticket queue in practice.
  • The leaky abstraction. The interface exposes the implementation, so every provider change breaks users.
  • The mandated platform. Adoption enforced rather than earned; teams route around it and you learn nothing about why.
  • The rebuilt wheel. A homegrown Postgres operator, a bespoke CI engine, a custom secrets store: effort spent below the line where your users care, at the cost of the interface work only you can do.

Exercises

paper and pen, mostly

Take the lab's tool list from make help and assign every layer to one or more of the paper's capability domains. Then mark which interfaces each capability is exposed through (portal, API, CLI) in this lab. Where the lab has no coverage (dev environments, messaging), say so; knowing where the gaps are is part of the exercise.

verify: check yourself against the table in the white paper's "Capabilities of platforms" section.

Write, exam style: why should a platform team not build its own Postgres operator?

verify: if your answer touches TVP, delegation to existing capability providers, and where the platform team's effort should go instead (the interface, the golden path, the integration), you have absorbed the paper.

Invent one platform API your own organization would use: a cache, a queue, a scheduled job, an internal endpoint. Write the ten-line manifest a developer would type. Then list, underneath, every real Kubernetes object it must expand into, and every field you deliberately did not expose.

verify: the not-exposed list is longer than the exposed one, and you can defend each omission with either "we choose it for them" or "they can escape-hatch to the underlying object". That defense is the platform product-management skill the paper is asking for.

The command prints the lab's inventory; there is nothing to grade in the output itself.

Take each tool make help names, put it in exactly one row of the capability table above, and then say which of the five maturity aspects the lab is weakest on. A tool you cannot place is either the wrong tool or a row you have not understood.

make -C "$REPO_ROOT" --no-print-directory help
outputcaptured 2026-09-13
$ make -C "$REPO_ROOT" --no-print-directory help
  help           Show this help
  host           Kernel limits, docker, thermal advice (run once, needs sudo)
  tools          Install or upgrade every CLI into ~/.local/bin
  refresh        Everything 'tools' does, plus updating an existing Gitea container
  up             Create the cluster: kind + Cilium + LB + metrics + VPA + registry
  gitea          Local git server, seeded repos, CoreDNS entry
  gitops         Argo CD, Argo Rollouts, Argo Workflows, Flux        [domain 2]
  cicd           Tekton Pipelines/Triggers/Dashboard, Trivy Operator [domain 2]
  api            Crossplane, CloudNativePG, kro, kubebuilder hints   [domain 3]
  obs            Prometheus, Grafana, OTel, Jaeger, Loki+Alloy, OpenCost    [domain 4]
  sec            Kyverno, Gatekeeper, sealed/external secrets, PSS   [domain 5]
  spire          SPIFFE/SPIRE workload identity                   [domain 5]
  mesh           Second cluster + Istio (MESH=linkerd works) + Flagger [domain 5]
  portal         Scaffold Backstage on the host                      [domain 3]
  core           Minimum useful lab (~4 GB)
  full           Everything on the main cluster (~14 GB)
  down-gitops    Remove Argo CD/Rollouts/Workflows/Flux from the cluster   [layer teardown]
  down-cicd      Remove Tekton and the Trivy operator                      [layer teardown]
  down-api       Remove Crossplane, CloudNativePG and kro                  [layer teardown]
  down-obs       Remove Prometheus, Grafana, Loki, Jaeger, OTel, OpenCost  [layer teardown]
  down-sec       Remove Kyverno, Gatekeeper, sealed/external secrets       [layer teardown]
  down-spire     Remove SPIFFE/SPIRE                                       [layer teardown]
  down-mesh      Delete the second cluster (Istio/Linkerd + Flagger)       [layer teardown]
  fix-cp-metrics Expose control-plane metrics on an existing cluster (Prometheus targets)
  validate       Functionally verify every layer (FAST=1 to skip probes)
  forward        Start a background port-forward for every UI
  forward-stop   Kill all port-forwards started by 'make forward'
  grade          Run a mock exam's grading block from its page (EXAM=1|2)
  study          Open the CNPE study console (curriculum) in a browser
  fonts          Re-cut assets/fonts from tools/fonts-src to the console's charset (needs fonttools, brotli)
  site           Stage the study console exactly as Pages publishes it, into ./_site
  browser        Browser-check the staged console like CI does (AREAS=sync,theme runs a subset)
  worker         Test the progress-sync Worker against a stub D1 (plain node, no deps)
  merge          Test the progress merge over plain objects (plain node, no deps; see also 'sim')
  sim            Test the quest's command interpreter over every battle scenario (plain node, no deps)
  ts             Compile curriculum/src/**/*.ts (the quest, merge.js, syntax.js) into curriculum/assets/ (needs typescript)
  ts-check       Fail if the committed compiled scripts no longer match their TypeScript
  syntax         Test the command-block highlighting over plain strings (plain node, no deps)
  typecheck      Type-check the console's JS via JSDoc and its TypeScript sources (needs typescript, @types/node, playwright resolvable)
  urls           Every UI, its URL/port-forward, and credentials
  status         Clusters, endpoints, unhealthy pods, host load
  break          Inject a random fault, then diagnose it under time pressure (DOMAIN=/FAULT= to scope)
  break-fix      Auto-diagnose and repair whatever 'make break' injected
  down           Delete both clusters (keeps git history + registry)
  nuke           Delete everything including Gitea data
  break-answer   Reveal the last injected fault
verify: every layer in the output lands in one row, you can defend the placements you found hard, and you can name the one capability the lab has no tool for at all. That gap is a legitimate exam answer.

Self-check

vocabulary, out loud
Define a platform in the paper's own terms.

An integrated collection of capabilities, defined and presented according to the needs of its users. The emphasis on users and on presentation (interfaces) is what distinguishes a platform from a pile of infrastructure.

What is the Thinnest Viable Platform, and what does it imply about build-vs-adopt?

The smallest layer that provides consistency and accelerates delivery. It implies adopting existing capabilities wherever they exist and spending your effort on interfaces, integration and golden paths: the parts nobody can buy for your organization.

A team can request a namespace through a form, and someone approves it within a day. Is that self-service?

No. The paper's definition requires fulfillment without a human in the loop. A form plus an approver is an interface over a ticket queue: operational maturity, not scalable. The fix is automated fulfillment with policy encoded in the API, and approval only where regulation demands it.

Name the DORA four and one platform-specific measure the paper adds.

Deployment frequency, lead time for changes, time to restore service, change failure rate. The paper adds measures like latency from request to fulfillment of a capability, time to first contribution, and adoption, all platform-specific because they measure the platform, not the applications on it.

Which platform attribute does a mandatory, non-composable platform violate, and why does it matter practically?

"Optional and composable". Practically: mandated platforms hide their own failures, because teams route around them quietly instead of filing the feedback that would improve them. Optionality is a feedback mechanism, not generosity.

List the five benefits the white paper says enterprises get from investing in platforms.

Reduce cognitive load on product teams and so accelerate delivery; improve reliability and resiliency by dedicating experts to platform capabilities; accelerate delivery by reusing and sharing tools and knowledge across teams; reduce security, regulatory and functional risk by governing capabilities, users, tools and processes; enable cost-effective use of public cloud and managed services by delegating implementation while keeping control of the user experience.

Name the five aspects and four levels of the Platform Engineering Maturity Model.

Aspects: Investment, Adoption, Interfaces, Operations, Measurement. Levels: Provisional, Operational, Scalable, Optimizing. Each aspect is scored on its own, and the model says the top level is not a goal in itself because each level demands more funding and time.

Which glossary role decides what the platform should do, and which builds it?

The platform product manager understands users, builds the roadmap and manages the team; the platform engineer develops and maintains interfaces and tools according to the product manager's requirements. Capability providers (internal teams or vendors) maintain the underlying capabilities; platform users are whoever runs software on the platform.

An organization mandates the platform for production releases and runs office hours because users cannot find features. Which aspect and level?

Adoption at level 2, Operational: extrinsic push (mandates and incentives), fragmented use, heavy reliance on help-desk style support. The adjacent target is level 3, intrinsic pull: documentation, ergonomic interfaces, shared roadmaps and feedback forums that make teams choose the platform on merit.

Docs to know your way around

study time, not exam time
  • tag-app-delivery.cncf.io/whitepapers/platforms/: the paper itself, plus its glossary (platform engineer vs platform product manager definitions have appeared in practice questions).
  • dora.dev: the four keys, one page.
  • Team Topologies: where "platform team" and "cognitive load" as terms of art come from; the paper leans on both.
  • tag-app-delivery.cncf.io/whitepapers/platform-eng-maturity-model: the model table and the per-level characteristics and example scenarios; the source is github.com/cncf/tag-app-delivery under platforms-maturity-model/v1.
  • github.com/cncf/tag-app-delivery, platforms-whitepaper/v1/index.md: the paper's markdown source, easier to search than the rendered site for exact wording of the attributes and the capability table.