This competency is listed explicitly in the official PDF and is the least practiced in the domain, because almost no home lab turns on API audit logging. This lab does, at cluster build.

needsmake up cicdmake sec

Orientation

competency 5.3 · audit trails and policy compliance

Three artefacts, three questions. The audit log answers "who did what to what, and did it work". SBOMs answer "what is inside the thing we shipped". Compliance reports answer "which controls are we failing right now". Being able to produce each on demand is the competency.

API server audit logs

policy · levels · stages · the event schema

The pipeline: a policy file tells the API server what to record and at what level; a backend (here, a log file on the control-plane node; in production often a webhook to a SIEM) receives one JSON event per stage of each request.

LevelRecordsUse for
Nonenothingnoise: leases, events, health endpoints
Metadatawho, what, when, verb, response codethe sensible catch-all
Request+ the request bodyRBAC and policy changes
RequestResponse+ the response bodythe highest-value resources only; it is enormous

Stages: RequestReceived, ResponseStarted (long-running watches), ResponseComplete, Panic. Most analysis uses ResponseComplete. Policy rules match in order, first hit wins (the same mental model as Alertmanager routes), so the file reads as: drop the noise, raise the sensitive, catch everything else at Metadata. Read kind/audit-policy.yaml in this repo and you will see exactly that.

Fields that answer real questions

FieldAnswers
user.username, user.groupswho (a person, or system:serviceaccount:ns:name)
verb, objectRef.{resource,namespace,name}did what, to what
responseStatus.codedid it work (403 = denied, 201 = created)
sourceIPs, userAgentfrom where, with which client
annotations["authorization.k8s.io/decision"]allow or forbid (the reason sits beside it in …/reason)
annotations["pod-security.kubernetes.io/audit-violations"]PSS violations from section 5.3, recorded quietly
stageTimestamp, auditIDwhen, and how to correlate the stages of one request
One operational failure the repo already documents

The apiserver is started with --audit-policy-file pointing into a mounted directory. If that file is missing, the API server does not start at all. An audit misconfiguration can therefore present as a completely dead cluster. It is the least intuitive control-plane failure in this curriculum.

What audit is and is not

It records API requests. It does not record what happened inside a container, what a pod did on the network, or anything that bypassed the API server. Pair it with Falco-style runtime detection, network flow logs (Hubble), and image scanning for the other layers, and say so if a scenario asks for "a complete audit trail".

Audit policy and backends, field by field

rule fields, ordering, the flags, what a good policy records

Rule fields

A Policy (audit.k8s.io/v1) has rules[] and a top-level omitStages. Each rule sets a level and narrows by any of users, userGroups, verbs, resources[] (each with group, resources such as pods or pods/exec, and optional resourceNames), namespaces, nonResourceURLs (/healthz*, /metrics, /version), plus a per-rule omitStages and omitManagedFields. A request is matched against rules top to bottom and the first match sets the level, so the file is ordered: drop noise, raise sensitive things, catch the rest.

apiVersion: audit.k8s.io/v1
kind: Policy
omitStages: [RequestReceived]          # one event per request, not two
rules:
- level: None                          # noise first
  users: ["system:kube-proxy"]
  verbs: ["watch"]
  resources: [{ group: "", resources: ["endpoints", "services", "endpointslices"] }]
- level: None
  nonResourceURLs: ["/healthz*", "/readyz*", "/livez*", "/metrics", "/version"]
- level: None
  resources: [{ group: "coordination.k8s.io", resources: ["leases"] }]
- level: Metadata                      # never the body: it holds the values
  resources: [{ group: "", resources: ["secrets", "configmaps"] },
              { group: "authentication.k8s.io", resources: ["tokenreviews"] }]
- level: Metadata
  resources: [{ group: "", resources: ["pods/exec", "pods/attach", "pods/portforward"] }]
- level: RequestResponse               # who changed access and policy, with the diff
  resources: [{ group: "rbac.authorization.k8s.io" },
              { group: "kyverno.io" }, { group: "policies.kyverno.io" },
              { group: "constraints.gatekeeper.sh" }, { group: "admissionregistration.k8s.io" }]
- level: Metadata                      # everything else
  omitStages: [RequestReceived]
  • Why those levels. Request/RequestResponse on Secrets writes secret values into the log; Metadata still records who read which Secret and whether it succeeded. pods/exec at Metadata answers "who shelled into production". RBAC and policy objects at RequestResponse give you the before/after for a permissions change. Request alone suits high-volume writes where the response body is redundant (a Deployment update).
  • Stages. RequestReceived fires before authorization, so it records requests that are later denied; most policies omit it and rely on ResponseComplete, which carries responseStatus.code. ResponseStarted exists only for long-running requests (watch, exec, port-forward). Panic is rare. Both events of one request share auditID.
  • Event fields you will grep for beyond the table above: level, stage, requestURI, impersonatedUser (present when --as was used; user is then the real caller), requestObject/responseObject at the higher levels, requestReceivedTimestamp. omitManagedFields: true strips the noisy managedFields block out of both bodies.

Backends and their flags

FlagBackendMeans
--audit-policy-fileboththe Policy above; if absent, nothing is logged; if the path is wrong, the API server does not start
--audit-log-pathlogJSON lines file (- for stdout); the flag that enables the log backend
--audit-log-maxage / -maxbackup / -maxsizelogrotation: days to keep, number of files, MB per file
--audit-log-formatlogjson (default) or legacy
--audit-log-modelogblocking (default), blocking-strict (a logging failure at RequestReceived fails the request), batch
--audit-webhook-config-filewebhooka kubeconfig naming the remote sink (a SIEM, Falco's k8saudit plugin, a Fluent Bit input)
--audit-webhook-modewebhookbatch (default), blocking, blocking-strict
--audit-webhook-batch-buffer-size / -max-size / -max-waitwebhook10000 events buffered, 400 per batch, 30s max wait; overflow drops events
--audit-webhook-initial-backoffwebhookretry delay after the first failed delivery (exponential after)

Static-pod control planes (kubeadm, kind) need the policy file and the log directory mounted as hostPath volumes into the kube-apiserver manifest; a missing mount is the "dead cluster" failure the callout above describes. On managed control planes you cannot set any of this; you consume the provider's audit stream instead, and the exam's equivalent question becomes "what would you look for in it".

Annotations other controls leave in the audit event

  • authorization.k8s.io/decision and authorization.k8s.io/reason: RBAC's verdict and the rule that produced it.
  • pod-security.kubernetes.io/audit-violations: the PSS audit label's findings, per pod.
  • validation.policy.admission.k8s.io/validation_failure: a ValidatingAdmissionPolicy binding with validationActions: [Audit] records here; the policy's own auditAnnotations add custom keys.
  • mutation.webhook.admission.k8s.io/round_0_index_N and patch.webhook.admission.k8s.io/...: which mutating webhooks touched the object and what they changed, when the level is Request or higher.
How this gets tested

Two kinds of task. Write or fix a policy: "record who reads Secrets without logging their contents, and drop health checks" is a Metadata rule on secrets plus a None rule on nonResourceURLs, in that order relative to the catch-all. Query the log: "which ServiceAccount deleted the Deployment" is a jq filter on verb, objectRef.resource, objectRef.name and user.username. Both are graded on the output, so run the query after the change.

Supply-chain paper trail: Trivy Operator

compliance as queryable API objects

The operator rescans continuously and materializes results as CRs, which is the platform move: compliance as queryable API objects rather than PDF attachments. That means kubectl, RBAC, GitOps and dashboards all work on your compliance data for free.

KindHolds
vulnerabilityreportsCVEs per workload image, with severity counts and fix versions
sbomreportsCycloneDX component inventory per image
configauditreportsmisconfigurations in the workload spec itself
exposedsecretreportscredentials found baked into images
rbacassessmentreportsrisky RBAC rules
clusterinfraassessmentreportsper-node control-plane and kubelet checks
clustercompliancereportsCIS, NSA and PSS rollups built from the above

make validate demands minimum counts of these, so a fresh cicd layer gives you a populated dataset to practice queries on.

SBOM vocabulary, because the words get tested

  • CycloneDX and SPDX are the two standard formats; Trivy emits both, and "which format" is a compatibility question, not a quality one.
  • An SBOM lists components and versions. It is not a vulnerability report; you join it against a CVE database, which is why an SBOM generated at build time stays useful after new CVEs are published.
  • The one-sentence purpose: when the next log4shell drops, you query your SBOMs for the package instead of rescanning the world.
  • Neighbors: provenance/attestation (how it was built; see section 5.6), signatures (who vouches for it). Three different documents, one trust story.

Runtime detection, SBOM tooling and compliance evidence

Falco, syft/grype/trivy, the Trivy Operator's report kinds, kube-bench, in-toto

Falco fills the gap the API audit leaves

The API server never sees a shell spawned inside a container, a file written under /etc, or an outbound connection to an unexpected host. Falco does: a kernel driver (modern eBPF by default) or a plugin (the k8saudit plugin consumes the audit webhook stream above, so one Falco can watch both layers) produces events, and rules match them. A rule has rule, desc, condition (a filter over event fields such as proc.name, container.image.repository, k8s.ns.name, fd.sip), output (a format string), priority (EMERGENCY to DEBUG) and tags; macro and list entries factor conditions. Rules ship in maturity tiers (stable rules are bundled; incubating and sandbox are opt-in) and are customized by override files rather than edits. Alerts go to stdout, a file, gRPC or HTTP; Falcosidekick fans them out to Slack, Loki, Alertmanager and the like. "Detect a shell in a production container" or "alert on writes below /etc" is a Falco question; "who created that pod" is an audit-log question.

SBOM and scanner tooling, one line each

  • Formats. SPDX (ISO/IEC 5962, tag-value or JSON) and CycloneDX (OWASP, JSON or XML); both list components with versions, licenses and hashes. trivy image --format cyclonedx|spdx-json, syft <image> -o cyclonedx-json. An SBOM is input to a scanner (trivy sbom sbom.json, grype sbom:sbom.json), which is why generating it at build time and rescanning later works.
  • Attaching it to the image. cosign attest --predicate sbom.json --type cyclonedx <image@digest> stores it as a signed in-toto attestation next to the image; admission policies can then require it (5.6).
  • VEX (Vulnerability Exploitability eXchange, OpenVEX or CSAF) is the vendor's or your statement that a CVE does not affect a product as shipped; trivy --vex <file> or --vex repo suppresses those findings with a recorded justification, which is the auditable alternative to a growing ignore file.

The Trivy Operator's report kinds

Kind (aquasecurity.github.io/v1alpha1)ScopeHolds
VulnerabilityReportnamespaced, one per container of a workload; named <kind>-<name>-<container>, owned by the ReplicaSet or controller so it is garbage-collected with itreport.summary.criticalCount|highCount|..., report.vulnerabilities[] with vulnerabilityID, installedVersion, fixedVersion, severity, primaryLink
SbomReport / ClusterSbomReportper containera CycloneDX BOM under report.components
ConfigAuditReport / ClusterConfigAuditReportper workload / per cluster objectmisconfiguration checks (AVD-KSV-* ids): privileged, missing limits, root user
ExposedSecretReportper containercredentials found in image layers
RbacAssessmentReport / ClusterRbacAssessmentReportper Role / ClusterRolerisky rules (wildcards, secrets access, escalate)
InfraAssessmentReport / ClusterInfraAssessmentReportper node componentkubelet, API server and etcd checks, the kube-bench territory
ClusterComplianceReportclusterspec.compliance.id in k8s-cis-1.23, k8s-nsa-1.0, k8s-pss-baseline-0.1, k8s-pss-restricted-0.1; spec.cron and spec.reportType: summary|all; status.summary.passCount/failCount and per-control results mapping to the check ids above

Reports carry labels trivy-operator.resource.kind, trivy-operator.resource.name, trivy-operator.container.name, so -l trivy-operator.resource.name=payments narrows a query without parsing names. Scans run as short-lived jobs in the operator's namespace (or the workload's, per OPERATOR_SCAN_JOB_... settings), which is why a NetworkPolicy or a registry pull secret can stop reports from appearing: the scan job is a pod like any other. Rescans happen on a TTL (OPERATOR_VULNERABILITY_SCANNER_REPORT_TTL) and when the image digest changes.

Benchmarks and attestations

  • CIS Kubernetes Benchmark checks the control plane and node configuration (API server flags, kubelet flags, file permissions) and is versioned separately from Kubernetes. kube-bench runs it as a Job with a target set (master, node, etcd, policies) and prints PASS/FAIL/WARN with the control number (for example 1.2.x API server); --benchmark picks the CIS version and managed-cluster variants (eks, gke, aks) exist. The Trivy Operator's k8s-cis compliance report covers the same ground as CRDs.
  • in-toto attestation is the envelope format everything else uses: a DSSE envelope (payloadType: application/vnd.in-toto+json, signatures) around a Statement (_type: https://in-toto.io/Statement/v1, subject[] with name and digest.sha256, predicateType, predicate). SLSA provenance, SBOMs, vulnerability scan results and VEX are all predicate types. "Attestation" in a task means one of these, signed, attached to an image digest.
  • An evidence pipeline, in one sentence per stage: build produces the image and its SBOM; scan gates on severity and writes the scan result as an attestation; sign the digest; store SBOM and provenance attestations in the registry beside the image; admission requires signature and attestations; the operator rescans continuously and the compliance reports roll it up. Each stage leaves a queryable artefact, which is what "generating audit trails" means for the supply chain.
Trap

A ClusterComplianceReport is only as current as its spec.cron and the underlying reports; a freshly changed cluster shows yesterday's numbers until the next run. Read status.updateTimestamp before quoting a pass count, and expect graders to check the report after your fix, not the fix itself.

Exercises

tick the dot when its check passes

It lives on the control-plane node, reachable through docker:

docker exec cnpe-control-plane tail -1 /var/log/kubernetes/audit.log | jq .
outputcaptured 2026-08-26
$ docker exec cnpe-control-plane tail -1 /var/log/kubernetes/audit.log | jq .
{
  "kind": "Event",
  "apiVersion": "audit.k8s.io/v1",
  "level": "Metadata",
  "auditID": "90a334b1-0a03-420e-a1bc-96d14fb79e67",
  "stage": "ResponseComplete",
  "requestURI": "/apis/wgpolicyk8s.io/v1alpha2/namespaces/spire/policyreports/6c7a1017-58fc-422a-8ba5-eed66dba4c27",
  "verb": "get",
  "user": {
    "username": "system:serviceaccount:kyverno:kyverno-reports-controller",
    "uid": "94675045-58e9-4c7d-a50a-0745ac4b4687",
    "groups": [
      "system:serviceaccounts",
      "system:serviceaccounts:kyverno",
      "system:authenticated"
    ],
    "extra": {
      "authentication.kubernetes.io/credential-id": [
        "JTI=48fd7eed-ebbd-47c4-b311-ef850241a1ec"
      ],
      "authentication.kubernetes.io/node-name": [
        "cnpe-worker2"
      ],
      "authentication.kubernetes.io/node-uid": [
        "3e493a08-e750-4c41-b117-9792d6b3d3e2"
      ],
      "authentication.kubernetes.io/pod-name": [
        "kyverno-reports-controller-7447456448-trwwg"
      ],
      "authentication.kubernetes.io/pod-uid": [
        "c2953876-66ee-4bd1-9d50-2ec02831230b"
      ]
    }
  },
  "sourceIPs": [
    "172.18.0.2"
  ],
  "userAgent": "reports-controller/v0.0.0 (linux/amd64) kubernetes/$Format",
  "objectRef": {
    "resource": "policyreports",
    "namespace": "spire",
    "name": "6c7a1017-58fc-422a-8ba5-eed66dba4c27",
    "apiGroup": "wgpolicyk8s.io",
    "apiVersion": "v1alpha2"
  },
  "responseStatus": {
    "metadata": {},
    "code": 200
  },
  "requestReceivedTimestamp": "2026-08-27T02:44:31.144541Z",
  "stageTimestamp": "2026-08-27T02:44:31.157282Z",
  "annotations": {
    "authorization.k8s.io/decision": "allow",
    "authorization.k8s.io/reason": "RBAC: allowed by ClusterRoleBinding \"kyverno:reports-controller\" of ClusterRole \"kyverno:reports-controller\" to ServiceAccount \"kyverno-reports-controller/kyverno\""
  }
}

Now answer three questions an auditor would ask, each with one jq line: who read any secret today (select(.objectRef.resource=="secrets" and .verb=="get"), print user and name); every delete that succeeded (select(.verb=="delete" and .responseStatus.code<300)); all denied requests (select(.annotations["authorization.k8s.io/decision"]=="forbid")). Then close the loop end to end: do something distinctive (kubectl -n team-a delete pod bare --ignore-not-found, or a kubectl auth can-i as dev-a from 5.1) and find your own action in the log within seconds.

verify: an audit trail you have personally queried for your own fingerprints is one you can be examined on.

With 5.3 done, grep the audit log for pod-security.kubernetes.io in annotations.

verify: the lazy pod's restricted violations were recorded by the namespace's audit: restricted label, timestamped and attributed. That one log line ties sections 5.2, 5.3 and 5.4 together, which is roughly how a real compliance program works.

Not "are there reports" but questions with answers:

kubectl get vulnerabilityreports -A -o json | jq -r '
  .items[] | [.metadata.namespace, .metadata.name,
  (.report.summary.criticalCount|tostring), (.report.summary.highCount|tostring)] | @tsv' | sort -t$'\t' -k3 -rn | head
outputcaptured 2026-08-26
$ kubectl get vulnerabilityreports -A -o json | jq -r '
  .items[] | [.metadata.namespace, .metadata.name,
  (.report.summary.criticalCount|tostring), (.report.summary.highCount|tostring)] | @tsv' | sort -t$'\t' -k3 -rn | head
default	pod-build-and-scan-skfx6-build-pod-step-build-and-push	25	233
default	pod-build-and-scan-skfx6-clone-pod-step-clone	19	113
tekton-pipelines-resolvers	replicaset-79d56b68cd	6	24
default	replicaset-example-7c5474d4c8-prometheus-example-app	4	37
default	pod-build-and-scan-skfx6-clone-pod-place-scripts	4	45
default	pod-build-and-scan-skfx6-build-pod-place-scripts	4	45
tekton-pipelines	replicaset-tekton-triggers-webhook-5ddd9c4cbd-webhook	3	10
tekton-pipelines	replicaset-tekton-pipelines-webhook-5568b6-webhook	3	10
tekton-pipelines	replicaset-7c789d7fd9	3	10
tekton-pipelines	replicaset-6985f46fcc	3	15

Pick the top image and drill in: which CVE, which package, is there a fix version (.report.vulnerabilities[] | select(.severity=="CRITICAL")).

verify: a ranked worst-offenders list, and one actionable package bump identified. That triage, from fleet view to one action, is the whole vulnerability-management job in miniature.

Fleet-side: kubectl get sbomreports -A | head, then extract one and count its components (kubectl get sbomreport <name> -n <ns> -o jsonpath='{.report.components.components}' | jq length). Artifact-side, the pipeline angle from 2.4: trivy image --format cyclonedx --output sbom.json ghcr.io/nginxinc/nginx-unprivileged:1.27-alpine and confirm both SBOMs speak the same CycloneDX.

verify: you can say what an SBOM is for in one sentence (when the next log4shell drops, you query your SBOMs for the package instead of rescanning the world).

kubectl get clustercompliancereports (CIS, NSA, PSS variants), then one in detail: kubectl get clustercompliancereport cis -o jsonpath='{.status.summary}', and find one failing control's ID and description in the full output.

verify: you can trace a failed control to the check behind it and say which earlier curriculum section would fix it. Most CIS findings in this lab trace back to sections 5.1–5.3.

An SBOM is only useful if you can scan it and compare it. Produce one with each tool, feed it back to the scanner, and then suppress a finding with VEX so the exception is a document rather than a flag.

trivy image --format spdx-json --output /tmp/sbom-trivy.json ghcr.io/stefanprodan/podinfo:6.7.1
jq '{name: .name, packages: (.packages | length)}' /tmp/sbom-trivy.json
mkdir -p "$HOME/.local/bin"; export PATH="$HOME/.local/bin:$PATH"
command -v syft >/dev/null || curl -sSfL https://raw.githubusercontent.com/anchore/syft/main/install.sh | sh -s -- -b "$HOME/.local/bin"
syft version | head -2
syft ghcr.io/stefanprodan/podinfo:6.7.1 -o spdx-json=/tmp/sbom-syft.json
jq '{name: .name, packages: (.packages | length)}' /tmp/sbom-syft.json
jq -r '.packages[].name' /tmp/sbom-trivy.json | sort -u > /tmp/p-trivy.txt
jq -r '.packages[].name' /tmp/sbom-syft.json | sort -u > /tmp/p-syft.txt
wc -l /tmp/p-trivy.txt /tmp/p-syft.txt
comm -3 /tmp/p-trivy.txt /tmp/p-syft.txt | head -10
trivy sbom /tmp/sbom-trivy.json --severity HIGH,CRITICAL | head -30
trivy image --severity HIGH,CRITICAL --format json ghcr.io/stefanprodan/podinfo:6.7.1 | jq -r '[.Results[].Vulnerabilities // [] | .[].VulnerabilityID] | unique | .[0:3][]'
CVE=$(trivy image --severity HIGH,CRITICAL --format json ghcr.io/stefanprodan/podinfo:6.7.1 | jq -r '[.Results[].Vulnerabilities // [] | .[].VulnerabilityID] | unique | .[0]')
echo "suppressing $CVE"
cat > /tmp/vex.json <<EOF
{
  "@context": "https://openvex.dev/ns/v0.2.0",
  "@id": "https://lab.local/vex/podinfo",
  "author": "platform team",
  "statements": [
    {
      "vulnerability": { "name": "${CVE}" },
      "products": [{ "@id": "pkg:oci/podinfo" }],
      "status": "not_affected",
      "justification": "vulnerable_code_not_in_execute_path"
    }
  ]
}
EOF
trivy image --vex /tmp/vex.json --severity HIGH,CRITICAL ghcr.io/stefanprodan/podinfo:6.7.1 | head -20
outputcaptured 2026-09-12
$ trivy image --format spdx-json --output /tmp/sbom-trivy.json ghcr.io/stefanprodan/podinfo:6.7.1
2026-09-13T13:01:03-04:00	INFO	"--format spdx-json" disables security scanning. Specify "--scanners vuln" explicitly if you want to include vulnerabilities in the "spdx-json" report.
2026-09-13T13:01:03-04:00	INFO	Detected OS	family="alpine" version="3.20.3"
2026-09-13T13:01:03-04:00	INFO	Number of language-specific files	num=2

📣 Notices:
  - Version 0.74.0 of Trivy is now available, current version is 0.73.0

To suppress version checks, run Trivy scans with the --skip-version-check flag
$ jq '{name: .name, packages: (.packages | length)}' /tmp/sbom-trivy.json
{
  "name": "ghcr.io/stefanprodan/podinfo:6.7.1",
  "packages": 115
}
$ mkdir -p "$HOME/.local/bin"; export PATH="$HOME/.local/bin:$PATH"
$ command -v syft >/dev/null || curl -sSfL https://raw.githubusercontent.com/anchore/syft/main/install.sh | sh -s -- -b "$HOME/.local/bin"
$ syft version | head -2
Application:   syft
Version:       1.51.1
$ syft ghcr.io/stefanprodan/podinfo:6.7.1 -o spdx-json=/tmp/sbom-syft.json
$ jq '{name: .name, packages: (.packages | length)}' /tmp/sbom-syft.json
{
  "name": "ghcr.io/stefanprodan/podinfo",
  "packages": 112
}
$ jq -r '.packages[].name' /tmp/sbom-trivy.json | sort -u > /tmp/p-trivy.txt
$ jq -r '.packages[].name' /tmp/sbom-syft.json | sort -u > /tmp/p-syft.txt
$ wc -l /tmp/p-trivy.txt /tmp/p-syft.txt
 103 /tmp/p-trivy.txt
 100 /tmp/p-syft.txt
 203 total
$ comm -3 /tmp/p-trivy.txt /tmp/p-syft.txt | head -10
alpine
	ghcr.io/stefanprodan/podinfo
ghcr.io/stefanprodan/podinfo:6.7.1
home/app/podinfo
usr/local/bin/podcli
$ trivy sbom /tmp/sbom-trivy.json --severity HIGH,CRITICAL | head -30
2026-09-13T13:01:08-04:00	INFO	[vuln] Vulnerability scanning is enabled
2026-09-13T13:01:08-04:00	INFO	Detected SBOM format	format="spdx-json"
2026-09-13T13:01:08-04:00	INFO	Detected OS	family="alpine" version="3.20.3"
2026-09-13T13:01:08-04:00	INFO	[alpine] Detecting vulnerabilities...	os_version="3.20" repository="" pkg_num=27
2026-09-13T13:01:08-04:00	INFO	Number of language-specific files	num=2
2026-09-13T13:01:08-04:00	INFO	[gobinary] Detecting vulnerabilities...
2026-09-13T13:01:08-04:00	WARN	Using severities from other vendors for some vulnerabilities. Read https://trivy.dev/docs/v0.73/guide/scanner/vulnerability#severity-selection for details.
2026-09-13T13:01:08-04:00	WARN	This OS version is no longer supported by the distribution	family="alpine" version="3.20.3"
2026-09-13T13:01:08-04:00	WARN	The vulnerability detection may be insufficient because security updates are not provided

Report Summary

┌──────────────────────────────────────┬──────────┬─────────────────┐
│                Target                │   Type   │ Vulnerabilities │
├──────────────────────────────────────┼──────────┼─────────────────┤
│ /tmp/sbom-trivy.json (alpine 3.20.3) │  alpine  │       21        │
├──────────────────────────────────────┼──────────┼─────────────────┤
│ home/app/podinfo                     │ gobinary │       36        │
├──────────────────────────────────────┼──────────┼─────────────────┤
│ usr/local/bin/podcli                 │ gobinary │       33        │
└──────────────────────────────────────┴──────────┴─────────────────┘
Legend:
- '-': Not scanned
- '0': Clean (no security findings detected)


/tmp/sbom-trivy.json (alpine 3.20.3)
====================================
Total: 21 (HIGH: 19, CRITICAL: 2)

┌────────────┬────────────────┬──────────┬────────┬───────────────────┬───────────────┬──────────────────────────────────────────────────────────────┐
│  Library   │ Vulnerability  │ Severity │ Status │ Installed Version │ Fixed Version │                            Title                             │
├────────────┼────────────────┼──────────┼────────┼───────────────────┼───────────────┼──────────────────────────────────────────────────────────────┤
│ libcrypto3 │ CVE-2026-31789 │ CRITICAL │ fixed  │ 3.3.2-r0          │ 3.3.7-r0      │ openssl: OpenSSL: Heap buffer overflow on 32-bit systems     │
│            │                │          │        │                   │               │ from large X.509 certificate...                              │
│            │                │          │        │                   │               │ https://avd.aquasec.com/nvd/cve-2026-31789                   │
│            ├────────────────┼──────────┤        │                   ├───────────────┼──────────────────────────────────────────────────────────────┤
│            │ CVE-2024-12797 │ HIGH     │        │                   │ 3.3.3-r0      │ openssl: RFC7250 handshakes with unauthenticated servers     │
│            │                │          │        │                   │               │ don't abort as expected                                      │
$ trivy image --severity HIGH,CRITICAL --format json ghcr.io/stefanprodan/podinfo:6.7.1 | jq -r '[.Results[].Vulnerabilities // [] | .[].VulnerabilityID] | unique | .[0:3][]'
2026-09-13T13:01:08-04:00	INFO	[vuln] Vulnerability scanning is enabled
2026-09-13T13:01:08-04:00	INFO	[secret] Secret scanning is enabled
2026-09-13T13:01:08-04:00	INFO	[secret] If your scanning is slow, please try '--scanners vuln' to disable secret scanning
2026-09-13T13:01:08-04:00	INFO	[secret] Please see https://trivy.dev/docs/v0.73/guide/scanner/secret#recommendation for faster secret detection
2026-09-13T13:01:08-04:00	INFO	Detected OS	family="alpine" version="3.20.3"
2026-09-13T13:01:08-04:00	INFO	[alpine] Detecting vulnerabilities...	os_version="3.20" repository="3.20" pkg_num=27
2026-09-13T13:01:08-04:00	INFO	Number of language-specific files	num=2
2026-09-13T13:01:08-04:00	INFO	[gobinary] Detecting vulnerabilities...
2026-09-13T13:01:08-04:00	WARN	Using severities from other vendors for some vulnerabilities. Read https://trivy.dev/docs/v0.73/guide/scanner/vulnerability#severity-selection for details.
2026-09-13T13:01:08-04:00	WARN	This OS version is no longer supported by the distribution	family="alpine" version="3.20.3"
2026-09-13T13:01:08-04:00	WARN	The vulnerability detection may be insufficient because security updates are not provided

📣 Notices:
  - Version 0.74.0 of Trivy is now available, current version is 0.73.0

To suppress version checks, run Trivy scans with the --skip-version-check flag

CVE-2024-12797
CVE-2024-45338
CVE-2025-15467
$ CVE=$(trivy image --severity HIGH,CRITICAL --format json ghcr.io/stefanprodan/podinfo:6.7.1 | jq -r '[.Results[].Vulnerabilities // [] | .[].VulnerabilityID] | unique | .[0]')
2026-09-13T13:01:08-04:00	INFO	[vuln] Vulnerability scanning is enabled
2026-09-13T13:01:08-04:00	INFO	[secret] Secret scanning is enabled
2026-09-13T13:01:08-04:00	INFO	[secret] If your scanning is slow, please try '--scanners vuln' to disable secret scanning
2026-09-13T13:01:08-04:00	INFO	[secret] Please see https://trivy.dev/docs/v0.73/guide/scanner/secret#recommendation for faster secret detection
2026-09-13T13:01:09-04:00	INFO	Detected OS	family="alpine" version="3.20.3"
2026-09-13T13:01:09-04:00	INFO	[alpine] Detecting vulnerabilities...	os_version="3.20" repository="3.20" pkg_num=27
2026-09-13T13:01:09-04:00	INFO	Number of language-specific files	num=2
2026-09-13T13:01:09-04:00	INFO	[gobinary] Detecting vulnerabilities...
2026-09-13T13:01:09-04:00	WARN	Using severities from other vendors for some vulnerabilities. Read https://trivy.dev/docs/v0.73/guide/scanner/vulnerability#severity-selection for details.
2026-09-13T13:01:09-04:00	WARN	This OS version is no longer supported by the distribution	family="alpine" version="3.20.3"
2026-09-13T13:01:09-04:00	WARN	The vulnerability detection may be insufficient because security updates are not provided

📣 Notices:
  - Version 0.74.0 of Trivy is now available, current version is 0.73.0

To suppress version checks, run Trivy scans with the --skip-version-check flag
$ echo "suppressing $CVE"
suppressing CVE-2024-12797
$ cat > /tmp/vex.json <<EOF
{
  "@context": "https://openvex.dev/ns/v0.2.0",
  "@id": "https://lab.local/vex/podinfo",
  "author": "platform team",
  "statements": [
    {
      "vulnerability": { "name": "${CVE}" },
      "products": [{ "@id": "pkg:oci/podinfo" }],
      "status": "not_affected",
      "justification": "vulnerable_code_not_in_execute_path"
    }
  ]
}
EOF
$ trivy image --vex /tmp/vex.json --severity HIGH,CRITICAL ghcr.io/stefanprodan/podinfo:6.7.1 | head -20
2026-09-13T13:01:09-04:00	INFO	[vuln] Vulnerability scanning is enabled
2026-09-13T13:01:09-04:00	INFO	[secret] Secret scanning is enabled
2026-09-13T13:01:09-04:00	INFO	[secret] If your scanning is slow, please try '--scanners vuln' to disable secret scanning
2026-09-13T13:01:09-04:00	INFO	[secret] Please see https://trivy.dev/docs/v0.73/guide/scanner/secret#recommendation for faster secret detection
2026-09-13T13:01:09-04:00	INFO	Detected OS	family="alpine" version="3.20.3"
2026-09-13T13:01:09-04:00	INFO	[alpine] Detecting vulnerabilities...	os_version="3.20" repository="3.20" pkg_num=27
2026-09-13T13:01:09-04:00	INFO	Number of language-specific files	num=2
2026-09-13T13:01:09-04:00	INFO	[gobinary] Detecting vulnerabilities...
2026-09-13T13:01:09-04:00	WARN	Using severities from other vendors for some vulnerabilities. Read https://trivy.dev/docs/v0.73/guide/scanner/vulnerability#severity-selection for details.
2026-09-13T13:01:09-04:00	WARN	This OS version is no longer supported by the distribution	family="alpine" version="3.20.3"
2026-09-13T13:01:09-04:00	WARN	The vulnerability detection may be insufficient because security updates are not provided
2026-09-13T13:01:09-04:00	INFO	Some vulnerabilities have been ignored/suppressed. Use the "--show-suppressed" flag to display them.

Report Summary

┌────────────────────────────────────────────────────┬──────────┬─────────────────┬─────────┐
│                       Target                       │   Type   │ Vulnerabilities │ Secrets │
├────────────────────────────────────────────────────┼──────────┼─────────────────┼─────────┤
│ ghcr.io/stefanprodan/podinfo:6.7.1 (alpine 3.20.3) │  alpine  │       19        │    -    │
├────────────────────────────────────────────────────┼──────────┼─────────────────┼─────────┤
│ home/app/podinfo                                   │ gobinary │       36        │    -    │
├────────────────────────────────────────────────────┼──────────┼─────────────────┼─────────┤
│ usr/local/bin/podcli                               │ gobinary │       33        │    -    │
└────────────────────────────────────────────────────┴──────────┴─────────────────┴─────────┘
Legend:
- '-': Not scanned
- '0': Clean (no security findings detected)


ghcr.io/stefanprodan/podinfo:6.7.1 (alpine 3.20.3)
==================================================
Total: 19 (HIGH: 17, CRITICAL: 2)
verify: two SBOM files exist, both list the alpine packages plus the Go module list, and comm shows where the two catalogers disagree. The counts are close but not identical, which is the honest result: each tool decides for itself what counts as a package. trivy sbom accepts either file as input. Then the VEX half: the same image rescanned with --vex reports 19 alpine findings where the plain scan reported 21, and trivy says so in a line of its own, "Some vulnerabilities have been ignored/suppressed". An exception you can diff, not a flag someone remembered to pass.

A ClusterComplianceReport is generated on a schedule, so a stale timestamp means you are reading yesterday's answer. Changing the cron is the supported way to force a rerun, and the timestamp proves it happened.

kubectl get clustercompliancereport
BEFORE=$(kubectl get clustercompliancereport k8s-cis-1.23 -o jsonpath='{.status.updateTimestamp}'); echo "before: $BEFORE"
kubectl get clustercompliancereport k8s-cis-1.23 -o jsonpath='{.spec.cron}{"\n"}'
kubectl patch clustercompliancereport k8s-cis-1.23 --type merge -p '{"spec":{"cron":"*/1 * * * *"}}'
for i in $(seq 1 24); do NOW=$(kubectl get clustercompliancereport k8s-cis-1.23 -o jsonpath='{.status.updateTimestamp}'); [ "$NOW" != "$BEFORE" ] && { echo "after:  $NOW (took ${i}0s)"; break; }; sleep 10; done
kubectl get clustercompliancereport k8s-cis-1.23 -o jsonpath='{.status.summary}' | jq
kubectl patch clustercompliancereport k8s-cis-1.23 --type merge -p '{"spec":{"cron":"0 */6 * * *"}}'
outputcaptured 2026-09-13
$ kubectl get clustercompliancereport
NAME                     AGE
k8s-cis-1.23             15h
k8s-nsa-1.0              15h
k8s-pss-baseline-0.1     15h
k8s-pss-restricted-0.1   15h
$ BEFORE=$(kubectl get clustercompliancereport k8s-cis-1.23 -o jsonpath='{.status.updateTimestamp}'); echo "before: $BEFORE"
before: 2026-09-13T00:00:00Z
$ kubectl get clustercompliancereport k8s-cis-1.23 -o jsonpath='{.spec.cron}{"\n"}'
0 */6 * * *
$ kubectl patch clustercompliancereport k8s-cis-1.23 --type merge -p '{"spec":{"cron":"*/1 * * * *"}}'
clustercompliancereport.aquasecurity.github.io/k8s-cis-1.23 patched
$ for i in $(seq 1 24); do NOW=$(kubectl get clustercompliancereport k8s-cis-1.23 -o jsonpath='{.status.updateTimestamp}'); [ "$NOW" != "$BEFORE" ] && { echo "after:  $NOW (took ${i}0s)"; break; }; sleep 10; done
after:  2026-09-13T10:49:49Z (took 10s)
$ kubectl get clustercompliancereport k8s-cis-1.23 -o jsonpath='{.status.summary}' | jq
{
  "failCount": 34,
  "passCount": 82
}
$ kubectl patch clustercompliancereport k8s-cis-1.23 --type merge -p '{"spec":{"cron":"0 */6 * * *"}}'
clustercompliancereport.aquasecurity.github.io/k8s-cis-1.23 patched
verify: the timestamp moves and the summary counts fail and pass totals.

CIS benchmark output is the format compliance conversations happen in: a control id, a verdict, and a remediation. Run it once so the ids mean something when a task quotes one at you.

kubectl apply -f - <<'EOF'
apiVersion: batch/v1
kind: Job
metadata: { name: kube-bench, namespace: default }
spec:
  template:
    spec:
      hostPID: true
      nodeSelector: { node-role.kubernetes.io/control-plane: "" }
      tolerations:
        - { key: node-role.kubernetes.io/control-plane, operator: Exists, effect: NoSchedule }
      containers:
        - name: kube-bench
          image: docker.io/aquasec/kube-bench:latest
          command: ["kube-bench", "run", "--targets", "master"]
          volumeMounts:
            - { name: etc-kubernetes, mountPath: /etc/kubernetes, readOnly: true }
            - { name: var-lib-etcd, mountPath: /var/lib/etcd, readOnly: true }
      restartPolicy: Never
      volumes:
        - { name: etc-kubernetes, hostPath: { path: /etc/kubernetes } }
        - { name: var-lib-etcd, hostPath: { path: /var/lib/etcd } }
  backoffLimit: 0
EOF
kubectl wait --for=condition=complete job/kube-bench --timeout=300s || true
kubectl logs job/kube-bench | grep -E '^\[(PASS|FAIL)\]' | head -10
kubectl logs job/kube-bench | tail -20
kubectl delete job kube-bench
outputcaptured 2026-09-12
$ kubectl apply -f - <<'EOF'
apiVersion: batch/v1
kind: Job
metadata: { name: kube-bench, namespace: default }
spec:
  template:
    spec:
      hostPID: true
      nodeSelector: { node-role.kubernetes.io/control-plane: "" }
      tolerations:
        - { key: node-role.kubernetes.io/control-plane, operator: Exists, effect: NoSchedule }
      containers:
        - name: kube-bench
          image: docker.io/aquasec/kube-bench:latest
          command: ["kube-bench", "run", "--targets", "master"]
          volumeMounts:
            - { name: etc-kubernetes, mountPath: /etc/kubernetes, readOnly: true }
            - { name: var-lib-etcd, mountPath: /var/lib/etcd, readOnly: true }
      restartPolicy: Never
      volumes:
        - { name: etc-kubernetes, hostPath: { path: /etc/kubernetes } }
        - { name: var-lib-etcd, hostPath: { path: /var/lib/etcd } }
  backoffLimit: 0
EOF
job.batch/kube-bench created
$ kubectl wait --for=condition=complete job/kube-bench --timeout=300s || true
job.batch/kube-bench condition met
$ kubectl logs job/kube-bench | grep -E '^\[(PASS|FAIL)\]' | head -10
[PASS] 1.1.1 Ensure that the API server pod specification file permissions are set to 600 or more restrictive (Automated)
[PASS] 1.1.2 Ensure that the API server pod specification file ownership is set to root:root (Automated)
[PASS] 1.1.3 Ensure that the controller manager pod specification file permissions are set to 600 or more restrictive (Automated)
[PASS] 1.1.4 Ensure that the controller manager pod specification file ownership is set to root:root (Automated)
[PASS] 1.1.5 Ensure that the scheduler pod specification file permissions are set to 600 or more restrictive (Automated)
[PASS] 1.1.6 Ensure that the scheduler pod specification file ownership is set to root:root (Automated)
[PASS] 1.1.7 Ensure that the etcd pod specification file permissions are set to 600 or more restrictive (Automated)
[PASS] 1.1.8 Ensure that the etcd pod specification file ownership is set to root:root (Automated)
[PASS] 1.1.11 Ensure that the etcd data directory permissions are set to 700 or more restrictive (Automated)
[FAIL] 1.1.12 Ensure that the etcd data directory ownership is set to etcd:etcd (Automated)
$ kubectl logs job/kube-bench | tail -20
1.4.1 Edit the Scheduler pod specification file /etc/kubernetes/manifests/kube-scheduler.yaml file
on the control plane node and set the below parameter.
--profiling=false

1.4.2 Edit the Scheduler pod specification file /etc/kubernetes/manifests/kube-scheduler.yaml
on the control plane node and ensure the correct value for the --bind-address parameter


== Summary master ==
38 checks PASS
11 checks FAIL
11 checks WARN
0 checks INFO

== Summary total ==
38 checks PASS
11 checks FAIL
11 checks WARN
0 checks INFO
$ kubectl delete job kube-bench
job.batch "kube-bench" deleted from default namespace
verify: you can quote one PASS and one FAIL with their control ids, and read the remediation text for the failure.

Self-check

answer before opening
Which audit level would you set for Secrets, and which for Events? Why?

Secrets: Metadata, never Request or RequestResponse, or you write secret values into a log file. Events: None; they are high-volume and low-value, and dropping them keeps the log readable. First-match-wins ordering makes those two rules the top of the file.

Find every request a specific ServiceAccount was denied. What does the query look like?

Filter on .user.username == "system:serviceaccount:<ns>:<name>" and either .responseStatus.code == 403 or .annotations["authorization.k8s.io/decision"] == "forbid", then print verb and objectRef. That is the audit-log form of the RBAC debugging you did in 5.1.

An SBOM and a vulnerability report: what is the difference and why keep both?

The SBOM is an inventory of components; the vulnerability report is that inventory joined against a CVE database at a point in time. Keep both because the inventory stays true while the CVE list changes daily; that is exactly why continuous rescanning (the operator) matters more than a scan at build time alone.

Why is "compliance as CRDs" a platform-engineering idea rather than a Trivy feature?

Because it makes compliance data a first-class API object: queryable with kubectl, protected by RBAC, watchable by controllers, graphable in Grafana, and reviewable in git. The alternative, reports as files in a bucket, cannot be joined to anything the platform already does.

Your cluster's API server will not start after an audit change. First hypothesis?

The audit policy file is missing or unparseable at the path given by --audit-policy-file (or the volume mount is wrong). The API server refuses to start rather than run unaudited: a safety choice that presents as a dead control plane.

Your audit policy puts the catch-all Metadata rule first and the Secrets rule after it. What is wrong, and what else should precede the catch-all?

First match wins, so the Secrets rule never applies and every request lands at Metadata (harmless for Secrets in this case, but the same mistake with a None rule after the catch-all means the noise is never dropped, and a RequestResponse rule after it never raises RBAC changes). Order: None for health URLs, leases and kube-proxy watches; Metadata for secrets, configmaps, tokenreviews and pods/exec; RequestResponse for RBAC and policy groups; then the catch-all.

Which layer catches each: a user reads a Secret; a process in a container starts /bin/sh; an image contains a package with a new CVE; a node's kubelet allows anonymous auth?

API audit log (Metadata on secrets); Falco syscall rules; the Trivy Operator's VulnerabilityReport on rescan (or a pipeline scan at build); kube-bench or the Trivy Operator's InfraAssessmentReport and CIS compliance report. Four different sensors, one compliance story; naming which sees what is the answer to "design a complete audit trail".

What is the difference between an SBOM attestation and a VEX document, and why keep both?

The SBOM lists what is inside the artefact; VEX states whether known vulnerabilities in those components actually affect it (not affected, affected, fixed, under investigation) with a justification. Scanners join the SBOM against CVE feeds and then apply VEX to suppress non-applicable findings with a recorded reason, which is auditable in a way a bare .trivyignore line is not.

Docs to know your way around

study time, not exam time
  • kubernetes.io: Auditing (policy levels, stages, and the event schema).
  • aquasecurity.github.io/trivy-operator: the CRD reference pages, one per report kind.
  • cyclonedx.org / spdx.dev: enough to know which is which.
  • Offline: jq against the log file, and kubectl explain clustercompliancereport.status.
  • kubernetes.io/docs/tasks/debug/debug-cluster/audit: rule fields, stages, and the full flag list for both backends; the Policy API reference under audit.k8s.io/v1 for omitManagedFields.
  • aquasecurity.github.io/trivy-operator/latest/docs/crds (one page per report kind) and /docs/compliance (the four compliance spec ids); aquasecurity.github.io/trivy/latest/docs/configuration/filtering for VEX and ignore files.
  • falco.org/docs/rules for rule syntax and github.com/in-toto/attestation/tree/main/spec for the Statement and DSSE layout.