Grafana is the pane of glass over both halves of this section: Prometheus metrics through dashboards, Loki logs through Explore. Credentials are genuinely admin/admin here; address from make urls.

needsmake up obs

Orientation

competency 4.1 · logging and actionable dashboards

The competency wording is "dashboards that provide actionable insight", which is a judgment claim, not a tooling claim. Two things make it real: dashboards shipped as code, and dashboards designed around a question someone actually asks during an incident.

Dashboards as code

the sidecar pattern

The exam-relevant fact about Grafana in a kube-prometheus-stack world: dashboards are provisioned, not clicked together. A sidecar container in the Grafana pod watches for ConfigMaps labeled grafana_dashboard: "1" and loads any dashboard JSON it finds inside. That is why the stack's forty-odd dashboards exist without anyone importing them, and it is how you ship a dashboard in a GitOps world: JSON in a ConfigMap in a git repo. Datasources provision the same way with grafana_datasource: "1".

git ──▶ ConfigMap (label grafana_dashboard=1) ──▶ sidecar writes /tmp/dashboards ──▶ Grafana loads it
                                                                                     │
   a dashboard clicked together in the UI lives only in Grafana's database           │
   and dies with the pod ────────────────────────────────────────────────────────────┘

What makes a dashboard actionable rather than decorative

  • Lead with the user-facing symptom. RED for services (Rate, Errors, Duration); USE for resources (Utilization, Saturation, Errors); the four golden signals (latency, traffic, errors, saturation) if you prefer that vocabulary. Know all three acronyms and which applies to what; it is a cheap question to set.
  • Template over a variable ($namespace, $pod) instead of hardcoding, so one dashboard serves every tenant. Variables come from label_values() queries, which is how the stock dashboards populate their dropdowns.
  • Every panel answers a question someone would ask during an incident. Forty panels of everything is a wall, not a dashboard.
  • Annotations and links. Deployment markers turn "when did this start" into a glance, and a panel link to the matching logs query turns a metric spike into a root cause.

Grafana specifics worth naming once: datasources (Prometheus, Loki, Jaeger here), panels with unit and threshold settings, alerting (Grafana has its own alert engine, deliberately unused in this lab; Prometheus rules are the exam's model), and folders + permissions for multi-team setups.

Loki and LogQL

index labels, not content

Loki indexes labels, not content: log lines are stored against a small label set (namespace, pod, container, app) and content matching happens at query time by scanning the selected streams. That is why it is cheap to run, and why every query must start from a label selector before filtering. It is also why label cardinality discipline matters even more than in Prometheus: a label per request ID would create a stream per request.

Collection here is Alloy, a single-replica deployment in monitoring that tails every pod through the kubelet API and pushes to Loki; Loki itself stores them as a StatefulSet next door. That is two hops, and make validate proves the pair works by demanding Loki actually returns streams. A running Loki with no shipper looks healthy and holds nothing.

LogQL pieceExampleDoes
stream selector{namespace="team-a"}required first step; picks streams
line filter|= "error" · != "/healthz" · |~ "5[0-9]{2}"substring / regex match on the line
parser| json · | logfmt · | patternextract fields into labels for filtering
label filter| status >= 500filter on extracted fields
metric querysum by (pod) (rate({namespace="x"} |= "error" [5m]))turns logs into a graphable series

Those five constructs cover exam depth. The one to internalize is the last: a metric query over logs is how you alert on something that only exists in text, and it is the bridge between this section and 4.2.

Terminal beats UI, often

stern <pattern> -n <ns> tails many pods at once with color-coded names, and kubectl logs --previous is the only way to read a crashed container's last words. Loki is for "what happened an hour ago across the fleet"; stern and kubectl are for "what is happening right now". Knowing which question you are asking picks the tool.

The three signals, and when logs are the right one

Metrics answer "how much / how often" cheaply over long windows. Logs answer "what exactly happened to this one thing", at higher cost per query. Traces (4.4) answer "where in the request path did it go wrong". An exam scenario that says "find the error message" is a log question; "how many failed" is a metric question; "which service is slow" is a trace question. Choosing wrong costs minutes.

Provisioning, field by field

datasource and dashboard files, sidecar knobs, the Grafana Operator

Grafana's own provisioning format

Files under provisioning/datasources/*.yaml and provisioning/dashboards/*.yaml start with apiVersion: 1. A datasource entry has name, type (prometheus, loki, jaeger, tempo), url, access: proxy, an optional stable uid (dashboards reference datasources by uid, so set it or every exported dashboard breaks on import), isDefault, editable: false to stop UI drift, and two maps: jsonData for settings (httpMethod: POST, timeInterval, Loki's derivedFields that turn a trace ID in a log line into a Tempo/Jaeger link, and alertmanagerUid) and secureJsonData for credentials, which Grafana encrypts at rest. deleteDatasources removes entries by name before applying. A dashboard provider entry names a folder, type: file and options.path, plus foldersFromFilesStructure, allowUiUpdates (default false: the file wins) and disableDeletion.

In kube-prometheus-stack the sidecar generates those provider files for you: ConfigMaps labeled grafana_dashboard: "1" land in the default folder unless annotated grafana_folder: <name>; ConfigMaps or Secrets labeled grafana_datasource: "1" are datasource files. The sidecar watches every namespace by default (sidecar.dashboards.searchNamespace), so a dashboard shipped with a team's Helm chart appears in the platform Grafana without anyone in monitoring touching anything.

The Grafana Operator, when the exam gives you CRDs

Kind (grafana.integreatly.org/v1beta1)HoldsField to watch
Grafanaan instance (or an external one via spec.external.url plus admin credentials)its labels are what every other kind selects with instanceSelector
GrafanaDashboardone dashboard from json, gzipJson, url, configMapRef, jsonnet or grafanaCom.idinstanceSelector, folder/folderRef, resyncPeriod, datasources[] to rewrite ${DS_PROMETHEUS} inputs
GrafanaDatasourceone datasource with the same fields as the provisioning file (datasource.jsonData, secureJsonData, valuesFrom for Secret refs)instanceSelector; allowCrossNamespaceImport: true on the Grafana when the datasource lives elsewhere
GrafanaFoldera folder, optionally with parentFolderRef and permissionsdashboards point at it by folderRef
GrafanaAlertRuleGroup / GrafanaContactPoint / GrafanaNotificationPolicy / GrafanaMuteTimingGrafana-managed alerting as codeirrelevant if Prometheus rules are the model, but a task may hand you one

The failure that recurs: a GrafanaDashboard whose instanceSelector.matchLabels matches no Grafana object is accepted and does nothing; its status stays empty. Same selector story as ServiceMonitors, one CRD family over. In every model, the JSON is the artefact: export with "export for sharing externally" off so datasource uids stay literal, keep id: null and a stable uid in the JSON, and template with label_values(kube_pod_info, namespace) variables rather than hardcoding a tenant.

Loki under the hood, and LogQL past the first five constructs

components, labels vs structured metadata, parsers, shippers, retention

Components and deployment modes

The write path is distributor (validates, hashes each stream to a ring position) → ingester (builds chunks in memory, flushes to object storage) with a quorum of replicas acknowledging. The read path is query-frontend (splits and queues) → query-scheduler → querier (asks ingesters for recent data, then the store). The compactor merges index files and applies retention; the ruler evaluates alerting and recording rules; the index gateway serves index lookups. Everything persists to one object store (S3, GCS, Azure, MinIO) with the TSDB index shipped alongside the chunks. Three deployment modes: single binary (-target=all), simple scalable (read, write, backend targets, the Helm default at scale) and full microservices. "Loki is up but returns nothing" in the SSD mode is usually the write path (shipper cannot push, ingester ring unhealthy) rather than the read path.

Labels, cardinality, structured metadata

Every unique label set is a stream, every stream has its own chunks and index entries, and Loki caps streams per tenant (max_streams_per_user) and labels per stream (15 by default). Labels should be low-cardinality and long-lived: namespace, app, container, level at most. Anything high-cardinality that you still want to filter on quickly (pod name, trace ID, request ID) belongs in structured metadata, per-line key-value pairs stored with the chunk and queryable with the same label-filter syntax, without creating streams. The OTLP endpoint (/otlp/v1/logs) maps resource attributes such as service.name and k8s.namespace.name to index labels and the rest to structured metadata automatically; the docs now recommend moving k8s.pod.name off the index label list.

LogQL you have not used yet

ConstructExampleNote
| pattern| pattern "<ip> - - <_> \"<method> <path> <_>\" <status>"fast field extraction for fixed-format lines; <_> skips
| regexp| regexp "(?P<user>\\w+) logged in"named groups become labels; slowest parser
| json a="b.c"extract only named pathsbare | json flattens everything, which explodes label counts on wide events
| line_format / | label_format| line_format "{{.method}} {{.status}}"Go templates; rewrite the line before further filters
| __error__=""drop lines the parser could not readparse failures are not dropped by default; they get an __error__ label
| unwrap durationquantile_over_time(0.95, {app="api"} | logfmt | unwrap duration [5m]) by (path)turns an extracted number into samples; unwrap duration_seconds(x) and bytes(x) convert units
count_over_time / bytes_rate / absent_over_timeabsent_over_time({job="backup"} |= "completed" [25h])the "the nightly job did not log success" alert
offsetcount_over_time({app="x"}[5m] offset 1w)inside the range brackets, unlike PromQL's position
| drop / | keep| keep status, pathtrim labels before an aggregation

Two syntax facts that cost minutes: label matchers in the stream selector are fully anchored regexes, line filters (|~) are not; and the log range must come immediately after the pipeline, so rate({a="b"} |= "x" [5m]) is right and rate({a="b"}[5m] |= "x") is a parse error.

Shippers, retention, rules

  • Promtail is end-of-life (LTS ended February 2026, EOL 2 March 2026). Alloy is the replacement: discovery.kubernetes → loki.source.kubernetes (or loki.source.file with a hostPath mount) → loki.process (stages: stage.json, stage.labels, stage.drop) → loki.write. alloy convert --source-format=promtail translates an old config. The OpenTelemetry Collector with a filelog receiver and an otlphttp exporter to Loki's /otlp endpoint, and Fluent Bit with its loki output, are the other two shippers a task may hand you; all three are DaemonSets that tail /var/log/pods and enrich with Kubernetes metadata.
  • Retention is the compactor's job: compactor.retention_enabled: true plus limits_config.retention_period (global) or per-tenant/per-stream retention_stream overrides. Without the compactor flag, data is kept forever regardless of the period.
  • The ruler takes Prometheus-format rule groups whose expr is LogQL, from local files or object storage (ruler.storage), and sends alerts to ruler.alertmanager_url; recording rules can remote_write the resulting series into Prometheus so metrics derived from logs join the normal dashboards.
How this gets tested

"Find the pod that logged X in the last hour" is a stream selector plus |=; "count errors per service over time" is sum by (app) (count_over_time(... |= "error" [5m])); "p95 latency from access logs" is logfmt or pattern, unwrap, quantile_over_time. Recognizing which of the three you are being asked for is most of the time saved.

Exercises

tick the dot when its check passes

In Grafana: Connections → Data sources. If Loki is not there, add it the provisioned way rather than the UI way:

kubectl -n monitoring apply -f - <<'EOF'
apiVersion: v1
kind: ConfigMap
metadata:
  name: loki-datasource
  labels: { grafana_datasource: "1" }
data:
  loki.yaml: |
    apiVersion: 1
    datasources:
      - name: Loki
        type: loki
        url: http://loki.monitoring.svc:3100
        access: proxy
EOF
outputcaptured 2026-08-26
$ kubectl -n monitoring apply -f - <<'EOF'
apiVersion: v1
kind: ConfigMap
metadata:
  name: loki-datasource
  labels: { grafana_datasource: "1" }
data:
  loki.yaml: |
    apiVersion: 1
    datasources:
      - name: Loki
        type: loki
        url: http://loki.monitoring.svc:3100
        access: proxy
EOF
configmap/loki-datasource created

The sidecar picks it up within a minute or so. If nothing arrives, bisect the two hops: shipper first (kubectl -n monitoring get deploy alloy, then its logs for push errors), Loki second.

verify: Explore → Loki → {namespace="monitoring"} returns log lines. If nothing returns, name which hop is broken.

Generate known lines, then find them (team-a comes from make sec; kubectl apply -f examples/multitenancy/team-a.yaml creates it standalone):

kubectl -n team-a run chatty --image=busybox:1.37 --restart=Never -- \
  sh -c 'for i in $(seq 60); do echo "level=error msg=payment_failed attempt=$i"; sleep 1; done'
outputcaptured 2026-08-26
$ kubectl -n team-a run chatty --image=busybox:1.37 --restart=Never -- \
  sh -c 'for i in $(seq 60); do echo "level=error msg=payment_failed attempt=$i"; sleep 1; done'
Warning: would violate PodSecurity "restricted:latest": allowPrivilegeEscalation != false (container "chatty" must set securityContext.allowPrivilegeEscalation=false), unrestricted capabilities (container "chatty" must set securityContext.capabilities.drop=["ALL"]), runAsNonRoot != true (pod or container "chatty" must set securityContext.runAsNonRoot=true), seccompProfile (pod or container "chatty" must set securityContext.seccompProfile.type to "RuntimeDefault" or "Localhost")
pod/chatty created

In Explore: {namespace="team-a"} |= "payment_failed", then the metric form sum(rate({namespace="team-a"} |= "payment_failed" [1m])).

verify: lines visible and the rate curve is about 1/s while the pod runs. Compare stern chatty -n team-a for the same lines live.

Build one panel in the UI first (New dashboard → add visualization → Prometheus → sum by (namespace) (rate(container_cpu_usage_seconds_total[5m]))), then export its JSON (share → Export → save JSON, set "export for sharing externally" off), and re-deliver it properly:

kubectl -n monitoring create configmap team-cpu-dashboard --from-file=dash.json=<your-file> \
  --dry-run=client -o yaml | kubectl label -f - --local --dry-run=client -o yaml grafana_dashboard=1 | kubectl apply -f -
outputcaptured 2026-08-26
$ kubectl -n monitoring create configmap team-cpu-dashboard --from-file=dash.json=dash.json \
  --dry-run=client -o yaml | kubectl label -f - --local --dry-run=client -o yaml grafana_dashboard=1 | kubectl apply -f -
configmap/team-cpu-dashboard created
verify: delete the hand-made dashboard in the UI, and the provisioned copy appears (default folder) and survives a Grafana pod delete, which the clicked one would not have. State the GitOps moral in one sentence.

Open "Kubernetes / Compute Resources / Namespace (Pods)" for team-a while the chatty pod runs. Answer: which panels are RED, which are USE, and which single panel would you keep if you could keep one during a pod-crash incident.

verify: there is no single right answer; having an answer with a reason is the skill the "actionable insights" wording points at.

A derived field turns a trace id in a log line into a link to the trace. It is a regex and a datasource uid in a provisioning file, and it is what joins the two halves of an investigation without a copy and paste.

# port 3000 on this host is Backstage, not Grafana: a curl there returns HTML and jq says 'Invalid numeric literal'
# and the derived field needs a Jaeger datasource with the uid it points at
kubectl -n monitoring delete cm loki-datasource --ignore-not-found
kubectl apply -f - <<'EOF'
apiVersion: v1
kind: ConfigMap
metadata:
  name: jaeger-datasource
  namespace: monitoring
  labels: { grafana_datasource: "1" }
data:
  jaeger.yaml: |
    apiVersion: 1
    datasources:
      - name: Jaeger
        type: jaeger
        uid: jaeger-lab
        access: proxy
        url: http://jaeger.tracing.svc:16686
EOF
kubectl -n monitoring get cm -l grafana_datasource=1 -o name
kubectl apply -f - <<'EOF'
apiVersion: v1
kind: ConfigMap
metadata:
  name: loki-datasource-linked
  namespace: monitoring
  labels: { grafana_datasource: "1" }
data:
  loki.yaml: |
    apiVersion: 1
    datasources:
      - name: Loki
        type: loki
        uid: loki-lab
        access: proxy
        url: http://loki.monitoring.svc:3100
        editable: false
        jsonData:
          derivedFields:
            - name: TraceID
              matcherRegex: "trace_id=(\\w+)"
              url: "$${__value.raw}"
              datasourceUid: jaeger-lab
EOF
kubectl -n monitoring wait --for=condition=available deploy/prometheus-grafana --timeout=180s
sleep 45
GF_PW=$(kubectl -n monitoring get secret prometheus-grafana -o jsonpath='{.data.admin-password}' | base64 -d)
kubectl -n monitoring port-forward --address 127.0.0.1 svc/prometheus-grafana 3300:80 & PF1=$!
sleep 5
curl -s -u admin:admin 127.0.0.1:3300/api/datasources | jq '[.[] | {name, uid, type}]'
curl -s -u admin:admin 127.0.0.1:3300/api/datasources/uid/loki-lab | jq '.jsonData'
kill $PF1
kubectl -n monitoring delete cm jaeger-datasource
kubectl -n monitoring delete cm loki-datasource-linked --ignore-not-found
outputcaptured 2026-09-12
$ # port 3000 on this host is Backstage, not Grafana: a curl there returns HTML and jq says 'Invalid numeric literal'
$ # and the derived field needs a Jaeger datasource with the uid it points at
$ kubectl -n monitoring delete cm loki-datasource --ignore-not-found
$ kubectl apply -f - <<'EOF'
apiVersion: v1
kind: ConfigMap
metadata:
  name: jaeger-datasource
  namespace: monitoring
  labels: { grafana_datasource: "1" }
data:
  jaeger.yaml: |
    apiVersion: 1
    datasources:
      - name: Jaeger
        type: jaeger
        uid: jaeger-lab
        access: proxy
        url: http://jaeger.tracing.svc:16686
EOF
configmap/jaeger-datasource created
$ kubectl -n monitoring get cm -l grafana_datasource=1 -o name
configmap/jaeger-datasource
configmap/loki-datasource-linked
configmap/prometheus-kube-prometheus-grafana-datasource
$ kubectl apply -f - <<'EOF'
apiVersion: v1
kind: ConfigMap
metadata:
  name: loki-datasource-linked
  namespace: monitoring
  labels: { grafana_datasource: "1" }
data:
  loki.yaml: |
    apiVersion: 1
    datasources:
      - name: Loki
        type: loki
        uid: loki-lab
        access: proxy
        url: http://loki.monitoring.svc:3100
        editable: false
        jsonData:
          derivedFields:
            - name: TraceID
              matcherRegex: "trace_id=(\\w+)"
              url: "$${__value.raw}"
              datasourceUid: jaeger-lab
EOF
configmap/loki-datasource-linked unchanged
$ kubectl -n monitoring wait --for=condition=available deploy/prometheus-grafana --timeout=180s
deployment.apps/prometheus-grafana condition met
$ sleep 45
$ GF_PW=$(kubectl -n monitoring get secret prometheus-grafana -o jsonpath='{.data.admin-password}' | base64 -d)
$ kubectl -n monitoring port-forward --address 127.0.0.1 svc/prometheus-grafana 3300:80 & PF1=$!
$ sleep 5
Forwarding from 127.0.0.1:3300 -> 3000
$ curl -s -u admin:admin 127.0.0.1:3300/api/datasources | jq '[.[] | {name, uid, type}]'
Handling connection for 3300
[
  {
    "name": "Alertmanager",
    "uid": "alertmanager",
    "type": "alertmanager"
  },
  {
    "name": "Jaeger",
    "uid": "jaeger-lab",
    "type": "jaeger"
  },
  {
    "name": "Loki",
    "uid": "loki-lab",
    "type": "loki"
  },
  {
    "name": "Prometheus",
    "uid": "prometheus",
    "type": "prometheus"
  }
]
$ curl -s -u admin:admin 127.0.0.1:3300/api/datasources/uid/loki-lab | jq '.jsonData'
Handling connection for 3300
{
  "derivedFields": [
    {
      "datasourceUid": "jaeger-lab",
      "matcherRegex": "trace_id=(\\w+)",
      "name": "TraceID",
      "url": "${__value.raw}"
    }
  ]
}
$ kill $PF1
$ kubectl -n monitoring delete cm jaeger-datasource
configmap "jaeger-datasource" deleted from monitoring namespace
$ kubectl -n monitoring delete cm loki-datasource-linked --ignore-not-found
configmap "loki-datasource-linked" deleted from monitoring namespace
verify: the datasource exists with the uid you chose and its jsonData carries the derived field. Open a log line in Explore and the TraceID appears as a link; if the link is dead, the target uid is wrong, not the regex.

A dashboard ConfigMap lands in the General folder unless you say otherwise, and General is where dashboards go to be lost. One annotation fixes it, and it is the difference between a dashboard people find and one they rebuild.

kubectl apply -f - <<'EOF'
apiVersion: v1
kind: ConfigMap
metadata:
  name: platform-overview
  namespace: monitoring
  labels: { grafana_dashboard: "1" }
  annotations: { grafana_folder: "Platform" }
data:
  platform-overview.json: |
    {
      "title": "Platform overview",
      "uid": "platform-overview",
      "schemaVersion": 39,
      "panels": [
        {"type": "timeseries", "title": "API server requests", "gridPos": {"h": 8, "w": 12, "x": 0, "y": 0},
         "targets": [{"expr": "sum(rate(apiserver_request_total[5m]))"}]}
      ]
    }
EOF
sleep 60
# port 3000 here is Backstage; Grafana answers on 3300
GF_PW=$(kubectl -n monitoring get secret prometheus-grafana -o jsonpath='{.data.admin-password}' | base64 -d)
kubectl -n monitoring port-forward --address 127.0.0.1 svc/prometheus-grafana 3300:80 & PF1=$!
sleep 5
curl -s -u admin:admin 127.0.0.1:3300/api/search?query=Platform%20overview | jq '.[] | {title, folderTitle, uid}'
# the sidecar only honours the annotation when it is told which one to read
kubectl -n monitoring get deploy prometheus-grafana -o jsonpath='{range .spec.template.spec.containers[?(@.name=="grafana-sc-dashboard")].env[*]}{.name}={.value}{"\n"}{end}' | grep -i -E 'folder|provider'
kill $PF1
kubectl -n monitoring delete cm platform-overview
outputcaptured 2026-09-13
$ kubectl apply -f - <<'EOF'
apiVersion: v1
kind: ConfigMap
metadata:
  name: platform-overview
  namespace: monitoring
  labels: { grafana_dashboard: "1" }
  annotations: { grafana_folder: "Platform" }
data:
  platform-overview.json: |
    {
      "title": "Platform overview",
      "uid": "platform-overview",
      "schemaVersion": 39,
      "panels": [
        {"type": "timeseries", "title": "API server requests", "gridPos": {"h": 8, "w": 12, "x": 0, "y": 0},
         "targets": [{"expr": "sum(rate(apiserver_request_total[5m]))"}]}
      ]
    }
EOF
configmap/platform-overview created
$ sleep 60
$ # port 3000 here is Backstage; Grafana answers on 3300
$ GF_PW=$(kubectl -n monitoring get secret prometheus-grafana -o jsonpath='{.data.admin-password}' | base64 -d)
$ kubectl -n monitoring port-forward --address 127.0.0.1 svc/prometheus-grafana 3300:80 & PF1=$!
$ sleep 5
Forwarding from 127.0.0.1:3300 -> 3000
$ curl -s -u admin:admin 127.0.0.1:3300/api/search?query=Platform%20overview | jq '.[] | {title, folderTitle, uid}'
Handling connection for 3300
{
  "title": "Platform overview",
  "folderTitle": null,
  "uid": "platform-overview"
}
$ # the sidecar only honours the annotation when it is told which one to read
$ kubectl -n monitoring get deploy prometheus-grafana -o jsonpath='{range .spec.template.spec.containers[?(@.name=="grafana-sc-dashboard")].env[*]}{.name}={.value}{"\n"}{end}' | grep -i -E 'folder|provider'
FOLDER=/tmp/dashboards
$ kill $PF1
$ kubectl -n monitoring delete cm platform-overview
configmap "platform-overview" deleted from monitoring namespace
verify: the dashboard is imported and /api/search finds it by title, but folderTitle is null: it landed at the top level, not in Platform. The annotation is not wrong, it is unread. The sidecar's environment carries only FOLDER=/tmp/dashboards and no FOLDER_ANNOTATION, so no annotation on any ConfigMap can place a dashboard in a folder on this install. Set sidecar.dashboards.folderAnnotation and sidecar.dashboards.provider.foldersFromFilesStructure in the chart values and the same ConfigMap lands in Platform. A convention that depends on a setting nobody enabled is the most common way dashboards end up in one flat list.

LogQL beyond a plain filter is where the marks are: parse a line into fields, turn a field into a number, notice an absence, and find the lines your parser could not handle.

kubectl -n monitoring port-forward svc/loki 3100:3100 & PF1=$!
sleep 5
curl -sG localhost:3100/loki/api/v1/query_range --data-urlencode 'query={namespace="team-a"} | pattern `<ip> - - <_> "<method> <uri> <_>" <status> <size>`' --data-urlencode 'limit=5' | jq '.data.result[0].values[0]'
curl -sG localhost:3100/loki/api/v1/query_range --data-urlencode 'query=quantile_over_time(0.95, {namespace="team-a"} | pattern `<_> <_> <_> <_> "<_> <_> <_>" <status> <size> ` | unwrap size [5m]) by (status)' | jq '.data.result'
curl -sG localhost:3100/loki/api/v1/query_range --data-urlencode 'query=absent_over_time({namespace="default", pod="does-not-exist"}[5m])' | jq '.data.result'
curl -sG localhost:3100/loki/api/v1/query_range --data-urlencode 'query={namespace="team-a"} | logfmt --strict | __error__ != ""' --data-urlencode 'limit=5' | jq '.data.result[0].stream'
kill $PF1
outputcaptured 2026-09-13
$ kubectl -n monitoring port-forward svc/loki 3100:3100 & PF1=$!
$ sleep 5
Forwarding from 127.0.0.1:3100 -> 3100
Forwarding from [::1]:3100 -> 3100
$ curl -sG localhost:3100/loki/api/v1/query_range --data-urlencode 'query={namespace="team-a"} | pattern `<ip> - - <_> "<method> <uri> <_>" <status> <size>`' --data-urlencode 'limit=5' | jq '.data.result[0].values[0]'
Handling connection for 3100
[
  "1789307349067886028",
  "10.244.2.200 - - [13/Sep/2026:13:49:09 +0000] \"GET / HTTP/1.1\" 200 615 \"-\" \"kube-probe/1.36\" \"-\"\n"
]
$ curl -sG localhost:3100/loki/api/v1/query_range --data-urlencode 'query=quantile_over_time(0.95, {namespace="team-a"} | pattern `<_> <_> <_> <_> "<_> <_> <_>" <status> <size> ` | unwrap size [5m]) by (status)' | jq '.data.result'
Handling connection for 3100
[
  {
    "metric": {
      "status": "200"
    },
    "values": [
      [
        1789303754,
        "615"
      ],
      [
        1789303768,
        "615"
      ],
      [
        1789303782,
        "615"
      ],
      [
        1789303796,
        "615"
      ],
      [
        1789303810,
        "615"
      ],
      [
        1789303824,
        "615"
      ],
      [
        1789303838,
        "615"
      ],
      [
        1789303852,
        "615"
      ],
      [
... 1002 more lines
$ curl -sG localhost:3100/loki/api/v1/query_range --data-urlencode 'query=absent_over_time({namespace="default", pod="does-not-exist"}[5m])' | jq '.data.result'
Handling connection for 3100
[
  {
    "metric": {
      "namespace": "default",
      "pod": "does-not-exist"
    },
    "values": [
      [
        1789303754,
        "1"
      ],
      [
        1789303768,
        "1"
      ],
      [
        1789303782,
        "1"
      ],
      [
        1789303796,
        "1"
      ],
      [
        1789303810,
        "1"
      ],
      [
        1789303824,
        "1"
      ],
      [
        1789303838,
        "1"
      ],
      [
        1789303852,
        "1"
      ],
... 1003 more lines
$ curl -sG localhost:3100/loki/api/v1/query_range --data-urlencode 'query={namespace="team-a"} | logfmt --strict | __error__ != ""' --data-urlencode 'limit=5' | jq '.data.result[0].stream'
Handling connection for 3100
{
  "__error__": "LogfmtParserErr",
  "__error_details__": "logfmt syntax error at pos 47 : unexpected '\"'",
  "app": "demo",
  "container": "web",
  "detected_level": "unknown",
  "instance": "team-a/staging-demo-66955f8974-l48bx:web",
  "job": "loki.source.kubernetes.pods",
  "namespace": "team-a",
  "pod": "staging-demo-66955f8974-l48bx",
  "service_name": "demo"
}
$ kill $PF1
verify: the pattern parser produces labeled fields, unwrap turns size into a number you can take a quantile of, absent_over_time returns a series for the pod that does not exist, and logfmt --strict plus __error__ != "" returns the nginx access lines, which are not logfmt at all. The --strict flag is the point: plain logfmt in Loki 3 skips what it cannot parse and sets no __error__, so the same query without it comes back empty and you conclude, wrongly, that every line parsed.

Alloy replaced promtail, and every migration task is the same: take the old config, run the converter, read what it produced. The command lives in the Alloy image.

cat > /tmp/promtail.yaml <<'EOF'
server:
  http_listen_port: 9080
clients:
  - url: http://loki.monitoring.svc:3100/loki/api/v1/push
scrape_configs:
  - job_name: kubernetes-pods
    kubernetes_sd_configs:
      - role: pod
    relabel_configs:
      - source_labels: [__meta_kubernetes_namespace]
        target_label: namespace
EOF
POD=$(kubectl -n monitoring get pod -l app.kubernetes.io/name=alloy -o jsonpath='{.items[0].metadata.name}')
kubectl -n monitoring cp /tmp/promtail.yaml "$POD":/tmp/promtail.yaml
kubectl -n monitoring exec "$POD" -- alloy convert --source-format=promtail /tmp/promtail.yaml
outputcaptured 2026-09-13
$ cat > /tmp/promtail.yaml <<'EOF'
server:
  http_listen_port: 9080
clients:
  - url: http://loki.monitoring.svc:3100/loki/api/v1/push
scrape_configs:
  - job_name: kubernetes-pods
    kubernetes_sd_configs:
      - role: pod
    relabel_configs:
      - source_labels: [__meta_kubernetes_namespace]
        target_label: namespace
EOF
$ POD=$(kubectl -n monitoring get pod -l app.kubernetes.io/name=alloy -o jsonpath='{.items[0].metadata.name}')
$ kubectl -n monitoring cp /tmp/promtail.yaml "$POD":/tmp/promtail.yaml
$ kubectl -n monitoring exec "$POD" -- alloy convert --source-format=promtail /tmp/promtail.yaml
discovery.kubernetes "kubernetes_pods" {
	role = "pod"

	selectors {
		role  = "pod"
		field = "spec.nodeName=" + coalesce(sys.env("HOSTNAME"), constants.hostname)
	}
}

discovery.relabel "kubernetes_pods" {
	targets = discovery.kubernetes.kubernetes_pods.targets

	rule {
		source_labels = ["__meta_kubernetes_namespace"]
		target_label  = "namespace"
	}
}

loki.source.file "kubernetes_pods" {
	targets    = discovery.relabel.kubernetes_pods.output
	forward_to = [loki.write.default.receiver]

	file_match {
		enabled = true
	}
	legacy_positions_file = "/var/log/positions.yaml"
}

loki.write "default" {
	endpoint {
		url = "http://loki.monitoring.svc:3100/loki/api/v1/push"
	}
	external_labels = {}
}
verify: the converter prints Alloy components: a discovery block, a relabel block, a source block and a write block, wired by references.

Retention in Loki is a compactor setting, and it is off by default in more distributions than people expect. Read it before you promise a week of logs to an audit.

kubectl -n monitoring get cm loki -o yaml | grep -A6 -E 'compactor|retention'
kubectl -n monitoring get cm loki -o jsonpath='{.data.config\.yaml}' | grep -E 'retention_enabled|retention_period|delete_request' || echo 'no retention keys in the rendered config'
kubectl -n monitoring get pvc -l app.kubernetes.io/name=loki
outputcaptured 2026-09-12
$ kubectl -n monitoring get cm loki -o yaml | grep -A6 -E 'compactor|retention'
      compactor_grpc_address: 'loki.monitoring.svc.cluster.local:9095'
      path_prefix: /var/loki
      replication_factor: 1
      storage:
        filesystem:
          chunks_directory: /var/loki/chunks
          rules_directory: /var/loki/rules
$ kubectl -n monitoring get cm loki -o jsonpath='{.data.config\.yaml}' | grep -E 'retention_enabled|retention_period|delete_request' || echo 'no retention keys in the rendered config'
no retention keys in the rendered config
$ kubectl -n monitoring get pvc -l app.kubernetes.io/name=loki
NAME             STATUS   VOLUME                                     CAPACITY   ACCESS MODES   STORAGECLASS   VOLUMEATTRIBUTESCLASS   AGE
storage-loki-0   Bound    pvc-4589502f-0ed6-4923-9993-473e891d6ebf   10Gi       RWO            standard       <unset>                 123m
verify: you can say whether retention_enabled is set on this install and what would delete old chunks if it is not. "The disk fills up" is a legitimate answer, and it is the one that matters for capacity planning.

Self-check

answer before opening
How does a dashboard get into Grafana in a GitOps platform, and why not click it?

JSON in a ConfigMap labeled grafana_dashboard: "1", delivered from git; the sidecar loads it. A clicked dashboard lives only in Grafana's database: unreviewable, unversioned, and gone with the pod unless persistence is configured.

Explain RED and USE, and which you would use for a queue worker.

RED = Rate, Errors, Duration, for request-driven services. USE = Utilization, Saturation, Errors, for resources. A queue worker is best served by both plus queue depth: USE for the worker pool, RED-ish for message processing, and saturation is really "is the backlog growing".

Loki returns nothing for {namespace="team-a"}. Two hops to check, in order.

The shipper first (Alloy running, not erroring on push, actually tailing that namespace), then Loki (ingester healthy, retention not eating the window, right tenant/org headers if multi-tenancy is on). A healthy Loki with no shipper looks fine and holds nothing, which is why make validate demands streams, not readiness.

Why is a per-request-ID label a bad idea in Loki?

Labels define streams, and Loki's index is per stream. A unique label value per request creates a stream per request: index explosion, slow queries, high memory. Put high-cardinality identifiers in the log line and filter or parse them at query time; that is what Loki's design is optimized for.

Turn "alert me when payments start failing" into a LogQL-based rule, conceptually.

A metric query over logs: sum(rate({namespace="payments"} |= "payment_failed" [5m])) > 0.1, evaluated by Loki's ruler (or mirrored into Prometheus via a recording pipeline). Same alerting mechanics as 4.2; the only difference is where the series comes from.

A dashboard imported from a ConfigMap shows "datasource not found" on every panel. What did the export get wrong, and what is the durable fix?

The JSON was exported "for sharing externally", so datasources are ${DS_PROMETHEUS} inputs (or reference a uid that does not exist here). Fix by provisioning the datasource with a stable uid and exporting with the sharing option off so the uid is literal, or with the Grafana Operator's GrafanaDashboard.spec.datasources[] mapping inputName to datasourceName.

Why is a trace ID a bad Loki label but a good candidate for structured metadata, and how do you still search it quickly?

A label defines a stream; a unique value per request means a stream per request, which blows up the index and chunk count. Structured metadata attaches key-values to individual lines without creating streams and remains filterable with label-filter syntax (| trace_id="abc"). Alloy and the OTLP ingest path put high-cardinality attributes there by default.

Your shipper is a Promtail DaemonSet from an older tutorial. Why is that a problem in 2026, and what replaces it?

Promtail reached end-of-life on 2 March 2026: no patches, no security fixes. Grafana Alloy replaces it (discovery.kubernetes, loki.source.kubernetes, loki.process, loki.write), and alloy convert --source-format=promtail migrates the configuration. An OpenTelemetry Collector with the filelog receiver or Fluent Bit are the vendor-neutral alternatives.

Docs to know your way around

study time, not exam time
  • grafana.com/docs: provisioning (datasources and dashboards), dashboard variables, and the LogQL reference.
  • The Explore UI's query builder doubles as LogQL documentation under time pressure.
  • Offline: stern --help, kubectl logs --previous, and the Grafana panel inspector (it shows the exact query and the raw response).
  • grafana.com/docs/grafana/latest/administration/provisioning: datasource and dashboard provider file formats (uid, jsonData, secureJsonData, foldersFromFilesStructure); grafana.github.io/grafana-operator/docs/api for the CRD fields.
  • grafana.com/docs/loki/latest/query/log_queries and /query/metric_queries: parsers, __error__, unwrap; /get-started/labels for the structured-metadata guidance.
  • grafana.com/docs/alloy/latest/set-up/migrate/from-promtail: the EOL statement and the component mapping.