Two systems, one handoff. Prometheus evaluates rules and decides what is true. Alertmanager receives firing alerts and decides who hears about it.

needsmake up obs

Orientation

competency 4.1 · alerting and notification routing

"Where do I configure X" questions are really "which side of the handoff is X" questions. Thresholds, durations and labels: Prometheus. Grouping, routing, silencing, inhibition, receivers: Alertmanager. Get that boundary right and the rest is syntax.

Prometheus                                       Alertmanager
   rule groups, evaluated every 30s              receives firing alerts via HTTP
   expr → Inactive → Pending (for:) → Firing ──▶ route → group → inhibit → silence → receiver
   labels decide routing                         group_wait / group_interval / repeat_interval
   annotations describe                          webhook · email · Slack · PagerDuty

Rules

PrometheusRule · states · recording rules

A PrometheusRule CRD holds groups of rules. An alerting rule is a PromQL expression plus:

  • for: how long the expression must hold before Pending becomes Firing. This is your flap filter, and picking it is a real decision: too short and you page on transients, too long and you notice late.
  • labels: routing material. severity, team, service. Whatever your routes match on must be produced here.
  • annotations: human material: summary, description, runbook_url. Templated with {{ $labels.x }} and {{ $value }}.
  • keep_firing_for is the opposite of for: keeps an alert firing briefly after the expression goes false, to damp flapping resolutions.

Recording rules are the other half of the CRD: they precompute an expensive expression into a new series on a schedule (job:http_errors:rate5m, by convention level:metric:operation). Dashboards and alerts then read a cheap series. If a scenario says "this dashboard takes 30 seconds to load", a recording rule is the expected answer.

The same selector story as ServiceMonitors

Rules are only evaluated if they match the Prometheus object's ruleSelector. This lab's is {} (everything); a stock install's is the release label. An unmatched rule produces no evaluation and no error. Read the selector before blaming the rule, and note that this is yet another place in domain 4 where a selector silently discards correct configuration.

Alert states, and where each is visible

Inactive → Pending (expression true, for not yet satisfied) → Firing. Pending alerts appear in Prometheus's Alerts page and nowhere else; they have not been sent. That matters when someone asks why nothing reached the receiver "even though the alert is showing". Add evaluation interval + for + group_wait to compute how long a fire genuinely takes; on a stock stack that is easily 2–3 minutes.

Add up the settings before you conclude that alerting is broken:

What makes an alert worth having

  • Symptom over cause. Alert on user-visible failure (error rate, latency, unavailability), not on every internal cause. Cause-based alerts multiply; symptom-based ones stay bounded.
  • Actionable. If nobody would do anything at 3am, it is a dashboard panel, not a page.
  • Alert on the SLO, if you can. Burn-rate alerting (fast burn on a short window, slow burn on a long one) is the modern form: it pages on error budget consumption instead of arbitrary thresholds. Even knowing the phrase "multi-window multi-burn-rate" signals you have read the SRE material.
  • Documented. A runbook_url annotation is the cheapest reliability improvement in this section.

Routing

tree · grouping · silence · inhibition

Note the order: routing happens first. The dispatcher matches an alert against the route tree, and the route it lands on decides the group (its group_by) and the timers; only then does the per-group pipeline run inhibition, silences, waiting and de-duplication before notifying. Reading it the other way round leads to wrong conclusions like "the silence should have stopped it from grouping".

Alertmanager's config is a tree of routes. Each alert enters at the root and descends to the most specific matching route; continue: true lets it match siblings too. A route names a receiver and grouping behavior:

SettingMeansTypical
group_bycollapse alerts sharing these labels into one notification[alertname, namespace]
group_waitwait before the first send, to batch siblings30s
group_intervalwait before sending new members of an existing group5m
repeat_intervalhow often to re-notify about an unresolved group4h–12h
matcherswhich alerts take this branchseverity="critical"

Silences mute matching alerts for a time window without touching config: create them in the UI or via amtool. They change notification, never truth: Prometheus still shows the alert firing. Inhibition suppresses alerts when a related, more severe one is already firing (node down inhibits everything on that node), matched by source_matchers, target_matchers and an equal label list. Distinguishing silence (temporary, human, targeted) from inhibition (permanent rule, relationship-based) is a fair exam question.

Verify the rendered truth

In this stack the live config is generated by the operator into a Secret, and the UI's Status page shows the rendered result. The operator also supports namespaced AlertmanagerConfig CRDs that merge into the tree; whether an Alertmanager picks them up depends on its alertmanagerConfigSelector. So verify pickup in the rendered config rather than trusting the apply. That habit transfers to every operator-managed config on the exam.

One special alert to recognize: Watchdog, which fires always, by design. It is a dead-man's switch: an external system watches for it and alerts when it stops arriving. That is how you detect that the alerting pipeline itself died. If a question asks why an always-firing alert is a feature, that is the answer.

Alertmanager configuration, field by field

matchers, time intervals, inhibition, AlertmanagerConfig scoping

Route fields

  • matchers is a list of strings with Prometheus selector syntax: severity="critical", team=~"platform|sre", namespace!="". The older match/match_re maps are deprecated but still parse. A child route with no matchers matches everything, which is a common way to accidentally swallow alerts meant for a later sibling.
  • continue (default false) lets an alert keep matching later siblings after this one; without it the first matching child wins. The root route must not have matchers and is the fallback receiver.
  • group_by inherits from the parent unless set; ['...'] disables grouping (one notification per alert). Defaults: group_wait: 30s, group_interval: 5m, repeat_interval: 4h. repeat_interval must be a multiple of group_interval to behave as you expect.
  • receiver names an entry in receivers[]; each receiver has one or more *_configs (webhook_configs, slack_configs, pagerduty_configs, email_configs, opsgenie_configs, msteams_configs). send_resolved: true is what makes "recovered" messages arrive.

Time-based muting

Top-level time_intervals (the older key mute_time_intervals is deprecated) defines named windows; each has a list of periods with times (start_time/end_time in HH:MM, end exclusive, 24:00 allowed), weekdays (['monday:friday']), days_of_month (['1:7', '-1']), months, years, and location (an IANA zone; UTC if omitted). A route then references them: mute_time_intervals: [weekends] suppresses notifications while the interval matches; active_time_intervals: [business_hours] suppresses them whenever it does not match. If both are set on a route, mute wins. Muted alerts still fire in Prometheus and still appear in the Alertmanager UI as active; only the notification is held, and it is sent when the window ends if the alert is still firing.

Inhibition and silences

  • An inhibit_rules[] entry has source_matchers (the alert that must be firing), target_matchers (the alerts to mute) and equal (labels that must carry the same value on both). A missing label and an empty label are treated as equal, so equal: [namespace] with a cluster-wide source inhibits every target that also lacks namespace. An alert matching both sides never inhibits itself.
  • Silences are matchers plus a start, an end and a comment, created in the UI, via amtool silence add alertname=X --duration 2h --comment ..., or by POSTing to /api/v2/silences. They are stored in Alertmanager's own data, not in the config, and are lost with the PVC if there is none. amtool config routes test --config.file am.yml severity=critical team=db prints which receiver a label set lands on; amtool check-config validates the file.

Operator-managed routing

Three ways to feed the Alertmanager object: spec.configSecret (a whole hand-written config), spec.alertmanagerConfiguration (one AlertmanagerConfig in the same namespace that becomes the global config), and namespaced AlertmanagerConfig objects selected by alertmanagerConfigSelector and alertmanagerConfigNamespaceSelector. The merged ones are scoped for safety: the operator adds a namespace="<its namespace>" matcher to every route and inhibition rule they contribute, so a team's AlertmanagerConfig can only route alerts from its own namespace. That is why a tenant's config "does not match" a cluster-wide alert, and why the field names inside the CRD are camelCase (groupBy, groupWait, repeatInterval, matchers[].name/value/matchType) while the rendered Secret uses the snake_case above. The rendered truth is in the Secret alertmanager-<name>-generated and on the UI's Status page.

Trap

Prometheus's own Alertmanager discovery is the other half of the handoff: the Prometheus object's spec.alerting.alertmanagers[] (namespace, name, port) must point at the Alertmanager Service. A stack where rules fire but Alertmanager shows nothing, and no silence or route explains it, usually has this pointer wrong or the Alertmanager Service port renamed. Status → Runtime & Build Information → Alertmanagers in the Prometheus UI lists what it discovered.

SLO burn-rate alerts and testing rules

the multi-window table, alert hygiene, promtool test rules

An SLO of 99.9% over 30 days gives an error budget of 0.1% of requests, about 43 minutes of full outage per month. Burn rate is "how many times faster than budget-neutral you are consuming it": burn rate 1 spends exactly the month's budget in a month. The SRE Workbook's standard ladder pairs a long window (decides) with a short window one twelfth its length (stops the alert once the problem is gone):

Burn rateLong windowShort windowBudget consumed at triggerAction
14.41h5m2% of the monthpage
66h30m5%page
31d2h10%ticket
13d6h10%ticket

For a 99.9% SLO the 1h rule reads error_ratio_1h > 14.4 * 0.001 and error_ratio_5m > 14.4 * 0.001, where each ratio is a recording rule such as sum(rate(http_requests_total{code=~"5.."}[1h])) / sum(rate(http_requests_total[1h])). Recording rules are not optional here: eight windows of raw division per service is exactly the slow-dashboard problem. Name them slo:sli_error:ratio_rate1h and friends, and put severity on the label so routing follows the table.

Alert hygiene that graders and reviewers look for

  • One severity taxonomy (critical pages, warning tickets, info dashboards) and routes that honor it; nothing pages without a runbook_url.
  • for long enough to survive a scrape gap (at least two scrape intervals), keep_firing_for when a symptom flaps at the threshold.
  • Templating in annotations: {{ $labels.namespace }}, {{ $value | humanize }}, {{ $value | humanizePercentage }}, {{ $externalLabels.cluster }}. Labels can be templated too, but a templated label that varies per evaluation creates a new alert identity each time and defeats grouping.
  • The ALERTS{alertname, alertstate="pending|firing"} series Prometheus writes for every rule is how you graph alert history and count noisy rules: count by (alertname) (changes(ALERTS[7d])).
  • Alertmanager's own health: alertmanager_notifications_failed_total by integration, alertmanager_alerts_invalid_total, and the Watchdog route to a dead man's switch receiver.

Test rules before you apply them

promtool check rules rules.yml catches syntax; promtool test rules test.yml runs rules against synthetic series. The test file names the rule files, an evaluation_interval, and test cases with input_series (a selector plus values in the compact 0+10x5 notation: start 0, add 10, five times), then alert_rule_test entries that state which alerts must be firing at eval_time with exact labels and annotations, or promql_expr_test entries with exp_samples. Rules written in a PrometheusRule test the same way once you extract spec.groups into a plain file.

rule_files: [rules.yml]
evaluation_interval: 1m
tests:
  - interval: 1m
    input_series:
      - series: 'up{job="payments", instance="a"}'
        values: '1 1 0 0 0 0 0'
    alert_rule_test:
      - alertname: PaymentsDown
        eval_time: 6m
        exp_alerts:
          - exp_labels: { severity: critical, job: payments, instance: a }
            exp_annotations: { summary: "payments target a is down" }
How this gets tested

Tasks read "alert when the error rate exceeds X for Y minutes, route severity="critical" to receiver Z, and do not page on weekends". That is a PrometheusRule with for and a severity label, a route with matchers and a receiver, and a time_intervals entry referenced by mute_time_intervals. Prove the routing with amtool config routes test or the rendered config, not by waiting for Saturday.

Exercises

tick the dot when its check passes

Alertmanager has no LoadBalancer here: kubectl -n monitoring port-forward svc/prometheus-kube-prometheus-alertmanager 9093:9093 and browse localhost:9093.

kubectl apply -f - <<'EOF'
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
  name: curriculum-drill
  namespace: monitoring
  labels: { release: prometheus }
spec:
  groups:
    - name: drill
      rules:
        - alert: TooManyExamplePods
          expr: count(kube_pod_info{namespace="default"}) > 0
          for: 1m
          labels: { severity: warning, team: platform }
          annotations:
            summary: "{{ $value }} pods in default"
            description: "Drill alert; fires whenever default has any pods."
EOF
outputcaptured 2026-08-26
$ kubectl apply -f - <<'EOF'
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
  name: curriculum-drill
  namespace: monitoring
  labels: { release: prometheus }
spec:
  groups:
    - name: drill
      rules:
        - alert: TooManyExamplePods
          expr: count(kube_pod_info{namespace="default"}) > 0
          for: 1m
          labels: { severity: warning, team: platform }
          annotations:
            summary: "{{ $value }} pods in default"
            description: "Drill alert; fires whenever default has any pods."
EOF
prometheusrule.monitoring.coreos.com/curriculum-drill created

Watch it walk the states: Prometheus UI → Alerts shows Pending, then Firing after the minute.

verify: it appears in the Alertmanager UI with your labels, grouped, and the annotation rendered the live value rather than the template text.

In the Alertmanager UI, create a silence matching alertname=TooManyExamplePods for 2 hours with a comment.

verify: the alert leaves the active view and lands under Silences, while Prometheus still shows it Firing. That split is the lesson: silencing changes notification, never truth.

Open Status in the Alertmanager UI and answer from the rendered config: what is the root receiver, what does group_by collapse on, which special route exists for the Watchdog alert (the stack ships one that fires always, as a dead-man's switch; know why that is a feature).

verify: you can trace where your drill alert's severity: warning lands in the tree and say which single config line you would change to route team: platform alerts to a new webhook receiver.

Deploy a trivial webhook sink (kubectl create deploy sink --image=mendhak/http-https-echo:31 plus a Service), add an AlertmanagerConfig in default routing team: platform to it, then verify pickup: does the rendered config in the UI now contain the route? If yes, kubectl logs deploy/sink shows the JSON payload when the drill alert fires. Read it once; its structure (groupLabels, commonAnnotations, alerts[]) is what every receiver integration parses.

verify: either the payload arrives, or you can name the selector on the Alertmanager object that would have to change. Clean up: delete the PrometheusRule; the silence expires on its own.

A silence is a manual action with an end time. A mute time interval is policy: this route does not page on weekends, forever. Tasks ask for the second and people reach for the first.

kubectl apply -f - <<'EOF'
apiVersion: monitoring.coreos.com/v1alpha1
kind: AlertmanagerConfig
metadata: { name: quiet-hours, namespace: default, labels: { release: prometheus } }
spec:
  route:
    groupBy: [alertname]
    receiver: blackhole
    muteTimeIntervals: [always-now]
  receivers:
    - name: blackhole
  muteTimeIntervals:
    - name: always-now
      timeIntervals:
        - times:
            - { startTime: "00:00", endTime: "24:00" }
EOF
sleep 45
kubectl -n monitoring get secret alertmanager-prometheus-kube-prometheus-alertmanager-generated -o jsonpath='{.data.alertmanager\.yaml\.gz}' | base64 -d | gunzip | sed -n '/mute_time_intervals/,/^[a-z]/p'
kubectl -n monitoring exec sts/alertmanager-prometheus-kube-prometheus-alertmanager -c alertmanager -- sh -c 'amtool config routes test --config.file=/etc/alertmanager/config_out/alertmanager.env.yaml alertname=TooManyExamplePods severity=warning namespace=default'
kubectl -n default delete alertmanagerconfig quiet-hours
outputcaptured 2026-09-13
$ kubectl apply -f - <<'EOF'
apiVersion: monitoring.coreos.com/v1alpha1
kind: AlertmanagerConfig
metadata: { name: quiet-hours, namespace: default, labels: { release: prometheus } }
spec:
  route:
    groupBy: [alertname]
    receiver: blackhole
    muteTimeIntervals: [always-now]
  receivers:
    - name: blackhole
  muteTimeIntervals:
    - name: always-now
      timeIntervals:
        - times:
            - { startTime: "00:00", endTime: "24:00" }
EOF
alertmanagerconfig.monitoring.coreos.com/quiet-hours created
$ sleep 45
$ kubectl -n monitoring get secret alertmanager-prometheus-kube-prometheus-alertmanager-generated -o jsonpath='{.data.alertmanager\.yaml\.gz}' | base64 -d | gunzip | sed -n '/mute_time_intervals/,/^[a-z]/p'
    mute_time_intervals:
    - default/quiet-hours/always-now
  - receiver: "null"
    matchers:
    - alertname = "Watchdog"
  group_wait: 30s
  group_interval: 5m
  repeat_interval: 12h
inhibit_rules:
mute_time_intervals:
- name: default/quiet-hours/always-now
  time_intervals:
  - times:
    - start_time: "00:00"
      end_time: "24:00"
templates:
$ kubectl -n monitoring exec sts/alertmanager-prometheus-kube-prometheus-alertmanager -c alertmanager -- sh -c 'amtool config routes test --config.file=/etc/alertmanager/config_out/alertmanager.env.yaml alertname=TooManyExamplePods severity=warning namespace=default'
default/quiet-hours/blackhole
$ kubectl -n default delete alertmanagerconfig quiet-hours
alertmanagerconfig.monitoring.coreos.com "quiet-hours" deleted from default namespace
verify: the rendered config contains your interval and the route that references it, and amtool config routes test names the receiver a matching alert would land on.

When the cluster is on fire you want one page, not forty. An inhibition rule suppresses warnings while a related critical is firing, matched on a shared label, and getting that equal list right is the entire skill.

kubectl apply -f - <<'EOF'
apiVersion: monitoring.coreos.com/v1alpha1
kind: AlertmanagerConfig
metadata: { name: inhibit-demo, namespace: default, labels: { release: prometheus } }
spec:
  route:
    groupBy: [alertname]
    receiver: blackhole
  receivers:
    - name: blackhole
  inhibitRules:
    - sourceMatch:
        - { name: alertname, value: DrillOutage }
      targetMatch:
        - { name: severity, value: warning }
      equal: [namespace]
EOF
sleep 120
kubectl -n monitoring get secret alertmanager-prometheus-kube-prometheus-alertmanager-generated -o jsonpath='{.data.alertmanager\.yaml\.gz}' | base64 -d | gunzip | grep -B4 -A4 DrillOutage
kubectl -n monitoring port-forward svc/prometheus-kube-prometheus-alertmanager 9093:9093 & PF1=$!
sleep 5
# Watchdog and the drill alert carry no namespace label, and the CRD adds namespace to both matchers,
# so nothing cluster-generated can ever match: inject a pair that can
curl -s -XPOST localhost:9093/api/v2/alerts -H 'Content-Type: application/json' -d '[{"labels":{"alertname":"DrillOutage","severity":"critical","namespace":"default"}},{"labels":{"alertname":"DrillNoise","severity":"warning","namespace":"default"}}]'
sleep 45
curl -s localhost:9093/api/v2/alerts | jq '[.[] | select(.labels.alertname|startswith("Drill")) | {alertname: .labels.alertname, state: .status.state, inhibitedBy: .status.inhibitedBy}]'
kill $PF1
kubectl -n default delete alertmanagerconfig inhibit-demo
outputcaptured 2026-09-13
$ kubectl apply -f - <<'EOF'
apiVersion: monitoring.coreos.com/v1alpha1
kind: AlertmanagerConfig
metadata: { name: inhibit-demo, namespace: default, labels: { release: prometheus } }
spec:
  route:
    groupBy: [alertname]
    receiver: blackhole
  receivers:
    - name: blackhole
  inhibitRules:
    - sourceMatch:
        - { name: alertname, value: DrillOutage }
      targetMatch:
        - { name: severity, value: warning }
      equal: [namespace]
EOF
alertmanagerconfig.monitoring.coreos.com/inhibit-demo created
$ sleep 120
$ kubectl -n monitoring get secret alertmanager-prometheus-kube-prometheus-alertmanager-generated -o jsonpath='{.data.alertmanager\.yaml\.gz}' | base64 -d | gunzip | grep -B4 -A4 DrillOutage
- target_matchers:
  - severity="warning"
  - namespace="default"
  source_matchers:
  - alertname="DrillOutage"
  - namespace="default"
  equal:
  - namespace
receivers:
$ kubectl -n monitoring port-forward svc/prometheus-kube-prometheus-alertmanager 9093:9093 & PF1=$!
$ sleep 5
Forwarding from 127.0.0.1:9093 -> 9093
Forwarding from [::1]:9093 -> 9093
$ # Watchdog and the drill alert carry no namespace label, and the CRD adds namespace to both matchers,
$ # so nothing cluster-generated can ever match: inject a pair that can
$ curl -s -XPOST localhost:9093/api/v2/alerts -H 'Content-Type: application/json' -d '[{"labels":{"alertname":"DrillOutage","severity":"critical","namespace":"default"}},{"labels":{"alertname":"DrillNoise","severity":"warning","namespace":"default"}}]'
Handling connection for 9093
$ sleep 45
$ curl -s localhost:9093/api/v2/alerts | jq '[.[] | select(.labels.alertname|startswith("Drill")) | {alertname: .labels.alertname, state: .status.state, inhibitedBy: .status.inhibitedBy}]'
Handling connection for 9093
[
  {
    "alertname": "DrillNoise",
    "state": "suppressed",
    "inhibitedBy": [
      "f3637fb5c13520d0"
    ]
  },
  {
    "alertname": "DrillOutage",
    "state": "active",
    "inhibitedBy": []
  }
]
$ kill $PF1
$ kubectl -n default delete alertmanagerconfig inhibit-demo
alertmanagerconfig.monitoring.coreos.com "inhibit-demo" deleted from default namespace
verify: the rendered config carries your rule with namespace="default" added to both matchers by the CRD, and DrillNoise comes back with inhibitedBy naming DrillOutage's fingerprint while DrillOutage itself has an empty list. Both injected alerts carry a namespace label on purpose: the stock Watchdog alert has none, so a rule that equals on namespace can never match it.

Alerting rules are code and can be tested without waiting for reality to cooperate. promtool test rules feeds synthetic series in and asserts which alerts fire, which is the only way to be sure about a for clause.

kubectl -n monitoring delete pod promtool-check --ignore-not-found
kubectl apply -f - <<'EOF'
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
  name: curriculum-drill
  namespace: monitoring
  labels: { release: prometheus }
spec:
  groups:
    - name: drill
      rules:
        - alert: TooManyExamplePods
          expr: count(kube_pod_info{namespace="default"}) > 0
          for: 1m
          labels: { severity: warning, team: platform }
          annotations:
            summary: "{{ $value }} pods in default"
            description: "Drill alert; fires whenever default has any pods."
EOF
kubectl -n monitoring get prometheusrule curriculum-drill -o jsonpath='{.spec}' | jq '{groups: .groups}' > /tmp/rules.json
python3 -c "import json,yaml; print(yaml.safe_dump(json.load(open('/tmp/rules.json'))))" > /tmp/rules.yml
cat /tmp/rules.yml
cat > /tmp/test.yml <<'EOF'
rule_files:
  - rules.yml
evaluation_interval: 1m
tests:
  - interval: 1m
    input_series:
      - series: 'kube_pod_info{namespace="default",pod="p1"}'
        values: '1+0x10'
    alert_rule_test:
      - eval_time: 30s
        alertname: TooManyExamplePods
        exp_alerts: []
      - eval_time: 5m
        alertname: TooManyExamplePods
        exp_alerts:
          - exp_labels:
              severity: warning
              team: platform
            exp_annotations:
              summary: "1 pods in default"
              description: "Drill alert; fires whenever default has any pods."
EOF
kubectl -n monitoring create configmap promtool-test --from-file=rules.yml=/tmp/rules.yml --from-file=test.yml=/tmp/test.yml --dry-run=client -o yaml | kubectl apply -f -
IMG=$(kubectl -n monitoring get sts prometheus-prometheus-kube-prometheus-prometheus -o jsonpath='{.spec.template.spec.containers[?(@.name=="prometheus")].image}'); echo "$IMG"
kubectl -n monitoring run promtool-check --rm -i --restart=Never --image="$IMG" --overrides="{\"spec\":{\"containers\":[{\"name\":\"promtool-check\",\"image\":\"$IMG\",\"workingDir\":\"/t\",\"command\":[\"promtool\",\"check\",\"rules\",\"rules.yml\"],\"volumeMounts\":[{\"name\":\"t\",\"mountPath\":\"/t\"}]}],\"volumes\":[{\"name\":\"t\",\"configMap\":{\"name\":\"promtool-test\"}}]}}"
kubectl -n monitoring run promtool-unit --rm -i --restart=Never --image="$IMG" --overrides="{\"spec\":{\"containers\":[{\"name\":\"promtool-unit\",\"image\":\"$IMG\",\"workingDir\":\"/t\",\"command\":[\"promtool\",\"test\",\"rules\",\"test.yml\"],\"volumeMounts\":[{\"name\":\"t\",\"mountPath\":\"/t\"}]}],\"volumes\":[{\"name\":\"t\",\"configMap\":{\"name\":\"promtool-test\"}}]}}"
kubectl -n monitoring delete configmap promtool-test
# otherwise TooManyExamplePods fires on this cluster forever
kubectl -n monitoring delete prometheusrule curriculum-drill
kubectl -n monitoring delete pod promtool-check --ignore-not-found
outputcaptured 2026-09-12
$ kubectl -n monitoring delete pod promtool-check --ignore-not-found
pod "promtool-check" deleted from monitoring namespace
$ kubectl apply -f - <<'EOF'
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
  name: curriculum-drill
  namespace: monitoring
  labels: { release: prometheus }
spec:
  groups:
    - name: drill
      rules:
        - alert: TooManyExamplePods
          expr: count(kube_pod_info{namespace="default"}) > 0
          for: 1m
          labels: { severity: warning, team: platform }
          annotations:
            summary: "{{ $value }} pods in default"
            description: "Drill alert; fires whenever default has any pods."
EOF
prometheusrule.monitoring.coreos.com/curriculum-drill created
$ kubectl -n monitoring get prometheusrule curriculum-drill -o jsonpath='{.spec}' | jq '{groups: .groups}' > /tmp/rules.json
$ python3 -c "import json,yaml; print(yaml.safe_dump(json.load(open('/tmp/rules.json'))))" > /tmp/rules.yml
$ cat /tmp/rules.yml
groups:
- name: drill
  rules:
  - alert: TooManyExamplePods
    annotations:
      description: Drill alert; fires whenever default has any pods.
      summary: '{{ $value }} pods in default'
    expr: count(kube_pod_info{namespace="default"}) > 0
    for: 1m
    labels:
      severity: warning
      team: platform
$ cat > /tmp/test.yml <<'EOF'
rule_files:
  - rules.yml
evaluation_interval: 1m
tests:
  - interval: 1m
    input_series:
      - series: 'kube_pod_info{namespace="default",pod="p1"}'
        values: '1+0x10'
    alert_rule_test:
      - eval_time: 30s
        alertname: TooManyExamplePods
        exp_alerts: []
      - eval_time: 5m
        alertname: TooManyExamplePods
        exp_alerts:
          - exp_labels:
              severity: warning
              team: platform
            exp_annotations:
              summary: "1 pods in default"
              description: "Drill alert; fires whenever default has any pods."
EOF
$ kubectl -n monitoring create configmap promtool-test --from-file=rules.yml=/tmp/rules.yml --from-file=test.yml=/tmp/test.yml --dry-run=client -o yaml | kubectl apply -f -
configmap/promtool-test created
$ IMG=$(kubectl -n monitoring get sts prometheus-prometheus-kube-prometheus-prometheus -o jsonpath='{.spec.template.spec.containers[?(@.name=="prometheus")].image}'); echo "$IMG"
quay.io/prometheus/prometheus:v3.14.0-distroless
$ kubectl -n monitoring run promtool-check --rm -i --restart=Never --image="$IMG" --overrides="{\"spec\":{\"containers\":[{\"name\":\"promtool-check\",\"image\":\"$IMG\",\"workingDir\":\"/t\",\"command\":[\"promtool\",\"check\",\"rules\",\"rules.yml\"],\"volumeMounts\":[{\"name\":\"t\",\"mountPath\":\"/t\"}]}],\"volumes\":[{\"name\":\"t\",\"configMap\":{\"name\":\"promtool-test\"}}]}}"
Checking rules.yml
  SUCCESS: 1 rules found

All commands and output from this session will be recorded in container logs, including credentials and sensitive information passed through the command prompt.
If you don't see a command prompt, try pressing enter.
warning: couldn't attach to pod/promtool-check, falling back to streaming logs: unable to upgrade connection: container promtool-check not found in pod promtool-check_monitoring
Checking rules.yml
  SUCCESS: 1 rules found

pod "promtool-check" deleted from monitoring namespace
$ kubectl -n monitoring run promtool-unit --rm -i --restart=Never --image="$IMG" --overrides="{\"spec\":{\"containers\":[{\"name\":\"promtool-unit\",\"image\":\"$IMG\",\"workingDir\":\"/t\",\"command\":[\"promtool\",\"test\",\"rules\",\"test.yml\"],\"volumeMounts\":[{\"name\":\"t\",\"mountPath\":\"/t\"}]}],\"volumes\":[{\"name\":\"t\",\"configMap\":{\"name\":\"promtool-test\"}}]}}"
  SUCCESS

All commands and output from this session will be recorded in container logs, including credentials and sensitive information passed through the command prompt.
If you don't see a command prompt, try pressing enter.
warning: couldn't attach to pod/promtool-unit, falling back to streaming logs: unable to upgrade connection: container promtool-unit not found in pod promtool-unit_monitoring
  SUCCESS

pod "promtool-unit" deleted from monitoring namespace
$ kubectl -n monitoring delete configmap promtool-test
configmap "promtool-test" deleted from monitoring namespace
$ # otherwise TooManyExamplePods fires on this cluster forever
$ kubectl -n monitoring delete prometheusrule curriculum-drill
prometheusrule.monitoring.coreos.com "curriculum-drill" deleted from monitoring namespace
$ kubectl -n monitoring delete pod promtool-check --ignore-not-found
verify: check rules passes and test rules reports SUCCESS, including the case at 30 seconds where the for clause has not elapsed. Change the for to 10m and watch the same test fail; that is the assertion earning its keep.

A good availability alert fires fast on a big burn and slow on a small one, and it does that with two windows over the same ratio. Build it on a recording rule so the alert expression stays readable.

kubectl apply -f - <<'EOF'
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata: { name: burn-rate, namespace: monitoring, labels: { release: prometheus } }
spec:
  groups:
    - name: sli
      interval: 30s
      rules:
        - record: job:request_error_ratio:rate5m
          expr: |
            sum(rate(apiserver_request_total{code=~"5.."}[5m]))
            /
            sum(rate(apiserver_request_total[5m]))
        - record: job:request_error_ratio:rate1h
          expr: |
            sum(rate(apiserver_request_total{code=~"5.."}[1h]))
            /
            sum(rate(apiserver_request_total[1h]))
        - alert: ErrorBudgetBurn
          expr: |
            job:request_error_ratio:rate5m > (14.4 * 0.001)
            and
            job:request_error_ratio:rate1h > (14.4 * 0.001)
          for: 2m
          labels: { severity: critical, team: platform }
          annotations:
            summary: "burning the error budget 14.4x faster than allowed"
EOF
sleep 120
kubectl -n monitoring port-forward svc/prometheus-kube-prometheus-prometheus 9090:9090 & PF1=$!
sleep 5
curl -sG localhost:9090/api/v1/query --data-urlencode 'query=job:request_error_ratio:rate5m' | jq '.data.result[0].value[1]'
curl -s localhost:9090/api/v1/rules | jq '.data.groups[] | select(.name=="sli") | .rules[] | {name: (.name // .alert), type, health, state}'
kill $PF1
kubectl -n monitoring delete prometheusrule burn-rate
outputcaptured 2026-09-12
$ kubectl apply -f - <<'EOF'
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata: { name: burn-rate, namespace: monitoring, labels: { release: prometheus } }
spec:
  groups:
    - name: sli
      interval: 30s
      rules:
        - record: job:request_error_ratio:rate5m
          expr: |
            sum(rate(apiserver_request_total{code=~"5.."}[5m]))
            /
            sum(rate(apiserver_request_total[5m]))
        - record: job:request_error_ratio:rate1h
          expr: |
            sum(rate(apiserver_request_total{code=~"5.."}[1h]))
            /
            sum(rate(apiserver_request_total[1h]))
        - alert: ErrorBudgetBurn
          expr: |
            job:request_error_ratio:rate5m > (14.4 * 0.001)
            and
            job:request_error_ratio:rate1h > (14.4 * 0.001)
          for: 2m
          labels: { severity: critical, team: platform }
          annotations:
            summary: "burning the error budget 14.4x faster than allowed"
EOF
prometheusrule.monitoring.coreos.com/burn-rate created
$ sleep 120
$ kubectl -n monitoring port-forward svc/prometheus-kube-prometheus-prometheus 9090:9090 & PF1=$!
$ sleep 5
Forwarding from 127.0.0.1:9090 -> 9090
Forwarding from [::1]:9090 -> 9090
$ curl -sG localhost:9090/api/v1/query --data-urlencode 'query=job:request_error_ratio:rate5m' | jq '.data.result[0].value[1]'
Handling connection for 9090
"0"
$ curl -s localhost:9090/api/v1/rules | jq '.data.groups[] | select(.name=="sli") | .rules[] | {name: (.name // .alert), type, health, state}'
Handling connection for 9090
{
  "name": "job:request_error_ratio:rate5m",
  "type": "recording",
  "health": "ok",
  "state": null
}
{
  "name": "job:request_error_ratio:rate1h",
  "type": "recording",
  "health": "ok",
  "state": null
}
{
  "name": "ErrorBudgetBurn",
  "type": "alerting",
  "health": "ok",
  "state": "inactive"
}
$ kill $PF1
$ kubectl -n monitoring delete prometheusrule burn-rate
prometheusrule.monitoring.coreos.com "burn-rate" deleted from monitoring namespace
verify: both recording rules are healthy and the alert is inactive on a quiet cluster.

The AlertmanagerConfig objects you write are merged into one file by the operator, with a namespace matcher injected into every route. Reading the rendered result is how you find out what your object actually became.

# nothing of yours is in the rendered file until an AlertmanagerConfig exists
kubectl apply -f - <<'EOF'
apiVersion: monitoring.coreos.com/v1alpha1
kind: AlertmanagerConfig
metadata: { name: read-me, namespace: default, labels: { release: prometheus } }
spec:
  route:
    groupBy: [alertname]
    receiver: mine
  receivers:
    - name: mine
EOF
sleep 60
kubectl -n monitoring get secret alertmanager-prometheus-kube-prometheus-alertmanager-generated -o jsonpath='{.data.alertmanager\.yaml\.gz}' | base64 -d | gunzip | head -60
kubectl -n monitoring get alertmanagerconfig -A
kubectl -n monitoring get alertmanager -o jsonpath='{.items[0].spec.alertmanagerConfigSelector}{"\n"}'
kubectl -n default delete alertmanagerconfig read-me
outputcaptured 2026-09-12
$ # nothing of yours is in the rendered file until an AlertmanagerConfig exists
$ kubectl apply -f - <<'EOF'
apiVersion: monitoring.coreos.com/v1alpha1
kind: AlertmanagerConfig
metadata: { name: read-me, namespace: default, labels: { release: prometheus } }
spec:
  route:
    groupBy: [alertname]
    receiver: mine
  receivers:
    - name: mine
EOF
alertmanagerconfig.monitoring.coreos.com/read-me created
$ sleep 60
$ kubectl -n monitoring get secret alertmanager-prometheus-kube-prometheus-alertmanager-generated -o jsonpath='{.data.alertmanager\.yaml\.gz}' | base64 -d | gunzip | head -60
global:
  resolve_timeout: 5m
route:
  receiver: "null"
  group_by:
  - namespace
  routes:
  - receiver: default/read-me/mine
    group_by:
    - alertname
    matchers:
    - namespace="default"
    continue: true
  - receiver: "null"
    matchers:
    - alertname = "Watchdog"
  group_wait: 30s
  group_interval: 5m
  repeat_interval: 12h
inhibit_rules:
- target_matchers:
  - severity =~ warning|info
  source_matchers:
  - severity = critical
  equal:
  - namespace
  - alertname
- target_matchers:
  - severity = info
  source_matchers:
  - severity = warning
  equal:
  - namespace
  - alertname
- target_matchers:
  - severity = info
  source_matchers:
  - alertname = InfoInhibitor
  equal:
  - namespace
... 7 more lines
$ kubectl -n monitoring get alertmanagerconfig -A
NAMESPACE   NAME      AGE
default     read-me   60s
$ kubectl -n monitoring get alertmanager -o jsonpath='{.items[0].spec.alertmanagerConfigSelector}{"\n"}'
{}
$ kubectl -n default delete alertmanagerconfig read-me
alertmanagerconfig.monitoring.coreos.com "read-me" deleted from default namespace
verify: the rendered file carries a route you never typed: your read-me config appears with namespace = "default" added to its matchers, and its receiver is renamed default/read-me/mine, prefixed with the namespace and the object name. That prefixing is why a receiver name in your object does not match the one in the UI. Without an AlertmanagerConfig of your own there is nothing of yours in the file, only the chart's null receiver and the stock inhibit rules.

Prometheus does not notify anyone; it posts to Alertmanager, and the discovery for that is in the Prometheus object rather than anywhere in your rules. On a cluster with alerts that never arrive, this is the second thing to check after the rule itself.

kubectl -n monitoring get prometheus -o jsonpath='{.items[0].spec.alerting}' | jq
kubectl -n monitoring port-forward svc/prometheus-kube-prometheus-prometheus 9090:9090 & PF1=$!
sleep 5
curl -s localhost:9090/api/v1/alertmanagers | jq '{active: [.data.activeAlertmanagers[].url], dropped: [.data.droppedAlertmanagers[].url]}'
kill $PF1
outputcaptured 2026-09-12
$ kubectl -n monitoring get prometheus -o jsonpath='{.items[0].spec.alerting}' | jq
{
  "alertmanagers": [
    {
      "apiVersion": "v2",
      "name": "prometheus-kube-prometheus-alertmanager",
      "namespace": "monitoring",
      "pathPrefix": "/",
      "port": "http-web"
    }
  ]
}
$ kubectl -n monitoring port-forward svc/prometheus-kube-prometheus-prometheus 9090:9090 & PF1=$!
$ sleep 5
Forwarding from 127.0.0.1:9090 -> 9090
Forwarding from [::1]:9090 -> 9090
$ curl -s localhost:9090/api/v1/alertmanagers | jq '{active: [.data.activeAlertmanagers[].url], dropped: [.data.droppedAlertmanagers[].url]}'
Handling connection for 9090
{
  "active": [
    "http://10.244.2.68:9093/api/v2/alerts"
  ],
  "dropped": [
    "http://10.244.2.242:7946/api/v2/alerts",
    "http://10.244.2.242:3100/api/v2/alerts",
    "http://10.244.2.242:9095/api/v2/alerts",
    "http://10.244.1.157:12345/api/v2/alerts",
    "http://10.244.2.3:10250/api/v2/alerts",
    "http://10.244.1.103:9090/api/v2/alerts",
    "http://10.244.1.103:8080/api/v2/alerts",
    "http://10.244.1.103:8081/api/v2/alerts",
    "http://10.244.2.192:9115/api/v2/alerts",
    "http://10.244.2.68:8080/api/v2/alerts",
    "http://10.244.2.68:9094/api/v2/alerts",
    "http://10.244.2.68:9094/api/v2/alerts",
    "http://10.244.2.68:8081/api/v2/alerts",
    "http://10.244.2.160:8080/api/v2/alerts",
    "http://172.18.0.3:9100/api/v2/alerts",
    "http://172.18.0.4:9100/api/v2/alerts",
    "http://172.18.0.5:9100/api/v2/alerts",
    "http://10.244.1.75:3000/api/v2/alerts",
    "http://10.244.1.75:9094/api/v2/alerts",
    "http://10.244.1.75:9094/api/v2/alerts",
    "http://10.244.1.75:6060/api/v2/alerts",
    "http://10.244.2.68:9094/api/v2/alerts",
    "http://10.244.2.68:9094/api/v2/alerts",
    "http://10.244.2.68:9093/api/v2/alerts",
    "http://10.244.2.68:8080/api/v2/alerts",
    "http://10.244.2.68:8081/api/v2/alerts",
    "http://10.244.1.103:9090/api/v2/alerts",
    "http://10.244.1.103:8080/api/v2/alerts",
    "http://10.244.1.103:8081/api/v2/alerts",
    "http://10.244.2.184:8080/api/v2/alerts",
    "http://10.244.2.184:8081/api/v2/alerts",
    "http://10.244.2.242:3100/api/v2/alerts",
    "http://10.244.2.160:8080/api/v2/alerts",
    "http://10.244.2.242:9095/api/v2/alerts",
... 6 more lines
$ kill $PF1
verify: the spec names the Alertmanager Service and port, and the API confirms at least one active endpoint. An empty active list with healthy rules is the whole diagnosis for "it says firing but nobody was paged".

Self-check

answer before opening
An alert shows Firing in Prometheus and nothing arrived. Three candidate causes?

A silence matches it; an inhibition rule suppresses it; or routing sent it to a receiver that is failing (check Alertmanager's own logs and its alertmanager_notifications_failed_total). A fourth, if it only just fired: group_wait has not elapsed.

Your rule applies cleanly and never evaluates. What did you not check?

The Prometheus object's ruleSelector (and its namespace selector). Unmatched rules are ignored silently, the same way ServiceMonitors fail, with the same first command: read the selector.

Silence versus inhibition: when do you use each?

Silence: temporary, human-initiated, targeted at a known maintenance or a known-noisy alert, with an expiry and a comment. Inhibition: a standing rule expressing a relationship: when the cause alert fires, suppress its downstream symptoms. Silences are operations; inhibitions are design.

How long, roughly, from "condition becomes true" to "notification sent" on a stock stack?

Up to one evaluation interval (30s) to notice, plus for (say 1–5m), plus group_wait (30s). So a couple of minutes minimum, which is why you wait before declaring an alerting pipeline broken, and why for: 15m on a page-worthy symptom is usually too slow.

What is the Watchdog alert for?

It fires permanently as a dead-man's switch: an external system expects to keep receiving it and alerts when it stops, catching the failure mode where Prometheus or Alertmanager itself dies and therefore cannot alert you about anything. Self-monitoring has to come from outside.

"Notify the database team only during business hours, page everyone else at any time." Which fields carry that?

A top-level time_intervals entry (times 09:00-17:00, weekdays monday:friday, a location), and on the database route active_time_intervals: [business_hours], which suppresses notifications outside the window. The other routes carry no interval. mute_time_intervals is the inverse (suppress while it matches); if both are present on one route, mute wins.

A team's AlertmanagerConfig in namespace team-a routes alertname="NodeDown" to their webhook, is selected, appears in the rendered config, and never fires. Why?

The operator scopes merged AlertmanagerConfigs with an added namespace="team-a" matcher on every route it contributes; NodeDown carries no such label, so the route can never match. Cluster-wide alerts must be routed from the global config (spec.configSecret or spec.alertmanagerConfiguration), where the namespace matcher is not enforced.

Explain the 14.4 in a burn-rate alert and why the rule has two windows.

14.4 is the burn rate at which 2% of a 30-day budget is consumed in one hour (30 days × 24 hours × 0.02 = 14.4). The long window (1h) decides that the burn is sustained; the short window (5m) must also exceed it so the alert stops as soon as the error rate falls, instead of firing for the rest of the hour on old data. Both conditions are joined with and.

How do you prove a new alerting rule fires under the right conditions without waiting for production to break?

promtool test rules with a test file: rule_files, evaluation_interval, input_series with synthetic values, and alert_rule_test cases asserting exp_alerts (labels and annotations) at a given eval_time. Pair it with promtool check rules for syntax, and amtool config routes test to prove where the labels route.

Docs to know your way around

study time, not exam time
  • prometheus.io: alerting rules, and the Alertmanager configuration page (route and inhibit_rule syntax).
  • prometheus-operator.dev: PrometheusRule and AlertmanagerConfig CRD references.
  • sre.google: the SRE workbook chapter on alerting on SLOs, for burn-rate vocabulary.
  • Offline: amtool config routes show / amtool config routes test if the binary is present, and the Alertmanager UI's Status page.
  • prometheus.io/docs/alerting/latest/configuration: the <route>, <inhibit_rule> and <time_interval> sections are the ones to have bookmarked; note which keys are deprecated.
  • prometheus.io/docs/prometheus/latest/configuration/unit_testing_rules: the promtool test rules file format; sre.google/workbook/alerting-on-slos for the burn-rate table.
  • prometheus-operator.dev/docs/developer/alerting: the three ways to configure the Alertmanager object and the namespace-scoping rule for merged AlertmanagerConfigs.