Two systems, one handoff. Prometheus evaluates rules and decides what is true. Alertmanager receives firing alerts and decides who hears about it.
make up obsOrientation
"Where do I configure X" questions are really "which side of the handoff is X" questions. Thresholds, durations and labels: Prometheus. Grouping, routing, silencing, inhibition, receivers: Alertmanager. Get that boundary right and the rest is syntax.
Prometheus Alertmanager rule groups, evaluated every 30s receives firing alerts via HTTP expr → Inactive → Pending (for:) → Firing ──▶ route → group → inhibit → silence → receiver labels decide routing group_wait / group_interval / repeat_interval annotations describe webhook · email · Slack · PagerDuty
Rules
A PrometheusRule CRD holds groups of rules. An alerting rule is a PromQL expression plus:
for: how long the expression must hold before Pending becomes Firing. This is your flap filter, and picking it is a real decision: too short and you page on transients, too long and you notice late.labels: routing material.severity,team,service. Whatever your routes match on must be produced here.annotations: human material:summary,description,runbook_url. Templated with{{ $labels.x }}and{{ $value }}.keep_firing_foris the opposite offor: keeps an alert firing briefly after the expression goes false, to damp flapping resolutions.
Recording rules are the other half of the CRD: they precompute an expensive expression into a new series on a schedule (job:http_errors:rate5m, by convention level:metric:operation). Dashboards and alerts then read a cheap series. If a scenario says "this dashboard takes 30 seconds to load", a recording rule is the expected answer.
Rules are only evaluated if they match the Prometheus object's ruleSelector. This lab's is {} (everything); a stock install's is the release label. An unmatched rule produces no evaluation and no error. Read the selector before blaming the rule, and note that this is yet another place in domain 4 where a selector silently discards correct configuration.
Inactive → Pending (expression true, for not yet satisfied) → Firing. Pending alerts appear in Prometheus's Alerts page and nowhere else; they have not been sent. That matters when someone asks why nothing reached the receiver "even though the alert is showing". Add evaluation interval + for + group_wait to compute how long a fire genuinely takes; on a stock stack that is easily 2–3 minutes.
Add up the settings before you conclude that alerting is broken:
What makes an alert worth having
- Symptom over cause. Alert on user-visible failure (error rate, latency, unavailability), not on every internal cause. Cause-based alerts multiply; symptom-based ones stay bounded.
- Actionable. If nobody would do anything at 3am, it is a dashboard panel, not a page.
- Alert on the SLO, if you can. Burn-rate alerting (fast burn on a short window, slow burn on a long one) is the modern form: it pages on error budget consumption instead of arbitrary thresholds. Even knowing the phrase "multi-window multi-burn-rate" signals you have read the SRE material.
- Documented. A
runbook_urlannotation is the cheapest reliability improvement in this section.
Routing
Note the order: routing happens first. The dispatcher matches an alert against the route tree, and the route it lands on decides the group (its group_by) and the timers; only then does the per-group pipeline run inhibition, silences, waiting and de-duplication before notifying. Reading it the other way round leads to wrong conclusions like "the silence should have stopped it from grouping".
Alertmanager's config is a tree of routes. Each alert enters at the root and descends to the most specific matching route; continue: true lets it match siblings too. A route names a receiver and grouping behavior:
| Setting | Means | Typical |
|---|---|---|
| group_by | collapse alerts sharing these labels into one notification | [alertname, namespace] |
| group_wait | wait before the first send, to batch siblings | 30s |
| group_interval | wait before sending new members of an existing group | 5m |
| repeat_interval | how often to re-notify about an unresolved group | 4h–12h |
| matchers | which alerts take this branch | severity="critical" |
Silences mute matching alerts for a time window without touching config: create them in the UI or via amtool. They change notification, never truth: Prometheus still shows the alert firing. Inhibition suppresses alerts when a related, more severe one is already firing (node down inhibits everything on that node), matched by source_matchers, target_matchers and an equal label list. Distinguishing silence (temporary, human, targeted) from inhibition (permanent rule, relationship-based) is a fair exam question.
In this stack the live config is generated by the operator into a Secret, and the UI's Status page shows the rendered result. The operator also supports namespaced AlertmanagerConfig CRDs that merge into the tree; whether an Alertmanager picks them up depends on its alertmanagerConfigSelector. So verify pickup in the rendered config rather than trusting the apply. That habit transfers to every operator-managed config on the exam.
One special alert to recognize: Watchdog, which fires always, by design. It is a dead-man's switch: an external system watches for it and alerts when it stops arriving. That is how you detect that the alerting pipeline itself died. If a question asks why an always-firing alert is a feature, that is the answer.
Alertmanager configuration, field by field
Route fields
matchersis a list of strings with Prometheus selector syntax:severity="critical",team=~"platform|sre",namespace!="". The oldermatch/match_remaps are deprecated but still parse. A child route with no matchers matches everything, which is a common way to accidentally swallow alerts meant for a later sibling.continue(defaultfalse) lets an alert keep matching later siblings after this one; without it the first matching child wins. The root route must not have matchers and is the fallback receiver.group_byinherits from the parent unless set;['...']disables grouping (one notification per alert). Defaults:group_wait: 30s,group_interval: 5m,repeat_interval: 4h.repeat_intervalmust be a multiple ofgroup_intervalto behave as you expect.receivernames an entry inreceivers[]; each receiver has one or more*_configs(webhook_configs,slack_configs,pagerduty_configs,email_configs,opsgenie_configs,msteams_configs).send_resolved: trueis what makes "recovered" messages arrive.
Time-based muting
Top-level time_intervals (the older key mute_time_intervals is deprecated) defines named windows; each has a list of periods with times (start_time/end_time in HH:MM, end exclusive, 24:00 allowed), weekdays (['monday:friday']), days_of_month (['1:7', '-1']), months, years, and location (an IANA zone; UTC if omitted). A route then references them: mute_time_intervals: [weekends] suppresses notifications while the interval matches; active_time_intervals: [business_hours] suppresses them whenever it does not match. If both are set on a route, mute wins. Muted alerts still fire in Prometheus and still appear in the Alertmanager UI as active; only the notification is held, and it is sent when the window ends if the alert is still firing.
Inhibition and silences
- An
inhibit_rules[]entry hassource_matchers(the alert that must be firing),target_matchers(the alerts to mute) andequal(labels that must carry the same value on both). A missing label and an empty label are treated as equal, soequal: [namespace]with a cluster-wide source inhibits every target that also lacksnamespace. An alert matching both sides never inhibits itself. - Silences are matchers plus a start, an end and a comment, created in the UI, via
amtool silence add alertname=X --duration 2h --comment ..., or by POSTing to/api/v2/silences. They are stored in Alertmanager's own data, not in the config, and are lost with the PVC if there is none.amtool config routes test --config.file am.yml severity=critical team=dbprints which receiver a label set lands on;amtool check-configvalidates the file.
Operator-managed routing
Three ways to feed the Alertmanager object: spec.configSecret (a whole hand-written config), spec.alertmanagerConfiguration (one AlertmanagerConfig in the same namespace that becomes the global config), and namespaced AlertmanagerConfig objects selected by alertmanagerConfigSelector and alertmanagerConfigNamespaceSelector. The merged ones are scoped for safety: the operator adds a namespace="<its namespace>" matcher to every route and inhibition rule they contribute, so a team's AlertmanagerConfig can only route alerts from its own namespace. That is why a tenant's config "does not match" a cluster-wide alert, and why the field names inside the CRD are camelCase (groupBy, groupWait, repeatInterval, matchers[].name/value/matchType) while the rendered Secret uses the snake_case above. The rendered truth is in the Secret alertmanager-<name>-generated and on the UI's Status page.
Prometheus's own Alertmanager discovery is the other half of the handoff: the Prometheus object's spec.alerting.alertmanagers[] (namespace, name, port) must point at the Alertmanager Service. A stack where rules fire but Alertmanager shows nothing, and no silence or route explains it, usually has this pointer wrong or the Alertmanager Service port renamed. Status → Runtime & Build Information → Alertmanagers in the Prometheus UI lists what it discovered.
SLO burn-rate alerts and testing rules
An SLO of 99.9% over 30 days gives an error budget of 0.1% of requests, about 43 minutes of full outage per month. Burn rate is "how many times faster than budget-neutral you are consuming it": burn rate 1 spends exactly the month's budget in a month. The SRE Workbook's standard ladder pairs a long window (decides) with a short window one twelfth its length (stops the alert once the problem is gone):
| Burn rate | Long window | Short window | Budget consumed at trigger | Action |
|---|---|---|---|---|
| 14.4 | 1h | 5m | 2% of the month | page |
| 6 | 6h | 30m | 5% | page |
| 3 | 1d | 2h | 10% | ticket |
| 1 | 3d | 6h | 10% | ticket |
For a 99.9% SLO the 1h rule reads error_ratio_1h > 14.4 * 0.001 and error_ratio_5m > 14.4 * 0.001, where each ratio is a recording rule such as sum(rate(http_requests_total{code=~"5.."}[1h])) / sum(rate(http_requests_total[1h])). Recording rules are not optional here: eight windows of raw division per service is exactly the slow-dashboard problem. Name them slo:sli_error:ratio_rate1h and friends, and put severity on the label so routing follows the table.
Alert hygiene that graders and reviewers look for
- One
severitytaxonomy (criticalpages,warningtickets,infodashboards) and routes that honor it; nothing pages without arunbook_url. forlong enough to survive a scrape gap (at least two scrape intervals),keep_firing_forwhen a symptom flaps at the threshold.- Templating in annotations:
{{ $labels.namespace }},{{ $value | humanize }},{{ $value | humanizePercentage }},{{ $externalLabels.cluster }}. Labels can be templated too, but a templated label that varies per evaluation creates a new alert identity each time and defeats grouping. - The
ALERTS{alertname, alertstate="pending|firing"}series Prometheus writes for every rule is how you graph alert history and count noisy rules:count by (alertname) (changes(ALERTS[7d])). - Alertmanager's own health:
alertmanager_notifications_failed_totalby integration,alertmanager_alerts_invalid_total, and the Watchdog route to a dead man's switch receiver.
Test rules before you apply them
promtool check rules rules.yml catches syntax; promtool test rules test.yml runs rules against synthetic series. The test file names the rule files, an evaluation_interval, and test cases with input_series (a selector plus values in the compact 0+10x5 notation: start 0, add 10, five times), then alert_rule_test entries that state which alerts must be firing at eval_time with exact labels and annotations, or promql_expr_test entries with exp_samples. Rules written in a PrometheusRule test the same way once you extract spec.groups into a plain file.
rule_files: [rules.yml]
evaluation_interval: 1m
tests:
- interval: 1m
input_series:
- series: 'up{job="payments", instance="a"}'
values: '1 1 0 0 0 0 0'
alert_rule_test:
- alertname: PaymentsDown
eval_time: 6m
exp_alerts:
- exp_labels: { severity: critical, job: payments, instance: a }
exp_annotations: { summary: "payments target a is down" }Tasks read "alert when the error rate exceeds X for Y minutes, route severity="critical" to receiver Z, and do not page on weekends". That is a PrometheusRule with for and a severity label, a route with matchers and a receiver, and a time_intervals entry referenced by mute_time_intervals. Prove the routing with amtool config routes test or the rendered config, not by waiting for Saturday.
Exercises
Alertmanager has no LoadBalancer here: kubectl -n monitoring port-forward svc/prometheus-kube-prometheus-alertmanager 9093:9093 and browse localhost:9093.
kubectl apply -f - <<'EOF'
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
name: curriculum-drill
namespace: monitoring
labels: { release: prometheus }
spec:
groups:
- name: drill
rules:
- alert: TooManyExamplePods
expr: count(kube_pod_info{namespace="default"}) > 0
for: 1m
labels: { severity: warning, team: platform }
annotations:
summary: "{{ $value }} pods in default"
description: "Drill alert; fires whenever default has any pods."
EOFoutputcaptured 2026-08-26
$ kubectl apply -f - <<'EOF'
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
name: curriculum-drill
namespace: monitoring
labels: { release: prometheus }
spec:
groups:
- name: drill
rules:
- alert: TooManyExamplePods
expr: count(kube_pod_info{namespace="default"}) > 0
for: 1m
labels: { severity: warning, team: platform }
annotations:
summary: "{{ $value }} pods in default"
description: "Drill alert; fires whenever default has any pods."
EOF
prometheusrule.monitoring.coreos.com/curriculum-drill createdWatch it walk the states: Prometheus UI → Alerts shows Pending, then Firing after the minute.
In the Alertmanager UI, create a silence matching alertname=TooManyExamplePods for 2 hours with a comment.
Open Status in the Alertmanager UI and answer from the rendered config: what is the root receiver, what does group_by collapse on, which special route exists for the Watchdog alert (the stack ships one that fires always, as a dead-man's switch; know why that is a feature).
severity: warning lands in the tree and say which single config line you would change to route team: platform alerts to a new webhook receiver.Deploy a trivial webhook sink (kubectl create deploy sink --image=mendhak/http-https-echo:31 plus a Service), add an AlertmanagerConfig in default routing team: platform to it, then verify pickup: does the rendered config in the UI now contain the route? If yes, kubectl logs deploy/sink shows the JSON payload when the drill alert fires. Read it once; its structure (groupLabels, commonAnnotations, alerts[]) is what every receiver integration parses.
A silence is a manual action with an end time. A mute time interval is policy: this route does not page on weekends, forever. Tasks ask for the second and people reach for the first.
kubectl apply -f - <<'EOF'
apiVersion: monitoring.coreos.com/v1alpha1
kind: AlertmanagerConfig
metadata: { name: quiet-hours, namespace: default, labels: { release: prometheus } }
spec:
route:
groupBy: [alertname]
receiver: blackhole
muteTimeIntervals: [always-now]
receivers:
- name: blackhole
muteTimeIntervals:
- name: always-now
timeIntervals:
- times:
- { startTime: "00:00", endTime: "24:00" }
EOF
sleep 45
kubectl -n monitoring get secret alertmanager-prometheus-kube-prometheus-alertmanager-generated -o jsonpath='{.data.alertmanager\.yaml\.gz}' | base64 -d | gunzip | sed -n '/mute_time_intervals/,/^[a-z]/p'
kubectl -n monitoring exec sts/alertmanager-prometheus-kube-prometheus-alertmanager -c alertmanager -- sh -c 'amtool config routes test --config.file=/etc/alertmanager/config_out/alertmanager.env.yaml alertname=TooManyExamplePods severity=warning namespace=default'
kubectl -n default delete alertmanagerconfig quiet-hoursoutputcaptured 2026-09-13
$ kubectl apply -f - <<'EOF'
apiVersion: monitoring.coreos.com/v1alpha1
kind: AlertmanagerConfig
metadata: { name: quiet-hours, namespace: default, labels: { release: prometheus } }
spec:
route:
groupBy: [alertname]
receiver: blackhole
muteTimeIntervals: [always-now]
receivers:
- name: blackhole
muteTimeIntervals:
- name: always-now
timeIntervals:
- times:
- { startTime: "00:00", endTime: "24:00" }
EOF
alertmanagerconfig.monitoring.coreos.com/quiet-hours created
$ sleep 45
$ kubectl -n monitoring get secret alertmanager-prometheus-kube-prometheus-alertmanager-generated -o jsonpath='{.data.alertmanager\.yaml\.gz}' | base64 -d | gunzip | sed -n '/mute_time_intervals/,/^[a-z]/p'
mute_time_intervals:
- default/quiet-hours/always-now
- receiver: "null"
matchers:
- alertname = "Watchdog"
group_wait: 30s
group_interval: 5m
repeat_interval: 12h
inhibit_rules:
mute_time_intervals:
- name: default/quiet-hours/always-now
time_intervals:
- times:
- start_time: "00:00"
end_time: "24:00"
templates:
$ kubectl -n monitoring exec sts/alertmanager-prometheus-kube-prometheus-alertmanager -c alertmanager -- sh -c 'amtool config routes test --config.file=/etc/alertmanager/config_out/alertmanager.env.yaml alertname=TooManyExamplePods severity=warning namespace=default'
default/quiet-hours/blackhole
$ kubectl -n default delete alertmanagerconfig quiet-hours
alertmanagerconfig.monitoring.coreos.com "quiet-hours" deleted from default namespaceamtool config routes test names the receiver a matching alert would land on.When the cluster is on fire you want one page, not forty. An inhibition rule suppresses warnings while a related critical is firing, matched on a shared label, and getting that equal list right is the entire skill.
kubectl apply -f - <<'EOF'
apiVersion: monitoring.coreos.com/v1alpha1
kind: AlertmanagerConfig
metadata: { name: inhibit-demo, namespace: default, labels: { release: prometheus } }
spec:
route:
groupBy: [alertname]
receiver: blackhole
receivers:
- name: blackhole
inhibitRules:
- sourceMatch:
- { name: alertname, value: DrillOutage }
targetMatch:
- { name: severity, value: warning }
equal: [namespace]
EOF
sleep 120
kubectl -n monitoring get secret alertmanager-prometheus-kube-prometheus-alertmanager-generated -o jsonpath='{.data.alertmanager\.yaml\.gz}' | base64 -d | gunzip | grep -B4 -A4 DrillOutage
kubectl -n monitoring port-forward svc/prometheus-kube-prometheus-alertmanager 9093:9093 & PF1=$!
sleep 5
# Watchdog and the drill alert carry no namespace label, and the CRD adds namespace to both matchers,
# so nothing cluster-generated can ever match: inject a pair that can
curl -s -XPOST localhost:9093/api/v2/alerts -H 'Content-Type: application/json' -d '[{"labels":{"alertname":"DrillOutage","severity":"critical","namespace":"default"}},{"labels":{"alertname":"DrillNoise","severity":"warning","namespace":"default"}}]'
sleep 45
curl -s localhost:9093/api/v2/alerts | jq '[.[] | select(.labels.alertname|startswith("Drill")) | {alertname: .labels.alertname, state: .status.state, inhibitedBy: .status.inhibitedBy}]'
kill $PF1
kubectl -n default delete alertmanagerconfig inhibit-demooutputcaptured 2026-09-13
$ kubectl apply -f - <<'EOF'
apiVersion: monitoring.coreos.com/v1alpha1
kind: AlertmanagerConfig
metadata: { name: inhibit-demo, namespace: default, labels: { release: prometheus } }
spec:
route:
groupBy: [alertname]
receiver: blackhole
receivers:
- name: blackhole
inhibitRules:
- sourceMatch:
- { name: alertname, value: DrillOutage }
targetMatch:
- { name: severity, value: warning }
equal: [namespace]
EOF
alertmanagerconfig.monitoring.coreos.com/inhibit-demo created
$ sleep 120
$ kubectl -n monitoring get secret alertmanager-prometheus-kube-prometheus-alertmanager-generated -o jsonpath='{.data.alertmanager\.yaml\.gz}' | base64 -d | gunzip | grep -B4 -A4 DrillOutage
- target_matchers:
- severity="warning"
- namespace="default"
source_matchers:
- alertname="DrillOutage"
- namespace="default"
equal:
- namespace
receivers:
$ kubectl -n monitoring port-forward svc/prometheus-kube-prometheus-alertmanager 9093:9093 & PF1=$!
$ sleep 5
Forwarding from 127.0.0.1:9093 -> 9093
Forwarding from [::1]:9093 -> 9093
$ # Watchdog and the drill alert carry no namespace label, and the CRD adds namespace to both matchers,
$ # so nothing cluster-generated can ever match: inject a pair that can
$ curl -s -XPOST localhost:9093/api/v2/alerts -H 'Content-Type: application/json' -d '[{"labels":{"alertname":"DrillOutage","severity":"critical","namespace":"default"}},{"labels":{"alertname":"DrillNoise","severity":"warning","namespace":"default"}}]'
Handling connection for 9093
$ sleep 45
$ curl -s localhost:9093/api/v2/alerts | jq '[.[] | select(.labels.alertname|startswith("Drill")) | {alertname: .labels.alertname, state: .status.state, inhibitedBy: .status.inhibitedBy}]'
Handling connection for 9093
[
{
"alertname": "DrillNoise",
"state": "suppressed",
"inhibitedBy": [
"f3637fb5c13520d0"
]
},
{
"alertname": "DrillOutage",
"state": "active",
"inhibitedBy": []
}
]
$ kill $PF1
$ kubectl -n default delete alertmanagerconfig inhibit-demo
alertmanagerconfig.monitoring.coreos.com "inhibit-demo" deleted from default namespacenamespace="default" added to both matchers by the CRD, and DrillNoise comes back with inhibitedBy naming DrillOutage's fingerprint while DrillOutage itself has an empty list. Both injected alerts carry a namespace label on purpose: the stock Watchdog alert has none, so a rule that equals on namespace can never match it.Alerting rules are code and can be tested without waiting for reality to cooperate. promtool test rules feeds synthetic series in and asserts which alerts fire, which is the only way to be sure about a for clause.
kubectl -n monitoring delete pod promtool-check --ignore-not-found
kubectl apply -f - <<'EOF'
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
name: curriculum-drill
namespace: monitoring
labels: { release: prometheus }
spec:
groups:
- name: drill
rules:
- alert: TooManyExamplePods
expr: count(kube_pod_info{namespace="default"}) > 0
for: 1m
labels: { severity: warning, team: platform }
annotations:
summary: "{{ $value }} pods in default"
description: "Drill alert; fires whenever default has any pods."
EOF
kubectl -n monitoring get prometheusrule curriculum-drill -o jsonpath='{.spec}' | jq '{groups: .groups}' > /tmp/rules.json
python3 -c "import json,yaml; print(yaml.safe_dump(json.load(open('/tmp/rules.json'))))" > /tmp/rules.yml
cat /tmp/rules.yml
cat > /tmp/test.yml <<'EOF'
rule_files:
- rules.yml
evaluation_interval: 1m
tests:
- interval: 1m
input_series:
- series: 'kube_pod_info{namespace="default",pod="p1"}'
values: '1+0x10'
alert_rule_test:
- eval_time: 30s
alertname: TooManyExamplePods
exp_alerts: []
- eval_time: 5m
alertname: TooManyExamplePods
exp_alerts:
- exp_labels:
severity: warning
team: platform
exp_annotations:
summary: "1 pods in default"
description: "Drill alert; fires whenever default has any pods."
EOF
kubectl -n monitoring create configmap promtool-test --from-file=rules.yml=/tmp/rules.yml --from-file=test.yml=/tmp/test.yml --dry-run=client -o yaml | kubectl apply -f -
IMG=$(kubectl -n monitoring get sts prometheus-prometheus-kube-prometheus-prometheus -o jsonpath='{.spec.template.spec.containers[?(@.name=="prometheus")].image}'); echo "$IMG"
kubectl -n monitoring run promtool-check --rm -i --restart=Never --image="$IMG" --overrides="{\"spec\":{\"containers\":[{\"name\":\"promtool-check\",\"image\":\"$IMG\",\"workingDir\":\"/t\",\"command\":[\"promtool\",\"check\",\"rules\",\"rules.yml\"],\"volumeMounts\":[{\"name\":\"t\",\"mountPath\":\"/t\"}]}],\"volumes\":[{\"name\":\"t\",\"configMap\":{\"name\":\"promtool-test\"}}]}}"
kubectl -n monitoring run promtool-unit --rm -i --restart=Never --image="$IMG" --overrides="{\"spec\":{\"containers\":[{\"name\":\"promtool-unit\",\"image\":\"$IMG\",\"workingDir\":\"/t\",\"command\":[\"promtool\",\"test\",\"rules\",\"test.yml\"],\"volumeMounts\":[{\"name\":\"t\",\"mountPath\":\"/t\"}]}],\"volumes\":[{\"name\":\"t\",\"configMap\":{\"name\":\"promtool-test\"}}]}}"
kubectl -n monitoring delete configmap promtool-test
# otherwise TooManyExamplePods fires on this cluster forever
kubectl -n monitoring delete prometheusrule curriculum-drill
kubectl -n monitoring delete pod promtool-check --ignore-not-foundoutputcaptured 2026-09-12
$ kubectl -n monitoring delete pod promtool-check --ignore-not-found
pod "promtool-check" deleted from monitoring namespace
$ kubectl apply -f - <<'EOF'
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
name: curriculum-drill
namespace: monitoring
labels: { release: prometheus }
spec:
groups:
- name: drill
rules:
- alert: TooManyExamplePods
expr: count(kube_pod_info{namespace="default"}) > 0
for: 1m
labels: { severity: warning, team: platform }
annotations:
summary: "{{ $value }} pods in default"
description: "Drill alert; fires whenever default has any pods."
EOF
prometheusrule.monitoring.coreos.com/curriculum-drill created
$ kubectl -n monitoring get prometheusrule curriculum-drill -o jsonpath='{.spec}' | jq '{groups: .groups}' > /tmp/rules.json
$ python3 -c "import json,yaml; print(yaml.safe_dump(json.load(open('/tmp/rules.json'))))" > /tmp/rules.yml
$ cat /tmp/rules.yml
groups:
- name: drill
rules:
- alert: TooManyExamplePods
annotations:
description: Drill alert; fires whenever default has any pods.
summary: '{{ $value }} pods in default'
expr: count(kube_pod_info{namespace="default"}) > 0
for: 1m
labels:
severity: warning
team: platform
$ cat > /tmp/test.yml <<'EOF'
rule_files:
- rules.yml
evaluation_interval: 1m
tests:
- interval: 1m
input_series:
- series: 'kube_pod_info{namespace="default",pod="p1"}'
values: '1+0x10'
alert_rule_test:
- eval_time: 30s
alertname: TooManyExamplePods
exp_alerts: []
- eval_time: 5m
alertname: TooManyExamplePods
exp_alerts:
- exp_labels:
severity: warning
team: platform
exp_annotations:
summary: "1 pods in default"
description: "Drill alert; fires whenever default has any pods."
EOF
$ kubectl -n monitoring create configmap promtool-test --from-file=rules.yml=/tmp/rules.yml --from-file=test.yml=/tmp/test.yml --dry-run=client -o yaml | kubectl apply -f -
configmap/promtool-test created
$ IMG=$(kubectl -n monitoring get sts prometheus-prometheus-kube-prometheus-prometheus -o jsonpath='{.spec.template.spec.containers[?(@.name=="prometheus")].image}'); echo "$IMG"
quay.io/prometheus/prometheus:v3.14.0-distroless
$ kubectl -n monitoring run promtool-check --rm -i --restart=Never --image="$IMG" --overrides="{\"spec\":{\"containers\":[{\"name\":\"promtool-check\",\"image\":\"$IMG\",\"workingDir\":\"/t\",\"command\":[\"promtool\",\"check\",\"rules\",\"rules.yml\"],\"volumeMounts\":[{\"name\":\"t\",\"mountPath\":\"/t\"}]}],\"volumes\":[{\"name\":\"t\",\"configMap\":{\"name\":\"promtool-test\"}}]}}"
Checking rules.yml
SUCCESS: 1 rules found
All commands and output from this session will be recorded in container logs, including credentials and sensitive information passed through the command prompt.
If you don't see a command prompt, try pressing enter.
warning: couldn't attach to pod/promtool-check, falling back to streaming logs: unable to upgrade connection: container promtool-check not found in pod promtool-check_monitoring
Checking rules.yml
SUCCESS: 1 rules found
pod "promtool-check" deleted from monitoring namespace
$ kubectl -n monitoring run promtool-unit --rm -i --restart=Never --image="$IMG" --overrides="{\"spec\":{\"containers\":[{\"name\":\"promtool-unit\",\"image\":\"$IMG\",\"workingDir\":\"/t\",\"command\":[\"promtool\",\"test\",\"rules\",\"test.yml\"],\"volumeMounts\":[{\"name\":\"t\",\"mountPath\":\"/t\"}]}],\"volumes\":[{\"name\":\"t\",\"configMap\":{\"name\":\"promtool-test\"}}]}}"
SUCCESS
All commands and output from this session will be recorded in container logs, including credentials and sensitive information passed through the command prompt.
If you don't see a command prompt, try pressing enter.
warning: couldn't attach to pod/promtool-unit, falling back to streaming logs: unable to upgrade connection: container promtool-unit not found in pod promtool-unit_monitoring
SUCCESS
pod "promtool-unit" deleted from monitoring namespace
$ kubectl -n monitoring delete configmap promtool-test
configmap "promtool-test" deleted from monitoring namespace
$ # otherwise TooManyExamplePods fires on this cluster forever
$ kubectl -n monitoring delete prometheusrule curriculum-drill
prometheusrule.monitoring.coreos.com "curriculum-drill" deleted from monitoring namespace
$ kubectl -n monitoring delete pod promtool-check --ignore-not-foundcheck rules passes and test rules reports SUCCESS, including the case at 30 seconds where the for clause has not elapsed. Change the for to 10m and watch the same test fail; that is the assertion earning its keep.A good availability alert fires fast on a big burn and slow on a small one, and it does that with two windows over the same ratio. Build it on a recording rule so the alert expression stays readable.
kubectl apply -f - <<'EOF'
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata: { name: burn-rate, namespace: monitoring, labels: { release: prometheus } }
spec:
groups:
- name: sli
interval: 30s
rules:
- record: job:request_error_ratio:rate5m
expr: |
sum(rate(apiserver_request_total{code=~"5.."}[5m]))
/
sum(rate(apiserver_request_total[5m]))
- record: job:request_error_ratio:rate1h
expr: |
sum(rate(apiserver_request_total{code=~"5.."}[1h]))
/
sum(rate(apiserver_request_total[1h]))
- alert: ErrorBudgetBurn
expr: |
job:request_error_ratio:rate5m > (14.4 * 0.001)
and
job:request_error_ratio:rate1h > (14.4 * 0.001)
for: 2m
labels: { severity: critical, team: platform }
annotations:
summary: "burning the error budget 14.4x faster than allowed"
EOF
sleep 120
kubectl -n monitoring port-forward svc/prometheus-kube-prometheus-prometheus 9090:9090 & PF1=$!
sleep 5
curl -sG localhost:9090/api/v1/query --data-urlencode 'query=job:request_error_ratio:rate5m' | jq '.data.result[0].value[1]'
curl -s localhost:9090/api/v1/rules | jq '.data.groups[] | select(.name=="sli") | .rules[] | {name: (.name // .alert), type, health, state}'
kill $PF1
kubectl -n monitoring delete prometheusrule burn-rateoutputcaptured 2026-09-12
$ kubectl apply -f - <<'EOF'
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata: { name: burn-rate, namespace: monitoring, labels: { release: prometheus } }
spec:
groups:
- name: sli
interval: 30s
rules:
- record: job:request_error_ratio:rate5m
expr: |
sum(rate(apiserver_request_total{code=~"5.."}[5m]))
/
sum(rate(apiserver_request_total[5m]))
- record: job:request_error_ratio:rate1h
expr: |
sum(rate(apiserver_request_total{code=~"5.."}[1h]))
/
sum(rate(apiserver_request_total[1h]))
- alert: ErrorBudgetBurn
expr: |
job:request_error_ratio:rate5m > (14.4 * 0.001)
and
job:request_error_ratio:rate1h > (14.4 * 0.001)
for: 2m
labels: { severity: critical, team: platform }
annotations:
summary: "burning the error budget 14.4x faster than allowed"
EOF
prometheusrule.monitoring.coreos.com/burn-rate created
$ sleep 120
$ kubectl -n monitoring port-forward svc/prometheus-kube-prometheus-prometheus 9090:9090 & PF1=$!
$ sleep 5
Forwarding from 127.0.0.1:9090 -> 9090
Forwarding from [::1]:9090 -> 9090
$ curl -sG localhost:9090/api/v1/query --data-urlencode 'query=job:request_error_ratio:rate5m' | jq '.data.result[0].value[1]'
Handling connection for 9090
"0"
$ curl -s localhost:9090/api/v1/rules | jq '.data.groups[] | select(.name=="sli") | .rules[] | {name: (.name // .alert), type, health, state}'
Handling connection for 9090
{
"name": "job:request_error_ratio:rate5m",
"type": "recording",
"health": "ok",
"state": null
}
{
"name": "job:request_error_ratio:rate1h",
"type": "recording",
"health": "ok",
"state": null
}
{
"name": "ErrorBudgetBurn",
"type": "alerting",
"health": "ok",
"state": "inactive"
}
$ kill $PF1
$ kubectl -n monitoring delete prometheusrule burn-rate
prometheusrule.monitoring.coreos.com "burn-rate" deleted from monitoring namespaceThe AlertmanagerConfig objects you write are merged into one file by the operator, with a namespace matcher injected into every route. Reading the rendered result is how you find out what your object actually became.
# nothing of yours is in the rendered file until an AlertmanagerConfig exists
kubectl apply -f - <<'EOF'
apiVersion: monitoring.coreos.com/v1alpha1
kind: AlertmanagerConfig
metadata: { name: read-me, namespace: default, labels: { release: prometheus } }
spec:
route:
groupBy: [alertname]
receiver: mine
receivers:
- name: mine
EOF
sleep 60
kubectl -n monitoring get secret alertmanager-prometheus-kube-prometheus-alertmanager-generated -o jsonpath='{.data.alertmanager\.yaml\.gz}' | base64 -d | gunzip | head -60
kubectl -n monitoring get alertmanagerconfig -A
kubectl -n monitoring get alertmanager -o jsonpath='{.items[0].spec.alertmanagerConfigSelector}{"\n"}'
kubectl -n default delete alertmanagerconfig read-meoutputcaptured 2026-09-12
$ # nothing of yours is in the rendered file until an AlertmanagerConfig exists
$ kubectl apply -f - <<'EOF'
apiVersion: monitoring.coreos.com/v1alpha1
kind: AlertmanagerConfig
metadata: { name: read-me, namespace: default, labels: { release: prometheus } }
spec:
route:
groupBy: [alertname]
receiver: mine
receivers:
- name: mine
EOF
alertmanagerconfig.monitoring.coreos.com/read-me created
$ sleep 60
$ kubectl -n monitoring get secret alertmanager-prometheus-kube-prometheus-alertmanager-generated -o jsonpath='{.data.alertmanager\.yaml\.gz}' | base64 -d | gunzip | head -60
global:
resolve_timeout: 5m
route:
receiver: "null"
group_by:
- namespace
routes:
- receiver: default/read-me/mine
group_by:
- alertname
matchers:
- namespace="default"
continue: true
- receiver: "null"
matchers:
- alertname = "Watchdog"
group_wait: 30s
group_interval: 5m
repeat_interval: 12h
inhibit_rules:
- target_matchers:
- severity =~ warning|info
source_matchers:
- severity = critical
equal:
- namespace
- alertname
- target_matchers:
- severity = info
source_matchers:
- severity = warning
equal:
- namespace
- alertname
- target_matchers:
- severity = info
source_matchers:
- alertname = InfoInhibitor
equal:
- namespace
... 7 more lines
$ kubectl -n monitoring get alertmanagerconfig -A
NAMESPACE NAME AGE
default read-me 60s
$ kubectl -n monitoring get alertmanager -o jsonpath='{.items[0].spec.alertmanagerConfigSelector}{"\n"}'
{}
$ kubectl -n default delete alertmanagerconfig read-me
alertmanagerconfig.monitoring.coreos.com "read-me" deleted from default namespaceread-me config appears with namespace = "default" added to its matchers, and its receiver is renamed default/read-me/mine, prefixed with the namespace and the object name. That prefixing is why a receiver name in your object does not match the one in the UI. Without an AlertmanagerConfig of your own there is nothing of yours in the file, only the chart's null receiver and the stock inhibit rules.Prometheus does not notify anyone; it posts to Alertmanager, and the discovery for that is in the Prometheus object rather than anywhere in your rules. On a cluster with alerts that never arrive, this is the second thing to check after the rule itself.
kubectl -n monitoring get prometheus -o jsonpath='{.items[0].spec.alerting}' | jq
kubectl -n monitoring port-forward svc/prometheus-kube-prometheus-prometheus 9090:9090 & PF1=$!
sleep 5
curl -s localhost:9090/api/v1/alertmanagers | jq '{active: [.data.activeAlertmanagers[].url], dropped: [.data.droppedAlertmanagers[].url]}'
kill $PF1outputcaptured 2026-09-12
$ kubectl -n monitoring get prometheus -o jsonpath='{.items[0].spec.alerting}' | jq
{
"alertmanagers": [
{
"apiVersion": "v2",
"name": "prometheus-kube-prometheus-alertmanager",
"namespace": "monitoring",
"pathPrefix": "/",
"port": "http-web"
}
]
}
$ kubectl -n monitoring port-forward svc/prometheus-kube-prometheus-prometheus 9090:9090 & PF1=$!
$ sleep 5
Forwarding from 127.0.0.1:9090 -> 9090
Forwarding from [::1]:9090 -> 9090
$ curl -s localhost:9090/api/v1/alertmanagers | jq '{active: [.data.activeAlertmanagers[].url], dropped: [.data.droppedAlertmanagers[].url]}'
Handling connection for 9090
{
"active": [
"http://10.244.2.68:9093/api/v2/alerts"
],
"dropped": [
"http://10.244.2.242:7946/api/v2/alerts",
"http://10.244.2.242:3100/api/v2/alerts",
"http://10.244.2.242:9095/api/v2/alerts",
"http://10.244.1.157:12345/api/v2/alerts",
"http://10.244.2.3:10250/api/v2/alerts",
"http://10.244.1.103:9090/api/v2/alerts",
"http://10.244.1.103:8080/api/v2/alerts",
"http://10.244.1.103:8081/api/v2/alerts",
"http://10.244.2.192:9115/api/v2/alerts",
"http://10.244.2.68:8080/api/v2/alerts",
"http://10.244.2.68:9094/api/v2/alerts",
"http://10.244.2.68:9094/api/v2/alerts",
"http://10.244.2.68:8081/api/v2/alerts",
"http://10.244.2.160:8080/api/v2/alerts",
"http://172.18.0.3:9100/api/v2/alerts",
"http://172.18.0.4:9100/api/v2/alerts",
"http://172.18.0.5:9100/api/v2/alerts",
"http://10.244.1.75:3000/api/v2/alerts",
"http://10.244.1.75:9094/api/v2/alerts",
"http://10.244.1.75:9094/api/v2/alerts",
"http://10.244.1.75:6060/api/v2/alerts",
"http://10.244.2.68:9094/api/v2/alerts",
"http://10.244.2.68:9094/api/v2/alerts",
"http://10.244.2.68:9093/api/v2/alerts",
"http://10.244.2.68:8080/api/v2/alerts",
"http://10.244.2.68:8081/api/v2/alerts",
"http://10.244.1.103:9090/api/v2/alerts",
"http://10.244.1.103:8080/api/v2/alerts",
"http://10.244.1.103:8081/api/v2/alerts",
"http://10.244.2.184:8080/api/v2/alerts",
"http://10.244.2.184:8081/api/v2/alerts",
"http://10.244.2.242:3100/api/v2/alerts",
"http://10.244.2.160:8080/api/v2/alerts",
"http://10.244.2.242:9095/api/v2/alerts",
... 6 more lines
$ kill $PF1Self-check
An alert shows Firing in Prometheus and nothing arrived. Three candidate causes?
A silence matches it; an inhibition rule suppresses it; or routing sent it to a receiver that is failing (check Alertmanager's own logs and its alertmanager_notifications_failed_total). A fourth, if it only just fired: group_wait has not elapsed.
Your rule applies cleanly and never evaluates. What did you not check?
The Prometheus object's ruleSelector (and its namespace selector). Unmatched rules are ignored silently, the same way ServiceMonitors fail, with the same first command: read the selector.
Silence versus inhibition: when do you use each?
Silence: temporary, human-initiated, targeted at a known maintenance or a known-noisy alert, with an expiry and a comment. Inhibition: a standing rule expressing a relationship: when the cause alert fires, suppress its downstream symptoms. Silences are operations; inhibitions are design.
How long, roughly, from "condition becomes true" to "notification sent" on a stock stack?
Up to one evaluation interval (30s) to notice, plus for (say 1–5m), plus group_wait (30s). So a couple of minutes minimum, which is why you wait before declaring an alerting pipeline broken, and why for: 15m on a page-worthy symptom is usually too slow.
What is the Watchdog alert for?
It fires permanently as a dead-man's switch: an external system expects to keep receiving it and alerts when it stops, catching the failure mode where Prometheus or Alertmanager itself dies and therefore cannot alert you about anything. Self-monitoring has to come from outside.
"Notify the database team only during business hours, page everyone else at any time." Which fields carry that?
A top-level time_intervals entry (times 09:00-17:00, weekdays monday:friday, a location), and on the database route active_time_intervals: [business_hours], which suppresses notifications outside the window. The other routes carry no interval. mute_time_intervals is the inverse (suppress while it matches); if both are present on one route, mute wins.
A team's AlertmanagerConfig in namespace team-a routes alertname="NodeDown" to their webhook, is selected, appears in the rendered config, and never fires. Why?
The operator scopes merged AlertmanagerConfigs with an added namespace="team-a" matcher on every route it contributes; NodeDown carries no such label, so the route can never match. Cluster-wide alerts must be routed from the global config (spec.configSecret or spec.alertmanagerConfiguration), where the namespace matcher is not enforced.
Explain the 14.4 in a burn-rate alert and why the rule has two windows.
14.4 is the burn rate at which 2% of a 30-day budget is consumed in one hour (30 days × 24 hours × 0.02 = 14.4). The long window (1h) decides that the burn is sustained; the short window (5m) must also exceed it so the alert stops as soon as the error rate falls, instead of firing for the rest of the hour on old data. Both conditions are joined with and.
How do you prove a new alerting rule fires under the right conditions without waiting for production to break?
promtool test rules with a test file: rule_files, evaluation_interval, input_series with synthetic values, and alert_rule_test cases asserting exp_alerts (labels and annotations) at a given eval_time. Pair it with promtool check rules for syntax, and amtool config routes test to prove where the labels route.
Docs to know your way around
- prometheus.io: alerting rules, and the Alertmanager configuration page (route and inhibit_rule syntax).
- prometheus-operator.dev: PrometheusRule and AlertmanagerConfig CRD references.
- sre.google: the SRE workbook chapter on alerting on SLOs, for burn-rate vocabulary.
- Offline:
amtool config routes show/amtool config routes testif the binary is present, and the Alertmanager UI's Status page. - prometheus.io/docs/alerting/latest/configuration: the
<route>,<inhibit_rule>and<time_interval>sections are the ones to have bookmarked; note which keys are deprecated. - prometheus.io/docs/prometheus/latest/configuration/unit_testing_rules: the
promtool test rulesfile format; sre.google/workbook/alerting-on-slos for the burn-rate table. - prometheus-operator.dev/docs/developer/alerting: the three ways to configure the Alertmanager object and the namespace-scoping rule for merged AlertmanagerConfigs.