Portal | Level: L2: Operations | Topics: Prometheus, Grafana | Domain: Observability
Lab Runtime 03 — Observability Target Down¶
Objective¶
Break the ServiceMonitor selector/labels so Prometheus loses the grokdevops scrape target, then fix the selector to restore the target.
Prerequisites¶
- Cluster running with the full stack deployed:
make deploy-all - Observability stack (Prometheus, Grafana) running in the
monitoringnamespace kubectlconfigured for the target cluster
Steps¶
-
Break it — Run
./break.sh. This patches the ServiceMonitor's label selector towrong-app-name, so Prometheus can no longer discover the grokdevops service as a scrape target. -
Observe — Wait ~60s for the next scrape cycle, then check Prometheus targets:
-
Fix it — Run
./fix.sh. This restores the selector togrokdevops. -
Verify — Run
./verify.shto confirm the ServiceMonitor selector is correct. Then check the Prometheus targets UI to confirm the target is back and UP.
Expected Observations¶
- The Prometheus
/targetspage shows the grokdevops target as DOWN or completely missing from the targets list. kubectl get servicemonitor -n monitoringstill shows the ServiceMonitor resource exists, but itsmatchLabelsselector no longer matches the grokdevops Service.- Prometheus configuration (viewable at
/configor via the operator-generated secret) no longer contains a scrape job for grokdevops.
Wrong Turns¶
- Restarting Prometheus pods — The scrape configuration is generated by the Prometheus operator from ServiceMonitor resources. Restarting Prometheus reloads the same (broken) generated config because the label mismatch persists.
- Checking the wrong namespace — The ServiceMonitor lives in
monitoringbut targets a Service ingrokdevops. Looking for problems only in one namespace misses half the picture. - Editing the Prometheus ConfigMap or Secret directly — The operator continuously reconciles the scrape config from ServiceMonitor CRDs. Manual edits to the generated config are overwritten on the next reconciliation cycle.
Minimal Explanation¶
The Prometheus operator watches ServiceMonitor custom resources. For each
ServiceMonitor, it matches the spec.selector.matchLabels against Service
labels in the target namespace. When a match is found, the operator
generates a scrape configuration block and injects it into the Prometheus
configuration Secret. Prometheus reloads this config automatically. When
the ServiceMonitor's selector is changed to wrong-app-name, no Service
matches, so the operator removes the scrape job from the generated config.
Prometheus reloads and the target disappears from its /targets page.
Transfer Pattern¶
- Adding new microservices without a ServiceMonitor: A team deploys a new service but forgets to create a ServiceMonitor, so Prometheus never scrapes it and there are no metrics or alerts.
- Label changes during Helm upgrade: Upgrading a Helm chart may change default labels on Services, breaking existing ServiceMonitor selectors silently.
See Also¶
training/library/runbooks/observability/prometheus-target-down.mdtraining/interview-scenarios/03-prometheus-target-down.md
Solution (spoilers)¶
See training/library/solutions/labs/lab-runtime-03.md for hints and explanation.
Teardown¶
Or reset the entire environment:
Wiki Navigation¶
Prerequisites¶
- Kubernetes Exercises (Quest Ladder) (CLI) (Exercise Set, L1)
Related Content¶
- Monitoring Fundamentals (Topic Pack, L1) — Grafana, Prometheus
- Monitoring Migration (Legacy to Modern) (Topic Pack, L2) — Grafana, Prometheus
- Observability Architecture (Reference, L2) — Grafana, Prometheus
- Observability Deep Dive (Topic Pack, L2) — Grafana, Prometheus
- Skillcheck: Observability (Assessment, L2) — Grafana, Prometheus
- Track: Observability (Reference, L2) — Grafana, Prometheus
- Adversarial Interview Gauntlet (30 sequences) (Scenario, L2) — Prometheus
- Alerting Rules (Topic Pack, L2) — Prometheus
- Alerting Rules Drills (Drill, L2) — Prometheus
- Capacity Planning (Topic Pack, L2) — Prometheus