Skip to content

Portal | Level: L2: Operations | Topics: Prometheus, Grafana | Domain: Observability

Lab Runtime 03 — Observability Target Down

Objective

Break the ServiceMonitor selector/labels so Prometheus loses the grokdevops scrape target, then fix the selector to restore the target.

Prerequisites

  • Cluster running with the full stack deployed: make deploy-all
  • Observability stack (Prometheus, Grafana) running in the monitoring namespace
  • kubectl configured for the target cluster

Steps

  1. Break it — Run ./break.sh. This patches the ServiceMonitor's label selector to wrong-app-name, so Prometheus can no longer discover the grokdevops service as a scrape target.

  2. Observe — Wait ~60s for the next scrape cycle, then check Prometheus targets:

    kubectl port-forward -n monitoring svc/kube-prometheus-stack-prometheus 9090:9090
    # Open http://localhost:9090/targets
    # The grokdevops target should be missing
    

  3. Fix it — Run ./fix.sh. This restores the selector to grokdevops.

  4. Verify — Run ./verify.sh to confirm the ServiceMonitor selector is correct. Then check the Prometheus targets UI to confirm the target is back and UP.

Expected Observations

  • The Prometheus /targets page shows the grokdevops target as DOWN or completely missing from the targets list.
  • kubectl get servicemonitor -n monitoring still shows the ServiceMonitor resource exists, but its matchLabels selector no longer matches the grokdevops Service.
  • Prometheus configuration (viewable at /config or via the operator-generated secret) no longer contains a scrape job for grokdevops.

Wrong Turns

  1. Restarting Prometheus pods — The scrape configuration is generated by the Prometheus operator from ServiceMonitor resources. Restarting Prometheus reloads the same (broken) generated config because the label mismatch persists.
  2. Checking the wrong namespace — The ServiceMonitor lives in monitoring but targets a Service in grokdevops. Looking for problems only in one namespace misses half the picture.
  3. Editing the Prometheus ConfigMap or Secret directly — The operator continuously reconciles the scrape config from ServiceMonitor CRDs. Manual edits to the generated config are overwritten on the next reconciliation cycle.

Minimal Explanation

The Prometheus operator watches ServiceMonitor custom resources. For each ServiceMonitor, it matches the spec.selector.matchLabels against Service labels in the target namespace. When a match is found, the operator generates a scrape configuration block and injects it into the Prometheus configuration Secret. Prometheus reloads this config automatically. When the ServiceMonitor's selector is changed to wrong-app-name, no Service matches, so the operator removes the scrape job from the generated config. Prometheus reloads and the target disappears from its /targets page.

Transfer Pattern

  • Adding new microservices without a ServiceMonitor: A team deploys a new service but forgets to create a ServiceMonitor, so Prometheus never scrapes it and there are no metrics or alerts.
  • Label changes during Helm upgrade: Upgrading a Helm chart may change default labels on Services, breaking existing ServiceMonitor selectors silently.

See Also

  • training/library/runbooks/observability/prometheus-target-down.md
  • training/interview-scenarios/03-prometheus-target-down.md

Solution (spoilers)

See training/library/solutions/labs/lab-runtime-03.md for hints and explanation.

Teardown

./teardown.sh

Or reset the entire environment:

make undeploy-all

Wiki Navigation

Prerequisites

  • Kubernetes Exercises (Quest Ladder) (CLI) (Exercise Set, L1)