Portal | Level: L2: Operations | Topics: Loki | Domain: Observability
Lab Runtime 04 — Loki No Logs¶
Objective¶
Break the promtail log pipeline so logs disappear from Loki/Grafana, then restore promtail and verify logs start flowing again.
Prerequisites¶
- Cluster running with the full stack deployed:
make deploy-all - Observability stack (Loki, Promtail, Grafana) running in the
monitoringnamespace kubectlconfigured for the target cluster
Steps¶
-
Break it — Run
./break.sh. This adds an impossiblenodeSelectorto the promtail DaemonSet, causing all promtail pods to terminate (no node matches the selector). New logs will stop flowing to Loki. -
Observe — Wait ~30s and check Grafana/Loki:
Also confirm promtail pods are gone: -
Fix it — Run
./fix.sh. This removes the brokennodeSelector, allowing promtail pods to schedule on all nodes again. -
Verify — Run
./verify.shto confirm promtail pods are running. Then check Grafana/Loki to see logs flowing again.
Expected Observations¶
kubectl get pods -n monitoring -l app.kubernetes.io/name=promtailshows no pods (0/0 DESIRED) because the DaemonSet cannot schedule anywhere.kubectl get daemonset -n monitoringshows promtail withDESIRED: 0,CURRENT: 0,READY: 0.- In Grafana, querying Loki with
{namespace="grokdevops"}shows that new log entries stop appearing (historical logs remain but nothing new arrives). - Events on the DaemonSet or its pods may show
FailedSchedulingwith a message about node selector not matching any nodes.
Wrong Turns¶
- Restarting Loki — Loki is the log storage/query engine and is working fine. The problem is that Promtail (the log collector) has no running pods, so no new logs are being shipped to Loki.
- Checking application logs directly with
kubectl logs— The application is still writing logs to stdout. The issue is that Promtail is not collecting them, sokubectl logsworks but Grafana/Loki does not show new entries. - Scaling Loki up or down — Loki's replica count is irrelevant. The pipeline is broken at the collection layer (Promtail), not at the storage layer.
Minimal Explanation¶
Promtail runs as a DaemonSet, meaning Kubernetes schedules one pod on
every node that matches the DaemonSet's nodeSelector and tolerations.
When the nodeSelector is set to an impossible value (a label no node has),
the DaemonSet controller computes zero desired pods because no nodes
qualify. With zero Promtail pods running, no agent reads container log
files from /var/log/pods on the nodes, and no logs are pushed to Loki.
The applications keep writing to stdout, and kubectl logs still works
(it reads directly from the kubelet), but the Loki pipeline is severed.
Transfer Pattern¶
- Node taints and label changes: Operations teams relabel or taint nodes (e.g., for dedicated workloads), inadvertently excluding DaemonSet pods like log collectors or monitoring agents.
- DaemonSet config drift: A manual patch or misconfigured Helm value adds an incorrect nodeSelector, silently killing log collection across the entire cluster.
See Also¶
training/library/runbooks/observability/loki_no_logs.mdtraining/interview-scenarios/04-loki-logs-disappeared.md
Solution (spoilers)¶
See training/library/solutions/labs/lab-runtime-04.md for hints and explanation.
Teardown¶
Or reset the entire environment:
Wiki Navigation¶
Prerequisites¶
- Kubernetes Exercises (Quest Ladder) (CLI) (Exercise Set, L1)
Related Content¶
- Incident Simulator (18 scenarios) (CLI) (Exercise Set, L2) — Loki
- Interview: Loki Logs Disappeared (Scenario, L2) — Loki
- Log Pipelines (Topic Pack, L2) — Loki
- LogQL Drills (Drill, L2) — Loki
- Loki Flashcards (CLI) (flashcard_deck, L1) — Loki
- Observability Architecture (Reference, L2) — Loki
- Observability Deep Dive (Topic Pack, L2) — Loki
- Observability Drills (Drill, L2) — Loki
- Runbook: Log Pipeline Backpressure / Logs Not Appearing (Runbook, L2) — Loki
- Runbook: Loki No Logs (Runbook, L2) — Loki