Skip to content

Portal | Level: L2: Operations | Topics: Loki | Domain: Observability

Lab Runtime 04 — Loki No Logs

Objective

Break the promtail log pipeline so logs disappear from Loki/Grafana, then restore promtail and verify logs start flowing again.

Prerequisites

  • Cluster running with the full stack deployed: make deploy-all
  • Observability stack (Loki, Promtail, Grafana) running in the monitoring namespace
  • kubectl configured for the target cluster

Steps

  1. Break it — Run ./break.sh. This adds an impossible nodeSelector to the promtail DaemonSet, causing all promtail pods to terminate (no node matches the selector). New logs will stop flowing to Loki.

  2. Observe — Wait ~30s and check Grafana/Loki:

    kubectl port-forward -n monitoring svc/kube-prometheus-stack-grafana 3000:80
    # Open http://localhost:3000 -> Explore -> Loki
    # Query: {namespace="grokdevops"}
    # New log entries should stop appearing
    
    Also confirm promtail pods are gone:
    kubectl get pods -n monitoring -l app.kubernetes.io/name=promtail
    

  3. Fix it — Run ./fix.sh. This removes the broken nodeSelector, allowing promtail pods to schedule on all nodes again.

  4. Verify — Run ./verify.sh to confirm promtail pods are running. Then check Grafana/Loki to see logs flowing again.

Expected Observations

  • kubectl get pods -n monitoring -l app.kubernetes.io/name=promtail shows no pods (0/0 DESIRED) because the DaemonSet cannot schedule anywhere.
  • kubectl get daemonset -n monitoring shows promtail with DESIRED: 0, CURRENT: 0, READY: 0.
  • In Grafana, querying Loki with {namespace="grokdevops"} shows that new log entries stop appearing (historical logs remain but nothing new arrives).
  • Events on the DaemonSet or its pods may show FailedScheduling with a message about node selector not matching any nodes.

Wrong Turns

  1. Restarting Loki — Loki is the log storage/query engine and is working fine. The problem is that Promtail (the log collector) has no running pods, so no new logs are being shipped to Loki.
  2. Checking application logs directly with kubectl logs — The application is still writing logs to stdout. The issue is that Promtail is not collecting them, so kubectl logs works but Grafana/Loki does not show new entries.
  3. Scaling Loki up or down — Loki's replica count is irrelevant. The pipeline is broken at the collection layer (Promtail), not at the storage layer.

Minimal Explanation

Promtail runs as a DaemonSet, meaning Kubernetes schedules one pod on every node that matches the DaemonSet's nodeSelector and tolerations. When the nodeSelector is set to an impossible value (a label no node has), the DaemonSet controller computes zero desired pods because no nodes qualify. With zero Promtail pods running, no agent reads container log files from /var/log/pods on the nodes, and no logs are pushed to Loki. The applications keep writing to stdout, and kubectl logs still works (it reads directly from the kubelet), but the Loki pipeline is severed.

Transfer Pattern

  • Node taints and label changes: Operations teams relabel or taint nodes (e.g., for dedicated workloads), inadvertently excluding DaemonSet pods like log collectors or monitoring agents.
  • DaemonSet config drift: A manual patch or misconfigured Helm value adds an incorrect nodeSelector, silently killing log collection across the entire cluster.

See Also

  • training/library/runbooks/observability/loki_no_logs.md
  • training/interview-scenarios/04-loki-logs-disappeared.md

Solution (spoilers)

See training/library/solutions/labs/lab-runtime-04.md for hints and explanation.

Teardown

./teardown.sh

Or reset the entire environment:

make undeploy-all

Wiki Navigation

Prerequisites

  • Kubernetes Exercises (Quest Ladder) (CLI) (Exercise Set, L1)