---
tags:
- k8s
- l1
- flashcard-deck
- k8s-troubleshooting
---
<!-- wiki:breadcrumb:start -->
[Portal](../../../../library/portal/index.md) | **Level:** [L1: Foundations](../../../../library/portal/levels.md) | **Topics:** [Kubernetes Debugging](../../../../library/portal/topics.md) | **Domain:** Kubernetes
<!-- wiki:breadcrumb:end -->

id	category	difficulty	tags	question	answer	source_path
k8s-tshoot/001	k8s-troubleshooting	easy	kubernetes, troubleshooting, crashloopbackoff	A pod is in CrashLoopBackOff. What does this status mean?	"The container starts, crashes, and Kubernetes restarts it with exponential backoff (10s, 20s, 40s, up to 5m). Common causes: missing config/secrets, unhandled exception at startup, OOMKilled, bad entrypoint command."\n\nRemember: CrashLoopBackOff = start-crash-retry with exponential backoff (10s→20s→40s→5min).\n\nGotcha: `kubectl logs pod --previous` shows crash reason. Common: missing config, OOM, wrong cmd.	training/interactive/exercises/k8s/crashloop
k8s-tshoot/002	k8s-troubleshooting	medium	kubernetes, troubleshooting, crashloopbackoff	How do you get the crash output from a CrashLoopBackOff pod?	"kubectl logs <pod> --previous shows stdout/stderr from the last crashed container. If the container exits too fast, kubectl describe pod <pod> Events section often reveals the reason (OOMKilled, exec format error, etc.)."\n\nRemember: CrashLoopBackOff = start-crash-retry with exponential backoff (10s→20s→40s→5min).\n\nGotcha: `kubectl logs pod --previous` shows crash reason. Common: missing config, OOM, wrong cmd.	training/interactive/exercises/k8s/crashloop
k8s-tshoot/003	k8s-troubleshooting	easy	kubernetes, troubleshooting, imagepullbackoff	A pod is stuck in ImagePullBackOff. What are the common causes?	"1) Image name or tag is wrong. 2) Image doesn't exist in the registry. 3) ImagePullSecret is missing or expired. 4) Private registry requires auth. 5) Network policy or firewall blocks registry access."\n\nRemember: ImagePullBackOff = can't pull image. Check: typo, auth, network, tag existence.\n\nGotcha: `kubectl describe pod` Events section has the exact error.	training/interactive/exercises/levels/level-03/k8s-imagepull
k8s-tshoot/004	k8s-troubleshooting	medium	kubernetes, troubleshooting, imagepullbackoff	How do you verify that an ImagePullSecret is correct?	"kubectl get secret <name> -o jsonpath='{.data.\\.dockerconfigjson}' | base64 -d to inspect credentials. Then test manually: docker login <registry> with those creds. Also check the secret is in the same namespace as the pod."\n\nRemember: K8s troubleshooting: Get→Describe→Logs→Exec. Mnemonic: "GDLE."\n\nGotcha: Always check Events with `kubectl describe` — they tell WHY, not just WHAT failed.	training/interactive/exercises/levels/level-03/k8s-imagepull
k8s-tshoot/005	k8s-troubleshooting	easy	kubernetes, troubleshooting, pending	A pod is stuck in Pending state. What do you check first?	"kubectl describe pod <pod> — look at Events. Common reasons: insufficient CPU/memory (no node fits requests), no matching nodeSelector/affinity, PVC not bound, taints without tolerations."\n\nRemember: Pending = can't schedule. Causes: no resources, taints, unbound PVC, no nodes.\n\nGotcha: `kubectl describe pod` shows FailedScheduling reason.	training/interactive/exercises/levels/level-04/k8s-pending
k8s-tshoot/006	k8s-troubleshooting	medium	kubernetes, troubleshooting, pending, scheduling	How do you determine if a pod is Pending due to resource pressure?	"kubectl describe nodes | grep -A 5 'Allocated resources' shows used vs allocatable. If requests exceed available, the scheduler can't place the pod. kubectl get events --field-selector reason=FailedScheduling confirms."\n\nRemember: Pending = can't schedule. Causes: no resources, taints, unbound PVC, no nodes.\n\nGotcha: `kubectl describe pod` shows FailedScheduling reason.	training/interactive/exercises/levels/level-04/k8s-pending
k8s-tshoot/007	k8s-troubleshooting	medium	kubernetes, troubleshooting, probes	A pod keeps restarting but logs show no errors. What could be wrong?	"Likely a misconfigured liveness probe. If the probe endpoint is wrong, too slow, or has a short timeout, kubelet kills the container as 'unhealthy'. Check: kubectl describe pod <pod> — look for 'Liveness probe failed' in Events."\n\nRemember: High restarts = crashing. `kubectl describe pod` for reason, `kubectl logs --previous` for logs.	training/interactive/exercises/k8s/probes
k8s-tshoot/008	k8s-troubleshooting	medium	kubernetes, troubleshooting, probes	What is the difference between liveness, readiness, and startup probes?	"Liveness: is the process alive? Failure = container restart. Readiness: can it serve traffic? Failure = removed from Service endpoints. Startup: is the app finished initializing? Blocks liveness/readiness until it passes. Use startup probes for slow-starting apps."\n\nRemember: K8s troubleshooting: Get→Describe→Logs→Exec. Mnemonic: "GDLE."\n\nGotcha: Always check Events with `kubectl describe` — they tell WHY, not just WHAT failed.	training/interactive/exercises/k8s/probes
k8s-tshoot/009	k8s-troubleshooting	easy	kubernetes, troubleshooting, events	How do you use kubectl events for troubleshooting?	"kubectl get events --sort-by=.lastTimestamp shows recent cluster events. kubectl get events --field-selector involvedObject.name=<pod> filters to one resource. Events expire after 1 hour by default — check quickly after an issue."\n\nRemember: Flow: Get→Describe→Logs→Exec. Mnemonic: "GDLE."	training/interactive/exercises/k8s/events
k8s-tshoot/010	k8s-troubleshooting	medium	kubernetes, troubleshooting, logs, describe	When do you use kubectl logs vs kubectl describe?	"Use logs to see application output (stdout/stderr). Use describe to see Kubernetes-level info: scheduling decisions, probe results, image pulls, resource limits, events. Start with describe for cluster issues, logs for app issues."\n\nExample: `kubectl logs pod --previous --tail=100` — last 100 lines from crashed container.\n\nGotcha: Logs lost on pod deletion. Set up log aggregation (Fluentd/Loki) for persistence.	training/interactive/exercises/k8s/events
k8s-tshoot/011	k8s-troubleshooting	medium	kubernetes, troubleshooting, services, endpoints	A Service exists but gets no traffic. How do you debug?	"kubectl get endpoints <svc> — if empty, the selector doesn't match any pod labels. Check: kubectl get pods --show-labels and compare with kubectl get svc <svc> -o yaml | grep selector. Also verify pods are Ready."\n\nRemember: Flow: Get→Describe→Logs→Exec. Mnemonic: "GDLE."	training/interactive/exercises/levels/level-28/k8s-endpoints
k8s-tshoot/012	k8s-troubleshooting	medium	kubernetes, troubleshooting, services	How do you test connectivity to a Service from inside the cluster?	"kubectl run tmp --image=busybox --rm -it -- wget -qO- http://<svc>.<ns>.svc.cluster.local:<port>. Or use kubectl exec into an existing pod. Check DNS resolution: nslookup <svc>.<ns>.svc.cluster.local."\n\nRemember: K8s troubleshooting: Get→Describe→Logs→Exec. Mnemonic: "GDLE."\n\nGotcha: Always check Events with `kubectl describe` — they tell WHY, not just WHAT failed.	training/interactive/exercises/levels/level-28/k8s-endpoints
k8s-tshoot/013	k8s-troubleshooting	medium	kubernetes, troubleshooting, dns	Pods can reach external IPs but internal DNS fails. What do you check?	"1) CoreDNS pods are running: kubectl get pods -n kube-system -l k8s-app=kube-dns. 2) CoreDNS service has endpoints. 3) Pod resolv.conf points to CoreDNS ClusterIP. 4) Check CoreDNS logs for errors. 5) NetworkPolicy may be blocking UDP/53."\n\nExample: `kubectl exec debug-pod -- nslookup kubernetes.default` verifies cluster DNS.	training/interactive/exercises/levels/level-23/k8s-dns
k8s-tshoot/014	k8s-troubleshooting	hard	kubernetes, troubleshooting, dns, ndots	DNS lookups are slow in pods. What is the ndots issue?	"Default ndots:5 means any name with fewer than 5 dots gets search domains appended first (e.g., api.example.com tries api.example.com.<ns>.svc.cluster.local before the real lookup). Fix: set dnsConfig.options ndots:2 in the pod spec or use FQDNs with trailing dots."\n\nExample: `kubectl exec debug-pod -- nslookup kubernetes.default` verifies cluster DNS.	training/interactive/exercises/levels/level-23/k8s-dns
k8s-tshoot/015	k8s-troubleshooting	medium	kubernetes, troubleshooting, taints, tolerations	Pods won't schedule on a specific node. How do you check taints?	"kubectl describe node <node> | grep Taints. Common taints: node.kubernetes.io/not-ready, node.kubernetes.io/memory-pressure, node.kubernetes.io/disk-pressure. Pods need matching tolerations in their spec to schedule on tainted nodes."\n\nRemember: K8s troubleshooting: Get→Describe→Logs→Exec. Mnemonic: "GDLE."\n\nGotcha: Always check Events with `kubectl describe` — they tell WHY, not just WHAT failed.	training/interactive/exercises/k8s/scheduling
k8s-tshoot/016	k8s-troubleshooting	medium	kubernetes, troubleshooting, node-pressure	A node shows NotReady status. How do you investigate?	"kubectl describe node <node> — check Conditions (MemoryPressure, DiskPressure, PIDPressure). SSH to the node: check kubelet logs (journalctl -u kubelet), disk space (df -h), memory (free -m). Common cause: kubelet can't reach the API server."\n\nRemember: K8s troubleshooting: Get→Describe→Logs→Exec. Mnemonic: "GDLE."\n\nGotcha: Always check Events with `kubectl describe` — they tell WHY, not just WHAT failed.	training/interactive/exercises/levels/level-45/k8s-node-affinity
k8s-tshoot/017	k8s-troubleshooting	medium	kubernetes, troubleshooting, pvc, storage	A PVC is stuck in Pending. What are the common causes?	"1) No StorageClass matches the request. 2) StorageClass provisioner can't create the volume (cloud API error, quota). 3) No PV available for static provisioning. 4) Access mode mismatch (ReadWriteMany not supported). Check: kubectl describe pvc <name> Events."\n\nRemember: K8s troubleshooting: Get→Describe→Logs→Exec. Mnemonic: "GDLE."\n\nGotcha: Always check Events with `kubectl describe` — they tell WHY, not just WHAT failed.	training/interactive/exercises/levels/level-35/k8s-storageclass
k8s-tshoot/018	k8s-troubleshooting	hard	kubernetes, troubleshooting, pvc, storage	A pod with a PVC can't start and shows a multi-attach error. What's wrong?	"The PV is ReadWriteOnce (RWO) and is already mounted on another node. This happens during rolling updates when old and new pods are on different nodes. Fix: use Recreate strategy instead of RollingUpdate, or switch to ReadWriteMany if the storage supports it."\n\nRemember: K8s troubleshooting: Get→Describe→Logs→Exec. Mnemonic: "GDLE."\n\nGotcha: Always check Events with `kubectl describe` — they tell WHY, not just WHAT failed.	training/interactive/exercises/levels/level-35/k8s-storageclass
k8s-tshoot/019	k8s-troubleshooting	medium	kubernetes, troubleshooting, rollout	A Deployment rollout is stuck. How do you diagnose it?	"kubectl rollout status deploy/<name> shows progress. kubectl describe deploy/<name> shows conditions. Common causes: new pods failing probes, insufficient quota, image pull errors. kubectl rollout undo deploy/<name> reverts to the last working revision."\n\nRemember: K8s troubleshooting: Get→Describe→Logs→Exec. Mnemonic: "GDLE."\n\nGotcha: Always check Events with `kubectl describe` — they tell WHY, not just WHAT failed.	training/interactive/exercises/levels/level-15/k8s-rollout
k8s-tshoot/020	k8s-troubleshooting	medium	kubernetes, troubleshooting, rollout	How do you check Deployment revision history and roll back to a specific version?	"kubectl rollout history deploy/<name> lists revisions. kubectl rollout history deploy/<name> --revision=2 shows details. kubectl rollout undo deploy/<name> --to-revision=2 rolls back. Revisions are stored in ReplicaSets — don't delete old ReplicaSets."\n\nRemember: K8s troubleshooting: Get→Describe→Logs→Exec. Mnemonic: "GDLE."\n\nGotcha: Always check Events with `kubectl describe` — they tell WHY, not just WHAT failed.	training/interactive/exercises/levels/level-15/k8s-rollout
k8s-tshoot/021	k8s-troubleshooting	easy	kubernetes, troubleshooting, labels, selectors	How do you find pods that match a particular label selector?	"kubectl get pods -l app=myapp,env=prod (comma = AND). kubectl get pods -l 'app in (myapp,otherapp)' (set-based). kubectl get pods --show-labels to see all labels. Mismatched labels are the top cause of Service/Deployment issues."\n\nRemember: K8s troubleshooting: Get→Describe→Logs→Exec. Mnemonic: "GDLE."\n\nGotcha: Always check Events with `kubectl describe` — they tell WHY, not just WHAT failed.	training/interactive/exercises/levels/level-05/k8s-labels
k8s-tshoot/022	k8s-troubleshooting	medium	kubernetes, troubleshooting, namespaces	A developer says their app can't reach a service. They're in different namespaces. What's the fix?	"Use the FQDN: <service>.<namespace>.svc.cluster.local. Short names only resolve within the same namespace. Also check: NetworkPolicy may restrict cross-namespace traffic. kubectl get netpol -n <ns> to verify."\n\nRemember: K8s troubleshooting: Get→Describe→Logs→Exec. Mnemonic: "GDLE."\n\nGotcha: Always check Events with `kubectl describe` — they tell WHY, not just WHAT failed.	training/interactive/exercises/levels/level-23/k8s-dns
k8s-tshoot/023	k8s-troubleshooting	hard	kubernetes, troubleshooting, oom	A pod is OOMKilled but the app's memory usage looks normal. What happened?	"Check memory limits vs actual usage: kubectl top pod <name>. The kernel OOM killer uses RSS (resident set size), which includes shared libraries and buffers. Java apps commonly exceed limits due to off-heap memory. Fix: increase limits or tune the runtime (e.g., -XX:MaxRAMPercentage for JVMs)."\n\nRemember: OOMKilled = exit 137 (128+SIGKILL). Exceeded memory limit. Increase or fix leak.\n\nGotcha: Check `kubectl describe pod` for Reason: OOMKilled in Last State.	training/interactive/exercises/k8s/oom
k8s-tshoot/024	k8s-troubleshooting	medium	kubernetes, troubleshooting, networking	A pod can reach the internet but not other pods. What do you check?	"1) CNI plugin is healthy (check kube-system pods). 2) NetworkPolicy blocking inter-pod traffic. 3) iptables rules on the node (kube-proxy issues). 4) Pod CIDR overlap with node network. Run kubectl exec <pod> -- ping <other-pod-ip> to confirm."\n\nRemember: K8s troubleshooting: Get→Describe→Logs→Exec. Mnemonic: "GDLE."\n\nGotcha: Always check Events with `kubectl describe` — they tell WHY, not just WHAT failed.	training/interactive/exercises/levels/level-25/k8s-networkpolicy
k8s-tshoot/025	k8s-troubleshooting	medium	kubernetes, troubleshooting, configmap, secrets	A pod starts but behaves incorrectly after a ConfigMap update. Why?	"ConfigMaps mounted as volumes update eventually (kubelet sync period, ~60s). But env vars from ConfigMaps are set at pod creation and never update. Fix: restart the pod (kubectl rollout restart deploy/<name>) or use a sidecar that watches for changes."\n\nRemember: K8s troubleshooting: Get→Describe→Logs→Exec. Mnemonic: "GDLE."\n\nGotcha: Always check Events with `kubectl describe` — they tell WHY, not just WHAT failed.	training/interactive/exercises/levels/level-36/k8s-configmap-key
k8s-troubleshooting/a1b2c3d4e5f6	k8s-troubleshooting	easy	oomkilled, kubernetes, exit-code	What exit code does a container terminated by OOMKilled have, and why?	Exit code 137 (128 + 9), because the Linux kernel sends SIGKILL (signal 9) when the process exceeds its memory cgroup limit.\n\nRemember: OOMKilled = exit 137 (128+SIGKILL). Exceeded memory limit. Increase or fix leak.\n\nGotcha: Check `kubectl describe pod` for Reason: OOMKilled in Last State.	training/library/topics/oomkilled/primer.md
k8s-troubleshooting/b2c3d4e5f6a7	k8s-troubleshooting	easy	oomkilled, kubernetes, diagnosis	What kubectl command shows whether a pod was OOMKilled, and what fields do you look for?	kubectl describe pod <name>. Look for Last State: Terminated, Reason: OOMKilled, Exit Code: 137 in the container status section.\n\nRemember: OOMKilled = exit 137 (128+SIGKILL). Exceeded memory limit. Increase or fix leak.\n\nGotcha: Check `kubectl describe pod` for Reason: OOMKilled in Last State.	training/library/topics/oomkilled/primer.md
k8s-troubleshooting/c3d4e5f6a7b8	k8s-troubleshooting	easy	oomkilled, kubernetes, resources	What is the difference between resources.requests.memory and resources.limits.memory in a Kubernetes pod spec?	requests.memory is the scheduling guarantee (the kubelet reserves this amount on the node). limits.memory is the hard ceiling enforced by the Linux cgroup — exceeding it triggers OOMKill.\n\nRemember: K8s troubleshooting: Get→Describe→Logs→Exec. Mnemonic: "GDLE."\n\nGotcha: Always check Events with `kubectl describe` — they tell WHY, not just WHAT failed.	training/library/topics/oomkilled/primer.md
k8s-troubleshooting/d4e5f6a7b8c9	k8s-troubleshooting	easy	oomkilled, kubernetes, qos	What are the three Kubernetes QoS classes and which is evicted first under memory pressure?	BestEffort (no requests/limits, evicted first), Burstable (requests < limits, evicted second), Guaranteed (requests == limits, evicted last).\n\nRemember: K8s troubleshooting: Get→Describe→Logs→Exec. Mnemonic: "GDLE."\n\nGotcha: Always check Events with `kubectl describe` — they tell WHY, not just WHAT failed.	training/library/topics/oomkilled/primer.md
k8s-troubleshooting/e5f6a7b8c9d0	k8s-troubleshooting	medium	oomkilled, kubernetes, jvm	Why does a Java application with -Xmx1g in a container limited to 512Mi get OOMKilled, and how do you fix it?	The JVM requests 1GB of heap from the OS, but the cgroup enforces a 512Mi ceiling and kills the process. Fix by using -XX:MaxRAMPercentage=75.0 so the JVM sizes its heap relative to the container's memory limit.\n\nRemember: OOMKilled = exit 137 (128+SIGKILL). Exceeded memory limit. Increase or fix leak.\n\nGotcha: Check `kubectl describe pod` for Reason: OOMKilled in Last State.	training/library/topics/oomkilled/primer.md
k8s-troubleshooting/f6a7b8c9d0e1	k8s-troubleshooting	medium	oomkilled, linux, kernel	What is oom_score_adj and how does Kubernetes use it to influence which process the OOM killer targets?	oom_score_adj (-1000 to 1000) adjusts a process's OOM kill priority. Kubernetes sets it by QoS class: Guaranteed gets -997 (killed last), BestEffort gets 1000 (killed first), and Burstable gets a scaled value in between.\n\nRemember: OOMKilled = exit 137 (128+SIGKILL). Exceeded memory limit. Increase or fix leak.\n\nGotcha: Check `kubectl describe pod` for Reason: OOMKilled in Last State.	training/library/topics/oomkilled/primer.md
k8s-troubleshooting/a7b8c9d0e1f2	k8s-troubleshooting	medium	oomkilled, kubernetes, monitoring	Which Prometheus metric should you use to predict OOMKill, and why not container_memory_usage_bytes?	Use container_memory_working_set_bytes because it excludes inactive file cache and reflects what the OOM killer actually evaluates. container_memory_usage_bytes includes reclaimable cache and overstates true pressure.\n\nRemember: OOMKilled = exit 137 (128+SIGKILL). Exceeded memory limit. Increase or fix leak.\n\nGotcha: Check `kubectl describe pod` for Reason: OOMKilled in Last State.	training/library/topics/oomkilled/primer.md
k8s-troubleshooting/b8c9d0e1f2a3	k8s-troubleshooting	medium	oomkilled, kubernetes, prevention	What is a LimitRange and how does it prevent OOMKilled caused by missing resource limits?	A LimitRange is a namespace-scoped object that sets default memory requests and limits for containers that do not specify their own. It ensures every pod has a cgroup ceiling, preventing unbounded memory consumption that causes node-level OOM.\n\nRemember: OOMKilled = exit 137 (128+SIGKILL). Exceeded memory limit. Increase or fix leak.\n\nGotcha: Check `kubectl describe pod` for Reason: OOMKilled in Last State.	training/library/topics/oomkilled/primer.md
k8s-troubleshooting/c9d0e1f2a3b4	k8s-troubleshooting	hard	oomkilled, linux, kernel, diagnosis	How do you distinguish a container-level OOM from a node-level OOM, and what commands reveal each?	Container-level: single pod affected, Exit Code 137, Reason OOMKilled in kubectl describe pod. Node-level: multiple pods affected, dmesg shows kernel OOM killer messages, kubelet logs show eviction activity, kubectl describe node shows MemoryPressure: True.\n\nRemember: OOMKilled = exit 137 (128+SIGKILL). Exceeded memory limit. Increase or fix leak.\n\nGotcha: Check `kubectl describe pod` for Reason: OOMKilled in Last State.	training/library/topics/oomkilled/primer.md
k8s-troubleshooting/d0e1f2a3b4c5	k8s-troubleshooting	hard	oomkilled, kubernetes, eviction	Explain the kubelet eviction thresholds for memory and how they interact with the kernel OOM killer.	The kubelet has --eviction-hard (e.g., memory.available<100Mi) and --eviction-soft thresholds. When available memory crosses the soft threshold for its grace period or hits the hard threshold, the kubelet evicts pods (BestEffort first). If eviction cannot free memory fast enough and available memory reaches zero, the kernel OOM killer fires as a last resort.\n\nRemember: OOMKilled = exit 137 (128+SIGKILL). Exceeded memory limit. Increase or fix leak.\n\nGotcha: Check `kubectl describe pod` for Reason: OOMKilled in Last State.	training/library/topics/oomkilled/primer.md
k8s-troubleshooting/e1f2a3b4c5d6	k8s-troubleshooting	hard	oomkilled, kubernetes, sidecars	A pod has a main container limited to 512Mi and an Istio sidecar limited to 256Mi. The main container is OOMKilled despite using only 400Mi. What is the likely cause?	Each container has its own cgroup and memory limit. If the main container is OOMKilled at 400Mi with a 512Mi limit, it may be counting shared memory (e.g., tmpfs mounts, emptyDir medium: Memory volumes) against the container's cgroup. Check for memory-backed volumes and sidecar memory consumption patterns with kubectl top pod --containers.\n\nRemember: OOMKilled = exit 137 (128+SIGKILL). Exceeded memory limit. Increase or fix leak.\n\nGotcha: Check `kubectl describe pod` for Reason: OOMKilled in Last State.	training/library/topics/oomkilled/primer.md
k8s-troubleshooting/f2a3b4c5d6e7	k8s-troubleshooting	hard	oomkilled, kubernetes, vpa, capacity	How does the Vertical Pod Autoscaler help prevent OOMKilled, and what are its risks?	VPA analyzes historical memory usage and recommends or automatically sets requests and limits. It provides lowerBound, target, and upperBound recommendations. Risks: in UpdateMode it restarts pods to apply new limits (disruption), it can undersize limits if load patterns are spiky, and it conflicts with HPA on the same resource — never use VPA and HPA both scaling on memory.\n\nRemember: OOMKilled = exit 137 (128+SIGKILL). Exceeded memory limit. Increase or fix leak.\n\nGotcha: Check `kubectl describe pod` for Reason: OOMKilled in Last State.	training/library/topics/oomkilled/primer.md
k8s-troubleshooting/5871df00	k8s-troubleshooting	easy	kubernetes, crashloopbackoff, troubleshooting	What does CrashLoopBackOff mean in Kubernetes?	CrashLoopBackOff is a status indicating the container started, crashed, and the kubelet is waiting with exponential backoff (10s, 20s, 40s... capped at 5 minutes) before restarting it again.\n\nRemember: CrashLoopBackOff = start-crash-retry with exponential backoff (10s→20s→40s→5min).\n\nGotcha: `kubectl logs pod --previous` shows crash reason. Common: missing config, OOM, wrong cmd.	training/library/topics/crashloopbackoff/primer.md
k8s-troubleshooting/72c49333	k8s-troubleshooting	easy	kubernetes, crashloopbackoff, kubectl	What kubectl command shows logs from a crashed container's previous run?	kubectl logs <pod-name> --previous. The --previous flag retrieves logs from the last terminated container instance.\n\nExample: `kubectl logs pod --previous --tail=100` — last 100 lines from crashed container.\n\nGotcha: Logs lost on pod deletion. Set up log aggregation (Fluentd/Loki) for persistence.	training/library/topics/crashloopbackoff/primer.md
k8s-troubleshooting/95b5023f	k8s-troubleshooting	easy	kubernetes, crashloopbackoff, exit-codes	A pod is in CrashLoopBackOff with exit code 137. What killed it and what do you check first?	Exit code 137 means SIGKILL — most commonly the kernel OOM killer. Run kubectl describe pod and look for Reason: OOMKilled in Last State, then check the container memory limit versus actual usage.\n\nRemember: CrashLoopBackOff = start-crash-retry with exponential backoff (10s→20s→40s→5min).\n\nGotcha: `kubectl logs pod --previous` shows crash reason. Common: missing config, OOM, wrong cmd.	training/library/topics/crashloopbackoff/primer.md
k8s-troubleshooting/4129f930	k8s-troubleshooting	easy	kubernetes, crashloopbackoff, diagnostics	What three kubectl commands form the basic CrashLoopBackOff diagnostic workflow?	1) kubectl get pods (see restart count and status), 2) kubectl describe pod (events, exit codes, last state), 3) kubectl logs --previous (see what the container printed before dying).\n\nRemember: CrashLoopBackOff = start-crash-retry with exponential backoff (10s→20s→40s→5min).\n\nGotcha: `kubectl logs pod --previous` shows crash reason. Common: missing config, OOM, wrong cmd.	training/library/topics/crashloopbackoff/primer.md
k8s-troubleshooting/f86911ac	k8s-troubleshooting	medium	kubernetes, crashloopbackoff, exit-codes	What is the difference between exit codes 126 and 127 in a container?	Exit code 126 means the entrypoint binary exists but cannot be executed (permission denied). Exit code 127 means the entrypoint binary does not exist in the image (command not found).\n\nRemember: K8s troubleshooting: Get→Describe→Logs→Exec. Mnemonic: "GDLE."\n\nGotcha: Always check Events with `kubectl describe` — they tell WHY, not just WHAT failed.	training/library/topics/crashloopbackoff/primer.md
k8s-troubleshooting/1eeb6d9f	k8s-troubleshooting	medium	kubernetes, crashloopbackoff, liveness-probes	How can a liveness probe cause CrashLoopBackOff, and how do you prevent it for slow-starting apps?	If initialDelaySeconds is too short, the liveness probe fails before the app finishes starting, causing Kubernetes to kill and restart the container repeatedly. Use a startupProbe with a high failureThreshold to give the app time to initialize before liveness checks begin.\n\nRemember: CrashLoopBackOff = start-crash-retry with exponential backoff (10s→20s→40s→5min).\n\nGotcha: `kubectl logs pod --previous` shows crash reason. Common: missing config, OOM, wrong cmd.	training/library/topics/crashloopbackoff/primer.md
k8s-troubleshooting/57398912	k8s-troubleshooting	medium	kubernetes, crashloopbackoff, oomkilled	A container keeps restarting with exit code 137. Describe your troubleshooting steps.	Run kubectl describe pod to confirm OOMKilled in the Last State reason. Check the container's memory limit in the pod spec. Use kubectl top pod or Prometheus metrics to see actual memory usage. Increase the memory limit or fix the memory leak in the application.\n\nRemember: Flow: Get→Describe→Logs→Exec. Mnemonic: "GDLE."	training/library/topics/crashloopbackoff/primer.md
k8s-troubleshooting/a02885f5	k8s-troubleshooting	medium	kubernetes, crashloopbackoff, init-containers	How do init containers help prevent CrashLoopBackOff caused by missing dependencies?	Init containers run before the main container and block startup until they succeed. You can use an init container to wait for a dependency (e.g., polling a database port with nc -z) so the main container only starts when its dependencies are actually ready.\n\nRemember: CrashLoopBackOff = start-crash-retry with exponential backoff (10s→20s→40s→5min).\n\nGotcha: `kubectl logs pod --previous` shows crash reason. Common: missing config, OOM, wrong cmd.	training/library/topics/crashloopbackoff/primer.md
k8s-troubleshooting/ddd184b7	k8s-troubleshooting	hard	kubernetes, crashloopbackoff, pid1, signals	What is the PID 1 problem in containers and how does it cause exit code 137 on pod termination?	In a container, the entrypoint runs as PID 1. If PID 1 does not handle SIGTERM, Kubernetes sends SIGTERM on shutdown, the process ignores it, Kubernetes waits the terminationGracePeriodSeconds (default 30s), then sends SIGKILL — resulting in exit code 137. Fix by using exec form in Dockerfile or a lightweight init system like tini.\n\nRemember: K8s troubleshooting: Get→Describe→Logs→Exec. Mnemonic: "GDLE."\n\nGotcha: Always check Events with `kubectl describe` — they tell WHY, not just WHAT failed.	training/library/topics/crashloopbackoff/primer.md
k8s-troubleshooting/1f0edee9	k8s-troubleshooting	hard	kubernetes, crashloopbackoff, debugging	How do you debug a CrashLoopBackOff when kubectl logs --previous shows no output?	Use kubectl debug -it <pod> --image=busybox --target=<container> to attach an ephemeral debug container sharing the pod's namespace. Alternatively, run a new pod with the same image but override the entrypoint to sleep (kubectl run debug --image=<image> --overrides=...) then exec in and manually run the entrypoint to observe the error interactively.\n\nRemember: CrashLoopBackOff = start-crash-retry with exponential backoff (10s→20s→40s→5min).\n\nGotcha: `kubectl logs pod --previous` shows crash reason. Common: missing config, OOM, wrong cmd.	training/library/topics/crashloopbackoff/primer.md
k8s-troubleshooting/8bf152b6	k8s-troubleshooting	hard	kubernetes, crashloopbackoff, differential-diagnosis	How do you distinguish CrashLoopBackOff from CreateContainerConfigError and ImagePullBackOff?	In CrashLoopBackOff the container started and ran before crashing — check application logs. In ImagePullBackOff the image could not be pulled (wrong tag, registry auth, network). In CreateContainerConfigError the container could not be configured (referenced ConfigMap or Secret does not exist). The key distinction is whether the container process ever executed.\n\nRemember: CrashLoopBackOff = start-crash-retry with exponential backoff (10s→20s→40s→5min).\n\nGotcha: `kubectl logs pod --previous` shows crash reason. Common: missing config, OOM, wrong cmd.	training/library/topics/crashloopbackoff/primer.md
k8s-troubleshooting/ee0a4388	k8s-troubleshooting	hard	kubernetes, crashloopbackoff, exit-code-zero	Why might a container with exit code 0 enter CrashLoopBackOff, and how do you fix it?	With the default restartPolicy: Always, Kubernetes restarts containers even on successful exit (code 0). If the container's process completes and exits cleanly, it will be restarted indefinitely. Fix by changing restartPolicy to OnFailure or Never (for Jobs/CronJobs), or redesign the container to run as a long-lived process that does not exit.\n\nRemember: CrashLoopBackOff = start-crash-retry with exponential backoff (10s→20s→40s→5min).\n\nGotcha: `kubectl logs pod --previous` shows crash reason. Common: missing config, OOM, wrong cmd.	training/library/topics/crashloopbackoff/primer.md

<!-- wiki:related:start -->
---

## Wiki Navigation

### Related Content

- Kubernetes Advanced Operations Flashcards *(CLI)* (flashcard_deck, L1) — Kubernetes Debugging
- [Kubernetes Debugging Playbook](../../../../library/topics/k8s-debugging-playbook/index.md) (Topic Pack, L2) — Kubernetes Debugging

<!-- wiki:related:end -->
