---
tags:
- k8s
- l1
- flashcard-deck
- k8s-advanced-ops
---
<!-- wiki:breadcrumb:start -->
[Portal](../../../../library/portal/index.md) | **Level:** [L1: Foundations](../../../../library/portal/levels.md) | **Topics:** [Kubernetes Debugging](../../../../library/portal/topics.md) | **Domain:** Kubernetes
<!-- wiki:breadcrumb:end -->

id	category	difficulty	tags	question	answer	source_path
k8s-adv/001	k8s-advanced-ops	medium	kubernetes, rollout, debugging	A Deployment rollout is stuck at 1/3 updated replicas. How do you diagnose it?	kubectl rollout status deployment/<name> shows the stall. kubectl describe deployment <name> reveals the reason under Conditions (e.g., ProgressDeadlineExceeded). Check the new ReplicaSet's pods with kubectl get rs then kubectl describe pod on the stuck pods — usually a crash, image pull failure, or resource quota exhaustion.	training/interactive/knowledge/data/cards/k8s-advanced-ops.tsv
k8s-adv/002	k8s-advanced-ops	hard	kubernetes, rollout, strategy	You run kubectl rollout undo but the previous version also had issues. How do you roll back to a specific revision?	kubectl rollout history deployment/<name> lists revisions with change-cause annotations. kubectl rollout undo deployment/<name> --to-revision=<N> targets a specific known-good revision. Always confirm with kubectl rollout status and check pod readiness before declaring recovery.	training/interactive/knowledge/data/cards/k8s-advanced-ops.tsv
k8s-adv/003	k8s-advanced-ops	medium	kubernetes, probes, liveness, troubleshooting	Pods restart every 90 seconds but application logs show no errors. What is the most likely cause?	A misconfigured liveness probe. The probe path may return non-200, the port may be wrong, or timeoutSeconds is too short for the endpoint. kubectl describe pod <pod> will show 'Liveness probe failed' events with the HTTP status or connection error. Fix the probe spec, not the app, if the app is actually healthy.	training/interactive/knowledge/data/cards/k8s-advanced-ops.tsv
k8s-adv/004	k8s-advanced-ops	hard	kubernetes, probes, startup	An app takes 120 seconds to initialize. Liveness probe kills it before startup completes. How do you fix this without removing the liveness probe?	Add a startup probe with a generous failureThreshold \n* periodSeconds window (e.g., failureThreshold: 30, periodSeconds: 5 = 150s). The startup probe blocks liveness and readiness probes until it succeeds. This keeps fast-restart protection while allowing slow cold starts.	training/interactive/knowledge/data/cards/k8s-advanced-ops.tsv
k8s-adv/005	k8s-advanced-ops	easy	kubernetes, events, troubleshooting	What is the standard three-command diagnostic flow when a pod is misbehaving?	1) kubectl describe pod <pod> — check Events, conditions, and container state. \n2) kubectl logs <pod> [-c container] [--previous] — read application stdout/stderr. \n3) kubectl get events --field-selector involvedObject.name=<pod> — see cluster-level events. This describe-logs-events flow covers 90% of initial triage.	training/interactive/knowledge/data/cards/k8s-advanced-ops.tsv
k8s-adv/006	k8s-advanced-ops	medium	kubernetes, scheduling, taints, tolerations	A new node is joined to the cluster but no pods schedule onto it. kubectl describe node shows a taint. What do you do?	Check the taint with kubectl describe node <node> | grep Taints. If the taint is intentional (e.g., dedicated=gpu:NoSchedule), add a matching toleration to pods that should run there. If accidental, remove it: kubectl taint nodes <node> dedicated=gpu:NoSchedule-. The trailing minus sign removes the taint.	training/interactive/knowledge/data/cards/k8s-advanced-ops.tsv
k8s-adv/007	k8s-advanced-ops	hard	kubernetes, taints, tolerations, master	Pods are Pending cluster-wide after a control plane upgrade. All worker nodes show a NoSchedule taint. What happened and how do you recover?	The upgrade likely re-applied node-role.kubernetes.io/control-plane:NoSchedule and may have incorrectly tainted workers. Verify with kubectl get nodes -o custom-columns=NAME:.metadata.name,TAINTS:.spec.taints. Remove errant taints from workers: kubectl taint nodes <worker> <key>:NoSchedule-. For control-plane nodes, add tolerations only for system-critical pods.	training/interactive/knowledge/data/cards/k8s-advanced-ops.tsv
k8s-adv/008	k8s-advanced-ops	medium	kubernetes, pvc, storage, troubleshooting	A PVC is stuck in Pending state. What are the common causes?	1) No PV matches the PVC's storageClassName, access mode, or capacity request. \n2) The StorageClass provisioner is misconfigured or not installed. \n3) The cloud provider hit a quota or zone availability limit. \n4) The PVC requests a mode (ReadWriteMany) the provisioner does not support. Check kubectl describe pvc <name> Events for the specific error.	training/interactive/knowledge/data/cards/k8s-advanced-ops.tsv
k8s-adv/009	k8s-advanced-ops	hard	kubernetes, pvc, resize	You need to expand a PVC from 10Gi to 50Gi but the pod won't restart. What is the process?	1) Verify the StorageClass has allowVolumeExpansion: true. \n2) Edit the PVC: kubectl patch pvc <name> -p '{\spec\":{\"resources\":{\"requests\":{\"storage\":\"50Gi\"}}}}'. \n3) For file-system volumes, the resize happens on the next pod mount — you must delete and recreate the pod (not the PVC). \n4) Check kubectl describe pvc for FileSystemResizePending condition."	training/interactive/knowledge/data/cards/k8s-advanced-ops.tsv
k8s-adv/010	k8s-advanced-ops	easy	kubernetes, selectors, service, troubleshooting	A Service has no endpoints even though pods are running. What do you check?	kubectl get endpoints <svc> shows empty. Compare kubectl get svc <svc> -o wide selector labels against kubectl get pods --show-labels. A mismatch between the Service selector and pod labels is the most common cause. Also verify pods are in the same namespace as the Service and that pods are Ready (readiness probe passing).	training/interactive/knowledge/data/cards/k8s-advanced-ops.tsv
k8s-adv/011	k8s-advanced-ops	medium	kubernetes, dns, service, troubleshooting	Pods can reach a Service by ClusterIP but DNS resolution for <svc>.<ns>.svc.cluster.local fails. What do you investigate?	1) Check CoreDNS pods are running: kubectl -n kube-system get pods -l k8s-app=kube-dns. \n2) Verify the pod's /etc/resolv.conf points to the kube-dns ClusterIP. \n3) Test from inside a pod: nslookup <svc>.<ns>.svc.cluster.local. \n4) Check CoreDNS logs for errors. \n5) Confirm no NetworkPolicy is blocking UDP/TCP 53 to kube-dns.	training/interactive/knowledge/data/cards/k8s-advanced-ops.tsv
k8s-adv/012	k8s-advanced-ops	medium	kubernetes, dns, ndots	DNS lookups from pods are slow, adding 5+ seconds of latency. What is the likely cause?	The default ndots:5 in pod resolv.conf causes the resolver to try multiple search domains before querying the absolute name. For external domains, each attempt times out against cluster DNS. Fix: set dnsConfig.options ndots:2 in the pod spec, or always use FQDNs with a trailing dot (e.g., api.example.com.) to bypass search expansion.	training/interactive/knowledge/data/cards/k8s-advanced-ops.tsv
k8s-adv/013	k8s-advanced-ops	medium	kubernetes, node-pressure, eviction	Pods are being evicted with the message 'The node was low on resource: ephemeral-storage'. What is happening?	Kubelet's eviction manager detected ephemeral storage usage exceeding the threshold (default ~85%). It evicts pods in order of priority and usage. Check node conditions: kubectl describe node <node> | grep -A 5 Conditions. Clean up: remove unused images (crictl rmi --prune), check for pods writing large files to emptyDir, and set resource limits on ephemeral-storage in pod specs.	training/interactive/knowledge/data/cards/k8s-advanced-ops.tsv
k8s-adv/014	k8s-advanced-ops	hard	kubernetes, node-pressure, memory	A node goes NotReady and multiple pods are evicted simultaneously. How do you investigate and prevent recurrence?	1) kubectl describe node — check Conditions for MemoryPressure, DiskPressure, PIDPressure. \n2) Check kubelet logs on the node: journalctl -u kubelet. \n3) For memory: identify pods without memory limits (they can consume unbounded memory). \n4) Prevent: set memory requests and limits on all pods, configure ResourceQuotas per namespace, and consider LimitRanges to enforce defaults.	training/interactive/knowledge/data/cards/k8s-advanced-ops.tsv
k8s-adv/015	k8s-advanced-ops	medium	kubernetes, resources, limits, oomkilled	A container is OOMKilled but the application memory profiler shows usage well below the limit. Why?	The OOM limit applies to the entire cgroup, not just heap. It includes RSS, page cache, tmpfs mounts, and child processes. Also, the JVM or runtime may allocate off-heap memory (NIO buffers, thread stacks). Check kubectl describe pod for the exact Last State OOMKilled exit code 137 and compare against the actual RSS with cat /sys/fs/cgroup/memory/memory.usage_in_bytes inside the container.	training/interactive/knowledge/data/cards/k8s-advanced-ops.tsv
k8s-adv/016	k8s-advanced-ops	easy	kubernetes, resources, requests, scheduling	What is the difference between resource requests and limits, and which one affects scheduling?	Requests are guaranteed resources the scheduler uses for bin-packing — a pod is scheduled only if a node has enough allocatable capacity for the request. Limits are the max the container can use; exceeding CPU limits causes throttling, exceeding memory limits causes OOMKill. \nBest practice: set requests close to actual usage and limits as a safety ceiling.	training/interactive/knowledge/data/cards/k8s-advanced-ops.tsv
k8s-adv/017	k8s-advanced-ops	hard	kubernetes, hpa, troubleshooting	HPA shows <unknown>/80% for CPU target and won't scale. What is wrong?	The HPA cannot read metrics. Common causes: \n1) metrics-server is not installed or not running. \n2) The target Deployment pods have no CPU requests set — HPA needs requests to compute utilization percentage. \n3) metrics-server cannot reach kubelets (firewall or certificate issue). Verify: kubectl top pods (should return data), and kubectl describe hpa <name> for Conditions and events.	training/interactive/knowledge/data/cards/k8s-advanced-ops.tsv
k8s-adv/018	k8s-advanced-ops	medium	kubernetes, hpa, scaling	HPA keeps scaling to max replicas even when average CPU is low. What could cause this?	1) One pod is spiking and the average is skewed by replica count. \n2) The metric source is wrong (using total CPU instead of per-pod average). \n3) Readiness probe failures cause fewer ready pods, inflating per-pod averages. \n4) A recent deploy created pods that are initializing and consuming startup CPU. Check kubectl describe hpa and kubectl top pods to correlate.	training/interactive/knowledge/data/cards/k8s-advanced-ops.tsv
k8s-adv/019	k8s-advanced-ops	medium	kubernetes, pdb, disruption	You try to drain a node but it hangs. kubectl drain shows 'evicting pod ... Cannot evict pod as it would violate the pod's disruption budget'. How do you proceed?	A PodDisruptionBudget (PDB) is blocking eviction because draining would reduce available replicas below minAvailable or above maxUnavailable. Options: \n1) Scale up the deployment first so draining one node stays within budget. \n2) Check if another node already has disrupted pods. \n3) As a last resort, delete the PDB temporarily (kubectl delete pdb <name>), drain, then recreate it.	training/interactive/knowledge/data/cards/k8s-advanced-ops.tsv
k8s-adv/020	k8s-advanced-ops	hard	kubernetes, pdb, rollout	A rolling update is stuck because the PDB minAvailable equals the replica count. Why is this a problem and how do you fix it?	If minAvailable equals replicas (e.g., 3/3), the controller cannot evict any old pod to make room for a new one — a deadlock. Fix: set minAvailable to replicas-1 (or use maxUnavailable: 1 instead). This allows the rolling update to terminate one old pod at a time while maintaining minimum availability.	training/interactive/knowledge/data/cards/k8s-advanced-ops.tsv
k8s-adv/021	k8s-advanced-ops	medium	kubernetes, init-containers, troubleshooting	A pod is stuck in Init:0/2 status. How do you debug it?	kubectl describe pod <pod> shows init container status and events. kubectl logs <pod> -c <init-container-name> shows the init container's output. Init containers run sequentially — 0/2 means the first init container has not completed. Common causes: waiting on a dependency (DNS, service, database), wrong command, or missing ConfigMap/Secret volume.	training/interactive/knowledge/data/cards/k8s-advanced-ops.tsv
k8s-adv/022	k8s-advanced-ops	hard	kubernetes, sidecar, pattern	You deploy an Envoy sidecar with your app container. After a rolling update, requests fail for a few seconds. What is the likely cause and how do you fix it?	The app container starts receiving traffic before the Envoy sidecar is ready (race condition). Fix: \n1) Use a startup/readiness probe on the sidecar. \n2) In Kubernetes 1.28+, use the native sidecar feature (restartPolicy: Always in initContainers) which guarantees sidecar readiness before the main container starts. \n3) Alternatively, add a postStart lifecycle hook that polls the sidecar's health endpoint.	training/interactive/knowledge/data/cards/k8s-advanced-ops.tsv
k8s-adv/023	k8s-advanced-ops	medium	kubernetes, namespace, troubleshooting	You create a resource but cannot find it with kubectl get. What namespace-related mistakes should you check?	1) The resource was created in a different namespace — always use -n <ns> or --all-namespaces. \n2) Your kubeconfig context has a default namespace set that differs from where the resource lives. \n3) The resource is cluster-scoped (e.g., ClusterRole, PV, Node) and does not appear with -n. Check with kubectl api-resources --namespaced=false. \n4) RBAC may hide resources you lack permission to list.	training/interactive/knowledge/data/cards/k8s-advanced-ops.tsv
k8s-adv/024	k8s-advanced-ops	medium	kubernetes, controller, reconciliation	You delete a pod managed by a Deployment but it immediately reappears. Why?	The Deployment's ReplicaSet controller continuously reconciles actual state to desired state. When you delete a pod, the controller detects the replica count is below spec.replicas and creates a replacement. To actually remove the workload, delete or scale down the Deployment (kubectl scale deployment <name> --replicas=0), not individual pods.	training/interactive/knowledge/data/cards/k8s-advanced-ops.tsv
k8s-adv/025	k8s-advanced-ops	hard	kubernetes, scheduling, affinity, topology	Pods from the same Deployment keep landing on the same node, causing a single point of failure. How do you spread them across nodes?	Use pod topology spread constraints: topologySpreadConstraints with topologyKey: kubernetes.io/hostname, maxSkew: 1, and whenUnsatisfiable: DoNotSchedule. Alternatively, use pod anti-affinity with requiredDuringSchedulingIgnoredDuringExecution matching the app label. Topology spread constraints are more flexible than anti-affinity because they allow fine-grained skew control across zones and nodes.	training/interactive/knowledge/data/cards/k8s-advanced-ops.tsv

<!-- wiki:related:start -->
---

## Wiki Navigation

### Related Content

- [Kubernetes Debugging Playbook](../../../../library/topics/k8s-debugging-playbook/index.md) (Topic Pack, L2) — Kubernetes Debugging
- Kubernetes Troubleshooting Flashcards *(CLI)* (flashcard_deck, L1) — Kubernetes Debugging

<!-- wiki:related:end -->
