---
tags:
- k8s
- l1
- flashcard-deck
- k8s-node-lifecycle
---
<!-- wiki:breadcrumb:start -->
[Portal](../../../../library/portal/index.md) | **Level:** [L1: Foundations](../../../../library/portal/levels.md) | **Topics:** [Node Lifecycle & Maintenance](../../../../library/portal/topics.md) | **Domain:** Kubernetes
<!-- wiki:breadcrumb:end -->

id	category	difficulty	tags	question	answer	source_path
k8s-node-lifecycle/8901c66badc4	k8s-node-lifecycle	easy	k8s, node, lifecycle	What is the lifecycle of a Kubernetes node?	Provision -> register (kubelet joins cluster) -> schedule workloads -> cordon (stop new pods) -> drain (move existing pods) -> decommission or upgrade -> re-register -> repeat. Nodes are treated as ephemeral — the Kubernetes model assumes they can be replaced.\n\nRemember: Node lifecycle: join→Ready→(cordon)→drain→remove. Monitor: `kubectl get nodes`.	training/library/topics/k8s-node-lifecycle/primer.md
k8s-node-lifecycle/47d97416cbff	k8s-node-lifecycle	easy	k8s, node, cordon-drain	What is the difference between cordoning and draining a node?	Cordoning marks a node as unschedulable — no new pods will be placed on it, but existing pods continue running. Draining goes further: it cordons the node AND evicts all existing pods, moving them to other nodes. Getting drain right means zero-downtime maintenance.\n\nRemember: Cordon=unschedulable, pods stay. "Caution tape." `kubectl uncordon` removes it.	training/library/topics/k8s-node-lifecycle/primer.md
k8s-node-lifecycle/18941afc46d4	k8s-node-lifecycle	easy	k8s, node, pdb	What is a PodDisruptionBudget (PDB) and why does it create tension with node drains?	A PDB specifies the minimum number (or percentage) of pods that must remain available during voluntary disruptions like node drains. It prevents drain from breaking applications by ensuring enough replicas stay running. However, PDBs are also what cause drains to get stuck — if a PDB cannot be satisfied (e.g., minAvailable equals replicas and no room to reschedule), the drain blocks indefinitely.\n\nRemember: Drain=cordon+evict. `--ignore-daemonsets --delete-emptydir-data` for stubborn pods.\n\nGotcha: DaemonSet pods can't be drained — that's why `--ignore-daemonsets` exists.	training/library/topics/k8s-node-lifecycle/primer.md
k8s-node-lifecycle/81e86ccd4f19	k8s-node-lifecycle	medium	k8s, node, taints-tolerations	How do taints and tolerations control pod scheduling on nodes?	Taints are applied to nodes to repel pods that don't explicitly tolerate them (e.g., key=gpu:NoSchedule). Tolerations are set on pods to allow scheduling on tainted nodes. Three taint effects: NoSchedule (prevent scheduling), PreferNoSchedule (soft preference), NoExecute (evict existing pods). This mechanism is used to dedicate nodes for specific workloads like GPU jobs or system components.\n\nRemember: Effects: NoSchedule(hard), PreferNoSchedule(soft), NoExecute(evict+block). Increasing severity.	training/library/topics/k8s-node-lifecycle/primer.md
k8s-node-lifecycle/0b63a7dde0c9	k8s-node-lifecycle	medium	k8s, node, conditions	What node conditions does Kubernetes monitor for health, and what happens when a node goes NotReady?	Kubernetes monitors conditions like MemoryPressure, DiskPressure, PIDPressure, NetworkUnavailable, and Ready. When a node's Ready condition becomes False or Unknown (kubelet stops reporting), the node controller waits for a grace period, then marks pods as Unknown status and begins evicting them to healthy nodes. This is the automated response to node failures.\n\nRemember: Node lifecycle: join→Ready→(cordon)→drain→remove. Monitor: `kubectl get nodes`.	training/library/topics/k8s-node-lifecycle/primer.md
k8s-node-lifecycle/3a10261de5f0	k8s-node-lifecycle	medium	k8s, node, autoscaling	What is the difference between manual node scaling and cluster autoscaling?	Manual scaling requires an operator to add or remove nodes. Cluster Autoscaler automatically adds nodes when pods are pending due to insufficient resources and removes underutilized nodes. The autoscaler respects PDBs during scale-down and checks that all pods on a node can be rescheduled elsewhere before removing it.\n\nRemember: Node lifecycle: join→Ready→(cordon)→drain→remove. Monitor: `kubectl get nodes`.	training/library/topics/k8s-node-lifecycle/primer.md
k8s-node-lifecycle/318eec5d4a5b	k8s-node-lifecycle	medium	k8s, node, upgrades	What is the standard approach for upgrading Kubernetes node versions?	Cordon the node, drain it (respecting PDBs), upgrade the kubelet and container runtime, then uncordon and allow pods to schedule back. In managed Kubernetes (EKS, GKE, AKS), this is often handled by rolling replacement of node groups — new nodes with the updated version are added while old nodes are cordoned, drained, and terminated.\n\nRemember: Node lifecycle: join→Ready→(cordon)→drain→remove. Monitor: `kubectl get nodes`.	training/library/topics/k8s-node-lifecycle/primer.md
k8s-node-lifecycle/d389c792d253	k8s-node-lifecycle	hard	k8s, node, stuck-drain	What causes a node drain to get stuck and how do you troubleshoot it?	Common causes: PDB cannot be satisfied (minAvailable equals replicas with no room to reschedule), pods with local storage (emptyDir) that can't be rescheduled, pods without a controller (bare pods not managed by a Deployment/ReplicaSet), or pods with long terminationGracePeriodSeconds. Troubleshoot by checking which pods remain, examining their PDBs, and using --ignore-daemonsets --delete-emptydir-data --force flags if safe.\n\nRemember: Drain=cordon+evict. `--ignore-daemonsets --delete-emptydir-data` for stubborn pods.\n\nGotcha: DaemonSet pods can't be drained — that's why `--ignore-daemonsets` exists.	training/library/topics/k8s-node-lifecycle/primer.md
k8s-node-lifecycle/03a23fcbb767	k8s-node-lifecycle	hard	k8s, node, kubelet	What are common causes of kubelet failures and how do they manifest?	Kubelet failures cause the node to go NotReady. Common causes: kubelet process crash (check systemctl status kubelet and journalctl -u kubelet), certificate expiration (kubelet can't authenticate to API server), disk pressure triggering evictions, container runtime failure (containerd/CRI-O down), or network partition isolating the node from the control plane. Each manifests differently in node conditions.\n\nRemember: Node lifecycle: join→Ready→(cordon)→drain→remove. Monitor: `kubectl get nodes`.	training/library/topics/k8s-node-lifecycle/primer.md
k8s-node-lifecycle/dbc21e9bdf2e	k8s-node-lifecycle	hard	k8s, node, zero-downtime	What combination of mechanisms ensures zero-downtime during node maintenance?	PDBs ensure minimum pod availability during drain. Pod anti-affinity spreads replicas across nodes so no single node holds all instances. PreStop hooks give pods time to finish in-flight requests before termination. The Cluster Autoscaler or surge capacity ensures replacement nodes exist before draining. Together: PDB + anti-affinity + graceful shutdown + capacity headroom = zero-downtime maintenance.\n\nRemember: Node lifecycle: join→Ready→(cordon)→drain→remove. Monitor: `kubectl get nodes`.	training/library/topics/k8s-node-lifecycle/primer.md

<!-- wiki:related:start -->
---

## Wiki Navigation

### Related Content

- [Case Study: DaemonSet Blocks Eviction](../../../../library/case-studies/kubernetes_ops/daemonset-blocks-eviction/README.md) (Case Study, L2) — Node Lifecycle & Maintenance
- [Kubernetes Node Lifecycle](../../../../library/topics/k8s-node-lifecycle/index.md) (Topic Pack, L2) — Node Lifecycle & Maintenance
- [Kubernetes Ops (Production)](../../../../library/topics/k8s-ops/index.md) (Topic Pack, L2) — Node Lifecycle & Maintenance
- [Node Maintenance](../../../../library/topics/node-maintenance/index.md) (Topic Pack, L1) — Node Lifecycle & Maintenance
- [Runbook: Node NotReady](../../../../library/runbooks/kubernetes/node-not-ready.md) (Runbook, L1) — Node Lifecycle & Maintenance
- [Skillcheck: Kubernetes Under the Covers](../../../../library/skillchecks/kubernetes.under.the.covers.md) (Assessment, L2) — Node Lifecycle & Maintenance

<!-- wiki:related:end -->
