Welcome to TiddlyWiki created by Jeremy Ruston; Copyright © 2004-2007 Jeremy Ruston, Copyright © 2007-2011 UnaMesa Association
A Kubernetes Ingress is an API object that defines routing rules for external HTTP/HTTPS traffic to reach internal Services. An ingress controller is the pod (or set of pods) that implements those rules: it watches the Kubernetes API for Ingress or Gateway API resources, dynamically generates proxy configuration (nginx.conf, Traefik rules, or equivalent), and routes incoming requests to the correct backend Service.
At the cluster edge, the ingress controller acts as a reverse proxy and centralized policy enforcement point. It terminates TLS, authenticates requests, applies rate limits, transforms requests, and collects logs and metrics — all before traffic reaches backend services. This frees individual services from embedding security, rate-limiting, and observability boilerplate.
The alternative — exposing each service via LoadBalancer or NodePort — is expensive (a cloud load balancer per service), operationally awkward, and insecure. A single ingress controller handles all external traffic, routing by host or path to the appropriate backend.
The typical layering is: external traffic → cloud load balancer or MetalLB → ingress controller pods → backend Services. Centralizing routing and security policy at the cluster edge rather than inside each service is a core pattern in production Kubernetes.
----
''Sources''
* <html><code>training/library/topics/api-gateways/primer.md</code></html>
* <html><code>training/interactive/knowledge/data/cards/k8s-services.tsv</code></html>
//Merged from 3 source atoms.//
kubeconfig is a YAML file (default: <html><code>~/.kube/config</code></html>) that stores multiple Kubernetes cluster endpoints, authentication credentials, and named contexts. Each context combines a cluster, user, and namespace triple. <html><code>kubectl</code></html> uses the current context—controlled by <html><code>kubectl config use-context</code></html> or the <html><code>--context</code></html> flag—to determine which cluster receives a request.
Multiple kubeconfig files can be merged at runtime by setting the <html><code>KUBECONFIG</code></html> environment variable to a colon-separated list of paths (e.g., <html><code>KUBECONFIG=f1:f2 kubectl config view --merge --flatten > merged</code></html>). The file can also be overridden per-command with <html><code>--kubeconfig</code></html>. Organizations can distribute kubeconfigs as Git-versioned files, letting users merge them locally and switch clusters without reconfiguring tools.
Tools like <html><code>kubectx</code></html> provide shortcuts for context switching; <html><code>kubectx -</code></html> toggles back to the previous context. Credentials stored in kubeconfig are unencrypted by default—use a secrets store or short-lived tokens for production access rather than long-lived static credentials in the file.
----
''Sources''
* <html><code>training/library/topics/cloud-deep-dive/multi-cluster.md</code></html>
* <html><code>training/interactive/knowledge/data/cards/k8s-ops.tsv</code></html>
//Merged from 2 source atoms.//
''Related atoms''
* [[You are managing multiple Kubernetes clusters. How do you quickly change between the cl…]]
* [[What is kubconfig? What do you use it for?]]
CrashLoopBackOff means a container starts, crashes immediately, and Kubernetes restarts it with exponential backoff (10s → 20s → 40s → up to 5 min). The pod is unexecutable while cycling.
''Systematic approach (GDLE: Get → Describe → Logs → Exec):''
''1. Logs first:''
<html><code>kubectl logs <pod> --previous</code></html> reads stdout/stderr from the last crashed container. If the container exits before writing anything, the problem is in the command or args.
''2. Describe for events:''
<html><code>kubectl describe pod <pod></code></html> — look for OOMKilled, image pull errors, failed mounts, probe failures. Exit codes: 1 = app error, 137 = SIGKILL/OOMKilled, 139 = segfault, 143 = SIGTERM, 126 = permission denied on binary.
''3. When logs are empty or container crashes too fast:''
* ''Sleep override:'' <html><code>kubectl run debug-app --image=<image>:<tag> --restart=Never --command -- sleep 3600</code></html>, then exec in to inspect filesystem, env, and network connectivity.
* ''Ephemeral debug container (K8s 1.23+):'' <html><code>kubectl debug -it <pod> --image=busybox --target=<container></code></html> attaches a debug image sharing the pod's process namespace and filesystems.
* ''Copy pod with override:'' <html><code>kubectl debug <pod> --copy-to=debug-pod --container=<container> -- sleep 3600</code></html> creates a modified copy that sleeps instead of crashing.
* ''Node-level via crictl:'' <html><code>kubectl debug node/<node> -it --image=busybox</code></html> → <html><code>chroot /host</code></html> → <html><code>crictl ps -a --name <container> -q | head -1</code></html> → <html><code>crictl logs <id></code></html>. Works even if the pod image has no shell.
Most crashes are caused by: missing config, wrong image tag, memory limit too low (OOMKilled), wrong entrypoint, or dependency not ready.
----
''Sources''
* <html><code>training/library/topics/containers-deep-dive/street_ops.md</code></html>
* <html><code>training/library/topics/crashloopbackoff/primer.md</code></html>
* <html><code>training/library/topics/crashloopbackoff/street_ops.md</code></html>
* <html><code>training/interactive/knowledge/data/cards/k8s-ops.tsv</code></html>
* <html><code>training/interactive/knowledge/data/cards/k8s-troubleshooting.tsv</code></html>
* <html><code>training/library/topics/k8s-ops/primer.md</code></html>
* <html><code>training/library/topics/k8s-pods-and-scheduling/street_ops.md</code></html>
* <html><code>training/library/topics/k8s-debugging-playbook/primer.md</code></html>
//Merged from 6 source atoms.//
''Related atoms''
* [[Why do container logs sometimes bring down Kubernetes nodes?]]
* [[Diagnosing a stuck Kubernetes Deployment rollout]]
Q: Why do container logs sometimes bring down Kubernetes nodes?
A: Uncontrolled log growth exhausts disk space or inodes, causing kubelet failure.
The cascade:
# Application logs to stdout/stderr
# Container runtime captures to JSON files
# Files grow without bound (misconfigured rotation)
# Disk fills OR inodes exhausted
# kubelet can't function (needs disk for pods)
# Node goes NotReady
# Pods rescheduled, bring their logs elsewhere
# Repeat on next node
Specific failure modes:
Example: <html><code>kubectl logs -f pod --previous</code></html> — follow live or show crashed container logs.
Remember: Multi-container: <html><code>--container=name</code></html>. Errors without it on multi-container pods.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-ops.tsv</code></html>
''Related atoms''
* [[Debugging CrashLoopBackOff pods: logs, debug containers, node inspection]]
* [[kubectl logs: view container logs in a Pod]]
Q: How do you prevent high memory usage in your Kubernetes cluster and possibly issues like memory leak and OOM?
A: Apply requests and limits, especially on third party applications (where the uncertainty is even bigger)
Example: Set in pod spec: <html><code>resources: {limits: {memory: 512Mi}, requests: {memory: 256Mi}}</code></html>.
Remember: "Requests = minimum guaranteed, Limits = maximum allowed." Think "Request a seat, Limit the legroom."
Gotcha: CPU is throttled; memory is OOMKilled. Over-limit CPU just slows down; over-limit memory kills the container.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-config.tsv</code></html>
''Related atoms''
* [[Pods without memory limits cause node-wide evictions and OOM kills]]
* [[What is a LimitRange and how does it prevent OOMKilled caused by missing resource limits?]]
CrashLoopBackOff is one of several failure states, each with a different root cause. The key distinction is whether the container process ever executed.
''CrashLoopBackOff'': The container started and ran, then crashed — check application logs and exit codes via <html><code>kubectl logs pod --previous</code></html>. Kubernetes retries with exponential backoff (10s → 20s → 40s → 5 min cap). Common causes: missing config, OOM kill, wrong entrypoint command.
''ImagePullBackOff'': The container image could not be pulled — wrong tag, missing registry auth, or unreachable registry. The container never starts.
''CreateContainerConfigError'': The pod spec references a ConfigMap or Secret that does not exist. The container never starts.
''RunContainerError'': The runtime (containerd, docker) cannot start the process — bad seccomp profile, missing device, or cgroup issue.
''ErrImageNeverPull'': The image does not exist locally and <html><code>imagePullPolicy: Never</code></html> forbids pulling it.
''Pending'': The pod has not been scheduled — no node has sufficient resources, or affinity/taint constraints prevent placement.
If the pod phase is <html><code>Running</code></html> and the container waiting reason is <html><code>CrashLoopBackOff</code></html>, the container did start — logs and exit codes are the correct diagnostic path. All other states above indicate the container never reached execution.
----
''Sources''
* <html><code>training/library/topics/crashloopbackoff/primer.md</code></html>
* <html><code>training/interactive/knowledge/data/cards/k8s-troubleshooting.tsv</code></html>
//Merged from 2 source atoms.//
''Related atoms''
* [[How do init containers help prevent CrashLoopBackOff caused by missing dependencies?]]
* [[What are the top 3 causes of CrashLoopBackOff?]]
Q: What three kubectl commands form the basic CrashLoopBackOff diagnostic workflow?
A: 1) kubectl get pods (see restart count and status), 2) kubectl describe pod (events, exit codes, last state), 3) kubectl logs --previous (see what the container printed before dying).
Remember: CrashLoopBackOff = start-crash-retry with exponential backoff (10s→20s→40s→5min).
Gotcha: <html><code>kubectl logs pod --previous</code></html> shows crash reason. Common: missing config, OOM, wrong cmd.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-troubleshooting.tsv</code></html>
''Related atoms''
* [[What is the standard three-command diagnostic flow when a pod is misbehaving?]]
<html><code>kubectl logs <pod> --previous</code></html> retrieves logs from the last terminated container instance, not the currently running (or re-starting) one. Without <html><code>--previous</code></html>, you get no output when the current container has crashed before logging anything — the flag is the single most effective first step for debugging CrashLoopBackOff. For multi-container pods, specify the target: <html><code>kubectl logs <pod> -c <container> --previous</code></html>. Combine with <html><code>--tail=100</code></html> to limit output: <html><code>kubectl logs <pod> --previous --tail=100</code></html>.
The previous logs reveal crash category at a glance: a startup failure shows an error during initialization; a runtime failure shows the app running then an error; no logs at all (exit 127) typically means the binary was not found. If you instead want to observe the current container as it starts, use <html><code>--follow --tail=100</code></html> without <html><code>--previous</code></html>, accepting that output may be incomplete if the container crashes immediately.
Logs are retained on the node until the pod is deleted or the kubelet reclaims the space — they are not durable. For persistence across pod deletion, set up a log aggregation pipeline (Fluentd, Loki, etc.).
----
''Sources''
* <html><code>training/library/topics/crashloopbackoff/street_ops.md</code></html>
* <html><code>training/library/topics/crashloopbackoff/footguns.md</code></html>
* <html><code>training/library/topics/crashloopbackoff/trivia.md</code></html>
* <html><code>training/interactive/knowledge/data/cards/k8s-troubleshooting.tsv</code></html>
//Merged from 3 source atoms.//
<html><code>kubectl rollout undo</code></html> reverts a deployment to the immediately-previous ReplicaSet, not necessarily a known-good one. If that prior version also contained a bug, or if multiple consecutive releases were broken, the rollback lands on a still-broken state and can extend an outage rather than resolve it.
The safe procedure:
# Run <html><code>kubectl rollout history deployment/<name></code></html> to list all revisions. Each entry shows a revision number and, if set, a <html><code>change-cause</code></html> annotation describing what was deployed.
# Identify a revision that is known to have been healthy.
# Target it explicitly: <html><code>kubectl rollout undo deployment/<name> --to-revision=<N></code></html>.
# Confirm recovery with <html><code>kubectl rollout status deployment/<name></code></html> and verify pod readiness before declaring the incident resolved.
This distinction is most consequential during rapid iteration or urgent patching, where developers may have pushed several unstable revisions in quick succession. A hasty <html><code>rollout undo</code></html> without inspecting history can cycle through multiple broken versions. Always confirm the rollback destination was actually stable before issuing the command.
----
''Sources''
* <html><code>training/library/topics/crashloopbackoff/street_ops.md</code></html>
* <html><code>training/interactive/knowledge/data/cards/k8s-advanced-ops.tsv</code></html>
//Merged from 2 source atoms.//
''Related atoms''
* [[Kubernetes Deployment rolling updates and rollbacks]]
Exit code 137 = 128 + 9 (SIGKILL). In container and Kubernetes contexts, 137 almost always means the kernel OOM killer forcibly terminated the process because it exceeded its memory limit. The application receives no signal handler opportunity and is killed instantly without warning.
Diagnosis: container logs will be empty or incomplete because the kill is external. Check <html><code>kubectl describe pod | grep 'Last State'</code></html> for <html><code>Reason: OOMKilled</code></html>. On the host, <html><code>dmesg</code></html> records the OOM killer's victim process.
Fix path: inspect the memory limit in the pod spec and compare against actual process memory usage. If the process consistently approaches the limit, increase it. If usage is well below the limit, suspect a memory leak.
A common trap: teams spend days auditing application code or adding logging, only to discover the fix was increasing the memory limit by 50 MB. When exit code 137 appears, start with the resource limit—not the application.
----
''Sources''
* <html><code>training/library/topics/crashloopbackoff/trivia.md</code></html>
* <html><code>training/interactive/knowledge/data/cards/k8s-ops.tsv</code></html>
//Merged from 2 source atoms.//
''Related atoms''
* [[Pod deletion: SIGTERM, grace period, and SIGKILL]]
When a liveness probe fails, kubelet sends SIGKILL to the container and restarts it immediately. The behavior is indistinguishable from a genuine crash: the pod enters CrashLoopBackOff, exit codes increment, and restart counts climb. Applications that need 60–300+ seconds to initialize (common in JVM languages, large Go binaries, or frameworks with heavy startup work) fail liveness checks before they ever become healthy — even if the application code is correct.
Kubernetes 1.18 introduced <html><code>startupProbe</code></html> to break this catch-22. The startup probe runs exclusively during initialization and tolerates failures until the app signals readiness; once it succeeds once, liveness and readiness probes activate. A generous configuration such as <html><code>failureThreshold: 60, periodSeconds: 5</code></html> grants 300 seconds of initialization budget without loosening the liveness probe's sensitivity to stuck processes.
Modern deployments should separate the three probe responsibilities: <html><code>startupProbe</code></html> for slow initialization, <html><code>livenessProbe</code></html> for detecting hung or deadlocked processes, and <html><code>readinessProbe</code></html> for temporarily blocking traffic. Conflating all three into a single liveness probe creates false crash-loops that appear as application bugs but are actually misconfigured observability.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
* <html><code>training/interactive/knowledge/data/cards/k8s-troubleshooting.tsv</code></html>
* <html><code>training/library/topics/crashloopbackoff/footguns.md</code></html>
* <html><code>training/library/topics/crashloopbackoff/trivia.md</code></html>
* <html><code>training/library/topics/crashloopbackoff/street_ops.md</code></html>
* <html><code>training/library/topics/k8s-ops/primer.md</code></html>
* <html><code>training/library/topics/k8s-ops/footguns.md</code></html>
//Merged from 2 source atoms.//
''Related atoms''
* [[Diagnose and fix failing Kubernetes liveness/readiness probes]]
PVCs are designed to outlive the pods and StatefulSets that use them — deleting a StatefulSet leaves its PVCs intact, preventing accidental data loss. When a PVC is eventually deleted, the <html><code>reclaimPolicy</code></html> of the underlying PersistentVolume determines what happens to the storage.
''Retain'': Preserves the PV and underlying cloud disk (EBS, GCE PD, etc.). The PV moves to <html><code>Released</code></html> status; its <html><code>claimRef</code></html> must be cleared manually before it can be reused with a new PVC. Manual cleanup is required to reclaim storage cost. Operators who deleted "orphaned" PVCs without realizing they held live data have caused data loss through this path.
''Delete'': Immediately destroys both the PV and underlying storage upon PVC deletion. This is the default <html><code>reclaimPolicy</code></html> in most StorageClasses. A teammate removing a PVC to clean up will permanently destroy database data with no recovery path.
''Recycle'' (performs <html><code>rm -rf /thevolume/*</code></html> to clean and reuse): Deprecated. Do not rely on it.
For databases and any stateful workload, create a dedicated StorageClass with <html><code>reclaimPolicy: Retain</code></html>. Use <html><code>Delete</code></html> only for caches and ephemeral workloads where automatic cleanup reduces operational burden.
StatefulSet scaling follows the same safety-first pattern: scaling down keeps all PVCs intact; scaling back up reattaches the same volumes. Data is preserved across deployments automatically.
----
''Sources''
* <html><code>training/library/topics/database-ops/primer.md</code></html>
* <html><code>training/library/topics/k8s-storage/footguns.md</code></html>
* <html><code>training/interactive/knowledge/data/cards/k8s-storage.tsv</code></html>
* <html><code>training/library/topics/k8s-storage/trivia.md</code></html>
* <html><code>training/library/topics/k8s-storage/primer.md</code></html>
* <html><code>training/library/topics/database-ops/footguns.md</code></html>
* <html><code>training/library/topics/k8s-storage/street_ops.md</code></html>
//Merged from 4 source atoms.//
''Related atoms''
* [[StatefulSets and volumeClaimTemplates: stable storage per replica]]
Q: A PVC is stuck in Pending state. Walk through your debugging process.
A: 1) kubectl describe pvc to check events. Common causes: no matching PV and no StorageClass provisioner (check SC exists and CSI driver pods are healthy), WaitForFirstConsumer mode (normal until a pod is scheduled), storageclass not found (typo in SC name), exceeded ResourceQuota, or requested capacity exceeds provider limits. 2) Check CSI driver pods in kube-system. 3) Check node availability zones vs volume binding mode.
Remember: Storage abstraction: StorageClass→PV→PVC→Pod mount. Dynamic provisioning automates PV.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-advanced-ops.tsv</code></html>
* <html><code>training/interactive/knowledge/data/cards/k8s-storage.tsv</code></html>
* <html><code>training/interactive/knowledge/data/cards/k8s-troubleshooting.tsv</code></html>
//Merged from 3 source atoms.//
''Related atoms''
* [[A pod with a PVC can't start and shows a multi-attach error. What's wrong?]]
* [[Diagnosing a stuck Kubernetes Deployment rollout]]
* [[Kubernetes dynamic provisioning and WaitForFirstConsumer zone awareness]]
Q: What is the difference between liveness and readiness probes in Kubernetes?
A: Both are health check mechanisms, but they serve different purposes:
Liveness Probe:
* Determines if the container is running
* If it fails, kubelet kills the container and restarts it
* Use case: Detect deadlocks, unresponsive applications
* "Is the application alive?"
Readiness Probe:
* Determines if the container is ready to receive traffic
* If it fails, Pod is removed from Service endpoints
* Container is NOT restarted
* Use case: Waiting for dependencies, warming caches
* "Is the application ready to serve requests?"
Best practices:
* Liveness: Simple check, shouldn't depend on external services
* Readiness: Can include dependency checks
* Don't make liveness check the same as readiness
* Set appropriate initialDelaySeconds for slow-starting apps
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
* <html><code>training/interactive/knowledge/data/cards/k8s-troubleshooting.tsv</code></html>
//Merged from 2 source atoms.//
''Related atoms''
* [[Startup probes replaced the initialDelaySeconds compromise in Kubernetes 1.18]]
* [[startupProbe prevents false crash-loops during slow initialization]]
Q: How should liveness and readiness probe endpoints differ in what they check, and why?
A: Liveness should only verify the process is alive and responsive (no dependency checks) — it answers "should I restart?" Readiness should check dependencies, cache warmth, and load — it answers "can I serve traffic?" Using the same endpoint for both causes dependency failures to trigger unnecessary restarts.
Remember: <html><code>kubectl explain <resource></code></html> for fields. <html><code>kubectl api-resources</code></html> for all resource types.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-ops.tsv</code></html>
''Related atoms''
* [[What happens when a readiness probe fails on a Kubernetes pod?]]
* [[How do readiness probes interact with rolling deployments?]]
* [[startupProbe prevents false crash-loops during slow initialization]]
Q: What is the constraint on successThreshold for liveness and startup probes, and why does it matter for readiness?
A: For liveness and startup probes, successThreshold must be 1 (Kubernetes ignores other values). Only readiness probes can require multiple consecutive successes before the pod is added back to endpoints. This matters because you may want a readiness probe to confirm stability (e.g., successThreshold=2) before resuming traffic after a transient failure.
Remember: <html><code>kubectl explain <resource></code></html> for fields. <html><code>kubectl api-resources</code></html> for all resource types.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-ops.tsv</code></html>
''Related atoms''
* [[How do readiness probes interact with rolling deployments?]]
* [[Startup probes replaced the initialDelaySeconds compromise in Kubernetes 1.18]]
Kubernetes pods have <html><code>ndots:5</code></html> set in their <html><code>/etc/resolv.conf</code></html> by default, causing any name with fewer than 5 dots to trigger searches with cluster-local suffixes appended. A lookup for <html><code>api.github.com</code></html> (2 dots) tries <html><code>api.github.com.default.svc.cluster.local</code></html>, <html><code>api.github.com.svc.cluster.local</code></html>, <html><code>api.github.com.cluster.local</code></html> before finally querying <html><code>api.github.com.</code></html> as an absolute FQDN — tripling the DNS load. In large clusters with high external DNS call volume (e.g., 1,000 calls/second), this generates 3,000 extra queries per second, overwhelming CoreDNS, causing latency spikes, and triggering cascading timeouts. Fix: for pods that primarily make external DNS queries, set <html><code>dnsConfig.options.ndots=2</code></html> in the pod spec. Do not set ndots to 0 or 1 cluster-wide — that breaks short-name Kubernetes service resolution within the cluster.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-advanced-ops.tsv</code></html>
* <html><code>training/interactive/knowledge/data/cards/k8s-troubleshooting.tsv</code></html>
* <html><code>training/library/topics/dns-deep-dive/footguns.md</code></html>
* <html><code>training/library/topics/dns-ops/street_ops.md</code></html>
* <html><code>training/library/topics/k8s-networking/trivia.md</code></html>
* <html><code>training/library/topics/k8s-networking/primer.md</code></html>
* <html><code>training/library/topics/k8s-networking/street_ops.md</code></html>
* <html><code>training/library/topics/k8s-services-and-ingress/footguns.md</code></html>
* <html><code>training/library/topics/k8s-services-and-ingress/primer.md</code></html>
* <html><code>training/library/topics/k8s-services-and-ingress/street_ops.md</code></html>
* <html><code>training/library/topics/dns-deep-dive/street_ops.md</code></html>
* <html><code>training/library/topics/dns-ops/footguns.md</code></html>
* <html><code>training/library/topics/k8s-networking/footguns.md</code></html>
* <html><code>training/library/topics/k8s-ops/street_ops.md</code></html>
* <html><code>training/library/topics/networking/footguns.md</code></html>
//Merged from 15 source atoms.//
''Related atoms''
* [[Pod reaches Service by ClusterIP but internal DNS fails: diagnosis]]
* [[Default-deny egress NetworkPolicies must explicitly allow DNS]]
Kubernetes automatically generates DNS names for services and pods following a hierarchical pattern: services resolve to <html><code><service>.<namespace>.svc.cluster.local</code></html>, returning a stable ClusterIP (or individual Pod IPs for headless services). Pods have an address-based name: <html><code><ip-dashed>.<namespace>.pod.cluster.local</code></html>. StatefulSet pods get predictable names: <html><code><pod-name>.<service>.<namespace>.svc.cluster.local</code></html>. Headless services (ClusterIP: None) return individual Pod IPs instead of a single virtual IP, enabling client-side load balancing and stateful ordering. All names resolve within the cluster via the CoreDNS or kube-dns addon. This built-in DNS layer eliminates the need for external service registries — Kubernetes' own API server is the source of truth for service endpoints. Applications inside pods can use short names (e.g., <html><code>api-server</code></html>) with implicit namespace search, or fully qualified names for cross-namespace discovery.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-concept-chain.tsv</code></html>
* <html><code>training/library/topics/dns-deep-dive/primer.md</code></html>
* <html><code>training/library/topics/k8s-networking/primer.md</code></html>
* <html><code>training/library/topics/k8s-services-and-ingress/primer.md</code></html>
//Merged from 4 source atoms.//
Q: How does Kubernetes handle DNS resolution for services and pods?
A: ''DNS Resolution in Kubernetes:''
* Pods are assigned a DNS name based on their name and namespace.
* The internal DNS service in Kubernetes resolves these names to corresponding pod IP addresses.
* Services also have DNS names based on their name and namespace, enabling easy service discovery.
* Kubernetes has an internal DNS service that automatically assigns DNS names to pods and services. This allows for easy and dynamic DNS-based service discovery within the cluster.
Remember: Services use label selectors to find pods. Change labels→change which pods get traffic.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-services.tsv</code></html>
''Related atoms''
* [[Pod IPs are ephemeral; use a Service for stable addressing]]
* [[Kubernetes Service: stable networking endpoint for pods]]
* [[Describe what happens when a container tries to connect with its corresponding Service …]]
Q: Describe the role of etcd in a Kubernetes cluster.
A: ''etcd:''
* Distributed, consistent, and highly available key-value store.
* Stores configuration data, state information, and metadata about the cluster.
* Provides a reliable and centralized data store for the entire Kubernetes control plane.
* etcd ensures consistency among the components of the Kubernetes control plane by serving as the primary data store. It holds critical information such as cluster configuration, node status, and metadata about pods, services, and other resources. Its distributed nature makes it resilient to failures, contributing to the overall reliability of the Kubernetes cluster.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
//Merged from 3 source atoms.//
''Related atoms''
* [[Explain the main components of Kubernetes architecture.]]
Cluster Autoscaler scales pre-defined node groups (ASGs/MIGs) up and down based on pending pods. It is constrained to predefined instance types and introduces configuration delays — scale-down waits approximately 10 minutes before removing underutilized nodes. It is portable across cloud providers and operationally simpler.
Karpenter (AWS) provisions individual, right-sized nodes on-demand based directly on pod resource requests, without pre-defined node groups. It independently selects instance type, capacity type (spot vs. on-demand), and node lifetime, making scaling decisions faster and more cost-efficient. Karpenter also consolidates underutilized nodes in real time.
Key tradeoffs: Karpenter delivers faster provisioning, greater flexibility, and more aggressive cost optimization, making it preferable for variable or cost-sensitive workloads. Cluster Autoscaler is simpler to operate and cloud-agnostic, making it the practical choice when portability or operational simplicity matters more than scaling speed.
----
''Sources''
* <html><code>training/library/topics/finops/primer.md</code></html>
* <html><code>training/interactive/knowledge/data/cards/k8s-concept-chain.tsv</code></html>
//Merged from 2 source atoms.//
''Related atoms''
* [[Manual node scaling vs Cluster Autoscaler in Kubernetes]]
* [[What is the purpose of Horizontal Pod Autoscaling in Kubernetes?]]
A Deployment rollout stalls when new pods fail to become Ready within <html><code>progressDeadlineSeconds</code></html> (default 600 s). <html><code>kubectl rollout status deployment/<name> -n <ns></code></html> will hang or report the stall. <html><code>kubectl describe deployment <name></code></html> exposes the reason under Conditions (e.g., <html><code>ProgressDeadlineExceeded</code></html>).
''Diagnosis steps (Get → Describe → Logs → Exec):''
# Find the new ReplicaSet: <html><code>kubectl get rs -n <ns> -l app=<name></code></html>, then <html><code>kubectl describe rs <new-rs></code></html>.
# Describe stuck pods: <html><code>kubectl describe pod <pod></code></html> — Events explain //why//, not just //what// failed.
# If pods are Pending, check node pressure: <html><code>kubectl describe node</code></html>.
# If pods are Running but not Ready, inspect readiness probe output: <html><code>kubectl logs <pod> -n <ns></code></html>.
''Common causes:''
* ''Readiness probe failure'' — pod is running but never passes readiness, so Kubernetes never marks it healthy and won't terminate the old pod. This is the most frequent cause.
* ''ImagePullBackOff'' — image tag doesn't exist or registry is unreachable.
* ''OOMKilled on startup'' — container exceeds its memory limit immediately.
* ''Resource quota exhaustion'' — no node has sufficient CPU/memory to schedule the new pod.
* ''PodDisruptionBudget (PDB)'' — blocks eviction of old pods, preventing progress.
* ''Missing dependencies'' — ConfigMap, Secret, or external service unavailable.
''Recovery:'' <html><code>kubectl rollout undo deploy/<name></code></html> reverts to the last working revision while the root cause is addressed.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-advanced-ops.tsv</code></html>
* <html><code>training/library/topics/k8s-ops/street_ops.md</code></html>
* <html><code>training/interactive/knowledge/data/cards/k8s-ops.tsv</code></html>
* <html><code>training/interactive/knowledge/data/cards/k8s-troubleshooting.tsv</code></html>
//Merged from 4 source atoms.//
''Related atoms''
* [[Debugging CrashLoopBackOff pods: logs, debug containers, node inspection]]
When a pod restarts repeatedly or enters NotReady, the most likely cause is a misconfigured liveness, readiness, or startup probe — not an application bug.
''Diagnose:'' <html><code>kubectl describe pod <pod></code></html> shows the probe spec and last failure message (e.g., HTTP status, connection error). <html><code>kubectl get events -n <ns> --field-selector involvedObject.name=<pod></code></html> shows probe failures with timestamps.
''Common misconfigurations:''
* <html><code>initialDelaySeconds</code></html> too short — the app isn't ready when the first probe fires; increase it.
* Wrong port or path — the probe's <html><code>containerPort</code></html> and <html><code>httpGet.path</code></html> don't match the app's actual endpoint; verify both.
* <html><code>timeoutSeconds</code></html> too short — the endpoint is slow; increase the timeout (must stay below <html><code>periodSeconds</code></html>).
* <html><code>periodSeconds</code></html> too aggressive — probing every 1 s on a 10 s-slow endpoint will always fail; increase the period.
''Key principle:'' Probe failures are almost always a configuration mismatch. If the application logs show no errors and the app is actually healthy, fix the probe spec, not the application.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-advanced-ops.tsv</code></html>
* <html><code>training/interactive/knowledge/data/cards/k8s-troubleshooting.tsv</code></html>
* <html><code>training/library/topics/k8s-ops/primer.md</code></html>
//Merged from 2 source atoms.//
''Related atoms''
* [[startupProbe prevents false crash-loops during slow initialization]]
Q: An app takes 120 seconds to initialize. Liveness probe kills it before startup completes. How do you fix this without removing the liveness probe?
A: Add a startup probe with a generous failureThreshold
* periodSeconds window (e.g., failureThreshold: 30, periodSeconds: 5 = 150s). The startup probe blocks liveness and readiness probes until it succeeds. This keeps fast-restart protection while allowing slow cold starts.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-advanced-ops.tsv</code></html>
''Related atoms''
* [[Startup probes replaced the initialDelaySeconds compromise in Kubernetes 1.18]]
* [[startupProbe prevents false crash-loops during slow initialization]]
Q: What is the standard three-command diagnostic flow when a pod is misbehaving?
A: 1) kubectl describe pod <pod> — check Events, conditions, and container state.
2) kubectl logs <pod> [-c container] [--previous] — read application stdout/stderr.
3) kubectl get events --field-selector involvedObject.name=<pod> — see cluster-level events. This describe-logs-events flow covers 90% of initial triage.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-advanced-ops.tsv</code></html>
''Related atoms''
* [[What three kubectl commands form the basic CrashLoopBackOff diagnostic workflow?]]
* [[Debugging a failing or non-starting pod in Kubernetes]]
Check taints with <html><code>kubectl describe node <node> | grep Taints</code></html>. Common system taints include <html><code>node.kubernetes.io/not-ready</code></html>, <html><code>node.kubernetes.io/memory-pressure</code></html>, and <html><code>node.kubernetes.io/disk-pressure</code></html>. Pods require matching tolerations in their spec to schedule on tainted nodes.
If the taint is intentional (e.g., <html><code>dedicated=gpu:NoSchedule</code></html> for GPU-only workloads), add a matching toleration to the pods that should run there. If the taint is accidental, remove it using the trailing-minus syntax: <html><code>kubectl taint nodes <node> dedicated=gpu:NoSchedule-</code></html>.
General K8s troubleshooting order: Get → Describe → Logs → Exec (mnemonic: GDLE). Always inspect Events via <html><code>kubectl describe</code></html> — they explain WHY scheduling failed, not just that it did.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-advanced-ops.tsv</code></html>
* <html><code>training/interactive/knowledge/data/cards/k8s-troubleshooting.tsv</code></html>
//Merged from 2 source atoms.//
''Related atoms''
* [[Diagnosing Kubernetes pods stuck in Pending state]]
Q: Pods are Pending cluster-wide after a control plane upgrade. All worker nodes show a NoSchedule taint. What happened and how do you recover?
A: The upgrade likely re-applied node-role.kubernetes.io/control-plane:NoSchedule and may have incorrectly tainted workers. Verify with kubectl get nodes -o custom-columns=NAME:.metadata.name,TAINTS:.spec.taints. Remove errant taints from workers: kubectl taint nodes <worker> <key>:NoSchedule-. For control-plane nodes, add tolerations only for system-critical pods.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-advanced-ops.tsv</code></html>
''Related atoms''
* [[Diagnosing pods that won't schedule on a node due to taints]]
Expanding a PVC requires the StorageClass to have <html><code>allowVolumeExpansion: true</code></html>. Without it, patching the PVC's storage request appears to succeed — the size field in the object changes — but the underlying volume does not grow and no error is surfaced. The pod eventually runs out of disk space. Always verify first: <html><code>kubectl get storageclass <name> -o jsonpath='{.allowVolumeExpansion}'</code></html>.
To expand: <html><code>kubectl patch pvc <name> -p '{"spec":{"resources":{"requests":{"storage":"50Gi"}}}}'</code></html>. For cloud-backed volumes, the block device grows in the background. However, the filesystem inside the container does not expand until the next mount — this is the two-phase nature of online expansion. If <html><code>kubectl get pvc</code></html> shows the new capacity but <html><code>df</code></html> inside the pod still reports the old size, the resize is pending. Delete and recreate the pod (not the PVC) to trigger a remount; Kubernetes automatically runs <html><code>resize2fs</code></html> (ext4) or <html><code>xfs_growfs</code></html> (XFS) at that point.
Verify completion: <html><code>kubectl describe pvc <name></code></html> or <html><code>kubectl get pvc <name> -o jsonpath='{.status.conditions}'</code></html>. A lingering <html><code>FileSystemResizePending</code></html> condition indicates the filesystem resize has not yet occurred and a pod bounce is still needed.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-advanced-ops.tsv</code></html>
* <html><code>training/library/topics/k8s-storage/footguns.md</code></html>
* <html><code>training/library/topics/k8s-storage/street_ops.md</code></html>
//Merged from 3 source atoms.//
The most common Kubernetes networking bug is a mismatch between a Service's <html><code>selector</code></html> and pod labels, causing zero endpoints to register. Traffic to the Service ClusterIP results in connection refused or timeout; no error appears on the Service object itself — both resources exist and look correct.
''Diagnosis''
# <html><code>kubectl get endpoints <svc></code></html> — if the list shows <html><code><none></code></html>, no pods matched the selector.
# Compare the Service selector against actual pod labels:
** <html><code>kubectl get svc <svc> -o jsonpath='{.spec.selector}'</code></html>
** <html><code>kubectl get pods --show-labels</code></html>
# Common mismatches: typos (<html><code>app: api-server</code></html> vs <html><code>app: api</code></html>), missing label keys, version labels that diverge over time.
# If labels match but endpoints are still empty, check pod readiness — unready pods are excluded from endpoints even when the selector matches. Inspect readiness probe logs.
# Verify <html><code>targetPort</code></html> matches the port the container actually listens on: <html><code>kubectl exec <pod> -- ss -tlnp</code></html>.
# Test connectivity from another pod using the Service's cluster-internal DNS name.
''Prevention''
Use <html><code>kubectl expose deployment <name></code></html> to auto-generate a Service with selectors derived directly from the Deployment's pod template, eliminating manual label transcription errors.
''Namespace gotcha''
Pods and Services must be in the same namespace; a cross-namespace selector silently produces empty endpoints.
''Debug flow mnemonic'': Get → Describe → Logs → Exec (GDLE). Start with <html><code>kubectl get endpoints</code></html> before reaching for Describe or Logs.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-advanced-ops.tsv</code></html>
* <html><code>training/library/topics/k8s-networking/footguns.md</code></html>
* <html><code>training/library/topics/k8s-services-and-ingress/footguns.md</code></html>
* <html><code>training/library/topics/k8s-services-and-ingress/street_ops.md</code></html>
* <html><code>training/interactive/knowledge/data/cards/k8s-troubleshooting.tsv</code></html>
//Merged from 5 source atoms.//
''Related atoms''
* [[How do you find pods that match a particular label selector?]]
* [[Users unable to reach an application running on a Pod on Kubernetes. What might be the …]]
Q: Pods are being evicted with the message 'The node was low on resource: ephemeral-storage'. What is happening?
A: Kubelet's eviction manager detected ephemeral storage usage exceeding the threshold (default ~85%). It evicts pods in order of priority and usage. Check node conditions: kubectl describe node <node> | grep -A 5 Conditions. Clean up: remove unused images (crictl rmi --prune), check for pods writing large files to emptyDir, and set resource limits on ephemeral-storage in pod specs.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-advanced-ops.tsv</code></html>
When a Kubernetes node runs low on memory, disk, or process IDs, the kubelet sets NodeCondition warnings — MemoryPressure, DiskPressure, or PIDPressure — and begins evicting pods.
''Diagnosis''
* <html><code>kubectl describe node <node></code></html> — inspect Conditions and their status.
* <html><code>kubectl top nodes</code></html> / <html><code>kubectl top pods -n <ns> --sort-by=memory</code></html> / <html><code>kubectl top pods -A --sort-by=cpu | head -20</code></html> — identify top resource consumers.
* <html><code>journalctl -u kubelet</code></html> on the node — check kubelet logs for pressure events.
* <html><code>df -h</code></html> — inspect disk usage for DiskPressure.
* <html><code>ps aux | wc -l</code></html> — check process count for PIDPressure.
''Eviction order (QoS priority)''
The kubelet evicts pods in ascending QoS order: (1) BestEffort pods (no requests or limits set), (2) Burstable pods exceeding their requests, (3) Burstable pods within their requests, (4) Guaranteed pods (requests == limits). This protects production-critical workloads at the expense of best-effort batch jobs.
''Default eviction thresholds''
* Memory: 100Mi available
* Disk: 10% available filesystem
* Image disk: 15% available
''Cleanup''
Evicted pods remain in Failed state. Remove them with <html><code>kubectl delete pods --field-selector status.phase=Failed</code></html>. Detect eviction by inspecting pod describe output for "The node was low on resource: memory" or by monitoring NodeConditions directly.
''Prevention''
* Set memory requests and limits on all pods; unbounded pods can trigger MemoryPressure.
* Configure ResourceQuotas per namespace to cap total consumption.
* Use LimitRanges to enforce default requests/limits on pods that omit them.
Persistent eviction indicates over-provisioning — the node is running more workloads than its capacity can sustain.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-advanced-ops.tsv</code></html>
* <html><code>training/library/topics/k8s-ops/primer.md</code></html>
* <html><code>training/library/topics/k8s-pods-and-scheduling/street_ops.md</code></html>
//Merged from 3 source atoms.//
''Related atoms''
* [[Pods without memory limits cause node-wide evictions and OOM kills]]
* [[How do you distinguish a container-level OOM from a node-level OOM, and what commands r…]]
Q: A container is OOMKilled but the application memory profiler shows usage well below the limit. Why?
A: The OOM limit applies to the entire cgroup, not just heap. It includes RSS, page cache, tmpfs mounts, and child processes. Also, the JVM or runtime may allocate off-heap memory (NIO buffers, thread stacks). Check kubectl describe pod for the exact Last State OOMKilled exit code 137 and compare against the actual RSS with cat /sys/fs/cgroup/memory/memory.usage_in_bytes inside the container.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-advanced-ops.tsv</code></html>
Q: What is the difference between resource requests and limits, and which one affects scheduling?
A: Requests are guaranteed resources the scheduler uses for bin-packing — a pod is scheduled only if a node has enough allocatable capacity for the request. Limits are the max the container can use; exceeding CPU limits causes throttling, exceeding memory limits causes OOMKill.
Best practice: set requests close to actual usage and limits as a safety ceiling.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-advanced-ops.tsv</code></html>
* <html><code>training/interactive/knowledge/data/cards/k8s-concept-chain.tsv</code></html>
//Merged from 2 source atoms.//
''Related atoms''
* [[What is the difference between resources.requests.memory and resources.limits.memory in…]]
* [[You would like to limit the number of resources being used in your cluster. For example…]]
Q: HPA shows <unknown>/80% for CPU target and won't scale. What is wrong?
A: The HPA cannot read metrics. Common causes:
1) metrics-server is not installed or not running.
2) The target Deployment pods have no CPU requests set — HPA needs requests to compute utilization percentage.
3) metrics-server cannot reach kubelets (firewall or certificate issue). Verify: kubectl top pods (should return data), and kubectl describe hpa <name> for Conditions and events.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-advanced-ops.tsv</code></html>
''Related atoms''
* [[What does HPA do and what problem remains after you enable it?]]
Q: HPA keeps scaling to max replicas even when average CPU is low. What could cause this?
A: 1) One pod is spiking and the average is skewed by replica count.
2) The metric source is wrong (using total CPU instead of per-pod average).
3) Readiness probe failures cause fewer ready pods, inflating per-pod averages.
4) A recent deploy created pods that are initializing and consuming startup CPU. Check kubectl describe hpa and kubectl top pods to correlate.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-advanced-ops.tsv</code></html>
''Related atoms''
* [[What does HPA do and what problem remains after you enable it?]]
* [[HPA requires resource requests to calculate utilization percentage]]
Q: A rolling update is stuck because the PDB minAvailable equals the replica count. Why is this a problem and how do you fix it?
A: If minAvailable equals replicas (e.g., 3/3), the controller cannot evict any old pod to make room for a new one — a deadlock. Fix: set minAvailable to replicas-1 (or use maxUnavailable: 1 instead). This allows the rolling update to terminate one old pod at a time while maintaining minimum availability.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-advanced-ops.tsv</code></html>
Q: A pod is stuck in Init:0/2 status. How do you debug it?
A: kubectl describe pod <pod> shows init container status and events. kubectl logs <pod> -c <init-container-name> shows the init container's output. Init containers run sequentially — 0/2 means the first init container has not completed. Common causes: waiting on a dependency (DNS, service, database), wrong command, or missing ConfigMap/Secret volume.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-advanced-ops.tsv</code></html>
''Related atoms''
* [[Diagnosing a stuck Kubernetes Deployment rollout]]
Q: You deploy an Envoy sidecar with your app container. After a rolling update, requests fail for a few seconds. What is the likely cause and how do you fix it?
A: The app container starts receiving traffic before the Envoy sidecar is ready (race condition). Fix:
1) Use a startup/readiness probe on the sidecar.
2) In Kubernetes 1.28+, use the native sidecar feature (restartPolicy: Always in initContainers) which guarantees sidecar readiness before the main container starts.
3) Alternatively, add a postStart lifecycle hook that polls the sidecar's health endpoint.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-advanced-ops.tsv</code></html>
Q: You create a resource but cannot find it with kubectl get. What namespace-related mistakes should you check?
A: 1) The resource was created in a different namespace — always use -n <ns> or --all-namespaces.
2) Your kubeconfig context has a default namespace set that differs from where the resource lives.
3) The resource is cluster-scoped (e.g., ClusterRole, PV, Node) and does not appear with -n. Check with kubectl api-resources --namespaced=false.
4) RBAC may hide resources you lack permission to list.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-advanced-ops.tsv</code></html>
''Related atoms''
* [[While namespaces do provide scope for resources, they are not isolating them]]
* [[Check how many namespaces are there]]
* [[Create a pod called "kartos" in the namespace dev. The pod should be using the "redis" …]]
A Deployment's ReplicaSet controller continuously reconciles actual state to desired state. When you delete an individual pod, the controller detects the replica count is below <html><code>spec.replicas</code></html> and immediately creates a replacement — the pod reappears. The same applies to manually deleting a ReplicaSet: the Deployment recreates it.
To actually remove a workload, target the Deployment itself:
* <html><code>kubectl delete deploy <name></code></html> — removes the Deployment and its managed ReplicaSets and Pods cleanly.
* <html><code>kubectl scale deployment <name> --replicas=0</code></html> — scales down to zero without deleting the Deployment object, preserving the spec for later re-use.
Add <html><code>--grace-period=0 --force</code></html> for immediate deletion, but only in dev/test environments; in production this can leave orphaned resources or skip graceful shutdown hooks.
The root principle: Kubernetes is declarative. You change desired state (the Deployment spec), and controllers converge actual state to match. Imperative actions on managed objects (pods, ReplicaSets) are overwritten on the next reconciliation loop.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-advanced-ops.tsv</code></html>
* <html><code>training/interactive/knowledge/data/cards/k8s-workloads.tsv</code></html>
//Merged from 2 source atoms.//
''Related atoms''
* [[How to delete a deployment?]]
Q: Pods from the same Deployment keep landing on the same node, causing a single point of failure. How do you spread them across nodes?
A: Use pod topology spread constraints: topologySpreadConstraints with topologyKey: kubernetes.io/hostname, maxSkew: 1, and whenUnsatisfiable: DoNotSchedule. Alternatively, use pod anti-affinity with requiredDuringSchedulingIgnoredDuringExecution matching the app label. Topology spread constraints are more flexible than anti-affinity because they allow fine-grained skew control across zones and nodes.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-advanced-ops.tsv</code></html>
Running a pod directly — not wrapped in a Deployment, StatefulSet, or DaemonSet — means no controller watches its lifecycle. When a bare pod crashes, it stays dead permanently: nothing restarts it, nothing replaces it, and users experience an outage until someone manually re-creates it. The pod is not rescheduled if its node goes down, and image updates must be applied manually.
Always use a Deployment to wrap pods, even for "temporary" or "one-off" workloads. A Deployment guarantees the desired replica count is maintained, automatically replaces crashed pods, and provides a controlled rollout path when the image is updated. StatefulSet and DaemonSet serve the same protective role for stateful and per-node workloads respectively. Running bare pods in production is a footgun: the system looks healthy until the first crash, at which point recovery requires manual intervention.
----
''Sources''
* <html><code>training/library/topics/k8s-concept-chain/footguns.md</code></html>
* <html><code>training/interactive/knowledge/data/cards/k8s-concept-chain.tsv</code></html>
//Merged from 2 source atoms.//
Pod IP addresses are assigned at creation and re-assigned on every restart or reschedule. Hardcoding a pod's IP—discovered via <html><code>kubectl get pod -o wide</code></html>—works only until the pod is evicted, restarted, or rescheduled, at which point it receives a new IP and traffic to the old address breaks. At scale, IPs change constantly and cannot be tracked manually.
The correct abstraction is a Kubernetes Service. A Service exposes a stable ClusterIP and a stable DNS name (e.g., <html><code>my-service.default.svc.cluster.local</code></html>) that resolves to the current set of healthy pod IPs. The Service selects its backing pods via label selectors, not by IP. Pods are ephemeral implementation details; the Service is the stable contract.
Never hardcode a pod IP in any configuration, environment variable, or application config. Always route inter-service traffic through a Service name. This holds whether communication is within the same namespace or across namespaces.
----
''Sources''
* <html><code>training/library/topics/k8s-concept-chain/footguns.md</code></html>
* <html><code>training/interactive/knowledge/data/cards/k8s-concept-chain.tsv</code></html>
//Merged from 2 source atoms.//
''Related atoms''
* [[How does Kubernetes handle DNS resolution for services and pods?]]
* [[ClusterIP: Kubernetes default internal-only Service type]]
Kubernetes Secrets are stored in etcd as base64-encoded plaintext by default. Base64 is a trivially-reversible encoding format, not a security mechanism — <html><code>echo cGFzc3dvcmQ= | base64 -d</code></html> instantly reveals the value. Anyone with direct etcd access can read every secret. Encryption at rest requires explicitly configuring an <html><code>EncryptionConfiguration</code></html> on the API server, which enables AES-CBC or AES-GCM encryption for Secret resources. Managed Kubernetes services (EKS, GKE, AKS) typically enable encryption at rest by default, which is a significant security advantage over self-managed clusters; self-managed deployments must configure it explicitly. This is a high-impact, commonly overlooked vulnerability because base64 encoding creates the illusion of obfuscation. If your threat model includes etcd compromise, encryption at rest is essential. Additional mitigations: restrict Secret access via RBAC, and consider an external secrets operator (e.g., External Secrets Operator, Vault Agent) that encrypts and audits all secret access.
----
''Sources''
* <html><code>training/library/topics/k8s-concept-chain/footguns.md</code></html>
* <html><code>training/interactive/knowledge/data/cards/k8s-concept-chain.tsv</code></html>
* <html><code>training/library/topics/k8s-ops/trivia.md</code></html>
//Merged from 2 source atoms.//
''Related atoms''
* [[True or False? storing data in a Secret component makes it automatically secured]]
* [[Role of kube-apiserver in Kubernetes]]
Pods with no <html><code>resources.limits.memory</code></html> (and no requests) receive QoS class BestEffort — the lowest tier. Under node memory pressure, the Kubernetes eviction manager kills pods in order: BestEffort first, then Burstable, then Guaranteed. If a BestEffort pod has a memory leak, it can consume all available node memory. Once exhausted, the kernel OOM killer fires and selects victims by OOM score — the leaking pod may or may not be the one killed, depending on scoring. Unrelated pods on the same node become collateral damage regardless of their own resource hygiene. A pod without limits has no cgroup ceiling, so no mechanism bounds the leak to a single pod. Even a generous limit is better than none: it creates a cgroup ceiling that contains damage to one pod rather than propagating node-wide. Fixes: (1) Always set both <html><code>resources.requests</code></html> and <html><code>resources.limits</code></html> for memory. (2) Start with requests = limits (Guaranteed QoS) and relax only after measuring real usage. (3) Apply a LimitRange object to enforce default limits namespace-wide. (4) Monitor node conditions with <html><code>kubectl describe node</code></html> and watch for MemoryPressure. The eviction cascade is the primary Kubernetes-layer mechanism; the kernel OOM killer is the fallback and is indiscriminate.
----
''Sources''
* <html><code>training/library/topics/k8s-concept-chain/footguns.md</code></html>
* <html><code>training/interactive/knowledge/data/cards/k8s-concept-chain.tsv</code></html>
* <html><code>training/interactive/knowledge/data/cards/k8s-config.tsv</code></html>
* <html><code>training/library/topics/oomkilled/footguns.md</code></html>
//Merged from 4 source atoms.//
''Related atoms''
* [[Kubernetes node resource pressure: diagnosis, eviction, and prevention]]
* [[How do you prevent high memory usage in your Kubernetes cluster and possibly issues lik…]]
Q: What problem does a Deployment solve that bare Pods cannot?
A: A Deployment ensures a desired number of replicas are always running. If a Pod dies, the Deployment controller creates a replacement automatically.
Remember: Deployment → ReplicaSet → Pods. The Deployment manages ReplicaSets for rolling updates.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-concept-chain.tsv</code></html>
''Related atoms''
* [[What is the relationship between Deployment, ReplicaSet, and Pod?]]
Q: What problem does Ingress solve that Services alone cannot?
A: With Services, each externally-exposed service needs its own LoadBalancer (one cloud LB each, at ~$18-20/month). Ingress provides L7 routing (hostname + path) so one load balancer can serve many services.
Remember: Ingress = one LB, many services, smart routing by host/path.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-concept-chain.tsv</code></html>
''Related atoms''
* [[What are some use cases for using Ingress?]]
* [[Why do Ingress resources do nothing by themselves?]]
Q: Why do Ingress resources do nothing by themselves?
A: Ingress is just a set of routing rules. An Ingress Controller (nginx, Traefik, AWS LB Controller) must be deployed to watch Ingress resources and actually configure traffic routing.
Gotcha: No Ingress Controller deployed = Ingress resources are completely inert. No errors, no warnings — just silence.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-concept-chain.tsv</code></html>
Q: What problem does a ConfigMap solve?
A: Hardcoding config inside the container image means rebuilding for every config change and risking wrong values per environment. A ConfigMap externalizes config so the same image runs in dev, staging, and prod with different settings injected at runtime.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-concept-chain.tsv</code></html>
''Related atoms''
* [[What is a ConfigMap, and how is it used in Kubernetes?]]
* [[How to use ConfigMaps?]]
Q: Why should you use a Secret instead of a ConfigMap for passwords?
A: ConfigMaps have no special access controls — anyone who can read ConfigMaps in the namespace sees the data. Secrets have separate RBAC controls and are meant for sensitive data like passwords, tokens, and TLS certs.
Gotcha: Secrets are base64-encoded, NOT encrypted. Enable etcd encryption at rest for real protection.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-concept-chain.tsv</code></html>
''Related atoms''
* [[True or False? Sensitive data, like credentials, should be stored in a ConfigMap]]
* [[Kubernetes Secrets: storage, access, and best practices]]
Q: What does HPA do and what problem remains after you enable it?
A: HPA (Horizontal Pod Autoscaler) watches metrics (CPU, memory, custom) and adjusts Deployment replica count automatically. But HPA only creates Pods — if nodes are full, new Pods sit in Pending state.
Remember: HPA scales Pods. Karpenter/Cluster Autoscaler scales nodes. You often need both.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-concept-chain.tsv</code></html>
''Related atoms''
* [[HPA keeps scaling to max replicas even when average CPU is low. What could cause this?]]
* [[What is the purpose of Horizontal Pod Autoscaling in Kubernetes?]]
Kubernetes assigns every pod one of three Quality of Service classes based on its container resource fields. The class determines eviction order when a node is under memory pressure.
''Guaranteed'' — every container has <html><code>requests == limits</code></html> for both CPU and memory. Evicted last.
''Burstable'' — at least one container has <html><code>requests < limits</code></html>, or only requests are set with no limits defined. Evicted second.
''BestEffort'' — no container sets any resource requests or limits. Evicted first.
QoS class is assigned automatically by the kubelet; it cannot be set directly in the pod spec. Under node memory pressure, the kubelet terminates BestEffort pods first (they consume resources with no declared bound), then Burstable pods (they may exceed their requested amount), and evicts Guaranteed pods only as a last resort.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-concept-chain.tsv</code></html>
* <html><code>training/interactive/knowledge/data/cards/k8s-troubleshooting.tsv</code></html>
//Merged from 2 source atoms.//
''Related atoms''
* [[What QoS classes are there?]]
* [[Kubernetes memory failure modes: eviction vs OOMKill]]
Q: What is the relationship between Deployment, ReplicaSet, and Pod?
A: Deployment manages ReplicaSets. ReplicaSet manages Pods. During a rolling update, the Deployment creates a new ReplicaSet (with the updated spec) and scales it up while scaling the old ReplicaSet down.
Remember: Deployment → ReplicaSet → Pods. Never edit ReplicaSets directly.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-concept-chain.tsv</code></html>
''Related atoms''
* [[What problem does a Deployment solve that bare Pods cannot?]]
* [[Kubernetes Deployment rolling updates and rollbacks]]
* [[Kubernetes ReplicaSet: maintaining stable pod replica count]]
Q: How does a Kubernetes Service find the right Pods to route traffic to?
A: The Service uses label selectors to match Pods. Matching Pods are added to the Service's Endpoints list. kube-proxy (or the CNI) programs iptables/IPVS rules to route Service IP traffic to endpoint Pod IPs.
Remember: No label match = no endpoints = Service returns connection refused.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-concept-chain.tsv</code></html>
''Related atoms''
* [[Kubernetes Service: stable networking endpoint for pods]]
* [[What is NodePort service type in Kubernetes?]]
Pods using ConfigMap values via environment variable injection receive those values only at creation time; they are never updated in a running pod. Pods using volume mounts will see the updated file after the kubelet sync period (default ~60 s), but the application must explicitly re-read the files to act on the change.
In either case, Kubernetes does not trigger a Pod restart when a ConfigMap changes. To propagate a ConfigMap update reliably, use <html><code>kubectl rollout restart deploy/<name></code></html> or encode the ConfigMap content as a checksum annotation on the Pod template so that a data change forces a rollout.
Troubleshooting a pod that misbehaves after a ConfigMap update: check whether the pod was restarted after the change (env var case) or whether the app re-reads its config files (volume case). Use <html><code>kubectl describe</code></html> to inspect Events, which surface the cause of failures, not just the symptom.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-concept-chain.tsv</code></html>
* <html><code>training/interactive/knowledge/data/cards/k8s-troubleshooting.tsv</code></html>
//Merged from 2 source atoms.//
''Related atoms''
* [[Explain how ConfigMap and Secret updates are handled in Kubernetes.]]
Q: What happens when you set replicas in a Deployment manifest and also use HPA?
A: Every time you <html><code>kubectl apply</code></html> the manifest, it resets the replica count to the hardcoded value, overriding HPA's scaling decisions.
Fix: remove the <html><code>replicas</code></html> field from the Deployment manifest when using HPA, or use server-side apply.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-concept-chain.tsv</code></html>
''Related atoms''
* [[How to scale a deployment to 8 replicas?]]
* [[What does HPA do and what problem remains after you enable it?]]
Q: Why do some teams set CPU requests but no CPU limits?
A: CPU limits cause kernel-level throttling (CFS quota) even when the node has spare CPU capacity. This creates unpredictable latency spikes. Memory limits should always be set (OOM is catastrophic), but CPU throttling is annoying rather than fatal.
Remember: CPU limit = throttle. Memory limit = OOMKill. Different failure modes require different strategies.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-concept-chain.tsv</code></html>
Q: Why is a PodDisruptionBudget important when using node autoscaling?
A: Without a PDB, Karpenter or Cluster Autoscaler can evict all replicas of a service from a node simultaneously during scale-down. A PDB (minAvailable or maxUnavailable) ensures the autoscaler respects application availability during node drains.
Remember: PDB protects against voluntary disruptions — node drains, autoscaler evictions, maintenance.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-concept-chain.tsv</code></html>
''Related atoms''
* [[How can PodDisruptionBudget conflict with HPA scale-down, and what is the best practice…]]
An Ingress resource defines HTTP/HTTPS routing rules (L7 only) but has no effect by itself — traffic does not route until a controller (NGINX, Traefik, HAProxy, etc.) is installed and actively watching for Ingress objects. You can apply an Ingress, see it report success, and still receive 404s from external traffic.
The three most likely causes of a 404 on a correctly-specified Ingress:
# ''No Ingress controller running.'' Ingress objects are inert without one. Verify with <html><code>kubectl get pods -n ingress-nginx</code></html> (or the relevant namespace).
# ''<html><code>ingressClassName</code></html> mismatch.'' The field must match the installed controller's IngressClass; a mismatch means the controller ignores the resource.
# ''Backend Service has no ready endpoints.'' Either the Service name/port in the Ingress spec is wrong, or the Service selector labels do not match the Pod labels. Verify with <html><code>kubectl get endpoints BACKEND_SVC</code></html>.
Debug sequence: <html><code>kubectl get pods -n ingress-nginx</code></html> → <html><code>kubectl describe ingress NAME</code></html> → <html><code>kubectl get endpoints BACKEND_SVC</code></html>.
Note: Ingress handles L7 (HTTP/HTTPS) only. For L4 traffic, use NodePort or LoadBalancer Services.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-concept-chain.tsv</code></html>
* <html><code>training/library/topics/k8s-concept-chain/footguns.md</code></html>
* <html><code>training/library/topics/k8s-networking/footguns.md</code></html>
* <html><code>training/interactive/knowledge/data/cards/k8s-networking.tsv</code></html>
//Merged from 3 source atoms.//
Pending means the scheduler examined every node and could not place the pod. The scheduler is never silent about why: it always appends an event to the pod explaining which filter rejected it. Always start with <html><code>kubectl describe pod <name></code></html> and read the Events section — it is the diagnosis, not a hint.
Top causes and their fixes:
* ''Insufficient CPU or memory'': reduce resource requests or add/scale nodes; check <html><code>kubectl describe nodes</code></html> for Allocated vs Allocatable.
* ''Taints with no matching tolerations'': add tolerations to the pod spec or remove the taint from the node.
* ''NodeSelector or affinity rules matching zero nodes'': fix node labels or relax affinity rules.
* ''PersistentVolumeClaim not yet bound'' (e.g., PV in a different availability zone, wrong StorageClass): fix the StorageClass or binding, or check that the PVC is not stuck.
* ''ResourceQuota limits reached'': inspect quota with <html><code>kubectl describe resourcequota</code></html> and adjust or raise limits.
* ''Scheduler not running'': verify with <html><code>kubectl get pods -n kube-system | grep scheduler</code></html>.
* ''Node autoscaler (Karpenter/CA) not provisioning'': check autoscaler logs and capacity limits.
If <html><code>kubectl get pods -o wide</code></html> shows no node assigned, the scheduler itself may be the failure point. Pending is always a cluster or scheduling configuration problem, never a pod-internal one.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-concept-chain.tsv</code></html>
* <html><code>training/interactive/knowledge/data/cards/k8s-troubleshooting.tsv</code></html>
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
* <html><code>training/library/topics/k8s-debugging-playbook/trivia.md</code></html>
* <html><code>training/interactive/knowledge/data/cards/k8s-ops.tsv</code></html>
* <html><code>training/library/topics/k8s-pods-and-scheduling/street_ops.md</code></html>
//Merged from 5 source atoms.//
''Related atoms''
* [[Diagnosing pods that won't schedule on a node due to taints]]
Q: What is the Kubernetes concept chain?
A: Each K8s abstraction exists because the previous layer has an unresolved problem:
Pod crashes → Deployment
IPs change → Service
Too many LBs → Ingress
Rules need engine → Ingress Controller
Config in image → ConfigMap
Passwords exposed → Secret
Manual scaling → HPA
Nodes full → Karpenter
Rogue resources → Requests & Limits
Remember: Every K8s concept is a solution to a specific problem. Learn the problems, and the solutions make sense.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-concept-chain.tsv</code></html>
''Related atoms''
* [[Kubernetes: definition, origin, and core features]]
* [[What problems does Kubernetes actually solve?]]
* [[What challenges do you anticipate when managing large-scale Kubernetes clusters, and ho…]]
Q: What is Resource Quota?
A: Resource quota provides constraints that limit aggregate resource consumption per namespace. It can limit the quantity of objects that can be created in a namespace by type, as well as the total amount of compute resources that may be consumed by resources in that namespace.
Example: <html><code>kubectl create quota my-quota --hard=pods=10,requests.cpu=4,requests.memory=8Gi -n dev</code></html> limits namespace resources.
Remember: Quotas are per-namespace. Think "Quota = Namespace Budget" — each team gets a spending limit.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-config.tsv</code></html>
''Related atoms''
* [[Explain why one would specify resource limits in regards to Pods]]
* [[What is the difference between resources.requests.memory and resources.limits.memory in…]]
Q: You would like to limit the number of resources being used in your cluster. For example no more than 4 replicasets, 2 services, etc. How would you achieve that?
A: Use ResourceQuotas to limit the total number of objects and resource consumption per namespace.
Example: <html><code>kubectl create quota my-quota --hard=replicasets=4,services=2,pods=10</code></html> in a namespace. ResourceQuotas also limit total CPU/memory: <html><code>requests.cpu=4,limits.memory=8Gi</code></html>.
Gotcha: without LimitRange, users can still create pods without resource requests, so combine both.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-config.tsv</code></html>
* <html><code>training/interactive/knowledge/data/cards/k8s-workloads.tsv</code></html>
//Merged from 2 source atoms.//
''Related atoms''
* [[Check if there are any limits on one of the pods in your cluster]]
* [[You have one Kubernetes cluster and multiple teams that would like to use it. You would…]]
Q: True or False? storing data in a Secret component makes it automatically secured
A: False. Some known security mechanisms like "encryption" aren't enabled by default.
Gotcha: K8s Secrets are base64-encoded, not encrypted at rest by default. Enable EncryptionConfiguration for etcd encryption.
Remember: "Base64 is a disguise, not a safe" — anyone with etcd access can decode without encryption at rest.
See also: External secret managers (Vault, AWS Secrets Manager) provide true encryption and rotation.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-config.tsv</code></html>
''Related atoms''
* [[Kubernetes Secrets are base64-encoded, not encrypted, in etcd by default]]
* [[True or False? Sensitive data, like credentials, should be stored in a ConfigMap]]
* [[Kubernetes Secrets: storage, access, and best practices]]
Q: What is the problem with the following Secret file:
A: Password isn't encrypted.
You should run something like this: <html><code>echo -n 'mySecretPassword' | base64</code></html> and paste the result to the file instead of using plain-text.
Remember: <html><code>echo -n</code></html> is critical — without <html><code>-n</code></html>, a trailing newline gets base64-encoded, causing auth failures.
Gotcha: <html><code>echo -n 'myPassword' | base64</code></html> vs <html><code>echo 'myPassword' | base64</code></html> produce different results. The newline matters!
Example: Compare: <html><code>echo -n 'pass' | base64</code></html> → <html><code>cGFzcw==</code></html> vs <html><code>echo 'pass' | base64</code></html> → <html><code>cGFzcwo=</code></html> (extra newline).
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-config.tsv</code></html>
Q: True or False? Memory is a compressible resource, meaning that when a container reach the memory limit, it will keep running
A: False. CPU is a compressible resource while memory is a non compressible resource - once a container reached the memory limit, it will be terminated.
Remember: CPU is compressible (throttled when over-limit), memory is NOT (OOMKilled). This distinction is critical.
Under the hood: Kernel CFS scheduler throttles CPU. OOM killer terminates memory hogs. Different enforcement mechanisms.
Gotcha: Over-provision memory to be safe (OOM=crash). CPU can be slightly under-provisioned (just slower).
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-config.tsv</code></html>
Kubernetes Secrets store sensitive information — passwords, SSH keys, API keys, certificates — as base64-encoded data in etcd, the cluster's distributed key-value store. Base64 is encoding, not encryption; etcd stores Secrets unencrypted by default. Enable encryption at rest for production clusters.
Secrets are structurally identical to ConfigMaps but carry slightly more restricted access semantics. Pods consume Secrets either by mounting them as volumes or injecting them as environment variables.
Access is governed by RBAC (Role-Based Access Control); only authorized users or workloads should be granted <html><code>get</code></html>/<html><code>list</code></html> permissions on Secret resources.
Creation: <html><code>kubectl create secret generic <name></code></html> accepts <html><code>--from-literal=key=value</code></html>, <html><code>--from-file=key=path</code></html>, and <html><code>--from-env-file=path</code></html>. Special characters in literal values require shell quoting — use single quotes: <html><code>--from-literal=pass='p@ss!'</code></html>.
Best practices:
* Enable etcd encryption at rest.
* Use RBAC to limit Secret access to the minimum necessary.
* Never bake secrets into container image layers.
* Rotate secrets regularly.
* Prefer dedicated secret-management backends (e.g., Vault, AWS Secrets Manager) with the Secrets Store CSI driver for production workloads that need stronger guarantees than native Kubernetes Secrets provide.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-config.tsv</code></html>
//Merged from 3 source atoms.//
''Related atoms''
* [[Explain how ConfigMap and Secret updates are handled in Kubernetes.]]
Q: True or False? Resource limits applied on a Pod level meaning, if limits is 2gb RAM and there are two container in a Pod that it's 1gb RAM each
A: False. It's per container and not per Pod.
Remember: K8s true/false questions test edge cases and defaults. Verify with <html><code>kubectl explain <resource></code></html>.
Gotcha: Use <html><code>kubectl explain <resource>.spec</code></html> to check field behavior directly from the CLI.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-config.tsv</code></html>
''Related atoms''
* [[What is the difference between resources.requests.memory and resources.limits.memory in…]]
* [[Check if there are any limits on one of the pods in your cluster]]
* [[A pod has a main container limited to 512Mi and an Istio sidecar limited to 256Mi. The …]]
Q: Explain why one would specify resource limits in regards to Pods
A: * You know how much RAM and/or CPU your app should be consuming and anything above that is not valid
* You would like to make sure that everyone can run their apps in the cluster and resources are not being solely used by one type of application
Remember: Use <html><code>kubectl explain <resource></code></html> to explore API fields interactively from the CLI.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-config.tsv</code></html>
''Related atoms''
* [[What is the difference between resources.requests.memory and resources.limits.memory in…]]
* [[What is Resource Quota?]]
Q: Run a pod called "yay2" with the image "python". Make sure it has resources request of 64Mi memory and 250m CPU and the limits are 128Mi memory and 500m CPU
A: <html><code>kubectl run yay2 --image=python --dry-run=client -o yaml > pod.yaml</code></html>
<html><code>vi pod.yaml</code></html>
<html><pre><code class="language-plaintext">spec:
containers:
- image: python
imagePullPolicy: Always
name: yay2
resources:
limits:
cpu: 500m
memory: 128Mi
requests:
cpu: 250m
memory: 64Mi</code></pre></html>
<html><code>kubectl apply -f pod.yaml</code></html>
Remember: Use <html><code>kubectl explain <resource></code></html> to explore API fields interactively from the CLI.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-config.tsv</code></html>
//Merged from 2 source atoms.//
''Related atoms''
* [[Check if there are any limits on one of the pods in your cluster]]
* [[Create a static pod with the image python that runs the command sleep 2017]]
Q: What QoS classes are there?
A: * Guaranteed
* Burstable
* BestEffort
Remember: QoS eviction order: BestEffort first, Burstable second, Guaranteed last. Mnemonic: "BBG."
Under the hood: Guaranteed=requests==limits for ALL containers. Burstable=at least one request. BestEffort=zero.
Gotcha: Miss one container's limits and the pod drops from Guaranteed to Burstable.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-config.tsv</code></html>
''Related atoms''
* [[Kubernetes QoS classes: criteria and eviction order]]
Q: How to commit secrets to Git and in general how to use encrypted secrets?
A: One possible process would be as follows:
# You create a Kubernetes secret (but don't commit it)
# You encrypt it using some 3rd party project (.e.g kubeseal)
# You apply the sealed/encrypted secret
# You commit the sealed secret to Git
# You deploy an application that requires the secret and it can be automatically decrypted by using for example a Bitnami Sealed secrets controller
Remember: <html><code>kubectl create secret generic</code></html> supports <html><code>--from-literal</code></html>, <html><code>--from-file</code></html>, <html><code>--from-env-file</code></html>.
Gotcha: Special chars need shell quoting. Use single quotes: <html><code>--from-literal=pass='p@ss!'</code></html>.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-config.tsv</code></html>
''Related atoms''
* [[Kubernetes Secrets: storage, access, and best practices]]
Q: How to create a Secret from a key and value?
A: <html><code>kubectl create secret generic some-secret --from-literal=password='donttellmypassword'</code></html>
Remember: <html><code>kubectl create secret generic</code></html> supports <html><code>--from-literal</code></html>, <html><code>--from-file</code></html>, <html><code>--from-env-file</code></html>.
Gotcha: Special chars need shell quoting. Use single quotes: <html><code>--from-literal=pass='p@ss!'</code></html>.
Example: <html><code>kubectl create secret generic ssh-key --from-file=ssh-privatekey=~/.ssh/id_rsa</code></html>
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-config.tsv</code></html>
//Merged from 2 source atoms.//
''Related atoms''
* [[How to commit secrets to Git and in general how to use encrypted secrets?]]
* [[Kubernetes Secrets: storage, access, and best practices]]
Q: What type: Opaque in a secret file means? What other types are there?
A: Opaque is the default type used for key-value pairs.
Remember: <html><code>kubectl create secret generic</code></html> supports <html><code>--from-literal</code></html>, <html><code>--from-file</code></html>, <html><code>--from-env-file</code></html>.
Gotcha: Special chars need shell quoting. Use single quotes: <html><code>--from-literal=pass='p@ss!'</code></html>.
Example: <html><code>kubectl create secret generic ssh-key --from-file=ssh-privatekey=~/.ssh/id_rsa</code></html>
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-config.tsv</code></html>
''Related atoms''
* [[How to commit secrets to Git and in general how to use encrypted secrets?]]
* [[Kubernetes Secrets: storage, access, and best practices]]
Q: What is a Kubernetes Secret and how does it store sensitive data?
A: An object for storing sensitive data like passwords or tokens.
Remember: <html><code>kubectl create secret generic</code></html> supports <html><code>--from-literal</code></html>, <html><code>--from-file</code></html>, <html><code>--from-env-file</code></html>.
Gotcha: Special chars need shell quoting. Use single quotes: <html><code>--from-literal=pass='p@ss!'</code></html>.
Example: <html><code>kubectl create secret generic ssh-key --from-file=ssh-privatekey=~/.ssh/id_rsa</code></html>
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-config.tsv</code></html>
''Related atoms''
* [[How to commit secrets to Git and in general how to use encrypted secrets?]]
Q: What is a ConfigMap, and how is it used in Kubernetes?
A: ''ConfigMap:''
* Kubernetes resource that stores configuration data in key-value pairs.
* Decouples configuration from application code.
* Can be used to store configuration files, command-line arguments, environment variables, etc.
* ConfigMaps allow for the separation of configuration from application logic, making it easier to manage and update configurations without modifying the application code.Applications can reference ConfigMaps, and changes to the ConfigMap are automatically reflected in the pods that reference it.
Example: <html><code>kubectl create configmap app-cfg --from-file=config.yaml --from-literal=LOG_LEVEL=debug</code></html>
Gotcha: ConfigMap updates don't auto-restart pods. Use Reloader or hash annotations for rolling updates.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-config.tsv</code></html>
//Merged from 4 source atoms.//
''Related atoms''
* [[Explain how you would manage configuration drift in a Kubernetes environment.]]
Q: Explain the concept of PodDisruptionBudget in Kubernetes.
A: * PodDisruptionBudget: PodDisruptionBudget is a resource in Kubernetes that defines policies for pod disruptions during voluntary disruptions (e.g., rolling updates).
* It limits the number of concurrently disrupted pods and ensures that a minimum number of replicas are available during disruptions.
* Helps prevent service disruption and ensures stability during maintenance activities.
* PodDisruptionBudgets are useful for controlling the impact of disruptions, reducing the risk of service degradation during planned maintenance or updates.
* They provide a balance between maintaining high availability and executing necessary maintenance tasks.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-config.tsv</code></html>
Q: True or False? Sensitive data, like credentials, should be stored in a ConfigMap
A: False. Use secret.
Remember: Sensitive data → Secrets, not ConfigMaps. ConfigMaps are plain text, visible to namespace users.
Gotcha: Even Secrets are only base64-encoded by default. Enable encryption at rest + RBAC.
See also: External Secrets Operator syncs from Vault/AWS into K8s Secrets automatically.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-config.tsv</code></html>
''Related atoms''
* [[True or False? storing data in a Secret component makes it automatically secured]]
* [[Why should you use a Secret instead of a ConfigMap for passwords?]]
Q: How to use ConfigMaps?
A: 1. Create it (from key&value, a file or an env file)
# Attach it. Mount a configmap as a volume
Remember: <html><code>--dry-run=client -o yaml</code></html> generates templates. Pipe to a file and customize.
Gotcha: <html><code>create</code></html> = imperative (fails if exists). <html><code>apply</code></html> = declarative (creates or updates). Production = <html><code>apply</code></html>.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-config.tsv</code></html>
''Related atoms''
* [[What is a ConfigMap, and how is it used in Kubernetes?]]
* [[What problem does a ConfigMap solve?]]
* [[How to create components in a namespace?]]
Q: Explain how ConfigMap and Secret updates are handled in Kubernetes.
A: * ConfigMap and Secret Updates: Changes to ConfigMaps or Secrets trigger updates in associated pods automatically.
* Pods referencing ConfigMaps or Secrets receive notifications about updates.
* Containers in the pod can watch for changes and adapt their configurations dynamically.
* ConfigMap and Secret updates are dynamically propagated to pods using them.
* Containers within pods can watch for changes and reconfigure themselves accordingly, ensuring that any modifications to configuration data are seamlessly applied.
Example: <html><code>kubectl create secret generic db-creds --from-literal=user=admin --from-literal=pass=s3cret</code></html>
Remember: Secrets are like ConfigMaps wearing sunglasses — same structure, base64-encoded, slightly more restricted access.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-config.tsv</code></html>
''Related atoms''
* [[ConfigMap updates do not automatically restart Pods]]
* [[Kubernetes Secrets: storage, access, and best practices]]
Q: What the master node is responsible for?
A: The master coordinates all the workflows in the cluster:
* Scheduling applications
* Managing desired state
* Rolling out new updates
Remember: Use <html><code>kubectl explain <resource></code></html> to explore API fields from the CLI.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
''Related atoms''
* [[Explain the main components of Kubernetes architecture.]]
* [[What is a node in Kubernetes?]]
* [[What is kubectl and how is it used to manage Kubernetes clusters?]]
Q: Explain the working of the master node in Kubernetes?
A: The master node dignifies the node that controls and manages the set of worker nodes. This kind resembles a cluster in Kubernetes. The nodes are responsible for the cluster management and the API used to configure and manage the resources within the collection. The master nodes of Kubernetes can run with Kubernetes itself, the asset of dedicated pods.
Remember: Use <html><code>kubectl explain <resource></code></html> to explore API fields from the CLI.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
''Related atoms''
* [[What is kubectl and how is it used to manage Kubernetes clusters?]]
* [[What does the node status contain?]]
* [[Explain the purpose of kubelet in the Kubernetes cluster.]]
Q: What process runs on Kubernetes Master Node?
A: The Kube-api server process runs on the master node and serves to scale the deployment of more instances.
Remember: Use <html><code>kubectl explain <resource></code></html> to explore API fields from the CLI.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
''Related atoms''
* [[Explain the main components of Kubernetes architecture.]]
* [[Kubernetes Control Plane: components and responsibilities]]
* [[What is a node in Kubernetes?]]
Q: What is the role of kube-apiserver aggregation layer in Kubernetes?
A: * kube-apiserver Aggregation Layer: The aggregation layer extends the kube-apiserver to support custom APIs and resources.
* It allows third-party APIs and controllers to be seamlessly integrated into the Kubernetes API.
* Facilitates the development and deployment of custom extensions by different organizations or vendors.
* The kube-apiserver aggregation layer enables the Kubernetes API to be extensible, supporting custom resources and APIs beyond the core Kubernetes objects.
* This extensibility is essential for accommodating a wide range of use cases and domain-specific requirements.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
''Related atoms''
* [[Discuss the role of kube-proxy in Kubernetes networking.]]
* [[Role of kube-apiserver in Kubernetes]]
* [[Explain the main components of Kubernetes architecture.]]
Q: What are the possible Pod phases?
A: * Running - The Pod bound to a node and at least one container is running
* Failed/Error - At least one container in the Pod terminated with a failure
* Succeeded - Every container in the Pod terminated with success
* Unknown - Pod's state could not be obtained
* Pending - Containers are not yet running (Perhaps images are still being downloaded or the pod wasn't scheduled yet)
Remember: Use <html><code>kubectl explain <resource></code></html> to explore API fields from the CLI.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
Controllers are control loops that continuously watch the state of a Kubernetes cluster and make or request changes to move the current state toward the desired state. This pattern is called reconciliation: compare desired state vs. actual state, then act on any delta.
The <html><code>kube-controller-manager</code></html> is the control-plane component that hosts and runs these controller processes. Each controller manages a specific aspect of cluster state and operates independently within the same binary.
Key built-in controllers:
* ''Node Controller'' – monitors node health. If a node becomes unreachable, it evacuates all pods running on it and updates the node status accordingly.
* ''Replication Controller'' – monitors pod replica counts. If the running count differs from the desired count, it creates or deletes pods to reconcile.
* ''Endpoints Controller'' – populates Endpoints objects as Services and Pods change.
Reconciliation example: a ReplicaSet specifies 3 replicas, but only 2 are running → the ReplicaSet controller creates 1 additional pod.
Explore API fields for any controller-managed resource with <html><code>kubectl explain <resource></code></html>.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
//Merged from 3 source atoms.//
''Related atoms''
* [[Role of kube-apiserver in Kubernetes]]
* [[Kubernetes ReplicaSet: maintaining stable pod replica count]]
Q: Which components can't be created within a namespace?
A: Volume and Node.
Example: <html><code>kubectl create ns staging && kubectl config set-context --current --namespace=staging</code></html>
Remember: Namespaces = logical isolation. Combine with ResourceQuotas, NetworkPolicies, RBAC.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
''Related atoms''
* [[What special namespaces are there by default when creating a Kubernetes cluster?]]
* [[How to get the name of the current namespace?]]
* [[Create a pod called "kartos" in the namespace dev. The pod should be using the "redis" …]]
Q: How to create components in a namespace?
A: One way is by specifying --namespace like this: <html><code>kubectl apply -f my_component.yaml --namespace=some-namespace</code></html>
Another way is by specifying it in the YAML itself:
<html><pre><code class="language-plaintext">apiVersion: v1
kind: ConfigMap
metadata:
name: some-configmap
namespace: some-namespace</code></pre></html>
and you can verify with: <html><code>kubectl get configmap -n some-namespace</code></html>
Example: <html><code>kubectl create ns staging && kubectl config set-context --current --namespace=staging</code></html>
Remember: Namespaces = logical isolation. Combine with ResourceQuotas, NetworkPolicies, RBAC.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
''Related atoms''
* [[You are looking for a Pod called "atreus". How to check in which namespace it runs?]]
* [[Create a pod called "kartos" in the namespace dev. The pod should be using the "redis" …]]
Q: Discuss the relationship between Kubernetes and container runtimes like Docker and containerd.
A: * Kubernetes and Container Runtimes: Kubernetes interacts with container runtimes to deploy and manage containers.
* Common container runtimes include Docker, containerd, and others.
* Kubernetes abstracts the container runtime, allowing flexibility in choosing the underlying technology.
Remember: Use <html><code>kubectl explain <resource></code></html> to explore API fields from the CLI.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
''Related atoms''
* [[Explain the main components of Kubernetes architecture.]]
Q: What happens to running pods if if you stop Kubelet on the worker nodes?
A: When you stop the kubelet service on a worker node, it will no longer be able to communicate with the Kubernetes API server. As a result, the node will be marked as NotReady and the pods running on that node will be marked as Unknown. The Kubernetes control plane will then attempt to reschedule the pods to other available nodes in the cluster.
Remember: kubelet = node agent. Takes PodSpecs, ensures containers are running and healthy.
Under the hood: kubelet talks to container runtime (containerd, CRI-O) via CRI interface.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
''Related atoms''
* [[True or False? Once a Pod is assisgned to a worker node, it will only run on that node,…]]
Q: Explain the purpose of kubelet in the Kubernetes cluster.
A: The kubelet is a node agent that runs on each node in a Kubernetes cluster and is responsible for ensuring that containers are running in a Pod as expected. It interacts with the container runtime (e.g., Docker) to manage the containers and communicates with the master node to receive pod specifications and report the status of pods.
''Key responsibilities of the kubelet include:''
* Pod Lifecycle Management: Ensures that the containers within a pod are running and healthy.
Remember: kubelet = node agent. Takes PodSpecs, ensures containers are running and healthy.
Under the hood: kubelet talks to container runtime (containerd, CRI-O) via CRI interface.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
//Merged from 4 source atoms.//
''Related atoms''
* [[Kubernetes Control Plane: components and responsibilities]]
* [[Kubernetes Static Pods: definition, use cases, and DaemonSet contrast]]
* [[What is a Kubernetes Cluster?]]
Q: Describe how would you delete a static Pod
A: Locate the static Pods directory (look at <html><code>staticPodPath</code></html> in kubelet configuration file).
Go to that directory and remove the manifest/definition of the staic Pod (<html><code>rm <STATIC_POD_PATH>/<POD_DEFINITION_FILE></code></html>)
Remember: <html><code>--dry-run=client -o yaml</code></html> generates manifests. <html><code>kubectl explain</code></html> shows field schemas.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
''Related atoms''
* [[Where static Pods manifests are located?]]
* [[How to identify which Pods are Static Pods?]]
* [[Create a static pod with the image python that runs the command sleep 2017]]
Q: How to get the name of the current namespace?
A: <html><code>kubectl config view | grep namespace</code></html>
Example: <html><code>kubectl create ns staging && kubectl config set-context --current --namespace=staging</code></html>
Remember: Namespaces = logical isolation. Combine with ResourceQuotas, NetworkPolicies, RBAC.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
''Related atoms''
* [[While namespaces do provide scope for resources, they are not isolating them]]
* [[How to get list of resources which are not bound to a specific namespace?]]
* [[Which service and in which namespace the following file is referencing?]]
Q: How to switch to another namespace? In other words how to change active namespace?
A: <html><code>kubectl config set-context --current --namespace=some-namespace</code></html> and validate with <html><code>kubectl config view --minify | grep namespace:</code></html>
OR
<html><code>kubens some-namespace</code></html>
Example: <html><code>kubectl create ns staging && kubectl config set-context --current --namespace=staging</code></html>
Remember: Namespaces = logical isolation. Combine with ResourceQuotas, NetworkPolicies, RBAC.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
''Related atoms''
* [[You are looking for a Pod called "atreus". How to check in which namespace it runs?]]
* [[Create a pod called "kartos" in the namespace dev. The pod should be using the "redis" …]]
* [[You are managing multiple Kubernetes clusters. How do you quickly change between the cl…]]
Q: How view all the pods running in all the namespaces?
A: <html><code>kubectl get pods --all-namespaces</code></html>
Example: <html><code>kubectl create ns staging && kubectl config set-context --current --namespace=staging</code></html>
Remember: Namespaces = logical isolation. Combine with ResourceQuotas, NetworkPolicies, RBAC.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
''Related atoms''
* [[Which command lists all Pods in a Kubernetes cluster?]]
* [[How to create components in a namespace?]]
* [[Create a pod called "kartos" in the namespace dev. The pod should be using the "redis" …]]
Q: Assuming you have multiple schedulers, how to know which scheduler was used for a given Pod?
A: Running <html><code>kubectl get events</code></html> you can see which scheduler was used.
Under the hood: Scheduler: filter (which nodes CAN?) then score (which is BEST?). "Filter then Score."
Remember: Scheduler watches for unassigned pods (no nodeName). It does NOT move running pods.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
''Related atoms''
* [[True or False? Using the node affinity type "preferredDuringSchedulingIgnoredDuringExec…]]
* [[True or False? The "Pending" phase means the Pod was not yet accepted by the Kubernetes…]]
* [[What is the job of the kube-scheduler?]]
Q: How to check to which worker node the pods were scheduled to? In other words, how to check on which node a certain Pod is running?
A: <html><code>kubectl get pods -o wide</code></html> shows pod details including the NODE column indicating which worker node each pod runs on. You can also use <html><code>kubectl describe pod <name></code></html> and look for the 'Node:' field. For filtering: <html><code>kubectl get pods --field-selector spec.nodeName=worker-01</code></html> lists pods on a specific node.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
''Related atoms''
* [[Which command lists all Pods in a Kubernetes cluster?]]
* [[How can you find out information on a Service related to a certain Pod if all you can u…]]
* [[How to schedule a pod on a node called "node1"?]]
Q: How can you find out information on a Service related to a certain Pod if all you can use is kubectl exec --
A: You can run <html><code>kubectl exec <POD_NAME> -- env</code></html> which will give you a couple environment variables related to the Service.
Variables such as <html><code>[SERVICE_NAME]_SERVICE_HOST</code></html>, <html><code>[SERVICE_NAME]_SERVICE_PORT</code></html>, ...
Example: <html><code>kubectl exec -it pod -- /bin/sh</code></html>. Multi-container: <html><code>--container=name</code></html>.
Gotcha: exec needs a running container. For crashed pods: <html><code>kubectl logs --previous</code></html>.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
''Related atoms''
* [[Debugging a failing or non-starting pod in Kubernetes]]
* [[kubectl logs: view container logs in a Pod]]
* [[How to get information on a certain service?]]
Multiple schedulers can run concurrently in a Kubernetes cluster. A custom scheduler is deployed as a Pod (or Deployment) invoking <html><code>kube-scheduler</code></html> with a distinct <html><code>--scheduler-name</code></html> flag and, when HA is needed, <html><code>--leader-elect=true</code></html>:
<html><pre><code class="language-yaml">spec:
containers:
- command:
- kube-scheduler
- --address=127.0.0.1
- --leader-elect=true
- --scheduler-name=some-custom-scheduler</code></pre></html>
To direct a specific Pod to that scheduler, set <html><code>schedulerName</code></html> in the Pod spec:
<html><pre><code class="language-yaml">spec:
schedulerName: some-custom-scheduler</code></pre></html>
Pods that omit <html><code>schedulerName</code></html> continue to use the default scheduler.
''How scheduling works (applies to all schedulers):'' Filter phase — determine which nodes //can// run the Pod (resource fit, taints, affinity). Score phase — rank eligible nodes to find the //best// one. "Filter then Score."
''Key constraint:'' The scheduler only acts on Pods with no <html><code>nodeName</code></html> set. It does not evict or relocate already-running Pods.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
//Merged from 2 source atoms.//
''Related atoms''
* [[How to schedule a pod on a node called "node1"?]]
* [[An engineer form your organization asked whether there is a way to prevent from Pods (w…]]
* [[Kubernetes DaemonSet: one pod per node]]
Q: How many containers can a pod contain?
A: A pod can include multiple containers but in most cases it would probably be one container per pod.
There are some patterns where it makes to run more than one container like the "side-car" pattern where you might want to perform logging or some other operation that is executed by another container running with your app container in the same Pod.
Remember: <html><code>--dry-run=client -o yaml</code></html> generates manifests. <html><code>kubectl explain</code></html> shows field schemas.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
''Related atoms''
* [[Why it's common to have only one container per Pod in most cases?]]
* [[What are your thoughts on "Pods are not meant to be created directly"?]]
* [[What kubectl describe pod [pod name] does? command does?]]
The kube-apiserver is the central component of the Kubernetes control plane and serves as the front door to the cluster. It exposes the Kubernetes API and acts as the single endpoint through which all components — kubectl, kubelet, scheduler, and controllers — communicate with the cluster.
Core responsibilities:
* Validates and processes RESTful API requests for all API objects (pods, services, replication controllers, etc.).
* Enforces authentication, authorization, and admission control before persisting state.
* Reads and writes cluster state to etcd, the distributed key-value backing store.
* Ensures consistency between the desired and actual state of the system.
Request flow through the API server follows the sequence: Authenticate → Authorize → Admit → Validate → etcd (mnemonic: AAAVE).
Because every control-plane interaction passes through it, the kube-apiserver is the critical coordination point for the entire Kubernetes architecture.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
//Merged from 2 source atoms.//
''Related atoms''
* [[Kubernetes controllers and kube-controller-manager]]
Q: Explain the main components of Kubernetes architecture.
A: The main components of Kubernetes architecture include:
''Master Node:''
* kube-apiserver: Serves as the API server for Kubernetes, handling communication with the cluster.
* etcd: Consistent and highly available key-value store for storing configuration data.
* kube-scheduler: Assigns work (pods) to nodes based on resource availability and constraints.
''Node (Minion) Nodes:''
* kubelet: Acts as the node agent, ensuring that containers are running on the node as expected.
Remember: Use <html><code>kubectl explain <resource></code></html> to explore API fields from the CLI.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
//Merged from 2 source atoms.//
''Related atoms''
* [[What is a Kubernetes Cluster?]]
* [[What process runs on Kubernetes Master Node?]]
* [[What the master node is responsible for?]]
Q: What happens when you run a Pod with kubectl?
A: 1. Kubectl sends a request to the API server (kube-apiserver) to create the Pod
## In the process the user gets authenticated and the request is being validated.
## etcd is being updated with the data
# The Scheduler detects that there is an unassigned Pod by monitoring the API server (kube-apiserver)
# The Scheduler chooses a node to assign the Pod to
## etcd is being updated with the information
# The Scheduler updates the API server about which node it chose
5.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
''Related atoms''
* [[True or False? Once a Pod is assisgned to a worker node, it will only run on that node,…]]
Annotations attach arbitrary non-identifying metadata to Kubernetes objects; clients such as tools and libraries can retrieve this metadata.
Key distinction from labels: labels identify and select objects — they can be used in selectors and queries. Annotations cannot be used to identify or select objects. They exist solely to carry supplemental metadata.
Annotation values may be small or large, structured or unstructured, and may include characters not permitted by label values.
Example:
<html><pre><code class="language-plaintext">kubectl annotate pod mypod description='Web server' build='v1.2.3'</code></pre></html>
Use annotations for data that tools, libraries, or humans need to read — build metadata, checksums, pointers to external systems, rollout context — where selection is not required. Use labels when you need grouping, filtering, or selector-based routing (e.g., Services, ReplicaSets).
Source: [[kubernetes.io — Annotations|https://kubernetes.io/docs/concepts/overview/working-with-objects/annotations/]]
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
//Merged from 2 source atoms.//
''Related atoms''
* [[What are Kubernetes selectors and how do they match resources?]]
* [[Kubernetes Service empty endpoints: label-selector mismatch diagnosis]]
kube-scheduler is the Kubernetes control plane component responsible for assigning pods to nodes. Its process has three stages:
# ''What to schedule?'' It parses the pod specification to understand resource requests, constraints, and metadata.
# ''Which node to schedule on?'' It evaluates candidate nodes against resource availability, affinity, anti-affinity rules, taints/tolerations, and other user-defined constraints to select the best fit.
# ''Bind.'' It writes the binding decision, associating the pod with the chosen node.
By considering these factors across all pending pods and available nodes, kube-scheduler maintains balance and efficiency in workload distribution. It ensures pods land only on nodes that satisfy their requirements, directly contributing to cluster health, performance, and resource utilization.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
//Merged from 2 source atoms.//
''Related atoms''
* [[Discuss the implications of pod sprawl and how to manage it effectively in Kubernetes.]]
Static Pods are managed directly by the kubelet daemon on a specific node, without the API server observing them. Unlike Pods managed by the control plane (e.g., via a Deployment), the kubelet itself watches each Static Pod and restarts it if it fails.
''Primary use case — Control Plane components'': kube-apiserver, kube-scheduler, kube-controller-manager, and etcd are run as Static Pods. They must operate on specific nodes and must remain available regardless of the state of other cluster components, making kubelet-direct management the correct mechanism.
''Contrast with DaemonSets'': kube-proxy is NOT a Static Pod — it runs as a DaemonSet. A DaemonSet schedules a Pod on every node in the cluster and is managed by the API server; a Static Pod is node-specific and kubelet-managed. Because kube-proxy has no requirement to run on one particular node — only the requirement to run on all nodes — it does not qualify as a Static Pod.
Tip: Use <html><code>kubectl explain <resource></code></html> to explore API fields from the CLI.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
//Merged from 3 source atoms.//
''Related atoms''
* [[Where static Pods manifests are located?]]
* [[Kubernetes DaemonSet: one pod per node]]
Q: What special namespaces are there by default when creating a Kubernetes cluster?
A: * default
* kube-system
* kube-public
* kube-node-lease
Remember: 4 default namespaces: default, kube-system, kube-public, kube-node-lease. Mnemonic: "DSPL."
Gotcha: Never run production workloads in <html><code>default</code></html> — no quotas, harder RBAC.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
''Related atoms''
* [[Which components can't be created within a namespace?]]
* [[Explain the concept of Namespaces in Kubernetes.]]
* [[How to get list of resources which are not bound to a specific namespace?]]
Q: An engineer form your organization told you he is interested only in seeing his team resources in Kubernetes. Instead, in reality, he sees resources of the whole organization, from multiple different teams. What Kubernetes concept can you use in order to deal with it?
A: Namespaces. See the following namespaces question and answer for more information.
Remember: Use <html><code>kubectl explain <resource></code></html> to explore API fields from the CLI.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
''Related atoms''
* [[True or False? With namespaces you can limit the resources consumed by the users/teams]]
* [[While namespaces do provide scope for resources, they are not isolating them]]
* [[What can you find in kube-system namespace?]]
Q: How to schedule a pod on a node called "node1"?
A: <html><code>k run some-pod --image=redix -o yaml --dry-run=client > pod.yaml</code></html>
<html><code>vi pod.yaml</code></html> and add:
<html><pre><code class="language-plaintext">spec:
nodeName: node1</code></pre></html>
<html><code>k apply -f pod.yaml</code></html>
Note: if you don't have a node1 in your cluster the Pod will be stuck on "Pending" state.
Remember: <html><code>--dry-run=client -o yaml</code></html> generates manifests. <html><code>kubectl explain</code></html> shows field schemas.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
''Related atoms''
* [[Kubernetes custom schedulers: deploy and use]]
* [[You applied a taint with k taint node minikube app=web:NoSchedule on the only node in y…]]
* [[Using node affinity, set a Pod to schedule on a node where the key is "region" and valu…]]
Q: You applied a taint with k taint node minikube app=web:NoSchedule on the only node in your cluster and then executed kubectl run some-pod --image=redis but the Pod is in pending state. How to fix it?
A: <html><code>kubectl edit po some-pod</code></html> and add the following
<html><pre><code class="language-plaintext"> - effect: NoSchedule
key: app
operator: Equal
value: web</code></pre></html>
Exit and save. The pod should be in Running state now.
Remember: Taint effects: NoSchedule (block), PreferNoSchedule (soft), NoExecute (evict+block).
Example: Add: <html><code>kubectl taint nodes n1 key=val:NoSchedule</code></html>. Remove: append <html><code>-</code></html> at end.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
//Merged from 2 source atoms.//
''Related atoms''
* [[How to schedule a pod on a node called "node1"?]]
* [[Kubernetes taints and tolerations: effects, syntax, and use cases]]
Q: Create a taint on one of the nodes in your cluster with key of "app" and value of "web" and effect of "NoSchedule". Verify it was applied
A: <html><code>k taint node minikube app=web:NoSchedule</code></html>
<html><code>k describe no minikube | grep -i taints</code></html>
Remember: Taint effects: NoSchedule (block), PreferNoSchedule (soft), NoExecute (evict+block).
Example: Add: <html><code>kubectl taint nodes n1 key=val:NoSchedule</code></html>. Remove: append <html><code>-</code></html> at end.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
//Merged from 2 source atoms.//
''Related atoms''
* [[An engineer form your organization asked whether there is a way to prevent from Pods (w…]]
* [[Kubernetes taints and tolerations: effects, syntax, and use cases]]
* [[Check what labels one of your nodes in the cluster has]]
Q: True or False? Once a Pod is assisgned to a worker node, it will only run on that node, even if it fails at some point and spins up a new Pod
A: True. Once the scheduler assigns a Pod to a node, it stays bound to that node for its lifetime. If the Pod fails, the controller (Deployment/ReplicaSet) creates a new Pod, which may be scheduled to a different node.
Remember: K8s true/false questions test defaults and edge cases. Verify: <html><code>kubectl explain</code></html>.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
''Related atoms''
* [[What happens when you run a Pod with kubectl?]]
* [[Kubernetes DaemonSet: one pod per node]]
* [[True or False? A single Pod can be split across multiple nodes]]
Q: What kubectl describe pod [pod name] does? command does?
A: Show details of a specific resource or group of resources.
Remember: <html><code>--dry-run=client -o yaml</code></html> generates manifests. <html><code>kubectl explain</code></html> shows field schemas.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
''Related atoms''
* [[Which command will list all the object types in a cluster?]]
* [[How can you find out information on a Service related to a certain Pod if all you can u…]]
* [[How many containers can a pod contain?]]
Q: What do you understand by Cloud controller manager?
A: With the help of cloud infrastructure technologies, you can run Kubernetes on them. In the context of Cloud Controller Manager, it is the control plane component that embeds the cloud-specific control logic. This process lets you link the cluster into the cloud provider's API and separates the elements that interact with the cloud platform from components that only interact with your cluster.
This also enables the cloud providers to release the features at a different pace compared to the main Kubernetes project. It is structured using a plugin mechanism and allows various cloud providers to integrate their platforms with Kubernetes.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
''Related atoms''
* [[Kubernetes Control Plane: components and responsibilities]]
* [[How does Kubernetes manage containerized applications?]]
* [[Kubernetes controllers and kube-controller-manager]]
A Pod is a group of one or more containers with shared storage, network resources, and a specification for how to run those containers. It is the smallest deployable unit of computing in Kubernetes — containers are not scheduled directly; Pods are.
Containers within the same Pod share a local network namespace and volumes, allowing them to communicate as if on the same machine while retaining a degree of isolation from containers in other Pods.
Key properties:
* Smallest schedulable unit (not a bare container).
* All containers in a Pod share the same IP address and port space.
* Pod IP is ephemeral and changes on restart; use a Service for a stable endpoint.
* Storage volumes mounted by a Pod are shared across all its containers.
Etymology: "Pod" derives from "pod of whales" — a group traveling together — reflecting the co-located, cooperative nature of its containers.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
//Merged from 3 source atoms.//
''Related atoms''
* [[Kubernetes Service: stable networking endpoint for pods]]
* [[Discuss the implications of pod sprawl and how to manage it effectively in Kubernetes.]]
Q: List all the pods with the label "env=prod"
A: <html><code>k get po -l env=prod</code></html>
To count them: <html><code>k get po -l env=prod --no-headers | wc -l</code></html>
Example: <html><code>kubectl label pod mypod app=web tier=frontend</code></html>
Remember: Labels = key-value pairs for selection. Services, Deployments, Jobs all use selectors.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
''Related atoms''
* [[Which command lists all Pods in a Kubernetes cluster?]]
* [[How to check how many Pods are ready as part of a replica set called "repli"?]]
* [[How to get information on a certain service?]]
Q: What does the node status contain?
A: The main components of a node status are:
* Address - Contains the hostname, external IP, and internal IP
* Condition - Describes the status of all running nodes
* Capacity - Describes the resources available on the node (CPU, memory, max pods)
* Info - Contains general information about the node (kernel version, Kubernetes version, container runtime details, OS)
Remember: Use <html><code>kubectl explain <resource></code></html> to explore API fields from the CLI.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
''Related atoms''
* [[What kubectl get componentstatus does?]]
* [[Which command lists all Pods in a Kubernetes cluster?]]
* [[What is a node in Kubernetes?]]
Q: Where static Pods manifests are located?
A: Most of the time it's in /etc/kubernetes/manifests but you can verify with <html><code>grep -i static /var/lib/kubelet/config.yaml</code></html> to locate the value of <html><code>statisPodsPath</code></html>.
It might be that your config is in different path. To verify run <html><code>ps -ef | grep kubelet</code></html> and see what is the value of --config argument of the process <html><code>/usr/bin/kubelet</code></html>
The key itself for defining the path of static Pods is <html><code>staticPodPath</code></html>. So if your config is in <html><code>/var/lib/kubelet/config.yaml</code></html> you can run <html><code>grep staticPodPath /var/lib/kubelet/config.yaml</code></html>.
Remember: Use <html><code>kubectl explain <resource></code></html> to explore API fields from the CLI.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
''Related atoms''
* [[Describe how would you delete a static Pod]]
* [[Kubernetes Static Pods: definition, use cases, and DaemonSet contrast]]
* [[How to identify which Pods are Static Pods?]]
Q: Why is etcd performance tied to disk latency more than CPU?
A: etcd uses Raft consensus which requires synchronous disk writes for every committed operation.
The bottleneck:
# Raft consensus protocol
** Every write must be persisted before acknowledgment
** Leader writes to WAL (Write-Ahead Log)
** Followers must also persist before ACK
** Quorum requires majority to persist
# fsync per commit
** etcd calls fsync() after every write batch
** fsync() = wait for disk confirmation
** Network latency + disk latency = total latency
** CPU sits idle waiting for disk
Remember: etcd is the cluster's "brain" — all state lives here. Back up with <html><code>etcdctl snapshot save</code></html>.
Fun fact: etcd uses Raft consensus. 3 nodes tolerate 1 failure; 5 tolerate 2.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
Q: What are Custom Resource Definitions (CRDs) in Kubernetes?
A: * Custom Resource Definitions (CRDs): CRDs extend the Kubernetes API to support custom resources defined by users.
* They allow users to define and use custom resource types beyond the core Kubernetes objects.
* CRDs are the foundation for creating custom controllers and operators.
* CRDs empower users to extend Kubernetes with domain-specific resources tailored to their applications. By defining custom resources, users can leverage the Kubernetes API and benefit from the same management capabilities for their specialized workloads.
Remember: Use <html><code>kubectl explain <resource></code></html> to explore API fields from the CLI.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
* <html><code>training/interactive/knowledge/data/cards/k8s-operators.tsv</code></html>
//Merged from 3 source atoms.//
''Related atoms''
* [[Kubernetes Operator: purpose and usage]]
* [[Explain the main components of Kubernetes architecture.]]
Q: True or False? Each Pod, when created, gets its own public IP address
A: False. Each Pod gets an IP address but an internal one and not publicly accessible.
To make a Pod externally accessible, we need to use an object called Service in Kubernetes.
Remember: K8s true/false questions test defaults and edge cases. Verify: <html><code>kubectl explain</code></html>.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
''Related atoms''
* [[True or False? By default, pods are isolated. This means they are unable to receive tra…]]
* [[What is the fundamental rule of the Kubernetes pod networking model?]]
* [[Pod IPs are ephemeral; use a Service for stable addressing]]
Q: True or False? When a namespace is deleted all resources in that namespace are not deleted but moved to another default namespace
A: False. When a namespace is deleted, the resources in that namespace are deleted as well.
Remember: 4 default namespaces: default, kube-system, kube-public, kube-node-lease. Mnemonic: "DSPL."
Gotcha: Never run production workloads in <html><code>default</code></html> — no quotas, harder RBAC.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
''Related atoms''
* [[How to get list of resources which are not bound to a specific namespace?]]
* [[What special namespaces are there by default when creating a Kubernetes cluster?]]
* [[True or False? With namespaces you can limit the resources consumed by the users/teams]]
Taints are applied to nodes and repel pods that do not explicitly tolerate them. Tolerations are set in pod specs and grant exceptions. Think of taints as stop signs; tolerations are permits. Each taint has a key, value, and effect.
''Effects:''
* <html><code>NoSchedule</code></html>: new pods without a matching toleration are not scheduled; existing pods are not evicted.
* <html><code>PreferNoSchedule</code></html>: scheduler avoids the node but will use it if no alternative exists.
* <html><code>NoExecute</code></html>: new pods are rejected and existing pods without a matching toleration are evicted; an optional <html><code>tolerationSeconds</code></html> field sets a grace period before eviction.
''Syntax:''
* Add: <html><code>kubectl taint nodes node1 gpu=true:NoSchedule</code></html>
* Remove: append <html><code>-</code></html> → <html><code>kubectl taint nodes node1 gpu=true:NoSchedule-</code></html>
* Apply at registration: <html><code>kubelet --register-with-taints</code></html>
''Built-in node-condition taints (applied automatically):''
<html><code>node.kubernetes.io/not-ready:NoExecute</code></html>, <html><code>node.kubernetes.io/unreachable:NoExecute</code></html>, <html><code>node.kubernetes.io/disk-pressure:NoSchedule</code></html>, <html><code>node.kubernetes.io/memory-pressure:NoSchedule</code></html>, <html><code>node.kubernetes.io/pid-pressure:NoSchedule</code></html>, <html><code>node.kubernetes.io/unschedulable:NoSchedule</code></html>.
''Operational use cases:''
* Dedicated nodes: taint GPU nodes so only GPU workloads (which tolerate the taint) land there.
* Maintenance: taint a node before draining to block new scheduling.
* Problematic nodes: apply <html><code>NoExecute</code></html> to evict all workloads that do not explicitly tolerate the condition.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
* <html><code>training/interactive/knowledge/data/cards/k8s-node-lifecycle.tsv</code></html>
* <html><code>training/library/topics/k8s-node-lifecycle/primer.md</code></html>
* <html><code>training/library/topics/k8s-node-lifecycle/street_ops.md</code></html>
//Merged from 3 source atoms.//
''Related atoms''
* [[You applied a taint with k taint node minikube app=web:NoSchedule on the only node in y…]]
* [[Kubernetes Static Pods: definition, use cases, and DaemonSet contrast]]
Q: How to identify which Pods are Static Pods?
A: There are several ways to identify static pods:
# ''Node name suffix'': Static pods have the node name appended as a suffix to their pod name. For example, a static pod from manifest <html><code>etcd.yaml</code></html> on node <html><code>master-01</code></html> would be named <html><code>etcd-master-01</code></html>.
# ''Owner reference'': Static pods have their <html><code>ownerReferences</code></html> field set to a Node object rather than a ReplicaSet, Deployment, or other controller.
Remember: <html><code>--dry-run=client -o yaml</code></html> generates manifests. <html><code>kubectl explain</code></html> shows field schemas.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
''Related atoms''
* [[What are your thoughts on "Pods are not meant to be created directly"?]]
* [[Describe how would you delete a static Pod]]
* [[Create a static pod with the image python that runs the command sleep 2017]]
Q: What use cases exist for running multiple containers in a single pod?
A: A web application with separate (= in their own containers) logging and monitoring components/adapters is one examples.
A CI/CD pipeline (using Tekton for example) can run multiple containers in one Pod if a Task contains multiple commands.
Remember: Use <html><code>kubectl explain <resource></code></html> to explore API fields from the CLI.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
''Related atoms''
* [[Why it's common to have only one container per Pod in most cases?]]
Q: Which resources are accessible from different namespaces?
A: Services (and a few other cluster-scoped resources like Nodes and PersistentVolumes). A Service in namespace A can be reached from namespace B via <html><code><svc>.<ns>.svc.cluster.local</code></html>.
Example: <html><code>kubectl create ns staging && kubectl config set-context --current --namespace=staging</code></html>
Remember: Namespaces = logical isolation. Combine with ResourceQuotas, NetworkPolicies, RBAC.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
''Related atoms''
* [[Which service and in which namespace the following file is referencing?]]
* [[You are looking for a Pod called "atreus". How to check in which namespace it runs?]]
Q: While namespaces do provide scope for resources, they are not isolating them
A: True. Try create two pods in two separate namespaces for example, and you'll see there is a connection between the two.
Example: <html><code>kubectl create ns staging && kubectl config set-context --current --namespace=staging</code></html>
Remember: Namespaces = logical isolation. Combine with ResourceQuotas, NetworkPolicies, RBAC.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
''Related atoms''
* [[How to get the name of the current namespace?]]
* [[True or False? With namespaces you can limit the resources consumed by the users/teams]]
Q: What can you find in kube-system namespace?
A: * Master and Kubectl processes
* System processes
Example: <html><code>kubectl create ns staging && kubectl config set-context --current --namespace=staging</code></html>
Remember: Namespaces = logical isolation. Combine with ResourceQuotas, NetworkPolicies, RBAC.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
''Related atoms''
* [[How to create components in a namespace?]]
* [[True or False? With namespaces you can limit the resources consumed by the users/teams]]
* [[Explain the concept of Namespaces in Kubernetes.]]
Q: Name the initial namespaces from which Kubernetes starts?
A: Kubernetes starts with three initial namespaces:
* default - The default namespace for objects with no other namespace
* kube-system - The namespace for objects created by the Kubernetes system
* kube-public - This namespace is readable by all users and is reserved for cluster usage
Example: <html><code>kubectl create ns staging && kubectl config set-context --current --namespace=staging</code></html>
Remember: Namespaces = logical isolation. Combine with ResourceQuotas, NetworkPolicies, RBAC.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
''Related atoms''
* [[Explain the concept of Namespaces in Kubernetes.]]
* [[Create a pod called "kartos" in the namespace dev. The pod should be using the "redis" …]]
Q: You have one Kubernetes cluster and multiple teams that would like to use it. You would like to limit the resources each team consumes in the cluster. Which Kubernetes concept would you use for that?
A: Namespaces will allow to limit resources and also make sure there are no collisions between teams when working in the cluster (like creating an app with the same name).
Remember: Use <html><code>kubectl explain <resource></code></html> to explore API fields from the CLI.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
''Related atoms''
* [[True or False? With namespaces you can limit the resources consumed by the users/teams]]
* [[Explain the concept of Namespaces in Kubernetes.]]
* [[You would like to limit the number of resources being used in your cluster. For example…]]
Q: Create a pod called "kartos" in the namespace dev. The pod should be using the "redis" image.
A: If the namespace doesn't exist already: <html><code>k create ns dev</code></html>
<html><code>k run kratos --image=redis -n dev</code></html>
Example: <html><code>kubectl create ns staging && kubectl config set-context --current --namespace=staging</code></html>
Remember: Namespaces = logical isolation. Combine with ResourceQuotas, NetworkPolicies, RBAC.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
''Related atoms''
* [[How to create components in a namespace?]]
* [[Name the initial namespaces from which Kubernetes starts?]]
* [[How view all the pods running in all the namespaces?]]
Q: You are looking for a Pod called "atreus". How to check in which namespace it runs?
A: <html><code>k get po -A | grep atreus</code></html>
Example: <html><code>kubectl create ns staging && kubectl config set-context --current --namespace=staging</code></html>
Remember: Namespaces = logical isolation. Combine with ResourceQuotas, NetworkPolicies, RBAC.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
''Related atoms''
* [[How to create components in a namespace?]]
* [[How to switch to another namespace? In other words how to change active namespace?]]
* [[How to get list of resources which are not bound to a specific namespace?]]
Q: Check how many namespaces are there
A: <html><code>k get ns --no-headers | wc -l</code></html>
Example: <html><code>kubectl create ns staging && kubectl config set-context --current --namespace=staging</code></html>
Remember: Namespaces = logical isolation. Combine with ResourceQuotas, NetworkPolicies, RBAC.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
//Merged from 3 source atoms.//
''Related atoms''
* [[Which service and in which namespace the following file is referencing?]]
* [[True or False? With namespaces you can limit the resources consumed by the users/teams]]
* [[A developer says their app can't reach a service. They're in different namespaces. What…]]
Q: How do you view logs for a container running in a Pod?
A: <html><code>kubectl logs POD_NAME</code></html> prints logs for a container in a pod. For multi-container pods, you must specify the target container: <html><code>kubectl logs POD_NAME --container=CONTAINER_NAME</code></html> (or <html><code>-c CONTAINER_NAME</code></html>). Omitting the container selector on a multi-container pod produces an error.
Useful flags:
* <html><code>-f</code></html> — follow (stream) log output in real time
* <html><code>--previous</code></html> — show logs from the previously terminated container instance (useful for crash diagnosis)
Example: <html><code>kubectl logs -f POD_NAME --previous</code></html>
Remember: single-container pods resolve automatically; multi-container pods require explicit <html><code>--container=name</code></html>.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
//Merged from 3 source atoms.//
''Related atoms''
* [[Why do container logs sometimes bring down Kubernetes nodes?]]
* [[Debugging a failing or non-starting pod in Kubernetes]]
* [[Debugging CrashLoopBackOff pods: logs, debug containers, node inspection]]
Q: What are Kubernetes selectors and how do they match resources?
A: [[Kubernetes.io|https://kubernetes.io/docs/concepts/overview/working-with-objects/labels/#label-selectors]]: "Unlike names and UIDs, labels do not provide uniqueness. In general, we expect many objects to carry the same label(s).
Via a label selector, the client/user can identify a set of objects. The label selector is the core grouping primitive in Kubernetes.
The API currently supports two types of selectors: equality-based and set-based. A label selector can be made of multiple requirements which are comma-separated. In the case of multiple requirements, all must be satisfied so the comma separator acts as a logical AND (&&) operator."
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
''Related atoms''
* [[Kubernetes Annotations vs Labels]]
Q: What kube-node-lease contains?
A: It holds information on heartbeats of nodes. Each node gets an object which holds information about its availability.
Remember: Use <html><code>kubectl explain <resource></code></html> to explore API fields from the CLI.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
''Related atoms''
* [[Explain the main components of Kubernetes architecture.]]
* [[What kube-public contains?]]
* [[What Kubernetes objects are there?]]
Q: True or False? By default, pods are isolated. This means they are unable to receive traffic from any source
A: False. By default, pods are non-isolated = pods accept traffic from any source.
Remember: K8s true/false questions test defaults and edge cases. Verify: <html><code>kubectl explain</code></html>.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
''Related atoms''
* [[A pod can reach the internet but not other pods. What do you check?]]
* [[True or False? Each Pod, when created, gets its own public IP address]]
* [[True or False? A single Pod can be split across multiple nodes]]
Labels are key/value pairs attached to Kubernetes objects (pods, nodes, services, etc.) at creation time or at any point thereafter. Each key must be unique per object. Labels specify identifying attributes meaningful to users but carry no inherent semantics to the core system.
Primary uses:
* ''Scheduling'': the scheduler can place pods with specific labels onto specific nodes.
* ''ReplicaSet tracking'': ReplicaSets use label selectors to identify which pods they manage and must scale.
* ''Service routing'': Services use selectors to find the pods they should route traffic to.
* ''Deployments and Jobs'' similarly rely on label selectors to track owned objects.
Selectors are queries that filter objects by their labels, enabling the creation of logical groups based on common attributes.
Examples:
<html><pre><code class="language-plaintext">kubectl label pod mypod app=web tier=frontend
kubectl get pods -l app=web,tier=frontend</code></pre></html>
Mental model: labels are sticky notes on objects; selectors are the search queries that find objects by those notes.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
//Merged from 3 source atoms.//
''Related atoms''
* [[Kubernetes Service: stable networking endpoint for pods]]
* [[Explain the main components of Kubernetes architecture.]]
Q: Why to use namespaces? What is the problem with using one default namespace?
A: When using the default namespace alone, it becomes hard over time to get an overview of all the applications you manage in your cluster. Namespaces make it easier to organize the applications into groups that makes sense, like a namespace of all the monitoring applications and a namespace for all the security applications, etc.
Namespaces can also be useful for managing Blue/Green environments where each namespace can include a different version of an app and also share resources that are in other namespaces (namespaces like logging,
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
''Related atoms''
* [[Explain the concept of Namespaces in Kubernetes.]]
Q: Discuss the role of kube-proxy in Kubernetes networking.
A: * kube-proxy Role: kube-proxy is responsible for maintaining network rules on nodes.
* It enables communication to and from pods and services.
* Implements load balancing for services with multiple pod instances.
* kube-proxy operates at the network layer, ensuring that communication between pods and services is correctly routed. It plays a crucial role in load balancing service traffic among available pod instances, contributing to the overall stability and performance of the Kubernetes networking model.
Remember: Use <html><code>kubectl explain <resource></code></html> to explore API fields from the CLI.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
//Merged from 4 source atoms.//
''Related atoms''
* [[ClusterIP: Kubernetes default internal-only Service type]]
Q: True or False? The "Pending" phase means the Pod was not yet accepted by the Kubernetes cluster so the scheduler can't run it unless it's accepted
A: False. "Pending" is after the Pod was accepted by the cluster, but the container can't run for different reasons like images not yet downloaded.
Under the hood: Scheduler: filter (which nodes CAN?) then score (which is BEST?). "Filter then Score."
Remember: Scheduler watches for unassigned pods (no nodeName). It does NOT move running pods.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
''Related atoms''
* [[True or False? Once a Pod is assisgned to a worker node, it will only run on that node,…]]
* [[How do you determine if a pod is Pending due to resource pressure?]]
Q: True or False? Using the node affinity type "preferredDuringSchedulingIgnoredDuringExecution" means the scheduler can't schedule unless the rule is met
A: False. The scheduler tries to find a node that meets the requirements/rules and if it doesn't it will schedule the Pod anyway.
Under the hood: Scheduler: filter (which nodes CAN?) then score (which is BEST?). "Filter then Score."
Remember: Scheduler watches for unassigned pods (no nodeName). It does NOT move running pods.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
''Related atoms''
* [[Kubernetes custom schedulers: deploy and use]]
* [[True or False? Once a Pod is assisgned to a worker node, it will only run on that node,…]]
Q: True or False? The scheduler is responsible for both deciding where a Pod will run and actually running it
A: False. While the scheduler is responsible for choosing the node on which the Pod will run, Kubelet is the one that actually runs the Pod.
Under the hood: Scheduler: filter (which nodes CAN?) then score (which is BEST?). "Filter then Score."
Remember: Scheduler watches for unassigned pods (no nodeName). It does NOT move running pods.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
''Related atoms''
* [[What is the job of the kube-scheduler?]]
* [[Kubernetes custom schedulers: deploy and use]]
* [[kube-scheduler: pod scheduling in the Kubernetes control plane]]
Q: What are the types of controller managers?
A: The primary controller managers that can run on the master node are:
* Node controller - Responsible for noticing and responding when nodes go down
* Replication controller - Responsible for maintaining the correct number of pods
* Endpoints controller - Populates the Endpoints object
* Service accounts controller - Creates default accounts for new namespaces
* Token controller - Creates tokens for the service accounts
* Namespace controller - Manages the lifecycle of namespaces
Remember: Controllers run reconciliation loops: desired state vs actual state → take action.
Example: ReplicaSet controller: 2 running, 3 desired → create 1. That's reconciliation.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
''Related atoms''
* [[Kubernetes controllers and kube-controller-manager]]
* [[What process is responsible for running and installing the different controllers?]]
* [[What do you understand by Cloud controller manager?]]
Q: Describe the Kubernetes API versioning strategy.
A: * API Versioning Strategy: Kubernetes follows a versioning scheme for its API to maintain compatibility and introduce new features.
* API versions are structured as "group/version."
* For example, "v1" represents the stable core API, while "apps/v1" may represent a group-specific version for certain resources.
* The versioning strategy allows Kubernetes to introduce enhancements without breaking existing deployments.
* It provides stability for core resources while enabling extensions and improvements through versioned groups.
Remember: Use <html><code>kubectl explain <resource></code></html> to explore API fields from the CLI.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
''Related atoms''
* [[How do you define a Kubernetes Deployment?]]
* [[What fields are mandatory with any Kubernetes object?]]
Q: What is a Kubernetes cluster and what are its components?
A: A Kubernetes cluster is a group of nodes (machines) managed by the Kubernetes control plane. It consists of: control plane components (API server, etcd, scheduler, controller-manager) and worker nodes running kubelet, kube-proxy, and a container runtime. Use <html><code>kubectl cluster-info</code></html> to view endpoints.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
''Related atoms''
* [[What is kubectl and how is it used to manage Kubernetes clusters?]]
* [[Kubernetes Control Plane: components and responsibilities]]
* [[Explain the purpose of kubelet in the Kubernetes cluster.]]
Q: What is a cluster of containers in Kubernetes?
A: A cluster of containers is a set of machine elements that are nodes. Clusters initiate specific routes so that the containers running on the nodes can communicate with each other. In Kubernetes, the container engine (not the server of the Kubernetes API) provides hosting for the API server.
Remember: Use <html><code>kubectl explain <resource></code></html> to explore API fields from the CLI.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
''Related atoms''
* [[Kubernetes Pod: the smallest deployable unit]]
* [[Explain the main components of Kubernetes architecture.]]
Q: What is a Kubernetes Cluster?
A: Red Hat Definition: "A Kubernetes cluster is a set of node machines for running containerized applications. If you’re running Kubernetes, you’re running a cluster.
At a minimum, a cluster contains a worker node and a master node."
Read more [[here|https://www.redhat.com/en/topics/containers/what-is-a-kubernetes-cluster]]
Remember: Use <html><code>kubectl explain <resource></code></html> to explore API fields from the CLI.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
''Related atoms''
* [[Explain the main components of Kubernetes architecture.]]
* [[Kubernetes Control Plane: components and responsibilities]]
* [[Explain the purpose of kubelet in the Kubernetes cluster.]]
Q: What is a node in Kubernetes?
A: A node is the smallest fundamental unit of computing hardware. It represents a single machine in a cluster, which could be a physical machine in a data center or a virtual machine from a cloud provider. Each machine can substitute any other machine in a Kubernetes cluster. The master in Kubernetes controls the nodes that have containers.
Remember: Use <html><code>kubectl explain <resource></code></html> to explore API fields from the CLI.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
//Merged from 3 source atoms.//
''Related atoms''
* [[Explain the main components of Kubernetes architecture.]]
* [[What does the node status contain?]]
* [[What the master node is responsible for?]]
Q: True or False? Every cluster must have 0 or more master nodes and at least 1 worker
A: False. A Kubernetes cluster consists of at least 1 master and can have 0 workers (although that wouldn't be very useful...)
Remember: K8s true/false questions test defaults and edge cases. Verify: <html><code>kubectl explain</code></html>.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
''Related atoms''
* [[True or False? A single Pod can be split across multiple nodes]]
* [[True or False? Once a Pod is assisgned to a worker node, it will only run on that node,…]]
* [[Explain the working of the master node in Kubernetes?]]
Q: Check what labels one of your nodes in the cluster has
A: <html><code>k get no minikube --show-labels</code></html>
Example: <html><code>kubectl label pod mypod app=web tier=frontend</code></html>
Remember: Labels = key-value pairs for selection. Services, Deployments, Jobs all use selectors.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
''Related atoms''
* [[Kubernetes Service empty endpoints: label-selector mismatch diagnosis]]
* [[You applied a taint with k taint node minikube app=web:NoSchedule on the only node in y…]]
The control plane (also called the master node) consists of components that manage the overall state of the cluster:
* ''API Server'' — the Kubernetes API endpoint; all cluster components communicate through it.
* ''Scheduler'' — assigns workloads (Pods) to worker nodes based on resource availability and constraints.
* ''Controller Manager'' — handles cluster maintenance tasks: enforcing replication counts, detecting node failures, reconciling desired vs. actual state, and more.
* ''etcd'' — a distributed key-value store that persists all cluster configuration and state.
The control plane does not run user workloads; it makes decisions and drives the cluster toward the desired state declared in the API. Use <html><code>kubectl explain <resource></code></html> to explore API fields from the CLI.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
//Merged from 2 source atoms.//
''Related atoms''
* [[Kubernetes DaemonSet: one pod per node]]
* [[What is a Kubernetes cluster and what are its components?]]
* [[Explain the purpose of kubelet in the Kubernetes cluster.]]
Q: True or False? By default there is no communication between two Pods in two different namespaces
A: False. By default two Pods in two different namespaces are able to communicate with each other.
Try it for yourself:
kubectl run test-prod -n prod --image ubuntu -- sleep 2000000000
kubectl run test-dev -n dev --image ubuntu -- sleep 2000000000
<html><code>k describe po test-prod -n prod</code></html> to get the IP of test-prod Pod.
Access dev Pod: <html><code>kubectl exec --stdin --tty test-dev -n dev -- /bin/bash</code></html>
And ping the IP of test-prod Pod you get earlier.You'll see that there is communication between the two pods, in two separate namespaces.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
''Related atoms''
* [[True or False? Each Pod, when created, gets its own public IP address]]
* [[A pod can reach the internet but not other pods. What do you check?]]
* [[Pod IPs are ephemeral; use a Service for stable addressing]]
Q: What are your thoughts on "Pods are not meant to be created directly"?
A: Pods are usually indeed not created directly. You'll notice that Pods are usually created as part of another entities such as Deployments or ReplicaSets.
If a Pod dies, Kubernetes will not bring it back. This is why it's more useful for example to define ReplicaSets that will make sure that a given number of Pods will always run, even after a certain Pod dies.
Remember: <html><code>--dry-run=client -o yaml</code></html> generates manifests. <html><code>kubectl explain</code></html> shows field schemas.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
''Related atoms''
* [[How many containers can a pod contain?]]
* [[How to identify which Pods are Static Pods?]]
* [[What is the difference between a Pod and a Deployment in Kubernetes?]]
Q: Explain the purpose of the following lines
A: They define a readiness probe where the Pod will not be marked as "Ready" before it will be possible to connect to port 2017 of the container. The first check/probe will start after 15 seconds from the moment the container started to run and will continue to run the check/probe every 20 seconds until it will manage to connect to the defined port.
Remember: Use <html><code>kubectl explain <resource></code></html> to explore API fields from the CLI.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
''Related atoms''
* [[What is the constraint on successThreshold for liveness and startup probes, and why doe…]]
* [[How do readiness probes interact with rolling deployments?]]
Q: Using node affinity, set a Pod to schedule on a node where the key is "region" and value is either "asia" or "emea"
A: <html><code>vi pod.yaml</code></html>
<html><pre><code class="language-yaml">affinity:
nodeAffinity:
requiredDuringSchedulingIgnoredDuringExecution:
nodeSelectorTerms:
- matchExpressions:
- key: region
operator: In
values:
- asia
- emea</code></pre></html>
Remember: Node affinity=where. Pod affinity=near whom. Pod anti-affinity=away from whom.
Gotcha: <html><code>required</code></html>=hard rule (must match). <html><code>preferred</code></html>=soft (best effort).
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
//Merged from 2 source atoms.//
''Related atoms''
* [[How to schedule a pod on a node called "node1"?]]
* [[True or False? Using the node affinity type "preferredDuringSchedulingIgnoredDuringExec…]]
* [[Kubernetes custom schedulers: deploy and use]]
Q: True or False? With namespaces you can limit the resources consumed by the users/teams
A: True. With namespaces you can limit CPU, RAM and storage usage.
Example: <html><code>kubectl create ns staging && kubectl config set-context --current --namespace=staging</code></html>
Remember: Namespaces = logical isolation. Combine with ResourceQuotas, NetworkPolicies, RBAC.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
''Related atoms''
* [[You have one Kubernetes cluster and multiple teams that would like to use it. You would…]]
* [[While namespaces do provide scope for resources, they are not isolating them]]
* [[What can you find in kube-system namespace?]]
Q: Check if there are taints on node "master"
A: <html><code>k describe no master | grep -i taints</code></html>
Remember: Taint effects: NoSchedule (block), PreferNoSchedule (soft), NoExecute (evict+block).
Example: Add: <html><code>kubectl taint nodes n1 key=val:NoSchedule</code></html>. Remove: append <html><code>-</code></html> at end.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
''Related atoms''
* [[You applied a taint with k taint node minikube app=web:NoSchedule on the only node in y…]]
* [[Kubernetes taints and tolerations: effects, syntax, and use cases]]
Q: Which service and in which namespace the following file is referencing?
A: It's referencing the service "samurai" in the namespace called "jack".
Example: <html><code>kubectl create ns staging && kubectl config set-context --current --namespace=staging</code></html>
Remember: Namespaces = logical isolation. Combine with ResourceQuotas, NetworkPolicies, RBAC.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
''Related atoms''
* [[Which resources are accessible from different namespaces?]]
* [[While namespaces do provide scope for resources, they are not isolating them]]
* [[Check how many namespaces are there]]
Q: What kubectl get componentstatus does?
A: Outputs the status of each of the control plane components.
Remember: Use <html><code>kubectl explain <resource></code></html> to explore API fields from the CLI.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
''Related atoms''
* [[What does the node status contain?]]
* [[Kubernetes Control Plane: components and responsibilities]]
* [[Describe shortly and in high-level, what happens when you run kubectl get nodes]]
Q: Why etcd? Why not some SQL or NoSQL database?
A: When chosen as the data store etcd was (and still is of course):
* Highly Available - you can deploy multiple nodes
* Fully Replicated - any node in etcd cluster is "primary" node and has full access to the data
* Consistent - reads return latest data
* Secured - supports both TLS and SSL
* Speed - high performance data store (10k writes per sec!)
Remember: etcd is the cluster's "brain" — all state lives here. Back up with <html><code>etcdctl snapshot save</code></html>.
Fun fact: etcd uses Raft consensus. 3 nodes tolerate 1 failure; 5 tolerate 2.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
''Related atoms''
* [[Why is etcd performance tied to disk latency more than CPU?]]
Q: How to delete all pods whose status is not "Running"?
A: <html><code>kubectl delete pods --field-selector=status.phase!=Running</code></html> removes all pods not in Running state. This catches Failed, Succeeded, Pending, and Unknown pods.
Gotcha: add <html><code>--dry-run=client</code></html> first to preview what would be deleted. Use <html><code>--field-selector=status.phase==Failed</code></html> to target only failed pods specifically.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
''Related atoms''
* [[Deleting a Deployment (not its pods) is the correct removal path]]
Q: An engineer form your organization asked whether there is a way to prevent from Pods (with cretain label) to be scheduled on one of the nodes in the cluster. Your reply is:
A: Yes, using taints, we could run the following command and it will prevent from all resources with label "app=web" to be scheduled on node1: <html><code>kubectl taint node node1 app=web:NoSchedule</code></html>
Example: <html><code>kubectl label pod mypod app=web tier=frontend</code></html>
Remember: Labels = key-value pairs for selection. Services, Deployments, Jobs all use selectors.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
''Related atoms''
* [[Kubernetes custom schedulers: deploy and use]]
* [[Diagnosing pods that won't schedule on a node due to taints]]
Q: True or False? A single Pod can be split across multiple nodes
A: False. A single Pod can run on a single node.
Remember: K8s true/false questions test defaults and edge cases. Verify: <html><code>kubectl explain</code></html>.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
''Related atoms''
* [[True or False? Every cluster must have 0 or more master nodes and at least 1 worker]]
* [[True or False? By default, pods are isolated. This means they are unable to receive tra…]]
* [[True or False? Once a Pod is assisgned to a worker node, it will only run on that node,…]]
Q: Why it's common to have only one container per Pod in most cases?
A: One reason is that it makes it harder to scale when you need to scale only one of the containers in a given Pod.
Remember: Use <html><code>kubectl explain <resource></code></html> to explore API fields from the CLI.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
''Related atoms''
* [[How many containers can a pod contain?]]
* [[What use cases exist for running multiple containers in a single pod?]]
* [[What is the purpose of Horizontal Pod Autoscaling in Kubernetes?]]
''Command:'' <html><code>kubectl delete pod <pod_name></code></html>
''Termination sequence:''
# <html><code>SIGTERM</code></html> is sent to the main process inside each container of the Pod.
# Each container is given a 30-second grace period to shut down gracefully.
# If the grace period expires, <html><code>SIGKILL</code></html> forcefully terminates the remaining processes and containers.
''Skip grace period:'' <html><code>kubectl delete pod <pod_name> --grace-period=0 --force</code></html> bypasses the wait and kills immediately.
''Cascade behavior:'' Deleting a Deployment also removes its ReplicaSets and all associated Pods.
''Tip:'' Use <html><code>kubectl explain <resource></code></html> to explore API fields (including <html><code>terminationGracePeriodSeconds</code></html>) from the CLI.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
* <html><code>training/interactive/knowledge/data/cards/k8s-ops.tsv</code></html>
//Merged from 2 source atoms.//
''Related atoms''
* [[Exit code 137 in containers means OOM kill, not app error]]
* [[What is the PID 1 problem in containers and how does it cause exit code 137 on pod term…]]
Q: Explain the concept of Namespaces in Kubernetes.
A: * Namespaces:
* Virtual clusters within a physical cluster, providing a way to partition resources.
* Used to organize and scope objects, preventing naming conflicts.
* Enable the isolation of resources, making it easier to manage and share clusters.
Namespaces allow multiple teams or applications to coexist within a shared Kubernetes cluster without interfering with each other. They provide a logical separation of resources, such as pods, services, and storage, making it possible to create isolated environments within a single physical cluster.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
//Merged from 4 source atoms.//
''Related atoms''
* [[Name the initial namespaces from which Kubernetes starts?]]
* [[Which resources are accessible from different namespaces?]]
* [[You have one Kubernetes cluster and multiple teams that would like to use it. You would…]]
Q: What Kubernetes objects are there?
A: * Pod
* Service
* ReplicationController
* ReplicaSet
* DaemonSet
* Namespace
* ConfigMap
...
Remember: Use <html><code>kubectl explain <resource></code></html> to explore API fields from the CLI.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
''Related atoms''
* [[What Kubernetes objects do you usually use when deploying applications in Kubernetes?]]
* [[Kubernetes Control Plane: components and responsibilities]]
* [[What kube-public contains?]]
Q: How to confirm a container is running after running the command kubectl run web --image nginxinc/nginx-unprivileged
A: * When you run <html><code>kubectl describe pods <POD_NAME></code></html> it will tell whether the container is running:
<html><code>Status: Running</code></html>
* Run a command inside the container: <html><code>kubectl exec web -- ls</code></html>
Remember: <html><code>--dry-run=client -o yaml</code></html> generates manifests. <html><code>kubectl explain</code></html> shows field schemas.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
''Related atoms''
* [[How can you find out information on a Service related to a certain Pod if all you can u…]]
* [[How do you verify that an ImagePullSecret is correct?]]
* [[How to verify a deployment was created?]]
Q: Create a static pod with the image python that runs the command sleep 2017
A: First change to the directory tracked by kubelet for creating static pod: <html><code>cd /etc/kubernetes/manifests</code></html> (you can verify path by reading kubelet conf file)
Now create the definition/manifest in that directory
<html><code>k run some-pod --image=python --command sleep 2017 --restart=Never --dry-run=client -o yaml > statuc-pod.yaml</code></html>
Remember: <html><code>--dry-run=client -o yaml</code></html> generates manifests. <html><code>kubectl explain</code></html> shows field schemas.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
''Related atoms''
* [[How to identify which Pods are Static Pods?]]
* [[Run a pod called "yay2" with the image "python". Make sure it has resources request of …]]
* [[How to schedule a pod on a node called "node1"?]]
Q: What kube-public contains?
A: * A configmap, which contains cluster information
* Publicly accessible data
Remember: Use <html><code>kubectl explain <resource></code></html> to explore API fields from the CLI.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
''Related atoms''
* [[What kube-node-lease contains?]]
* [[What Kubernetes objects are there?]]
* [[Explain the main components of Kubernetes architecture.]]
Q: What are all the phases/steps of a control loop?
A: - Observe - identify the cluster current state
* Diff - Identify whether a diff exists between current state and desired state
* Act - Bring current cluster state to the desired state (basically reach a state where there is no diff)
Remember: Use <html><code>kubectl explain <resource></code></html> to explore API fields from the CLI.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
''Related atoms''
* [[Kubernetes Operator: components and control loop pattern]]
* [[What is the reconciliation loop in a Kubernetes operator?]]
* [[Kubernetes Control Plane: components and responsibilities]]
Q: What process is responsible for running and installing the different controllers?
A: The kube-controller-manager runs all built-in controllers (ReplicaSet, Deployment, Job, Node, ServiceAccount, etc.) as goroutines in a single process. Each controller watches the API server for its resource type and reconciles desired vs actual state. If kube-controller-manager crashes, no new reconciliation happens until it restarts.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
''Related atoms''
* [[Kubernetes controllers and kube-controller-manager]]
* [[How does Kubernetes manage containerized applications?]]
* [[What are the types of controller managers?]]
Q: How to get list of resources which are not bound to a specific namespace?
A: kubectl api-resources --namespaced=false
Remember: 4 default namespaces: default, kube-system, kube-public, kube-node-lease. Mnemonic: "DSPL."
Gotcha: Never run production workloads in <html><code>default</code></html> — no quotas, harder RBAC.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
//Merged from 4 source atoms.//
''Related atoms''
* [[How to get the name of the current namespace?]]
* [[You are looking for a Pod called "atreus". How to check in which namespace it runs?]]
* [[While namespaces do provide scope for resources, they are not isolating them]]
Q: What are the components of a worker node (aka data plane)?
A: * Kubelet - an agent responsible for node communication with the master.
* Kube-proxy - load balancing traffic between app components
* Container runtime - the engine runs the containers (Podman, Docker, ...)
Remember: Use <html><code>kubectl explain <resource></code></html> to explore API fields from the CLI.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-core.tsv</code></html>
''Related atoms''
* [[Explain the main components of Kubernetes architecture.]]
* [[What is a node in Kubernetes?]]
* [[What is a Kubernetes cluster and what are its components?]]
Q: How do you verify that an ImagePullSecret is correct?
A: kubectl get secret <name> -o jsonpath='{.data.\\.dockerconfigjson}' | base64 -d to inspect credentials. Then test manually: docker login <registry> with those creds. Also check the secret is in the same namespace as the pod.
Remember: K8s troubleshooting: Get→Describe→Logs→Exec. Mnemonic: "GDLE."
Gotcha: Always check Events with <html><code>kubectl describe</code></html> — they tell WHY, not just WHAT failed.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-troubleshooting.tsv</code></html>
''Related atoms''
* [[A pod is stuck in ImagePullBackOff. What are the common causes?]]
* [[Debugging a failing or non-starting pod in Kubernetes]]
Q: Check if there are any limits on one of the pods in your cluster
A: <html><code>kubectl describe pod <POD_NAME> | grep -i limits</code></html> shows resource limits (CPU, memory) set on the pod's containers. You can also use <html><code>kubectl get pod <name> -o jsonpath='{.spec.containers[*].resources}'</code></html> for structured output.
Gotcha: pods without limits can consume unbounded resources and affect other workloads on the same node.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-ops.tsv</code></html>
Q: A container keeps restarting with exit code 137. Describe your troubleshooting steps.
A: Run kubectl describe pod to confirm OOMKilled in the Last State reason. Check the container's memory limit in the pod spec. Use kubectl top pod or Prometheus metrics to see actual memory usage. Increase the memory limit or fix the memory leak in the application.
Remember: Flow: Get→Describe→Logs→Exec. Mnemonic: "GDLE."
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-troubleshooting.tsv</code></html>
''Related atoms''
* [[A pod has a main container limited to 512Mi and an Istio sidecar limited to 256Mi. The …]]
* [[A pod is OOMKilled but the app's memory usage looks normal. What happened?]]
An operator should set <html><code>ownerReferences</code></html> on every child resource it creates. This enables Kubernetes garbage collection: when the parent custom resource is deleted, Kubernetes automatically deletes all owned children. Without owner references, deleted CRs leave orphaned child resources behind.
Critical constraint: <html><code>ownerReferences</code></html> must only be set on resources the operator itself creates. Setting them on resources owned by another team or system causes cascading deletion—when the parent CR is deleted, Kubernetes garbage-collects those foreign resources too, removing workloads unintentionally.
<html><code>ownerReferences</code></html> encodes ownership (and therefore deletion responsibility), not association. For resources that are related to a CR but must outlive it, or that belong to another system, use labels and selectors instead. Labels express association without triggering garbage collection.
Summary rule: create resource → set ownerReference. Watch or reference a resource you did not create → use labels/selectors only.
----
''Sources''
* <html><code>training/library/topics/k8s-ecosystem/footguns.md</code></html>
* <html><code>training/interactive/knowledge/data/cards/k8s-operators.tsv</code></html>
//Merged from 2 source atoms.//
''Related atoms''
* [[What are Kubernetes selectors and how do they match resources?]]
Q: What challenges do you anticipate when managing large-scale Kubernetes clusters, and how would you address them?
A: * Challenges in Large-Scale Clusters: Resource Scaling: Ensuring adequate resources for a growing number of pods.
* Network Complexity: Handling increased network traffic and potential bottlenecks.
* Cluster Monitoring: Implementing effective monitoring and logging at scale.
* Configuration Management: Managing configurations consistently across a large number of nodes.
Fun fact: Kubernetes (K8s) is Greek for "helmsman." The 8 = letters between K and s.
Remember: K8s promise: declarative desired state + reconciliation loops = self-healing.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-general.tsv</code></html>
''Related atoms''
* [[Why do Kubernetes clusters fail at scale?]]
* [[What are the challenges in managing stateful applications in Kubernetes?]]
* [[What Kubernetes objects do you usually use when deploying applications in Kubernetes?]]
Q: What problems does Kubernetes actually solve?
A: Kubernetes solves operational problems, not application problems:
''Scheduling'': Places containers on nodes based on resources, constraints, affinity rules.
''Self-healing'': Restarts failed containers, replaces unhealthy pods, reschedules when nodes die.
''Service discovery'': Internal DNS, load balancing, service endpoints - apps find each other.
''Declarative state'': You describe desired state; Kubernetes converges to it.
''Rollouts'': Controlled deployments with rollback capability.
''What it does NOT fix'':
* Bad application architecture
* Poor observability
* Security vulnerabilities in code
* Stateful application complexity
Kubernetes is powerful infrastructure, not a magic wand for bad apps.
Fun fact: Kubernetes (K8s) is Greek for "helmsman." The 8 = letters between K and s.
Remember: K8s promise: declarative desired state + reconciliation loops = self-healing.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-general.tsv</code></html>
''Related atoms''
* [[How does Kubernetes manage containerized applications?]]
* [[When or why NOT to use Kubernetes?]]
Q: What actions or operations you consider as best practices when it comes to Kubernetes?
A: - Always make sure Kubernetes YAML files are valid. Applying automated checks and pipelines is recommended.
** Always specify requests and limits to prevent situation where containers are using the entire cluster memory which may lead to OOM issue
** Specify labels to logically group Pods, Deployments, etc. Use labels to identify the type of the application for example, among other things
Fun fact: Kubernetes (K8s) is Greek for "helmsman." The 8 = letters between K and s.
Remember: K8s promise: declarative desired state + reconciliation loops = self-healing.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-general.tsv</code></html>
''Related atoms''
* [[When or why NOT to use Kubernetes?]]
* [[Kubernetes: definition, origin, and core features]]
* [[How do you define a Kubernetes Deployment?]]
Q: What are federated clusters?
A: The aggregation of multiple clusters that treat them as a single logical cluster refers to cluster federation. In this, multiple clusters may be managed as a single cluster. They stay with the assistance of federated groups. Also, users can create various clusters within the data center or cloud and use the federation to control or manage them in one place.
You can perform cluster federation by doing the following:
* Cross cluster that provides the ability to have DNS and Load Balancer with backend from the participating clusters
* Users can sync resources across different clusters in order to deploy the same deployment set across the various clusters
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-general.tsv</code></html>
''Related atoms''
* [[What is a cluster of containers in Kubernetes?]]
Q: Tell me about your Kubernetes experience.
A: I've supported Kubernetes clusters from an operations standpoint — troubleshooting pods, reviewing logs, debugging failed deployments, managing YAML manifests, and handling upgrades. I'm strong with kubectl workflows and understanding how workloads behave at the node and container level. I haven't acted as a cluster architect, but I'm very effective on the operations and reliability side.
Remember: K8s design: declare desired state, controllers reconcile. "Desired vs Actual → Action."
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-general.tsv</code></html>
''Related atoms''
* [[You encounter a performance issue in a Kubernetes cluster. How do you diagnose and reso…]]
* [[Kubernetes: definition, origin, and core features]]
* [[Kubernetes Control Plane: components and responsibilities]]
Q: When or why NOT to use Kubernetes?
A: - If you manage low level infrastructure or baremetals, Kubernetes is probably not what you need or want
** If you are a small team (like less than 20 engineers) running less than a dozen of containers, Kubernetes might be an overkill (even if you need scale, rolling out updates, etc.). You might still enjoy the benefits of using managed Kubernetes, but you definitely want to think about it carefully before making a decision on whether to adopt it.
Remember: K8s design: declare desired state, controllers reconcile. "Desired vs Actual → Action."
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-general.tsv</code></html>
''Related atoms''
* [[What actions or operations you consider as best practices when it comes to Kubernetes?]]
* [[How does Kubernetes manage containerized applications?]]
* [[What problems does Kubernetes actually solve?]]
Q: You have a microservices-based application. How would you deploy and manage it in Kubernetes?
A: ''Microservices Deployment in Kubernetes:''
* Containerize Services: Package each microservice into a container.
* Define Kubernetes Resources: Create Deployment or StatefulSet for each microservice.
* Service Discovery: Use Kubernetes Services for inter-microservice communication.
* Configurations: Utilize ConfigMaps and Secrets for configuration management.
* Horizontal Scaling: Leverage Kubernetes autoscaling for dynamic workload adjustments.
* Health Checks: Implement readiness and liveness probes for robust application health monitoring.
* Kubernetes simplifies the deployment and management of microservices by providing abstractions for containers, services, and dynamic scaling, along with features like ConfigMaps and Secrets for configuration management.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-general.tsv</code></html>
''Related atoms''
* [[How do you define a Kubernetes Deployment?]]
* [[Discuss the considerations for migrating an application from a monolithic architecture …]]
* [[Kubernetes: definition, origin, and core features]]
Q: Explain how you would monitor and scale a critical production application in Kubernetes.
A: **Monitoring:* •
• Use monitoring tools like Prometheus, Grafana, or Kubernetes-native solutions.
• Set up alerts based on key metrics, including resource utilization and application health.
• Monitor pod and node status, and track events.
**Scaling:* •
• Utilize Horizontal Pod Autoscaling (HPA) based on metrics like CPU or custom metrics.
• Consider Vertical Pod Autoscaling for adjusting resource limits dynamically.
• Implement Cluster Autoscaler for scaling the node pool based on demand.
• Effective monitoring involves selecting appropriate tools and setting up alerts.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-ops.tsv</code></html>
''Related atoms''
* [[Kubernetes monitoring solutions overview]]
* [[Manual node scaling vs Cluster Autoscaler in Kubernetes]]
Q: What is orchestration when it comes to software and DevOps?
A: Orchestration refers to the integration of multiple services that allows them to automate processes or synchronize information in a timely fashion. Say, for example, you have six or seven microservices for an application to run. If you place them in separate containers, this would inevitably create obstacles for communication. Orchestration would help in such a situation by enabling all services in individual containers to work seamlessly to accomplish a single goal.
Remember: K8s design: declare desired state, controllers reconcile. "Desired vs Actual → Action."
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-general.tsv</code></html>
Kubernetes (K8s) is an open-source container orchestration system that automates the deployment, scaling, and management of containerized applications across a cluster of machines.
''Name origin:'' Greek for "helmsman." The abbreviation K8s encodes the 8 letters between K and s. Created at Google, derived from their internal Borg system (managing production workloads since 2003); donated to the CNCF in 2015.
''Core features:''
* Automates manual processes: decides where to place containers and how to launch them.
* Manages multiple clusters simultaneously.
* Provides additional services: container management, security, networking, and storage.
* Self-monitors health of nodes and containers (self-healing).
* Supports both horizontal and vertical scaling, quickly and easily.
''Fundamental promise:'' declarative desired state + reconciliation loops = self-healing. Users declare what they want; Kubernetes continuously reconciles actual state toward that declaration.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-general.tsv</code></html>
//Merged from 2 source atoms.//
''Related atoms''
* [[What actions or operations you consider as best practices when it comes to Kubernetes?]]
* [[What are some of Kubernetes features?]]
* [[You have a microservices-based application. How would you deploy and manage it in Kuber…]]
Q: What is Istio? What is it used for?
A: Istio is an open source service mesh that helps organizations run distributed, microservices-based apps anywhere. Istio enables organizations to secure, connect, and monitor microservices, so they can modernize their enterprise apps more swiftly and securely.
Remember: K8s design: declare desired state, controllers reconcile. "Desired vs Actual → Action."
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-general.tsv</code></html>
Q: Why does Kubernetes stress Linux more than VMs?
A: Kubernetes uses Linux primitives heavily, creating unique pressure points:
''cgroups'': Every container has resource limits enforced by cgroups. Hundreds of containers = hundreds of cgroup hierarchies.
''Namespaces'': Network, PID, mount namespaces per container. Context switching overhead.
''iptables/nftables'': Service networking creates massive rule chains. Every service adds rules. Conntrack tables fill up.
''Overlay networking'': CNI plugins add network overhead. VXLAN/GENEVE encapsulation, bridge networking.
Remember: K8s design: declare desired state, controllers reconcile. "Desired vs Actual → Action."
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-general.tsv</code></html>
''Related atoms''
* [[Why do Kubernetes clusters fail at scale?]]
* [[When or why NOT to use Kubernetes?]]
* [[What problems does Kubernetes actually solve?]]
Q: What is the Google Container Engine?
A: The Google Container Engine (GKE - Google Kubernetes Engine) is an open-source management platform tailor-made for Docker containers and clusters to provide support for the clusters that run in Google public cloud services.
Remember: K8s design: declare desired state, controllers reconcile. "Desired vs Actual → Action."
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-general.tsv</code></html>
''Related atoms''
* [[How does Kubernetes integrate with cloud providers like AWS, Azure, and GCP?]]
* [[What is a cluster of containers in Kubernetes?]]
Q: What are the main differences between Docker Swarm and Kubernetes?
A: Docker Swarm is Docker's native, open-source container orchestration platform that is used to cluster and schedule Docker containers. Swarm differs from Kubernetes in the following ways:
* Docker Swarm is more convenient to set up but doesn't have a robust cluster, while Kubernetes is more complicated to set up but the benefit of having the assurance of a robust cluster
* Docker Swarm can't do auto-scaling (as can Kubernetes); however, Docker scaling is five times faster than Kubernetes
Fun fact: Kubernetes (K8s) is Greek for "helmsman." The 8 = letters between K and s.
Remember: K8s promise: declarative desired state + reconciliation loops = self-healing.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-general.tsv</code></html>
''Related atoms''
* [[How are Kubernetes and Docker related?]]
Q: What is the difference between deploying applications on hosts and containers?
A: Deploying Applications on hosts consist of an architecture that has an operating system. The operating system will have a kernel that holds various libraries installed on the operating system needed for an application.
Whereas container host refers to the system that runs the containerized processes. This kind is isolated from the other applications; therefore, the applications must have the necessary libraries. The binaries are separated from the rest of the system and cannot infringe any other application.
Remember: K8s design: declare desired state, controllers reconcile. "Desired vs Actual → Action."
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-general.tsv</code></html>
''Related atoms''
* [[How does Kubernetes manage containerized applications?]]
* [[How are Kubernetes and Docker related?]]
Q: What is Minikube and when would you use it?
A: With the help of Minikube, users can run Kubernetes locally. This process lets the user run a single-node Kubernetes cluster on your personal computer, including Windows, macOS, and Linux PCs. With this, users can try out Kubernetes and also use it for daily development work.
Remember: K8s design: declare desired state, controllers reconcile. "Desired vs Actual → Action."
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-general.tsv</code></html>
''Related atoms''
* [[When or why NOT to use Kubernetes?]]
* [[Tell me about your Kubernetes experience.]]
* [[Kubernetes: definition, origin, and core features]]
Q: How does Kubernetes manage containerized applications?
A: Kubernetes manages containerized applications through a declarative configuration model and a set of controllers. Users describe the desired state of their applications using YAML or JSON files, and Kubernetes controllers continuously work to maintain that desired state. Key components include Deployments, Services, and other abstractions that define and control the application's behavior.
* Declarative Configuration: Users define the desired state of their applications and infrastructure using configuration files.
Remember: K8s design: declare desired state, controllers reconcile. "Desired vs Actual → Action."
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-general.tsv</code></html>
''Related atoms''
* [[When or why NOT to use Kubernetes?]]
* [[What are the challenges in managing stateful applications in Kubernetes?]]
* [[What is the difference between deploying applications on hosts and containers?]]
Q: What fields are mandatory with any Kubernetes object?
A: metadata, kind and apiVersion
Fun fact: Kubernetes (K8s) is Greek for "helmsman." The 8 = letters between K and s.
Remember: K8s promise: declarative desired state + reconciliation loops = self-healing.
Remember: Every K8s manifest needs at least: apiVersion, kind, metadata. These three fields are mandatory.
Gotcha: apiVersion changes between K8s versions. Use <html><code>kubectl api-resources</code></html> to find the current group/version for each resource.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-general.tsv</code></html>
''Related atoms''
* [[What Kubernetes objects do you usually use when deploying applications in Kubernetes?]]
* [[Describe the Kubernetes API versioning strategy.]]
* [[What actions or operations you consider as best practices when it comes to Kubernetes?]]
Q: Discuss the considerations for migrating an application from a monolithic architecture to Kubernetes.
A: ''Considerations for Migration:''
* Containerization: Break the monolith into smaller, containerized services.
* Data Migration: Plan for migrating and managing data in a microservices environment.
* Service Dependencies: Understand and manage dependencies between services.
* Networking: Design and implement a robust network architecture.
* Scalability: Leverage Kubernetes scaling features for individual services.
* Monitoring and Logging: Implement effective monitoring and logging for improved observability.
Remember: K8s design: declare desired state, controllers reconcile. "Desired vs Actual → Action."
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-general.tsv</code></html>
''Related atoms''
* [[You have a microservices-based application. How would you deploy and manage it in Kuber…]]
* [[Kubernetes: definition, origin, and core features]]
* [[What problems does Kubernetes actually solve?]]
Q: What Kubernetes objects do you usually use when deploying applications in Kubernetes?
A: * Deployment - creates the Pods () and watches them
* Service: route traffic to Pods internally
* Ingress: route traffic from outside the cluster
Fun fact: Kubernetes (K8s) is Greek for "helmsman." The 8 = letters between K and s.
Remember: K8s promise: declarative desired state + reconciliation loops = self-healing.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-general.tsv</code></html>
''Related atoms''
* [[What are some of Kubernetes features?]]
* [[What Kubernetes objects are there?]]
* [[What fields are mandatory with any Kubernetes object?]]
Q: Why do Kubernetes clusters fail at scale?
A: Scale exposes design assumptions:
''Poor resource requests'': Without proper requests/limits, scheduler makes bad decisions. Nodes get overcommitted.
''etcd overload'': etcd is the bottleneck. Too many objects, frequent updates, large secrets overwhelm it.
''Bad networking assumptions'': CNI plugins have limits. Service mesh overhead. Network policies complexity.
''Logging storms'': Container logs to stdout without limits fills disks, kills kubelet.
''Control plane sizing'': Default control plane can't handle thousands of nodes/pods.
''API server rate limiting'': Too many controllers/operators hammering API.
''Fix'': Right-size resources, monitor etcd, limit log retention, use appropriate CNI, scale control plane.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-general.tsv</code></html>
''Related atoms''
* [[What challenges do you anticipate when managing large-scale Kubernetes clusters, and ho…]]
* [[You encounter a performance issue in a Kubernetes cluster. How do you diagnose and reso…]]
Q: Discuss the differences between OpenShift and vanilla Kubernetes.
A: * OpenShift vs. Vanilla Kubernetes: OpenShift is a Kubernetes distribution with additional features, including developer and operations tools.
* OpenShift has built-in security features, integrated CI/CD pipelines, and a developer-friendly web console.
* Vanilla Kubernetes is the upstream project, while OpenShift is a product built on top of Kubernetes.
Remember: K8s design: declare desired state, controllers reconcile. "Desired vs Actual → Action."
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-general.tsv</code></html>
''Related atoms''
* [[When or why NOT to use Kubernetes?]]
* [[How are Kubernetes and Docker related?]]
* [[How does Kubernetes manage containerized applications?]]
Q: How does Kubernetes integrate with cloud providers like AWS, Azure, and GCP?
A: * Kubernetes and Cloud Providers: Cloud providers offer managed Kubernetes services (EKS for AWS, AKS for Azure, GKE for GCP).
* These services simplify cluster management, scaling, and integration with other cloud services.
* Kubernetes itself is cloud-agnostic, running on any infrastructure.
Remember: K8s design: declare desired state, controllers reconcile. "Desired vs Actual → Action."
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-general.tsv</code></html>
''Related atoms''
* [[How does Kubernetes manage containerized applications?]]
* [[Tell me about your Kubernetes experience.]]
Q: How are Kubernetes and Docker related?
A: Docker is an open-source platform used to handle software development. Its main benefit is that it packages the settings and dependencies that the software/application needs to run into a container, which allows for portability and several other advantages. Kubernetes allows for the manual linking and orchestration of several containers, running on multiple hosts that have been created using Docker.
Remember: K8s design: declare desired state, controllers reconcile. "Desired vs Actual → Action."
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-general.tsv</code></html>
''Related atoms''
* [[What is the difference between deploying applications on hosts and containers?]]
* [[Kubernetes: definition, origin, and core features]]
Q: What does being cloud-native mean?
A: The term cloud native refers to the concept of building and running applications to take advantage of the distributed computing offered by the cloud delivery model.
Remember: K8s design: declare desired state, controllers reconcile. "Desired vs Actual → Action."
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-general.tsv</code></html>
''Related atoms''
* [[How does Kubernetes integrate with cloud providers like AWS, Azure, and GCP?]]
* [[Discuss the considerations for migrating an application from a monolithic architecture …]]
* [[Kubernetes: definition, origin, and core features]]
Q: What are some of Kubernetes features?
A: - Self-Healing: Kubernetes uses health checks to monitor containers and run certain actions upon failure or other type of events, like restarting the container
** Load Balancing: Kubernetes can split and/or balance requests to applications running in the cluster, based on the state of the Pods running the application
** Operators: Kubernetes packaged applications that can use the API of the cluster to update its state and trigger actions based on events and application state changes
** Automated Rollout: Gradual updates roll out to
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-general.tsv</code></html>
''Related atoms''
* [[What Kubernetes objects do you usually use when deploying applications in Kubernetes?]]
* [[Kubernetes: definition, origin, and core features]]
* [[How do you define a Kubernetes Deployment?]]
The Container Network Interface (CNI) standard delegates pod-to-pod networking to pluggable implementations. Without a CNI plugin, pods on different nodes cannot communicate.
Calico provides BGP-based routing with native or overlay modes and enforces NetworkPolicies; it scales well for clusters of 1000+ nodes.
Cilium uses eBPF programs in the Linux kernel for networking, load balancing, and security policies. It can replace kube-proxy entirely, delivering lower latency and deep network flow visibility. Also scales to 1000+ nodes.
Flannel is the simplest option — VXLAN overlay only — with no NetworkPolicy enforcement. Suitable for small clusters; pair with Calico's policy engine if segmentation is needed.
Selection criteria: scale (Cilium and Calico for large clusters; Flannel for small), whether NetworkPolicy enforcement is required, and whether you want to manage eBPF programs (Cilium) versus traditional routing daemons (Calico, Flannel). Mnemonic: CCF — Calico (policy), Cilium (eBPF), Flannel (simple).
----
''Sources''
* <html><code>training/library/topics/k8s-networking/primer.md</code></html>
* <html><code>training/interactive/knowledge/data/cards/k8s-networking.tsv</code></html>
//Merged from 2 source atoms.//
Q: What is NodePort service type in Kubernetes?
A: The NodePort service is the most fundamental way to get external traffic directly to your service. It opens a specific port on all Nodes and forwards any traffic sent to this port to the service. NodePort exposes the service on each Node's IP at a static port (the NodePort).
Remember: NodePort: 30000-32767. "30K to 32K." Opens on ALL nodes.
Gotcha: NodePort on every node, even without target pods. Traffic forwarded internally.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-services.tsv</code></html>
''Related atoms''
* [[What the following command does?]]
* [[kubectl expose: create a Service from a workload resource]]
Q: Explain the differences between ClusterIP, NodePort, and LoadBalancer service types.
A: ''Service Types in Kubernetes:''
• ClusterIP:
• Exposes the service on an internal IP within the cluster.
• Suitable for intra-cluster communication.
• NodePort:
• Exposes the service on a static port on each node's IP.
• Makes the service accessible externally.
• LoadBalancer:
• Requests an external load balancer to manage the service.
• Useful for exposing services externally in cloud environments.
Different service types provide varying levels of accessibility. projects/knowledge/interview/kubernetes/350-explain-the-differences-between-clusterip-nodeport.txt
Remember: ClusterIP = default, internal only. "Cluster Internal Protocol."
Under the hood: kube-proxy programs iptables/IPVS to load-balance to pods.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-concept-chain.tsv</code></html>
* <html><code>training/interactive/knowledge/data/cards/k8s-networking.tsv</code></html>
* <html><code>training/interactive/knowledge/data/cards/k8s-services.tsv</code></html>
//Merged from 6 source atoms.//
''Related atoms''
* [[Service type LoadBalancer provisions cloud load balancer]]
By default, all pod-to-pod traffic in Kubernetes is unrestricted. The moment the first NetworkPolicy is applied to a namespace, the namespace enters whitelist mode: all traffic not explicitly allowed by a policy rule is denied. There is no explicit deny primitive — policies are purely additive allow lists.
A <html><code>podSelector</code></html> field targets the pods the policy protects. An empty <html><code>podSelector: {}</code></html> matches all pods in the namespace. Combined with <html><code>policyTypes: [Ingress]</code></html> and an empty <html><code>ingress</code></html> list, this produces a default-deny ingress baseline that blocks all inbound traffic:
<html><pre><code class="language-yaml">apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: deny-all-ingress
spec:
podSelector: {}
ingress: []</code></pre></html>
Rules permit traffic from specific source pods (via <html><code>podSelector</code></html>), namespaces (via <html><code>namespaceSelector</code></html>), or CIDR ranges (via <html><code>ipBlock</code></html>), scoped to specified ports and protocols.
''Common operational traps:''
* ''DNS egress'': Teams adding a default-deny egress policy often forget to allow UDP/TCP port 53 to <html><code>kube-system</code></html>, breaking all name resolution. Pods can still reach services by IP but cannot resolve DNS names.
* ''Label reuse'': A label such as <html><code>role: frontend</code></html> shared across unrelated workloads causes those pods to match rules in other policies, granting unintended access. Selector labels should be scoped to the exact workload they target.
K8s security layers by concern: RBAC (who), NetworkPolicy (what traffic), PSA (how pods run), encryption (data at rest/transit).
----
''Sources''
* <html><code>training/library/topics/k8s-networking/primer.md</code></html>
* <html><code>training/library/topics/k8s-networking/street_ops.md</code></html>
* <html><code>training/interactive/knowledge/data/cards/k8s-security.tsv</code></html>
//Merged from 3 source atoms.//
''Related atoms''
* [[True or False? If no network policies are applied to a pod, then no connections to or f…]]
A default-deny egress NetworkPolicy blocks all outbound traffic, including DNS resolution. Without an explicit egress rule allowing UDP and TCP port 53 to the CoreDNS service in kube-system, every pod in the namespace loses the ability to resolve Service names. The failure is silent: connections time out waiting for DNS resolution rather than producing an immediate error, making the cluster appear broadly broken — service discovery fails, database connections drop, and the root cause is non-obvious.
Fix: always include an egress rule permitting traffic to the kube-dns Service in kube-system on port 53 (both UDP and TCP) whenever restricting egress.
Under the hood: pod <html><code>/etc/resolv.conf</code></html> is configured to point at CoreDNS, which auto-discovers Services and resolves them as <html><code><svc>.<ns>.svc.cluster.local</code></html>. Blocking port 53 severs this path entirely.
Additional consideration: if cross-namespace traffic is required — for example, Prometheus scraping metrics from application pods — include a <html><code>namespaceSelector</code></html> ingress rule to allow traffic from the monitoring namespace. Both omissions (DNS egress and cross-namespace ingress) are among the most common NetworkPolicy mistakes.
----
''Sources''
* <html><code>training/library/topics/k8s-networking/primer.md</code></html>
* <html><code>training/library/topics/k8s-networking/anti_primer.md</code></html>
* <html><code>training/library/topics/k8s-services-and-ingress/footguns.md</code></html>
* <html><code>training/library/topics/k8s-services-and-ingress/primer.md</code></html>
* <html><code>training/library/topics/k8s-networking/footguns.md</code></html>
* <html><code>training/library/topics/networking/footguns.md</code></html>
* <html><code>training/interactive/knowledge/data/cards/k8s-networking.tsv</code></html>
//Merged from 2 source atoms.//
''Related atoms''
* [[Kubernetes ndots:5 causes excessive external DNS queries]]
Setting <html><code>clusterIP: None</code></html> creates a headless Service that omits the virtual IP layer entirely. DNS queries return actual pod IP addresses as A records, updated dynamically as pods are created or destroyed. This makes headless Services the right choice for StatefulSets where clients must address a specific pod by stable hostname—for example, <html><code>postgres-0.db-headless.svc.cluster.local</code></html>—rather than routing through a shared virtual IP.
Headless Services do not load-balance. Most DNS clients pick the first returned IP and cache it, so all traffic concentrates on one pod while the remaining replicas sit idle. For load-balanced access to stateless workloads, use a regular ClusterIP Service instead. If an application must use a headless Service for distribution, it must implement client-side round-robin across all returned IPs, or use a client library that supports it.
Key invariants: headless = <html><code>clusterIP: None</code></html>; DNS returns individual pod IPs, not a VIP; designed for stateful workloads (databases, message brokers) needing per-pod addressing; not a substitute for ClusterIP-based load balancing.
----
''Sources''
* <html><code>training/library/topics/k8s-networking/trivia.md</code></html>
* <html><code>training/library/topics/k8s-services-and-ingress/primer.md</code></html>
* <html><code>training/library/topics/k8s-services-and-ingress/footguns.md</code></html>
* <html><code>training/interactive/knowledge/data/cards/k8s-networking.tsv</code></html>
* <html><code>training/interactive/knowledge/data/cards/k8s-services.tsv</code></html>
//Merged from 3 source atoms.//
''Related atoms''
* [[ClusterIP: Kubernetes default internal-only Service type]]
CNI (Container Network Interface) is a minimal specification for configuring network interfaces of container workloads. A CNI plugin is an executable that accepts a JSON config and a network namespace path, sets up networking—assigning IP addresses, configuring routes, and adding iptables rules—and returns the assigned IP. Kubelet calls CNI plugins to set up and tear down pod network namespaces, enabling cross-node pod communication. Without a CNI plugin, pods on different nodes cannot communicate.
The spec is intentionally brief: its narrow interface allows fundamentally different architectures—overlay tunnels (Flannel), BGP routing, and eBPF dataplanes (Cilium)—to coexist and evolve independently without requiring Kubernetes core changes. Calico layers network policy enforcement on top of routing. CNI provides a standardized contract between container runtimes (Docker, containerd) and network plugins, decoupling runtime implementation from networking strategy. This simplicity lets implementations improve without touching Kubernetes core.
----
''Sources''
* <html><code>training/library/topics/k8s-networking/trivia.md</code></html>
* <html><code>training/interactive/knowledge/data/cards/k8s-networking.tsv</code></html>
* <html><code>training/interactive/knowledge/data/cards/k8s-services.tsv</code></html>
//Merged from 3 source atoms.//
Q: What is the fundamental rule of the Kubernetes pod networking model?
A: Every pod gets its own IP address, and all pods can communicate with all other pods without NAT (flat network model).
Remember: K8s networking rule: every pod gets a unique IP. Pod-to-pod across nodes without NAT.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-networking.tsv</code></html>
''Related atoms''
* [[True or False? Each Pod, when created, gets its own public IP address]]
Q: What DNS record format does Kubernetes create for a ClusterIP Service?
A: <service-name>.<namespace>.svc.cluster.local — for example, backend-api.production.svc.cluster.local.
Remember: K8s DNS: <html><code><svc>.<ns>.svc.cluster.local</code></html>. CoreDNS runs in kube-system.
Under the hood: CoreDNS auto-discovers Services. Pod /etc/resolv.conf points to it.
Fun fact: CoreDNS replaced kube-dns in K8s 1.13 (2018). It is written in Go and uses a plugin architecture.
Remember: K8s DNS format: <svc>.<ns>.svc.cluster.local. Within the same namespace, just <svc> works thanks to search domains.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-networking.tsv</code></html>
''Related atoms''
* [[Kubernetes ndots:5 causes excessive external DNS queries]]
Q: What is the key difference between kube-proxy iptables mode and IPVS mode?
A: iptables mode rewrites the full rule chain on every Service change (O(n) updates, slow at scale), while IPVS uses kernel-level hash tables for O(1) lookup performance and better scalability beyond ~5,000 Services.
Remember: K8s networking rule: every pod gets a unique IP. Pod-to-pod across nodes without NAT.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-networking.tsv</code></html>
Q: What happens when you apply a NetworkPolicy with an empty podSelector and policyTypes: [Ingress] but no ingress rules?
A: It creates a default-deny-ingress policy that blocks ALL inbound traffic to every pod in the namespace, since the empty podSelector matches all pods and no ingress rules means nothing is allowed.
Remember: Ingress routes HTTP/HTTPS by hostname/path. L7 only. Use NodePort/LB for L4.
Gotcha: Without an Ingress Controller deployed, Ingress resources have no effect.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-networking.tsv</code></html>
''Related atoms''
* [[Ingress 404: controller missing, wrong ingressClassName, or no endpoints]]
* [[True or False? If no network policies are applied to a pod, then no connections to or f…]]
When a pod can reach a Service by ClusterIP but not by DNS name, the root cause is almost always in the DNS resolution path, not the network itself.
''Diagnosis steps:''
# Inspect the pod's resolver config: <html><code>kubectl exec -it <pod> -- cat /etc/resolv.conf</code></html> — verify the nameserver points to CoreDNS ClusterIP and search domains include <html><code><ns>.svc.cluster.local</code></html>.
# Test basic cluster DNS: <html><code>kubectl exec <pod> -- nslookup kubernetes.default</code></html> — failure here confirms DNS, not routing.
# Confirm CoreDNS pods are running: <html><code>kubectl get pods -n kube-system -l k8s-app=kube-dns</code></html>.
# Confirm CoreDNS Service has endpoints: <html><code>kubectl get endpoints -n kube-system kube-dns</code></html>.
# Check CoreDNS logs for upstream or parsing errors: <html><code>kubectl logs -n kube-system -l k8s-app=kube-dns</code></html>.
# Check NetworkPolicy: a policy blocking UDP/TCP port 53 to CoreDNS will silently break DNS while leaving direct IP traffic unaffected.
''Key facts:''
* Kubernetes DNS FQDN format: <html><code><svc>.<namespace>.svc.cluster.local</code></html>.
* CoreDNS runs in the <html><code>kube-system</code></html> namespace and auto-discovers Services via the API.
* Each pod's <html><code>/etc/resolv.conf</code></html> is injected by kubelet to point at CoreDNS ClusterIP.
* External IPs being reachable while internal DNS fails isolates the fault to CoreDNS or DNS config, not the CNI or kube-proxy.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-advanced-ops.tsv</code></html>
* <html><code>training/interactive/knowledge/data/cards/k8s-networking.tsv</code></html>
* <html><code>training/interactive/knowledge/data/cards/k8s-troubleshooting.tsv</code></html>
//Merged from 2 source atoms.//
''Related atoms''
* [[Users unable to reach an application running on a Pod on Kubernetes. What might be the …]]
* [[A pod can reach the internet but not other pods. What do you check?]]
Q: How do you capture network traffic inside a running pod without modifying its image?
A: Use an ephemeral debug container (K8s 1.23+): kubectl debug -it problem-pod --image=nicolaka/netshoot --target=app-container -- tcpdump -i eth0 -nn port 8080. This attaches a debug container to the pod's network namespace without restarting the pod.
Remember: K8s networking rule: every pod gets a unique IP. Pod-to-pod across nodes without NAT.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-networking.tsv</code></html>
kubectl drain hangs indefinitely when a PodDisruptionBudget prevents eviction. A deployment has one replica and a PDB requiring minAvailable: 1. Draining the node would leave zero replicas, violating the budget. The drain command cannot proceed. Your maintenance window expires while the drain waits forever. A team sets up a maintenance window, starts the drain, and after 45 minutes discovers the drain is still running — they missed the budget check.
Before draining a node, query all PodDisruptionBudgets: kubectl get pdb --all-namespaces. For each PDB, verify the affected deployment has enough replicas that evicting one or more pods does not violate the budget. If a PDB blocks the drain, scale up the deployment first, then drain, then scale back down. Alternatively, temporarily remove or increase the PDB threshold if it's safe to do so. Always check PDBs before draining — it's a one-line check that saves a blown maintenance window.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-advanced-ops.tsv</code></html>
* <html><code>training/library/topics/k8s-node-lifecycle/footguns.md</code></html>
* <html><code>training/library/topics/k8s-node-lifecycle/primer.md</code></html>
* <html><code>training/library/topics/k8s-ops/l3-node-lifecycle.md</code></html>
* <html><code>training/library/topics/node-maintenance/primer.md</code></html>
* <html><code>training/library/topics/node-maintenance/footguns.md</code></html>
//Merged from 6 source atoms.//
''Related atoms''
* [[What causes a node drain to get stuck and how do you troubleshoot it?]]
Q: What node conditions does Kubernetes monitor for health, and what happens when a node goes NotReady?
A: Kubernetes monitors conditions like MemoryPressure, DiskPressure, PIDPressure, NetworkUnavailable, and Ready. When a node's Ready condition becomes False or Unknown (kubelet stops reporting), the node controller waits for a grace period, then marks pods as Unknown status and begins evicting them to healthy nodes. This is the automated response to node failures.
Remember: Node lifecycle: join→Ready→(cordon)→drain→remove. Monitor: <html><code>kubectl get nodes</code></html>.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-node-lifecycle.tsv</code></html>
''Related atoms''
* [[What are common causes of kubelet failures and how do they manifest?]]
* [[What combination of mechanisms ensures zero-downtime during node maintenance?]]
Nodes can be added manually (cloud: new VM via provider API; bare metal: new hardware + <html><code>kubeadm join</code></html>) or removed via cordon, drain, and <html><code>kubectl delete node</code></html>. Cluster Autoscaler automates this: it watches for pods stuck in Pending due to insufficient resources and provisions new nodes to schedule them, then removes underutilized nodes that fall below a resource threshold (default: 50% utilization for 10 minutes).
Scale-down constraints: the autoscaler respects PodDisruptionBudgets — if evicting a node would violate a PDB, the node stays. Pods using local storage also block scale-down. DaemonSet-only nodes are not removed; the autoscaler recognizes they carry no reschedulable workloads. Mark pods tolerant of eviction with annotation <html><code>cluster-autoscaler.kubernetes.io/safe-to-evict: "true"</code></html>.
Node lifecycle: join → Ready → (cordon) → drain → remove. Diagnose scale-down blocking reasons with <html><code>kubectl describe configmap cluster-autoscaler-status -n kube-system</code></html>. Monitor node state with <html><code>kubectl get nodes</code></html>.
----
''Sources''
* <html><code>training/library/topics/k8s-node-lifecycle/street_ops.md</code></html>
* <html><code>training/interactive/knowledge/data/cards/k8s-node-lifecycle.tsv</code></html>
//Merged from 2 source atoms.//
''Related atoms''
* [[What combination of mechanisms ensures zero-downtime during node maintenance?]]
* [[Karpenter vs Cluster Autoscaler: key differences]]
Upgrade nodes one at a time: cordon the node, drain it (respecting PodDisruptionBudgets), upgrade the kubelet and container runtime packages (<html><code>apt-get install kubelet=<version> kubectl=<version></code></html>), restart kubelet (<html><code>systemctl daemon-reload && systemctl restart kubelet</code></html>), verify the node version, uncordon, and confirm pod rescheduling before proceeding to the next node.
In managed Kubernetes (EKS, GKE, AKS), this is typically handled by rolling node-group replacement: new nodes at the updated version join, old nodes are cordoned, drained, and terminated.
Key constraints and rules:
* Never skip minor versions — upgrade 1.28 → 1.29 → 1.30, not 1.28 → 1.30.
* Upgrade the control plane before workers; the control plane must be at least as new as any worker node.
* Test in staging first. Use a canary approach: upgrade one node, observe for ~an hour, then proceed in batches.
* Check for removed APIs in the target version that may break existing workloads.
Rollback: cordon and drain the broken node, downgrade kubelet to the previous version, restart, then uncordon.
Behavior notes:
* Restarting kubelet does not restart running containers — they continue executing.
* On restart, kubelet re-syncs state with the API server. If kubelet was down longer than the pod-eviction-timeout, evicted pods may have rescheduled elsewhere, temporarily creating duplicates.
Node lifecycle summary: join → Ready → (cordon) → drain → remove. Monitor with <html><code>kubectl get nodes</code></html>.
----
''Sources''
* <html><code>training/library/topics/k8s-node-lifecycle/street_ops.md</code></html>
* <html><code>training/interactive/knowledge/data/cards/k8s-node-lifecycle.tsv</code></html>
//Merged from 2 source atoms.//
When a node shows NotReady, begin with <html><code>kubectl describe node <node></code></html> and inspect the Conditions section for MemoryPressure, DiskPressure, or PIDPressure flags, then read the Events — they report //why//, not just //what//. Troubleshooting order: Get → Describe → Logs → Exec (GDLE).
SSH to the node and verify the kubelet is running (<html><code>systemctl status kubelet</code></html>); if not, check logs (<html><code>journalctl -u kubelet</code></html>) and restart. Confirm the container runtime is also running, since kubelet depends on it. Test API connectivity: <html><code>kubectl logs</code></html> against a pod on that node should succeed.
Check for resource exhaustion: memory (identify and kill the hog or add capacity), disk (clean images with <html><code>crictl rmi --prune</code></html>, remove old logs), or PID exhaustion (find the process leak). Kubelet logs commonly reveal certificate expiration, CNI plugin failures, or — after a kubelet upgrade — certificate mismatches, CNI incompatibilities, or container runtime version mismatches.
If the node went NotReady following a kubelet upgrade, roll back to the previous version via the package manager (e.g., <html><code>apt-get install -y kubelet=1.28.5-00</code></html>), then reload and restart the service; the node will rejoin once authentication succeeds and the CNI plugin runs cleanly. If rollback does not resolve it, inspect per-pod logs under <html><code>/var/log/pods/</code></html> to isolate the failing component.
----
''Sources''
* <html><code>training/library/topics/k8s-node-lifecycle/street_ops.md</code></html>
* <html><code>training/library/topics/k8s-ops/l3-node-lifecycle.md</code></html>
* <html><code>training/interactive/knowledge/data/cards/k8s-troubleshooting.tsv</code></html>
//Merged from 3 source atoms.//
Q: What is the lifecycle of a Kubernetes node?
A: Provision -> register (kubelet joins cluster) -> schedule workloads -> cordon (stop new pods) -> drain (move existing pods) -> decommission or upgrade -> re-register -> repeat. Nodes are treated as ephemeral — the Kubernetes model assumes they can be replaced.
Remember: Node lifecycle: join→Ready→(cordon)→drain→remove. Monitor: <html><code>kubectl get nodes</code></html>.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-node-lifecycle.tsv</code></html>
''Related atoms''
* [[What node conditions does Kubernetes monitor for health, and what happens when a node g…]]
* [[What combination of mechanisms ensures zero-downtime during node maintenance?]]
Q: What is the difference between cordoning and draining a node?
A: Cordoning marks a node as unschedulable — no new pods will be placed on it, but existing pods continue running. Draining goes further: it cordons the node AND evicts all existing pods, moving them to other nodes. Getting drain right means zero-downtime maintenance.
Remember: Cordon=unschedulable, pods stay. "Caution tape." <html><code>kubectl uncordon</code></html> removes it.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-node-lifecycle.tsv</code></html>
''Related atoms''
* [[What is a PodDisruptionBudget (PDB) and why does it create tension with node drains?]]
Q: What is a PodDisruptionBudget (PDB) and why does it create tension with node drains?
A: A PDB specifies the minimum number (or percentage) of pods that must remain available during voluntary disruptions like node drains. It prevents drain from breaking applications by ensuring enough replicas stay running. However, PDBs are also what cause drains to get stuck — if a PDB cannot be satisfied (e.g., minAvailable equals replicas and no room to reschedule), the drain blocks indefinitely.
Remember: Drain=cordon+evict. <html><code>--ignore-daemonsets --delete-emptydir-data</code></html> for stubborn pods.
Gotcha: DaemonSet pods can't be drained — that's why <html><code>--ignore-daemonsets</code></html> exists.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-node-lifecycle.tsv</code></html>
Q: What causes a node drain to get stuck and how do you troubleshoot it?
A: Common causes: PDB cannot be satisfied (minAvailable equals replicas with no room to reschedule), pods with local storage (emptyDir) that can't be rescheduled, pods without a controller (bare pods not managed by a Deployment/ReplicaSet), or pods with long terminationGracePeriodSeconds. Troubleshoot by checking which pods remain, examining their PDBs, and using --ignore-daemonsets --delete-emptydir-data --force flags if safe.
Remember: Drain=cordon+evict. <html><code>--ignore-daemonsets --delete-emptydir-data</code></html> for stubborn pods.
Gotcha: DaemonSet pods can't be drained — that's why <html><code>--ignore-daemonsets</code></html> exists.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-node-lifecycle.tsv</code></html>
''Related atoms''
* [[PodDisruptionBudgets block kubectl drain indefinitely]]
Q: What are common causes of kubelet failures and how do they manifest?
A: Kubelet failures cause the node to go NotReady. Common causes: kubelet process crash (check systemctl status kubelet and journalctl -u kubelet), certificate expiration (kubelet can't authenticate to API server), disk pressure triggering evictions, container runtime failure (containerd/CRI-O down), or network partition isolating the node from the control plane. Each manifests differently in node conditions.
Remember: Node lifecycle: join→Ready→(cordon)→drain→remove. Monitor: <html><code>kubectl get nodes</code></html>.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-node-lifecycle.tsv</code></html>
''Related atoms''
* [[What node conditions does Kubernetes monitor for health, and what happens when a node g…]]
Q: What combination of mechanisms ensures zero-downtime during node maintenance?
A: PDBs ensure minimum pod availability during drain. Pod anti-affinity spreads replicas across nodes so no single node holds all instances. PreStop hooks give pods time to finish in-flight requests before termination. The Cluster Autoscaler or surge capacity ensures replacement nodes exist before draining. Together: PDB + anti-affinity + graceful shutdown + capacity headroom = zero-downtime maintenance.
Remember: Node lifecycle: join→Ready→(cordon)→drain→remove. Monitor: <html><code>kubectl get nodes</code></html>.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-node-lifecycle.tsv</code></html>
''Related atoms''
* [[Manual node scaling vs Cluster Autoscaler in Kubernetes]]
* [[What node conditions does Kubernetes monitor for health, and what happens when a node g…]]
Q: What is the reconciliation loop in a Kubernetes operator?
A: A control loop where the controller: 1) watches for changes to custom resources, 2) compares desired state (CR spec) with actual state (cluster), 3) creates/updates/deletes resources to match desired state, 4) updates CR status, and 5) repeats.
Remember: Operator = custom controller + CRD. Encodes ops knowledge as code. "Robot SRE."
Example: PostgreSQL Operator handles failover, backups, scaling — automating DBA tasks.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-operators.tsv</code></html>
''Related atoms''
* [[Why must the Reconcile function in a Kubernetes operator be idempotent?]]
* [[In a Go-based operator using Kubebuilder, what does returning ctrl.Result{RequeueAfter:…]]
A Kubernetes Operator is a custom controller paired with a Custom Resource Definition (CRD) that extends Kubernetes to manage complex, stateful applications. It watches custom resources and continuously reconciles actual cluster state to the desired state declared in those resources.
Operators encode domain-specific operational knowledge — scaling, backup, failover, upgrades — directly into software, replacing manual runbooks with automated lifecycle management. They are sometimes described as a "robot SRE" because they replicate the judgment a human operator would apply.
Common usage patterns:
* Deploy an Operator via a CRD that defines application-specific fields (<html><code>kubectl explain <resource></code></html> inspects those fields; <html><code>kubectl api-resources</code></html> lists all available resource types).
* The controller loop detects drift between desired and actual state and acts to correct it.
Example: a PostgreSQL Operator handles primary-replica failover, scheduled backups, and read-replica scaling, automating tasks that would otherwise fall to a DBA.
Key identity: Operator = custom controller + CRD. The CRD declares the schema; the controller provides the behavior.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-operators.tsv</code></html>
* <html><code>training/interactive/knowledge/data/cards/k8s-ops.tsv</code></html>
//Merged from 2 source atoms.//
''Related atoms''
* [[What are Custom Resource Definitions (CRDs) in Kubernetes?]]
* [[Describe in detail what is the Operator Lifecycle Manager]]
Q: Name three operator-building frameworks and when you would choose each.
A: Kubebuilder (Go, production operators), Operator SDK (Go/Ansible/Helm, Red Hat ecosystem), and Kopf (Python, quick prototypes or Python shops). Kubebuilder and Operator SDK are for production use; Kopf is for lower complexity and faster development.
Remember: Operator SDK: Helm (easy), Ansible (medium), Go (powerful). Choose by complexity.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-operators.tsv</code></html>
''Related atoms''
* [[What components the Operator Framework consists of?]]
* [[What is the Operator Framework?]]
Q: What are the five levels of the operator maturity model?
A: Level 1: Basic install (Helm wrapper). Level 2: Seamless upgrades (rolling updates, version migration). Level 3: Full lifecycle (backup, restore, scaling). Level 4: Deep insights (metrics, alerts, log analysis). Level 5: Auto-pilot (auto-scaling, auto-tuning, self-healing).
Remember: Operator = custom controller + CRD. Encodes ops knowledge as code. "Robot SRE."
Example: PostgreSQL Operator handles failover, backups, scaling — automating DBA tasks.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-operators.tsv</code></html>
''Related atoms''
* [[Describe in detail what is the Operator Lifecycle Manager]]
* [[Name three operator-building frameworks and when you would choose each.]]
Q: What are finalizers in the context of Kubernetes operators, and why are they needed?
A: Finalizers let an operator run cleanup logic before a CR is deleted. When deletion is requested, the operator detects the deletion timestamp, runs cleanup (e.g., take a final backup), removes the finalizer, and then Kubernetes completes the deletion. Without finalizers, external resources may be left behind.
Remember: Operator = custom controller + CRD. Encodes ops knowledge as code. "Robot SRE."
Example: PostgreSQL Operator handles failover, backups, scaling — automating DBA tasks.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-operators.tsv</code></html>
''Related atoms''
* [[Kubernetes Operator: purpose and usage]]
* [[What is the reconciliation loop in a Kubernetes operator?]]
* [[Why do we need Operators?]]
Q: In a Go-based operator using Kubebuilder, what does returning ctrl.Result{RequeueAfter: 30 * time.Second} from the Reconcile function do?
A: It tells the controller to re-run reconciliation for this resource after 30 seconds, even if no changes are detected. This is useful for polling external state or ensuring periodic consistency checks. Without a bounded requeue, the operator only reconciles on watch events.
Remember: Operator = custom controller + CRD. Encodes ops knowledge as code. "Robot SRE."
Example: PostgreSQL Operator handles failover, backups, scaling — automating DBA tasks.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-operators.tsv</code></html>
''Related atoms''
* [[What is the reconciliation loop in a Kubernetes operator?]]
* [[Why must the Reconcile function in a Kubernetes operator be idempotent?]]
* [[What are finalizers in the context of Kubernetes operators, and why are they needed?]]
Q: Why should operators use the status subresource for status updates instead of updating the entire CR?
A: The status subresource allows updating status independently from spec, preventing conflicts where a status update could overwrite spec changes made by users. It also enables RBAC separation so controllers can update status without permission to modify spec.
Remember: Operators follow watch→diff→act. Reconcile custom resources like built-in controllers.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-operators.tsv</code></html>
Q: Why must the Reconcile function in a Kubernetes operator be idempotent?
A: The reconciliation loop may be called multiple times for the same state due to watch events, requeues, or controller restarts. If Reconcile is not idempotent, it may create duplicate resources, send duplicate notifications, or produce inconsistent state. Every reconciliation must produce the same result regardless of how many times it runs.
Remember: Operators follow watch→diff→act. Reconcile custom resources like built-in controllers.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-operators.tsv</code></html>
''Related atoms''
* [[What is the reconciliation loop in a Kubernetes operator?]]
* [[In a Go-based operator using Kubebuilder, what does returning ctrl.Result{RequeueAfter:…]]
The Horizontal Pod Autoscaler scales replicas based on metrics like CPU or memory utilization. HPA computes utilization as <html><code>currentUsage / request</code></html> — a percentage of the requested amount. If pods have no <html><code>resources.requests.cpu</code></html> defined, there is no denominator: HPA cannot compute a percentage and reports <html><code><unknown>/70%</code></html>. Without a known utilization value, HPA cannot make scaling decisions and remains inactive.
The fix is to always define <html><code>resources.requests.cpu</code></html> on any pod targeted by HPA. The request is not a hard limit — it is a statement to both the scheduler and the autoscaler about the expected resource baseline. HPA uses this baseline to contextualize observed usage: 50mCPU on a 100mCPU request is light load (50%), but on a 10mCPU request it would represent 500% overload. Without a request, the metric is undefined and HPA is non-functional.
Quick reference: <html><code>kubectl explain pod.spec.containers.resources</code></html> for field docs; <html><code>kubectl api-resources</code></html> for all resource types.
----
''Sources''
* <html><code>training/library/topics/k8s-ops/footguns.md</code></html>
* <html><code>training/library/topics/k8s-ops/primer.md</code></html>
* <html><code>training/interactive/knowledge/data/cards/k8s-ops.tsv</code></html>
//Merged from 2 source atoms.//
''Related atoms''
* [[HPA keeps scaling to max replicas even when average CPU is low. What could cause this?]]
Using memory utilization as the primary scaling metric in HPA causes replicas to scale up under load but never scale down. Many applications with garbage-collected runtimes (JVM, Python) allocate memory on startup and hold it even when idle — the heap does not shrink and the process keeps the allocation. Once HPA scales up based on rising memory usage, the newly added replicas also allocate and hold memory, preventing utilization from dropping enough to trigger scale-down. The cluster ends up over-provisioned and wasteful.
CPU is a better primary scaling metric: it is instantaneous and reversible — when load drops, CPU usage drops immediately, allowing HPA to scale down normally. Memory should serve as a secondary metric for safety (scale up if memory reaches a dangerous threshold) or run in recommendation-only mode.
For workloads with genuine memory scaling requirements (caching layers, in-memory databases), accept that scale-down will be slow or manual, or use per-pod memory limits and eviction policies instead of relying on HPA memory targets.
----
''Sources''
* <html><code>training/library/topics/k8s-ops/footguns.md</code></html>
* <html><code>training/interactive/knowledge/data/cards/k8s-ops.tsv</code></html>
//Merged from 2 source atoms.//
''Related atoms''
* [[What does HPA do and what problem remains after you enable it?]]
* [[HPA keeps scaling to max replicas even when average CPU is low. What could cause this?]]
Horizontal Pod Autoscaler (HPA) scales replica count based on per-pod utilization (<html><code>usage / request</code></html>). Vertical Pod Autoscaler (VPA) adjusts resource requests and limits on individual pods. When both target the same metric — typically CPU — they conflict in a destabilizing feedback loop: VPA increasing a pod's CPU request lowers the utilization ratio, signaling HPA to scale down; HPA scaling out changes per-pod utilization, which in turn triggers further VPA adjustments. The result is an unstable oscillation in both replica count and resource configuration.
Mitigation options:
* Use VPA for memory-based optimization only, leaving CPU requests as a stable baseline for HPA.
* Run VPA in recommendation-only mode (<html><code>updatePolicy.updateMode: 'off'</code></html>) and apply its CPU recommendations manually.
* Never allow both controllers to actively adjust the same metric simultaneously.
Key fields: <html><code>kubectl explain vpa.spec.updatePolicy</code></html>; <html><code>kubectl api-resources</code></html> to confirm VPA CRDs are installed.
----
''Sources''
* <html><code>training/library/topics/k8s-ops/footguns.md</code></html>
* <html><code>training/interactive/knowledge/data/cards/k8s-ops.tsv</code></html>
//Merged from 2 source atoms.//
''Related atoms''
* [[HPA keeps scaling to max replicas even when average CPU is low. What could cause this?]]
* [[What is the purpose of Horizontal Pod Autoscaling in Kubernetes?]]
Liveness probes are meant to detect dead or hung processes and trigger a container restart. Adding dependency checks — such as making a database query inside the liveness endpoint — turns the probe into a system-wide health check rather than a per-container aliveness check. When the database becomes temporarily unavailable, the database connection times out, all pods in the deployment fail liveness simultaneously, and Kubernetes kills and restarts them all at once. On restart, every pod immediately attempts to reconnect to the same downed database — a thundering herd that overwhelms the database further, extends the outage, and causes the restart cycle to repeat. Readiness probes are the correct place to check dependencies: a failing readiness probe removes the pod from the load balancer without triggering a restart, allowing the database to recover without the added chaos of simultaneous reconnection storms. Liveness should only verify that the process itself is alive — a fast, local check such as returning HTTP 200 with no I/O. With this separation, a database outage causes pods to stop receiving traffic (readiness fails) but not to restart (liveness still passes), dramatically reducing recovery time.
----
''Sources''
* <html><code>training/library/topics/k8s-ops/footguns.md</code></html>
* <html><code>training/library/topics/k8s-ops/primer.md</code></html>
* <html><code>training/interactive/knowledge/data/cards/k8s-ops.tsv</code></html>
//Merged from 2 source atoms.//
''Related atoms''
* [[Startup probes replaced the initialDelaySeconds compromise in Kubernetes 1.18]]
Full JVM garbage collection pauses can stop the JVM for several seconds. If <html><code>timeoutSeconds</code></html> is set too low (e.g., 1s) and a GC pause exceeds that threshold, the probe times out and counts as a failure. Three consecutive timeouts trigger a pod restart — even though the application is not actually stuck or unresponsive. GC pauses are normal and temporary, making this failure mode particularly dangerous.
Mitigations:
* Set <html><code>timeoutSeconds</code></html> to at least 5 seconds, or higher if production GC pauses warrant it.
* Use a <html><code>startupProbe</code></html> for slow JVM boot sequences so the liveness probe does not fire prematurely.
* Tune GC settings (e.g., pause-time targets, heap sizing) to reduce worst-case pause durations.
* Monitor GC behavior separately — heap usage and pause duration metrics help calibrate the threshold.
The same principle applies to any runtime with stop-the-world pauses: Go's GC, Python's GIL, or Rust's allocator. The rule generalizes: set <html><code>timeoutSeconds</code></html> to comfortably exceed your worst-case observed pause duration, not just your average.
----
''Sources''
* <html><code>training/library/topics/k8s-ops/footguns.md</code></html>
* <html><code>training/interactive/knowledge/data/cards/k8s-ops.tsv</code></html>
//Merged from 2 source atoms.//
''Related atoms''
* [[startupProbe prevents false crash-loops during slow initialization]]
Q: A pod is OOMKilled but the app's memory usage looks normal. What happened?
A: Check memory limits vs actual usage: kubectl top pod <name>. The kernel OOM killer uses RSS (resident set size), which includes shared libraries and buffers. Java apps commonly exceed limits due to off-heap memory. Fix: increase limits or tune the runtime (e.g., -XX:MaxRAMPercentage for JVMs).
Remember: OOMKilled = exit 137 (128+SIGKILL). Exceeded memory limit. Increase or fix leak.
Gotcha: Check <html><code>kubectl describe pod</code></html> for Reason: OOMKilled in Last State.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-troubleshooting.tsv</code></html>
''Related atoms''
* [[A container keeps restarting with exit code 137. Describe your troubleshooting steps.]]
Q: What kubectl command shows whether a pod was OOMKilled, and what fields do you look for?
A: kubectl describe pod <name>. Look for Last State: Terminated, Reason: OOMKilled, Exit Code: 137 in the container status section.
Remember: OOMKilled = exit 137 (128+SIGKILL). Exceeded memory limit. Increase or fix leak.
Gotcha: Check <html><code>kubectl describe pod</code></html> for Reason: OOMKilled in Last State.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-troubleshooting.tsv</code></html>
''Related atoms''
* [[A pod has a main container limited to 512Mi and an Istio sidecar limited to 256Mi. The …]]
* [[Debugging CrashLoopBackOff pods: logs, debug containers, node inspection]]
Q: How do you distinguish a container-level OOM from a node-level OOM, and what commands reveal each?
A: Container-level: single pod affected, Exit Code 137, Reason OOMKilled in kubectl describe pod. Node-level: multiple pods affected, dmesg shows kernel OOM killer messages, kubelet logs show eviction activity, kubectl describe node shows MemoryPressure: True.
Remember: OOMKilled = exit 137 (128+SIGKILL). Exceeded memory limit. Increase or fix leak.
Gotcha: Check <html><code>kubectl describe pod</code></html> for Reason: OOMKilled in Last State.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-troubleshooting.tsv</code></html>
''Related atoms''
* [[Which Prometheus metric should you use to predict OOMKill, and why not container_memory…]]
* [[Pods without memory limits cause node-wide evictions and OOM kills]]
* [[Kubernetes node resource pressure: diagnosis, eviction, and prevention]]
Q: A pod has a main container limited to 512Mi and an Istio sidecar limited to 256Mi. The main container is OOMKilled despite using only 400Mi. What is the likely cause?
A: Each container has its own cgroup and memory limit. If the main container is OOMKilled at 400Mi with a 512Mi limit, it may be counting shared memory (e.g., tmpfs mounts, emptyDir medium: Memory volumes) against the container's cgroup. Check for memory-backed volumes and sidecar memory consumption patterns with kubectl top pod --containers.
Remember: OOMKilled = exit 137 (128+SIGKILL). Exceeded memory limit. Increase or fix leak.
Gotcha: Check <html><code>kubectl describe pod</code></html> for Reason: OOMKilled in Last State.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-troubleshooting.tsv</code></html>
''Related atoms''
* [[What kubectl command shows whether a pod was OOMKilled, and what fields do you look for?]]
* [[A container keeps restarting with exit code 137. Describe your troubleshooting steps.]]
Q: What is the purpose of Horizontal Pod Autoscaling in Kubernetes?
A: ''Horizontal Pod Autoscaling (HPA):''
* Automatically adjusts the number of pod replicas based on observed metrics, such as CPU utilization or custom metrics.
* Ensures optimal resource utilization and responsiveness.
* Helps maintain a balance between application performance and resource efficiency.
// HPA allows Kubernetes to dynamically scale the number of pod replicas in response to changing demand. By automatically adjusting the replica count based on specified metrics, // HPA ensures that applications can efficiently handle varying workloads without manual intervention.
Remember: Workload hierarchy: Deployment/StatefulSet/DaemonSet→ReplicaSet→Pod.
Gotcha: Use <html><code>kubectl explain <resource></code></html> for field reference without leaving the CLI.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-workloads.tsv</code></html>
//Merged from 2 source atoms.//
''Related atoms''
* [[Discuss the implications of pod sprawl and how to manage it effectively in Kubernetes.]]
* [[HPA and VPA on the same metric create conflicting feedback loops]]
* [[Scale a ReplicaSet with kubectl scale rs]]
HPA (Horizontal Pod Autoscaler) requires <html><code>minReplicas >= 1</code></html> and cannot scale to zero: with zero pods there are no resource metrics to evaluate for scale-up decisions. KEDA (Kubernetes Event-Driven Autoscaling), a CNCF project, removes both limitations. It scales on Kafka consumer lag, RabbitMQ queue depth, Prometheus queries, HTTP requests per second, and 60+ other external event sources — and it can scale deployments all the way to zero replicas, triggering scale-up from zero based on external signals such as queue depth or incoming HTTP traffic. Knative Serving and custom controllers are alternative approaches to scale-to-zero. KEDA is the most widely adopted solution and integrates with the existing HPA API. It is ideal for event-driven workloads such as batch processors, scheduled tasks, and serverless-style on-demand scaling where idle replicas waste cost.
----
''Sources''
* <html><code>training/library/topics/k8s-ops/trivia.md</code></html>
* <html><code>training/interactive/knowledge/data/cards/k8s-ops.tsv</code></html>
//Merged from 2 source atoms.//
Before Kubernetes 1.18 (March 2020), the only tool for slow-starting containers was <html><code>initialDelaySeconds</code></html> on liveness probes. This forced a fragile tradeoff: to survive a worst-case 120-second startup, operators had to set <html><code>initialDelaySeconds: 120</code></html>. That same delay then applied during normal operation, meaning a truly dead container could go undetected for up to 120 seconds. The failure modes cut both ways: if the app started faster than expected, crash detection was still delayed by the full window; if it started slower, the liveness probe would kill the container during boot before it ever finished initializing. Startup probes, introduced in Kubernetes 1.18, broke this coupling. A startup probe runs first and gates liveness/readiness probes until the container signals it has finished starting. Once the startup probe succeeds, liveness probes can run with aggressive intervals and thresholds without penalizing startup time. This separates the concern of "is the app done initializing" from "is the running app healthy."
----
''Sources''
* <html><code>training/library/topics/k8s-ops/trivia.md</code></html>
* <html><code>training/interactive/knowledge/data/cards/k8s-ops.tsv</code></html>
//Merged from 2 source atoms.//
''Related atoms''
* [[What is the difference between liveness and readiness probes in Kubernetes?]]
* [[Dependency checks in liveness probes cause cascading restarts]]
Q: Do you have experience with deploying a Kubernetes cluster? If so, can you describe the process in high-level?
A: 1. Create multiple instances you will use as Kubernetes nodes/workers. Create also an instance to act as the Master. The instances can be provisioned in a cloud or they can be virtual machines on bare metal hosts.
# Provision a certificate authority that will be used to generate TLS certificates for the different components of a Kubernetes cluster (kubelet, etcd, ...)
## Generate a certificate and private key for the different components
# Generate kubeconfigs so the different clients of Kubernetes can locate the API servers and authenticate.
# Generate encryption key that will be used for encrypting the cluster data
# Create an etcd cluster
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-ops.tsv</code></html>
''Related atoms''
* [[What security best practices do you follow in regards to the Kubernetes cluster?]]
* [[You have a microservices-based application. How would you deploy and manage it in Kuber…]]
* [[Explain the main components of Kubernetes architecture.]]
Q: After running kubectl run database --image mongo you see the status is "CrashLoopBackOff". What could possibly went wrong and what do you do to confirm?
A: CrashLoopBackOff means the Pod is starting, crashing, starting...and so it repeats itself.
There are many different reasons to get this error - lack of permissions, init-container misconfiguration, persistent volume connection issue, etc.
One of the ways to check why it happened is to run <html><code>kubectl describe po <POD_NAME></code></html> and having a look at the exit code
<html><pre><code class="language-plaintext"> Last State: Terminated
Reason: Error
Exit Code: 100</code></pre></html>
Another way to check what's going on, is to run <html><code>kubectl logs <POD_NAME></code></html>. This will provide us with the logs from the containers running in that Pod.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-ops.tsv</code></html>
Q: Would you use Helm, Go or something else for creating an Operator?
A: Depends on the scope and maturity of the Operator. If it mainly covers installation and upgrades, Helm might be enough. If you want to go for Lifecycle management, insights and auto-pilot, this is where you'd probably use Go.
Remember: Helm = K8s package manager. Chart=package, Release=instance, Repo=store.
Example: <html><code>helm install my-app bitnami/nginx --set service.type=LoadBalancer</code></html>
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-ops.tsv</code></html>
''Related atoms''
* [[What is Helm, and how is it used in Kubernetes?]]
* [[What are some use cases for using Helm template file?]]
Q: Why there is no such command in Kubernetes? kubectl get containers
A: Because a container is not a Kubernetes object. The smallest object unit in Kubernetes is a Pod. In a single Pod you can find one or more containers.
Remember: <html><code>kubectl explain <resource></code></html> for fields. <html><code>kubectl api-resources</code></html> for all resource types.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-ops.tsv</code></html>
''Related atoms''
* [[Which command lists all Pods in a Kubernetes cluster?]]
* [[What Kubernetes objects are there?]]
* [[What kubectl describe pod [pod name] does? command does?]]
Q: How would you approach version upgrades of Kubernetes in a production environment?
A: **Version Upgrades in Production:* •
• Conduct thorough testing in a staging environment before production.
• Follow Kubernetes documentation and release notes for upgrade procedures.
• Use tools like kubeadm for streamlined upgrade processes.
• Ensure backups and have a rollback plan in case of issues. projects/knowledge/interview/kubernetes/366-how-would-you-approach-version-upgrades-of-kuberne.txt
Remember: Version skew: kubelet ≤1 minor behind API server. Upgrade control plane first.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-ops.tsv</code></html>
''Related atoms''
* [[What rollout/deployment strategies are you familiar with?]]
Helm is a package manager for Kubernetes. It bundles related YAML manifests — Secrets, Services, ConfigMaps, Deployments, and others — into a single distributable unit called a Chart, instead of requiring each component to be authored and applied individually.
The analogy to OS package managers is direct: Helm is to Kubernetes what dnf is to Fedora/RHEL or apt is to Ubuntu.
Core use cases:
# ''Multi-cluster deployments.'' When running parallel clusters (prod, dev, staging), one Chart can be installed uniformly across all of them rather than re-applying separate YAML files per cluster.
# ''Reducing repetitive YAML authorship.'' Deploying a complex application requires coordinating many resource types. Helm collapses that into a single <html><code>helm install</code></html> invocation.
# ''Community sharing.'' Charts can be published to public or private repositories, letting teams and companies distribute application deployment configurations for others to consume in their own clusters.
Helm's value is therefore both operational (consistent, repeatable deployments across environments) and collaborative (reusable, shareable packaging of Kubernetes workloads).
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-ops.tsv</code></html>
//Merged from 2 source atoms.//
''Related atoms''
* [[What are some use cases for using Helm template file?]]
Q: What components the Operator Framework consists of?
A: 1. Operator SDK - allows developers to build operators
# Operator Lifecycle Manager - helps to install, update and generally manage the lifecycle of all operators
# Operator Metering - Enables usage reporting for operators that provide specialized services
4.
Remember: <html><code>kubectl explain <resource></code></html> for fields. <html><code>kubectl api-resources</code></html> for all resource types.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-ops.tsv</code></html>
''Related atoms''
* [[Name three operator-building frameworks and when you would choose each.]]
* [[Are there any tools, projects you are using for building Operators?]]
Q: Describe in detail what is the Operator Lifecycle Manager
A: It's part of the Operator Framework, used for managing the lifecycle of operators. It basically extends Kubernetes so a user can use a declarative way to manage operators (installation, upgrade, ...).
Remember: <html><code>kubectl explain <resource></code></html> for fields. <html><code>kubectl api-resources</code></html> for all resource types.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-ops.tsv</code></html>
''Related atoms''
* [[Kubernetes Operator: purpose and usage]]
* [[Are there any tools, projects you are using for building Operators?]]
Q: What is the Operator Framework?
A: open source toolkit used to manage k8s native applications, called operators, in an automated and efficient way.
Remember: <html><code>kubectl explain <resource></code></html> for fields. <html><code>kubectl api-resources</code></html> for all resource types.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-ops.tsv</code></html>
''Related atoms''
* [[Name three operator-building frameworks and when you would choose each.]]
* [[Kubernetes Operator: purpose and usage]]
Q: You are managing multiple Kubernetes clusters. How do you quickly change between the clusters using kubectl?
A: <html><code>kubectl config use-context <context-name></code></html> switches between clusters by changing the active context in your kubeconfig. List available contexts with <html><code>kubectl config get-contexts</code></html>. Each context combines a cluster, user, and namespace. Consider using <html><code>kubectx</code></html> for faster switching in multi-cluster environments.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-ops.tsv</code></html>
''Related atoms''
* [[kubeconfig: composable context and credential store for kubectl]]
* [[How to switch to another namespace? In other words how to change active namespace?]]
* [[What can you find in kube-system namespace?]]
Q: What are some use cases for using Helm template file?
A: * Deploy the same application across multiple different environments
* CI/CD
Remember: Helm = K8s package manager. Chart=package, Release=instance, Repo=store.
Example: <html><code>helm install my-app bitnami/nginx --set service.type=LoadBalancer</code></html>
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-ops.tsv</code></html>
''Related atoms''
* [[What is Helm, and how is it used in Kubernetes?]]
* [[Helm: Kubernetes package manager for bundling and distributing YAML]]
* [[Explain the Helm Chart Directory Structure]]
Q: It is said that Helm is also Templating Engine. What does it mean?
A: It is useful for scenarios where you have multiple applications and all are similar, so there are minor differences in their configuration files and most values are the same. With Helm you can define a common blueprint for all of them and the values that are not fixed and change can be placeholders. This is called a template file and it looks similar to the following
<html><pre><code class="language-plaintext">apiVersion: v1
kind: Pod
metadata:
name: {[ .Values.name ]}
spec:
containers:
- name: {{ .Values.container.name }}
image: {{ .Values.container.image }}
port: {{ .Values.container.port }}</code></pre></html>
The values themselves will in separate file:
<html><pre><code class="language-plaintext">name: some-app
container:
name: some-app-container
image: some-app-image
port: 1991</code></pre></html>
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-ops.tsv</code></html>
''Related atoms''
* [[What is Helm, and how is it used in Kubernetes?]]
* [[Explain the Helm Chart Directory Structure]]
A Kubernetes Operator is composed of two parts:
# ''CRD (Custom Resource Definition)'' — A user- or developer-defined Kubernetes resource that extends the Kubernetes API beyond built-in types (Deployment, Pod, Service, etc.). Operators register one or more CRDs to represent the application's desired state.
# ''Controller'' — A custom control loop that runs against the CRD. It watches for changes in the application state (current vs. desired) and reconciles them, following the same watch-reconcile loop pattern Kubernetes uses internally for native resources.
Together, a CRD + controller pair lets an Operator encode domain-specific operational knowledge (upgrades, backups, failover) as automation inside the cluster.
Useful commands:
* <html><code>kubectl explain <resource></code></html> — inspect fields of any resource, including custom ones.
* <html><code>kubectl api-resources</code></html> — list all registered resource types, including CRDs.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-ops.tsv</code></html>
//Merged from 2 source atoms.//
''Related atoms''
* [[Kubernetes Control Plane: components and responsibilities]]
* [[Kubernetes controllers and kube-controller-manager]]
Q: How Helm supports release management?
A: Helm allows you to upgrade, remove and rollback to previous versions of charts. In version 2 of Helm it was with what is known as "Tiller". In version 3, it was removed due to security concerns.
Remember: Helm = K8s package manager. Chart=package, Release=instance, Repo=store.
Example: <html><code>helm install my-app bitnami/nginx --set service.type=LoadBalancer</code></html>
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-ops.tsv</code></html>
''Related atoms''
* [[Explain the Helm Chart Directory Structure]]
* [[How to upgrade a release?]]
Q: Run a command to view all nodes of the cluster
A: <html><code>kubectl get nodes</code></html>
Note: You might want to create an alias (<html><code>alias k=kubectl</code></html>) and get used to <html><code>k get no</code></html>
Remember: <html><code>kubectl explain <resource></code></html> for fields. <html><code>kubectl api-resources</code></html> for all resource types.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-ops.tsv</code></html>
''Related atoms''
* [[Which command will list all the object types in a cluster?]]
* [[Which command lists all Pods in a Kubernetes cluster?]]
* [[What kube-public contains?]]
Common Kubernetes monitoring solutions:
* ''Prometheus + Grafana'': Most popular open-source stack. Prometheus scrapes metrics via ServiceMonitors; Grafana provides dashboards. Alertmanager handles alerts.
* ''metrics-server'': Lightweight, in-memory, open-source. Powers <html><code>kubectl top</code></html> but provides no persistence or alerting.
* ''Datadog / New Relic / Dynatrace'': Commercial SaaS platforms with auto-discovery, APM, and built-in dashboards.
* ''ELK / Loki'': Log aggregation — Elasticsearch or Loki + Grafana for unified metrics and logs.
Key things to monitor: node resources, pod status, API server latency, etcd health, PersistentVolume usage, and network policies.
Useful commands: <html><code>kubectl explain <resource></code></html> for field definitions; <html><code>kubectl api-resources</code></html> for all available resource types.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-ops.tsv</code></html>
//Merged from 2 source atoms.//
''Related atoms''
* [[Explain how you would monitor and scale a critical production application in Kubernetes.]]
Q: How do you list deployed releases?
A: <html><code>helm ls</code></html> or <html><code>helm list</code></html> shows all deployed Helm releases in the current namespace. Add <html><code>--all-namespaces</code></html> or <html><code>-A</code></html> to see releases across all namespaces. The output shows NAME, NAMESPACE, REVISION, STATUS, CHART, and APP VERSION. Use <html><code>--filter <regex></code></html> to search for specific releases.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-ops.tsv</code></html>
''Related atoms''
* [[How to view revision history for a certain release?]]
* [[How to upgrade a release?]]
* [[How Helm supports release management?]]
Q: How to display the resources usages of pods?
A: <html><code>kubectl top pod</code></html> shows real-time CPU and memory usage per pod. Requires metrics-server to be installed in the cluster. Add <html><code>--containers</code></html> to see per-container metrics within pods. Use <html><code>--sort-by=cpu</code></html> or <html><code>--sort-by=memory</code></html> to find resource-hungry pods quickly.
Gotcha: if metrics-server is not installed, this command returns an error.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-ops.tsv</code></html>
''Related atoms''
* [[Check if there are any limits on one of the pods in your cluster]]
* [[List all the pods with the label "env=prod"]]
* [[Which command lists all Pods in a Kubernetes cluster?]]
Q: Why do we need Operators?
A: The process of managing stateful applications in Kubernetes isn't as straightforward as managing stateless applications where reaching the desired status and upgrades are both handled the same way for every replica. In stateful applications, upgrading each replica might require different handling due to the stateful nature of the app, each replica might be in a different status. As a result, we often need a human operator to manage stateful applications. Kubernetes Operator is suppose to assist with this.
This also help with automating a standard process on multiple Kubernetes clusters
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-ops.tsv</code></html>
''Related atoms''
* [[Kubernetes Operator: purpose and usage]]
* [[How does Kubernetes manage containerized applications?]]
Q: Explain the need for Kustomize by describing actual use cases
A: * You have an helm chart of an application used by multiple teams in your organization and there is a requirement to add annotation to the app specifying the name of the of team owning the app
* Without Kustomize you would need to copy the files (chart template in this case) and modify it to include the specific annotations we need
* With Kustomize you don't need to copy the entire repo or files
* You are asked to apply a change/patch to some app without modifying the original files of the app
* With Kustomize you can define kustomization.yml file that defines these customizations so you don't need to touch the original app files
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-ops.tsv</code></html>
''Related atoms''
* [[Describe in high-level how Kustomize works]]
Q: What is Heapster in Kubernetes?
A: Heapster is a performance monitoring and metrics collection system for data collected by the Kubelet. This aggregator is natively supported and runs like any other pod within a Kubernetes cluster, which allows it to discover and query usage data from all nodes within the cluster. Note: Heapster has been deprecated and replaced by the Metrics Server.
Remember: <html><code>kubectl explain <resource></code></html> for fields. <html><code>kubectl api-resources</code></html> for all resource types.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-ops.tsv</code></html>
''Related atoms''
* [[A container is OOMKilled but the application memory profiler shows usage well below the…]]
* [[What is kubectl and how is it used to manage Kubernetes clusters?]]
Q: Explain how you would manage configuration drift in a Kubernetes environment.
A: Managing Configuration Drift:
* Regularly audit configurations using tools like kube-score.
* Use version control for configuration files to track changes.
* Implement GitOps practices for declarative cluster configuration.
* Leverage Helm or Kustomize for consistent application configuration.
Remember: <html><code>kubectl explain <resource></code></html> for fields. <html><code>kubectl api-resources</code></html> for all resource types.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-ops.tsv</code></html>
''Related atoms''
* [[You have a microservices-based application. How would you deploy and manage it in Kuber…]]
* [[Explain how you would monitor and scale a critical production application in Kubernetes.]]
* [[How do you define a Kubernetes Deployment?]]
Q: What is kubconfig? What do you use it for?
A: A kubeconfig file is a file used to configure access to Kubernetes when used in conjunction with the kubectl commandline tool (or other clients).
Use kubeconfig files to organize information about clusters, users, namespaces, and authentication mechanisms.
Remember: <html><code>kubectl explain <resource></code></html> for fields. <html><code>kubectl api-resources</code></html> for all resource types.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-ops.tsv</code></html>
''Related atoms''
* [[Explain the main components of Kubernetes architecture.]]
* [[kubeconfig: composable context and credential store for kubectl]]
* [[What kube-public contains?]]
Q: Create a list of all nodes in JSON format and store it in a file called "some_nodes.json"
A: <html><code>kubectl get nodes -o json > some_nodes.json</code></html> exports all node objects in JSON format. Use <html><code>-o jsonpath='{.items[*].metadata.name}'</code></html> for just the names. Other output formats: <html><code>-o yaml</code></html>, <html><code>-o wide</code></html> (extra columns), <html><code>-o name</code></html> (just resource names). Pipe to <html><code>jq</code></html> for filtering: <html><code>kubectl get nodes -o json | jq '.items[].status.conditions'</code></html>.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-ops.tsv</code></html>
''Related atoms''
* [[Run a command to view all nodes of the cluster]]
Q: After creating a service, how to check it was created?
A: <html><code>kubectl get svc</code></html> lists all services in the current namespace showing NAME, TYPE, CLUSTER-IP, EXTERNAL-IP, and PORT(S). Add <html><code>-o wide</code></html> for selector details or <html><code>--all-namespaces</code></html> for cluster-wide view. Use <html><code>kubectl describe svc <name></code></html> for endpoint details and to verify pods are being targeted correctly.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-ops.tsv</code></html>
''Related atoms''
* [[How can you find out information on a Service related to a certain Pod if all you can u…]]
* [[You are looking for a Pod called "atreus". How to check in which namespace it runs?]]
Q: Is it possible to override values in values.yaml file when installing a chart?
A: Yes. You can pass another values file:
<html><code>helm install --values=override-values.yaml [CHART_NAME]</code></html>
Or directly on the command line: <html><code>helm install --set some_key=some_value</code></html>
Remember: <html><code>kubectl explain <resource></code></html> for fields. <html><code>kubectl api-resources</code></html> for all resource types.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-ops.tsv</code></html>
''Related atoms''
* [[Explain the Helm Chart Directory Structure]]
Q: Which command lists all Pods in a Kubernetes cluster?
A: kubectl get pods (for current namespace) or kubectl get pods -A (for all namespaces) will list all pods and their status.
Remember: <html><code>kubectl explain <resource></code></html> for fields. <html><code>kubectl api-resources</code></html> for all resource types.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-ops.tsv</code></html>
''Related atoms''
* [[How view all the pods running in all the namespaces?]]
* [[Why there is no such command in Kubernetes? kubectl get containers]]
* [[List all the pods with the label "env=prod"]]
Q: Which command will list all the object types in a cluster?
A: <html><code>kubectl api-resources</code></html> lists all resource types (pods, services, deployments, etc.) available in the cluster. Shows NAME, SHORTNAMES, APIVERSION, NAMESPACED, and KIND columns. Use <html><code>--namespaced=true</code></html> to filter namespaced resources only. Combine with <html><code>kubectl explain <resource></code></html> to explore field schemas for any listed resource.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-ops.tsv</code></html>
''Related atoms''
* [[What kubectl describe pod [pod name] does? command does?]]
* [[Run a command to view all nodes of the cluster]]
* [[Why there is no such command in Kubernetes? kubectl get containers]]
Q: What the following output of kubectl get rs means?
A: The replicaset <html><code>web</code></html> has 2 replicas. It seems that the containers inside the Pod(s) are not yet running since the value of READY is 0. It might be normal since it takes time for some containers to start running and it might be due to an error. Running <html><code>kubectl describe po POD_NAME</code></html> or <html><code>kubectl logs POD_NAME</code></html> can give us more information.
Remember: <html><code>kubectl explain <resource></code></html> for fields. <html><code>kubectl api-resources</code></html> for all resource types.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-ops.tsv</code></html>
''Related atoms''
* [[What does the "ErrImagePull" status of a Pod means?]]
* [[What are the possible Pod phases?]]
* [[ReplicaSets are running the moment the user executed the command to create them (like k…]]
Standard debug flow: ''Get → Describe → Logs → Exec'' (mnemonic: GDLE).
# ''kubectl get pods'' — check current status and RESTARTS count.
# ''kubectl describe pod <pod-name>'' — read Events for probe failures, image pull errors, and scheduling issues; check <html><code>lastState</code></html> for exit codes (e.g., 137 = OOMKilled).
# ''kubectl logs <pod-name>'' / ''kubectl logs --previous'' — inspect stdout/stderr from the running or most recently crashed container.
# ''kubectl exec'' into the pod and manually curl the probe endpoint to verify it returns the expected status code.
Additional checks:
* ''Resource constraints'': confirm requests and limits are set correctly; OOMKilled (exit 137) indicates the container exceeded its memory limit.
* ''Health probes'': verify readiness/liveness probe port matches the container port and the target path/endpoint exists.
* ''Configuration'': review referenced ConfigMaps and Secrets for missing keys or mount errors.
* ''Image availability'': ensure the container image tag exists and the node can pull it (check imagePullPolicy and registry credentials).
* ''Networking'': inspect NetworkPolicies and DNS connectivity if the probe or app depends on external services.
* ''Node health'': check whether the node itself is under pressure (disk, memory, CPU) using <html><code>kubectl describe node</code></html>.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-ops.tsv</code></html>
//Merged from 3 source atoms.//
Kubernetes service networking relies on conntrack (connection tracking) to handle NAT translations. kube-proxy uses iptables or IPVS to load-balance traffic to ClusterIP/NodePort services; every connection is NAT'd to a pod, and the kernel must record that translation so return packets can route back correctly. Each active connection consumes one conntrack entry.
The kernel's conntrack table has a fixed maximum size, defaulting to 65,536 entries (sometimes 131,072). On a busy cluster — where every Service × Pod combination generates NAT entries — this table fills quickly. When it exhausts, new connections are silently dropped in the kernel: no error surfaces to the application, only unexplained timeouts or intermittent 503s. The failure affects the entire cluster, not just a single service.
Diagnosis: inspect <html><code>nf_conntrack_count</code></html> vs <html><code>nf_conntrack_max</code></html> on nodes. Enable kernel logging to catch drop events (disabled by default).
Remediation:
* Raise the limit: <html><code>sysctl -w net.netfilter.nf_conntrack_max=262144</code></html> (tune to expected peak concurrency; production clusters with NodePort-heavy workloads may need 1M+).
* Persist via <html><code>/etc/sysctl.d/</code></html>.
* Monitor <html><code>nf_conntrack_count</code></html> as an alerting metric.
Real-world example: intermittent NodePort 503 errors on a loaded cluster traced to pod-to-service NAT exhausting the default table; increasing <html><code>nf_conntrack_max</code></html> to 1,000,000 resolved it.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-ops.tsv</code></html>
* <html><code>training/interactive/knowledge/data/cards/k8s-services.tsv</code></html>
* <html><code>training/library/topics/nat/footguns.md</code></html>
//Merged from 3 source atoms.//
Q: Describe in high-level how Kustomize works
A: 1. You add kustomization.yml file in the folder of the app you would like to customize.
## You define the customizations you would like to perform
# You run <html><code>kustomize build APP_PATH</code></html> where your kustomization.yml also resides
Remember: <html><code>kubectl explain <resource></code></html> for fields. <html><code>kubectl api-resources</code></html> for all resource types.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-ops.tsv</code></html>
''Related atoms''
* [[Explain the need for Kustomize by describing actual use cases]]
* [[Explain the main components of Kubernetes architecture.]]
Q: What is kubectl and how is it used to manage Kubernetes clusters?
A: Kubectl is a CLI (command-line interface) that is used to run commands against Kubernetes clusters. As such, it controls the Kubernetes cluster manager through different create and manage commands on the Kubernetes components. It is the primary tool for interacting with Kubernetes clusters.
Remember: <html><code>kubectl explain <resource></code></html> for fields. <html><code>kubectl api-resources</code></html> for all resource types.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-ops.tsv</code></html>
''Related atoms''
* [[What is a Kubernetes cluster and what are its components?]]
* [[Explain the working of the master node in Kubernetes?]]
* [[What the master node is responsible for?]]
Q: Perhaps a general question but, you suspect one of the pods is having issues, you don't know what exactly. What do you do?
A: Start by inspecting the pods status. we can use the command <html><code>kubectl get pods</code></html> (--all-namespaces for pods in system namespace)
If we see "Error" status, we can keep debugging by running the command <html><code>kubectl describe pod [name]</code></html>. In case we still don't see anything useful we can try stern for log tailing.
In case we find out there was a temporary issue with the pod or the system, we can try restarting the pod with the following <html><code>kubectl scale deployment [name] --replicas=0</code></html>
Setting the replicas to 0 will shut down the process. Now start it with <html><code>kubectl scale deployment [name] --replicas=1</code></html>
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-ops.tsv</code></html>
''Related atoms''
* [[Why do container logs sometimes bring down Kubernetes nodes?]]
Q: How to execute the command "ls" in an existing pod?
A: <html><code>kubectl exec some-pod -it -- ls</code></html> runs the <html><code>ls</code></html> command inside a running pod's container. For a shell: <html><code>kubectl exec -it pod-name -- /bin/sh</code></html>. In multi-container pods, specify the container: <html><code>kubectl exec -it pod-name --container=sidecar -- /bin/sh</code></html>.
Gotcha: exec requires the container to have the binary installed — distroless images may lack common tools.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-ops.tsv</code></html>
''Related atoms''
* [[How can you find out information on a Service related to a certain Pod if all you can u…]]
Q: What openshift-operator-lifecycle-manager namespace includes?
A: It includes:
* catalog-operator - Resolving and installing ClusterServiceVersions the resource they specify.
* olm-operator - Deploys applications defined by ClusterServiceVersion resource
Remember: <html><code>kubectl explain <resource></code></html> for fields. <html><code>kubectl api-resources</code></html> for all resource types.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-ops.tsv</code></html>
''Related atoms''
* [[Describe in detail what is the Operator Lifecycle Manager]]
* [[What special namespaces are there by default when creating a Kubernetes cluster?]]
Q: How to view revision history for a certain release?
A: <html><code>helm history RELEASE_NAME</code></html> shows the revision history including REVISION number, STATUS, CHART version, and DESCRIPTION. This lets you see what changed between deployments and identify which revision to rollback to.
Example: <html><code>helm history my-app</code></html> then <html><code>helm rollback my-app 3</code></html> to revert to revision 3.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-ops.tsv</code></html>
''Related atoms''
* [[Check Deployment rollout history and roll back to a specific revision]]
* [[How do you list deployed releases?]]
Q: Describe shortly and in high-level, what happens when you run kubectl get nodes
A: 1. Your user is getting authenticated
# Request is validated by the kube-apiserver
# Data is retrieved from etcd
Remember: get=summary list, describe=detailed+events. "get=glance, describe=deep dive."
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-ops.tsv</code></html>
''Related atoms''
* [[Role of kube-apiserver in Kubernetes]]
* [[What happens when you run a Pod with kubectl?]]
Q: What is Container resource monitoring?
A: Container resource monitoring refers to the activity that collects the metrics and tracks the health of containerized applications and microservices environments. It helps to improve health and performance and also makes sure that they operate smoothly. Common tools include Prometheus, Grafana, and the Kubernetes Metrics Server.
Remember: <html><code>kubectl explain <resource></code></html> for fields. <html><code>kubectl api-resources</code></html> for all resource types.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-ops.tsv</code></html>
''Related atoms''
* [[Kubernetes monitoring solutions overview]]
Q: What the following command does?
A: It exposes a ReplicaSet by creating a service called 'replicaset-svc'. The exposed port is 2017 (this is the port used by the application) and the service type is NodePort which means it will be reachable externally.
Remember: <html><code>kubectl explain <resource></code></html> for fields. <html><code>kubectl api-resources</code></html> for all resource types.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-ops.tsv</code></html>
''Related atoms''
* [[True or False? the target port, in the case of running the following command, will be e…]]
* [[What is NodePort service type in Kubernetes?]]
* [[ClusterIP: Kubernetes default internal-only Service type]]
Q: How to upgrade a release?
A: <html><code>helm upgrade RELEASE_NAME CHART_NAME</code></html> updates a deployed release with new chart values or a new chart version. Add <html><code>--set key=value</code></html> for inline overrides or <html><code>-f values.yaml</code></html> for file-based config. Use <html><code>--dry-run</code></html> to preview changes before applying.
Gotcha: always pin chart versions in production to avoid unexpected upgrades.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-ops.tsv</code></html>
''Related atoms''
* [[How Helm supports release management?]]
Q: What is Helm, and how is it used in Kubernetes?
A: * Helm: Helm is a package manager for Kubernetes applications.
* It simplifies the deployment and management of Kubernetes applications by packaging them into charts.
* Charts are pre-configured Kubernetes resource definitions that can be easily deployed and versioned.
* Helm allows users to define, install, and upgrade even the most complex Kubernetes applications with a single command.
* Charts encapsulate all the required Kubernetes resources and configurations, making it easier to share and reproduce application deployments across different environments.
Remember: Helm = K8s package manager. Chart=package, Release=instance, Repo=store.
Example: <html><code>helm install my-app bitnami/nginx --set service.type=LoadBalancer</code></html>
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-ops.tsv</code></html>
//Merged from 3 source atoms.//
''Related atoms''
* [[What are some use cases for using Helm template file?]]
Q: Explain the Helm Chart Directory Structure
A: someChart/ -> the name of the chart
Chart.yaml -> meta information on the chart
values.yaml -> values for template files
charts/ -> chart dependencies
templates/ -> templates files :)
Remember: Helm = K8s package manager. Chart=package, Release=instance, Repo=store.
Example: <html><code>helm install my-app bitnami/nginx --set service.type=LoadBalancer</code></html>
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-ops.tsv</code></html>
//Merged from 2 source atoms.//
''Related atoms''
* [[How Helm supports release management?]]
* [[What are some use cases for using Helm template file?]]
Q: You encounter a performance issue in a Kubernetes cluster. How do you diagnose and resolve it?
A: **Diagnosing and Resolving Performance Issues:* •
• Resource Utilization: Check CPU, memory, and storage usage for nodes and pods.
• Logs and Events: Analyze container logs and Kubernetes events for anomalies.
• Network: Examine network policies, traffic, and potential bottlenecks.
• Pod Placement: Review node placement and resource allocation for pods.
• Kubernetes Components: Inspect the health and performance of Kubernetes control plane components.
• Application Code: Review application code for performance bottlenecks.
• Scaling: Consider scaling resources based on demand.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-ops.tsv</code></html>
Q: Are there any tools, projects you are using for building Operators?
A: This one is based more on a personal experience and taste...
* Operator Framework
* Kubebuilder
* Controller Runtime
...
Remember: <html><code>kubectl explain <resource></code></html> for fields. <html><code>kubectl api-resources</code></html> for all resource types.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-ops.tsv</code></html>
''Related atoms''
* [[Describe in detail what is the Operator Lifecycle Manager]]
* [[What components the Operator Framework consists of?]]
* [[Name three operator-building frameworks and when you would choose each.]]
Q: What does the "ErrImagePull" status of a Pod means?
A: It wasn't able to pull the image specified for running the container(s). This can happen if the client didn't authenticated for example.
More details can be obtained with <html><code>kubectl describe po <POD_NAME></code></html>.
Remember: <html><code>kubectl explain <resource></code></html> for fields. <html><code>kubectl api-resources</code></html> for all resource types.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-ops.tsv</code></html>
''Related atoms''
* [[What the following output of kubectl get rs means?]]
* [[What are the possible Pod phases?]]
* [[A pod is stuck in ImagePullBackOff. What are the common causes?]]
Q: How do you search for charts?
A: <html><code>helm search hub <keyword></code></html> searches Artifact Hub for charts across all repositories. <html><code>helm search repo <keyword></code></html> searches only locally-added repositories.
Example: <html><code>helm search hub prometheus</code></html> finds Prometheus charts from multiple publishers. Add <html><code>--max-col-width 80</code></html> for readable descriptions and <html><code>--version <semver></code></html> to filter by chart version.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-ops.tsv</code></html>
''Related atoms''
* [[Explain the Helm Chart Directory Structure]]
* [[How do you list deployed releases?]]
* [[What is Helm, and how is it used in Kubernetes?]]
Q: Users unable to reach an application running on a Pod on Kubernetes. What might be the issue and how to check?
A: Troubleshoot Kubernetes application connectivity layer by layer:
# ''Pod status'': <html><code>kubectl get pods</code></html> / <html><code>kubectl describe pod</code></html> — check for CrashLoopBackOff, Pending, ImagePullBackOff
# ''Pod health'': <html><code>kubectl exec -it pod -- curl localhost:PORT</code></html> — verify app responds inside container
# ''Service & endpoints'': <html><code>kubectl get svc</code></html> / <html><code>kubectl get endpoints</code></html> — confirm selector matches pod labels
# ''Network policies'': Check if ingress/egress rules are blocking traffic
# ''DNS'': <html><code>kubectl exec -- nslookup service-name</code></html> — verify CoreDNS resolution
# ''Ingress/LB'': Check ingress rules, TLS config, and controller logs
# ''External access'': Verify NodePort, LoadBalancer provisioning, or DNS pointing to cluster
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-ops.tsv</code></html>
''Related atoms''
* [[How do you test connectivity to a Service from inside the cluster?]]
* [[Pod reaches Service by ClusterIP but internal DNS fails: diagnosis]]
Q: A pod can reach the internet but not other pods. What do you check?
A: 1) CNI plugin is healthy (check kube-system pods). 2) NetworkPolicy blocking inter-pod traffic. 3) iptables rules on the node (kube-proxy issues). 4) Pod CIDR overlap with node network. Run kubectl exec <pod> -- ping <other-pod-ip> to confirm.
Remember: K8s troubleshooting: Get→Describe→Logs→Exec. Mnemonic: "GDLE."
Gotcha: Always check Events with <html><code>kubectl describe</code></html> — they tell WHY, not just WHAT failed.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-troubleshooting.tsv</code></html>
''Related atoms''
* [[Pod reaches Service by ClusterIP but internal DNS fails: diagnosis]]
* [[How do you test connectivity to a Service from inside the cluster?]]
Q: What happens when a readiness probe fails on a Kubernetes pod?
A: The pod is removed from service endpoints, so it stops receiving traffic. Unlike a liveness probe failure (which restarts the container), a readiness failure just takes the pod out of the load balancer. The pod keeps running. See: training/interactive/runtime-labs/lab-runtime-01-rollout-probe-failure/
Remember: <html><code>kubectl explain <resource></code></html> for fields. <html><code>kubectl api-resources</code></html> for all resource types.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-ops.tsv</code></html>
//Merged from 2 source atoms.//
''Related atoms''
* [[How should liveness and readiness probe endpoints differ in what they check, and why?]]
* [[Debugging a failing or non-starting pod in Kubernetes]]
Q: How do readiness probes interact with rolling deployments?
A: New pods must pass their readiness probe before receiving traffic and before old pods are terminated. If new pods never become ready (e.g., broken config), the rollout stalls and old pods continue serving — preventing a bad deploy from causing downtime.
Remember: <html><code>kubectl explain <resource></code></html> for fields. <html><code>kubectl api-resources</code></html> for all resource types.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-ops.tsv</code></html>
''Related atoms''
* [[What is the constraint on successThreshold for liveness and startup probes, and why doe…]]
* [[How does Kubernetes handle rolling updates with zero downtime?]]
* [[How does rolling deployment work in Kubernetes?]]
Q: What does '<unknown>/50%' mean in HPA status?
A: It means the HPA cannot read CPU metrics. Usually because metrics-server is not installed, not healthy, or the deployment doesn't have CPU resource requests defined. HPA needs requests to calculate percentage-based utilization.
Remember: <html><code>kubectl explain <resource></code></html> for fields. <html><code>kubectl api-resources</code></html> for all resource types.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-ops.tsv</code></html>
''Related atoms''
* [[What are the prerequisites for HPA to work?]]
* [[What formula does the HPA use to compute the desired replica count?]]
* [[What are the four metric types supported by HPA v2 and when would you use each?]]
Q: What are the prerequisites for HPA to work?
A: 1) metrics-server must be installed and healthy; 2) The target deployment must have resource requests (at minimum CPU requests for CPU-based scaling); 3) The metrics API must be accessible (/apis/metrics.k8s.io/v1beta1).
Remember: <html><code>kubectl explain <resource></code></html> for fields. <html><code>kubectl api-resources</code></html> for all resource types.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-ops.tsv</code></html>
''Related atoms''
* [[What does '<unknown>/50%' mean in HPA status?]]
* [[When HPA is configured with multiple metrics, how does it decide the replica count?]]
Q: What are the four metric types supported by HPA v2 and when would you use each?
A: Resource (CPU/memory from metrics-server), Pods (per-pod app metrics like RPS via custom metrics adapter), External (cloud service metrics like SQS queue depth via external adapter), and Object (metrics from a specific Kubernetes object like Ingress RPS via custom adapter).
Remember: <html><code>kubectl explain <resource></code></html> for fields. <html><code>kubectl api-resources</code></html> for all resource types.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-ops.tsv</code></html>
''Related atoms''
* [[When HPA is configured with multiple metrics, how does it decide the replica count?]]
* [[HPA requires resource requests to calculate utilization percentage]]
* [[Explain why one would specify resource limits in regards to Pods]]
Q: How do you right-size memory limits for a Kubernetes deployment?
A: 1) Set requests based on steady-state usage (observe via kubectl top or Prometheus); 2) Set limits 1.5-2x requests to allow for spikes; 3) Monitor with container_memory_working_set_bytes metric; 4) Set alerts at 80% of limit. Never set limits = requests for variable workloads.
Remember: <html><code>kubectl explain <resource></code></html> for fields. <html><code>kubectl api-resources</code></html> for all resource types.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-ops.tsv</code></html>
Q: What are the top 3 causes of CrashLoopBackOff?
A: 1) Application error on startup (bad config, missing env var, import error); 2) OOMKilled (memory limit too low); 3) Liveness probe failing too aggressively (app healthy but probe times out). Always check 'kubectl logs --previous' to see the crash reason.
Remember: <html><code>kubectl explain <resource></code></html> for fields. <html><code>kubectl api-resources</code></html> for all resource types.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-ops.tsv</code></html>
''Related atoms''
* [[Distinguish CrashLoopBackOff from other pod failure states]]
Q: What's the difference between ImagePullBackOff and ErrImagePull?
A: ErrImagePull is the first failure to pull an image. ImagePullBackOff means Kubernetes tried, failed, and is now waiting with exponential backoff before retrying. Common causes: wrong tag, private registry without credentials, or image not imported into k3s.
Remember: <html><code>kubectl explain <resource></code></html> for fields. <html><code>kubectl api-resources</code></html> for all resource types.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-ops.tsv</code></html>
''Related atoms''
* [[Distinguish CrashLoopBackOff from other pod failure states]]
Q: What is the default HPA scale-down stabilization window?
A: 300 seconds (5 minutes). The controller looks back over this window and picks the highest (most conservative) replica count recommendation to prevent flapping.
Example: <html><code>kubectl scale deployment web --replicas=5</code></html>. HPA for automatic scaling.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-ops.tsv</code></html>
''Related atoms''
* [[What does HPA do and what problem remains after you enable it?]]
Q: How do you verify that metrics-server is running and providing data?
A: Run kubectl top nodes and kubectl top pods. If they return metrics, the server is working. Also check kubectl get apiservices | grep metrics and kubectl -n kube-system get pods -l k8s-app=metrics-server.
Remember: <html><code>kubectl explain <resource></code></html> for fields. <html><code>kubectl api-resources</code></html> for all resource types.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-ops.tsv</code></html>
''Related atoms''
* [[What are the four metric types supported by HPA v2 and when would you use each?]]
* [[Kubernetes monitoring solutions overview]]
Q: What formula does the HPA use to compute the desired replica count?
A: desiredReplicas = ceil(currentReplicas * (currentMetricValue / desiredMetricValue)). For example, if current CPU is 90% and target is 70%, the scale factor is 90/70 = 1.28, so replicas increase by roughly 28%.
Remember: <html><code>kubectl explain <resource></code></html> for fields. <html><code>kubectl api-resources</code></html> for all resource types.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-ops.tsv</code></html>
''Related atoms''
* [[When HPA is configured with multiple metrics, how does it decide the replica count?]]
* [[HPA requires resource requests to calculate utilization percentage]]
Q: When HPA is configured with multiple metrics, how does it decide the replica count?
A: It evaluates each metric independently and takes the maximum desired replica count across all metrics. The most demanding metric wins. This means combining metrics with very different response characteristics can lead to unexpected scaling behavior.
Remember: <html><code>kubectl explain <resource></code></html> for fields. <html><code>kubectl api-resources</code></html> for all resource types.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-ops.tsv</code></html>
''Related atoms''
* [[What are the four metric types supported by HPA v2 and when would you use each?]]
* [[What formula does the HPA use to compute the desired replica count?]]
* [[What are the prerequisites for HPA to work?]]
Q: Explain the selectPolicy field in HPA behavior and how Max vs Min affect scaling aggressiveness.
A: selectPolicy determines which policy to apply when multiple policies are defined. Max picks whichever policy allows the largest change (most aggressive scaling). Min picks the smallest change (most conservative). Disabled prevents scaling in that direction entirely. For example, with both a Percent(100%) and Pods(5) policy on scaleUp with selectPolicy: Max, the HPA uses whichever allows adding more pods.
Remember: <html><code>kubectl explain <resource></code></html> for fields. <html><code>kubectl api-resources</code></html> for all resource types.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-ops.tsv</code></html>
''Related atoms''
* [[What is the purpose of Horizontal Pod Autoscaling in Kubernetes?]]
* [[HPA and VPA on the same metric create conflicting feedback loops]]
* [[How can PodDisruptionBudget conflict with HPA scale-down, and what is the best practice…]]
Q: How can PodDisruptionBudget conflict with HPA scale-down, and what is the best practice to avoid it?
A: PDB enforces minAvailable during voluntary disruptions. HPA sets the desired replica count, but if scaling down would violate PDB constraints during node drains or spot terminations, evictions are blocked. The HPA controller itself does not check PDB. Best practice: set HPA minReplicas to at least what PDB requires as minimum available.
Example: <html><code>kubectl scale deployment web --replicas=5</code></html>. HPA for automatic scaling.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-ops.tsv</code></html>
''Related atoms''
* [[How to scale a deployment to 8 replicas?]]
Q: What does a Kubernetes liveness probe determine?
A: Whether the container is still alive. If the liveness probe fails, Kubernetes kills and restarts the container.
Remember: <html><code>kubectl explain <resource></code></html> for fields. <html><code>kubectl api-resources</code></html> for all resource types.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-ops.tsv</code></html>
''Related atoms''
* [[Dependency checks in liveness probes cause cascading restarts]]
* [[What is the constraint on successThreshold for liveness and startup probes, and why doe…]]
Q: What are the four probe mechanisms Kubernetes supports?
A: httpGet (HTTP GET returning 2xx/3xx), tcpSocket (TCP port is open), exec (command exits 0), and grpc (gRPC health check, Kubernetes 1.24+).
Remember: <html><code>kubectl explain <resource></code></html> for fields. <html><code>kubectl api-resources</code></html> for all resource types.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-ops.tsv</code></html>
''Related atoms''
* [[What is the purpose of a startup probe?]]
* [[Explain the main components of Kubernetes architecture.]]
Q: What is the purpose of a startup probe?
A: It tells Kubernetes the container is still booting. While the startup probe is running, liveness and readiness probes are disabled. Once it succeeds, it never runs again and the other probes take over.
Remember: <html><code>kubectl explain <resource></code></html> for fields. <html><code>kubectl api-resources</code></html> for all resource types.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-ops.tsv</code></html>
''Related atoms''
* [[Startup probes replaced the initialDelaySeconds compromise in Kubernetes 1.18]]
* [[startupProbe prevents false crash-loops during slow initialization]]
* [[What is the constraint on successThreshold for liveness and startup probes, and why doe…]]
Q: What does failureThreshold control, and how does it interact with periodSeconds?
A: failureThreshold is the number of consecutive probe failures before Kubernetes takes action. Combined with periodSeconds, it sets the detection window: e.g., failureThreshold=3 and periodSeconds=10 means 30 seconds of failures before restart or endpoint removal.
Remember: <html><code>kubectl explain <resource></code></html> for fields. <html><code>kubectl api-resources</code></html> for all resource types.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-ops.tsv</code></html>
''Related atoms''
* [[What is the constraint on successThreshold for liveness and startup probes, and why doe…]]
* [[What happens when a readiness probe fails on a Kubernetes pod?]]
Q: A pod is stuck in ImagePullBackOff. What are the common causes?
A: 1) Image name or tag is wrong. 2) Image doesn't exist in the registry. 3) ImagePullSecret is missing or expired. 4) Private registry requires auth. 5) Network policy or firewall blocks registry access.
Remember: ImagePullBackOff = can't pull image. Check: typo, auth, network, tag existence.
Gotcha: <html><code>kubectl describe pod</code></html> Events section has the exact error.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-troubleshooting.tsv</code></html>
''Related atoms''
* [[How do you verify that an ImagePullSecret is correct?]]
In Kubernetes RBAC, a resource and its subresources are treated as distinct API objects. Granting <html><code>get</code></html> on <html><code>pods</code></html> does not automatically grant <html><code>get</code></html> on <html><code>pods/log</code></html>, <html><code>pods/exec</code></html>, <html><code>pods/portforward</code></html>, <html><code>pods/attach</code></html>, or any other subresource. Users with apparent pod access will receive 403 Forbidden when attempting <html><code>kubectl logs</code></html> or <html><code>kubectl exec</code></html> if the subresource rule is missing. The same pattern applies to <html><code>deployments/scale</code></html> and <html><code>statefulsets/scale</code></html> — the parent resource grant does not cascade. This is one of the most common mistakes when writing RBAC rules for the first time, because it is natural to assume a resource rule covers everything under that resource. When authoring a Role or ClusterRole, enumerate every subresource the subject needs explicitly. Verify with <html><code>kubectl auth can-i get pods/log -n <namespace></code></html> before deploying the role. Reminder: RBAC involves exactly four object kinds — Role, ClusterRole, RoleBinding, ClusterRoleBinding.
----
''Sources''
* <html><code>training/library/topics/k8s-rbac/footguns.md</code></html>
* <html><code>training/interactive/knowledge/data/cards/k8s-rbac.tsv</code></html>
* <html><code>training/library/topics/k8s-rbac/primer.md</code></html>
//Merged from 2 source atoms.//
By default, every pod receives a mounted service account token at <html><code>/var/run/secrets/kubernetes.io/serviceaccount/token</code></html>. If the pod never calls the Kubernetes API — for example, a stateless web application that only queries a database — the token is still mounted and exposed to any file-read vulnerability in the pod. An attacker who gains code execution can read the token and use it to call the Kubernetes API from outside the cluster, impersonating the service account. This increases the blast radius of any pod vulnerability. The mitigation is to explicitly set <html><code>automountServiceAccountToken: false</code></html> on the ServiceAccount definition, which prevents the token from being mounted by default. Then, only on pods that genuinely need API access, set <html><code>automountServiceAccountToken: true</code></html> in the pod spec to override the account setting. This ensures tokens are only mounted where they are needed, minimizing exposure.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-rbac.tsv</code></html>
* <html><code>training/library/topics/k8s-rbac/footguns.md</code></html>
* <html><code>training/library/topics/k8s-rbac/primer.md</code></html>
* <html><code>training/library/topics/k8s-rbac/street_ops.md</code></html>
//Merged from 4 source atoms.//
A Role grants permissions to resources within a single namespace — pods, services, configmaps within that namespace only. Its rules specify API groups, resource names, and verbs. A ClusterRole grants permissions cluster-wide and uses the same rule syntax; scope is the only structural difference. ClusterRole is required in two scenarios: accessing cluster-scoped resources (nodes, namespaces, persistent volumes) that exist outside any namespace, or granting the same permissions uniformly across all namespaces. A ClusterRole can also be referenced by a namespace-scoped RoleBinding, which constrains its effect to that namespace alone — a common reuse pattern for shared utility roles that many namespaces need without repeating the permission definition. A ClusterRoleBinding applies the ClusterRole globally across the entire cluster. Summary of scope rules: Role = namespaced permissions; ClusterRole = cluster-wide permissions; RoleBinding = namespaced grant; ClusterRoleBinding = global grant. Gotcha: ClusterRoles referenced by RoleBindings are scoped to that namespace, not the whole cluster.
----
''Sources''
* <html><code>training/library/topics/k8s-rbac/primer.md</code></html>
* <html><code>training/interactive/knowledge/data/cards/k8s-rbac.tsv</code></html>
* <html><code>training/interactive/knowledge/data/cards/k8s-security.tsv</code></html>
//Merged from 3 source atoms.//
''Related atoms''
* [[Aggregated ClusterRoles compose permissions via label selectors]]
A RoleBinding grants permissions within one namespace and can reference either a Role or a ClusterRole. A ClusterRoleBinding grants permissions cluster-wide. The critical insight: the scope is determined by the binding type, not the role type. A ClusterRole + RoleBinding grants cluster-wide permissions within a single namespace. A ClusterRole + ClusterRoleBinding grants those permissions everywhere. This flexibility allows you to define a single ClusterRole for, say, "view all resources," then use RoleBindings in namespace A to grant it there and ClusterRoleBindings to grant it globally, without duplicating the rule definitions.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-rbac.tsv</code></html>
* <html><code>training/library/topics/k8s-rbac/primer.md</code></html>
//Merged from 2 source atoms.//
''Related atoms''
* [[Aggregated ClusterRoles compose permissions via label selectors]]
Aggregation composes multiple ClusterRoles into a single effective role using label selectors. A ClusterRole labeled <html><code>rbac.authorization.k8s.io/aggregate-to-view: "true"</code></html> has its rules automatically merged into the built-in <html><code>view</code></html> ClusterRole; equivalent labels exist for <html><code>edit</code></html> and <html><code>admin</code></html>. The built-in <html><code>admin</code></html>, <html><code>edit</code></html>, and <html><code>view</code></html> roles are themselves aggregations — they pull permissions from dozens of smaller labeled roles rather than listing rules directly.
The primary use case is CRD installation: when you add a custom resource, create a corresponding ClusterRole with the appropriate aggregate label, and cluster operators automatically receive the intended access level through their existing role bindings — no manual edits to built-in roles required. Editing built-in roles directly is an anti-pattern; aggregation is the sanctioned extension mechanism.
Scope reminder: Role and RoleBinding are namespace-scoped; ClusterRole and ClusterRoleBinding are cluster-wide. Aggregated ClusterRoles are always cluster-scoped, though they can be bound within a namespace via a namespaced RoleBinding.
----
''Sources''
* <html><code>training/library/topics/k8s-rbac/primer.md</code></html>
* <html><code>training/interactive/knowledge/data/cards/k8s-rbac.tsv</code></html>
//Merged from 2 source atoms.//
''Related atoms''
* [[Role is namespace-scoped; ClusterRole is cluster-wide]]
ClusterRoles with <html><code>apiGroups: ["*"]</code></html>, <html><code>resources: ["*"]</code></html>, or <html><code>verbs: ["*"]</code></html> grant unrestricted access to every current and future API resource in the cluster, including secrets, RBAC objects, and node operations. If a pod bearing these permissions is compromised, the attacker gains full cluster-admin access and can perform lateral movement across namespaces — for example, reading database credentials stored in secrets belonging to other namespaces.
Wildcard permissions must never be granted to workload service accounts. <html><code>cluster-admin</code></html> is reserved for emergency administrative access and trusted operators only.
The correct alternative is explicit enumeration of only the minimum required apiGroups (e.g., <html><code>"apps"</code></html>, <html><code>""</code></html>), resources (e.g., <html><code>"deployments"</code></html>, <html><code>"pods"</code></html>), and verbs (e.g., <html><code>"get"</code></html>, <html><code>"list"</code></html>, <html><code>"create"</code></html>). Use separate roles for read vs. write access. Bind at the narrowest scope: prefer a namespaced <html><code>RoleBinding</code></html> over a cluster-wide <html><code>ClusterRoleBinding</code></html>.
Key scoping reminder: <html><code>Role</code></html> and <html><code>RoleBinding</code></html> are namespace-scoped; <html><code>ClusterRole</code></html> and <html><code>ClusterRoleBinding</code></html> are cluster-wide. Overly broad bindings compound the risk of wildcard rules by removing namespace isolation entirely.
----
''Sources''
* <html><code>training/library/topics/k8s-rbac/primer.md</code></html>
* <html><code>training/interactive/knowledge/data/cards/k8s-rbac.tsv</code></html>
//Merged from 2 source atoms.//
Q: How do you use kubectl auth can-i to debug RBAC permissions?
A: kubectl auth can-i <verb> <resource> -n <namespace> checks your own permissions. Use --as=system:serviceaccount:<ns>:<sa> to impersonate another identity. kubectl auth can-i --list -n <namespace> shows all permissions for the current user in that namespace. This is the primary RBAC debugging tool.
Remember: RBAC = 4 objects: Role, ClusterRole, RoleBinding, ClusterRoleBinding. "2 Roles + 2 Bindings."
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-rbac.tsv</code></html>
Kubernetes' first authorization model was ABAC (Attribute-Based Access Control), which required a static policy file on each API server and a restart to change permissions. RBAC (Role-Based Access Control) was introduced as an alternative in Kubernetes 1.6 (March 2017) and became the default authorization mode in version 1.8. ABAC still exists as a flag (<html><code>--authorization-mode=ABAC</code></html>) for backward compatibility, but it is functionally deprecated — virtually no production clusters use it. RBAC's dynamic, object-based approach is far more suitable for multi-tenant and growing clusters.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-security.tsv</code></html>
* <html><code>training/library/topics/k8s-rbac/trivia.md</code></html>
* <html><code>training/library/topics/k8s-rbac/primer.md</code></html>
//Merged from 4 source atoms.//
Q: What are the standard RBAC verbs in Kubernetes and what API operations do they map to?
A: get (GET single resource), list (GET collection), watch (GET streaming), create (POST), update (PUT full replace), patch (PATCH partial modify), delete (DELETE single), deletecollection (DELETE multiple). Special verbs include bind, escalate, and impersonate.
Remember: RBAC = 4 objects: Role, ClusterRole, RoleBinding, ClusterRoleBinding. "2 Roles + 2 Bindings."
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-rbac.tsv</code></html>
Q: Describe a least-privilege RBAC pattern for a CI/CD deployer service account.
A: Create a namespace-scoped Role with only the verbs and resources needed: get/list/create/update/patch on deployments (apps group), services, and configmaps. Bind it with a RoleBinding to a dedicated ServiceAccount in the CI namespace. Never use ClusterRoleBinding. Never grant delete on namespaces or access to secrets unless specifically required.
Remember: Every namespace gets <html><code>default</code></html> SA. Best practice: dedicated SA per workload.
Gotcha: K8s 1.24+: SA tokens no longer auto-mounted as secrets. Use TokenRequest API.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-rbac.tsv</code></html>
''Related atoms''
* [[Wildcard RBAC rules grant unrestricted access and enable privilege escalation]]
Q: What are the security risks of using the default ServiceAccount and how do you audit for them?
A: The default SA starts with no permissions, but Helm charts or cluster operators may bind roles to it, meaning every pod in that namespace inherits those permissions silently. Audit with: kubectl get rolebindings,clusterrolebindings -A -o json | jq for subjects with name default. Fix by creating dedicated SAs per workload and ensuring default has no bindings beyond discovery.
Remember: Every namespace gets <html><code>default</code></html> SA. Best practice: dedicated SA per workload.
Gotcha: K8s 1.24+: SA tokens no longer auto-mounted as secrets. Use TokenRequest API.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-rbac.tsv</code></html>
''Related atoms''
* [[How to list Service Accounts?]]
* [[Describe a least-privilege RBAC pattern for a CI/CD deployer service account.]]
Q: Explain the escalate and bind verbs. Why are they dangerous?
A: The escalate verb allows a subject to modify a Role or ClusterRole to include permissions they do not already hold — bypassing the normal RBAC escalation prevention. The bind verb allows creating RoleBindings that reference roles the subject could not otherwise grant. Together they enable privilege escalation. Never grant these verbs unless the subject genuinely manages RBAC for the cluster.
Remember: RBAC = 4 objects: Role, ClusterRole, RoleBinding, ClusterRoleBinding. "2 Roles + 2 Bindings."
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-rbac.tsv</code></html>
Q: A developer reports they cannot exec into pods despite having pod access. Walk through your debugging process.
A: 1) Check what they can do: kubectl auth can-i create pods/exec -n <ns> --as=<user>. 2) pods/exec is a subresource separate from pods. The role must explicitly include pods/exec with the create verb. 3) Check their RoleBindings: kubectl get rolebindings -n <ns> -o json and inspect roleRef. 4) Inspect the referenced Role for pods/exec rules. 5) Fix by adding a rule for resources: ["pods/exec"] with verbs: ["create"]. 6) Verify with auth can-i.
Remember: RBAC = 4 objects: Role, ClusterRole, RoleBinding, ClusterRoleBinding. "2 Roles + 2 Bindings."
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-rbac.tsv</code></html>
Q: True or False? If no network policies are applied to a pod, then no connections to or from it are allowed
A: False. By default, pods are non-isolated — all ingress and egress traffic is allowed. Network policies only take effect when explicitly applied, and they use a whitelist model: once any policy selects a pod, all traffic not explicitly allowed is denied.
Gotcha: NetworkPolicy requires a CNI plugin that supports it (Calico, Cilium), not all do.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-security.tsv</code></html>
''Related atoms''
* [[Kubernetes NetworkPolicy: whitelist semantics and operational traps]]
* [[What are some use cases for using Network Policies?]]
Q: How does Kubernetes manage security, and what are some best practices?
A: ''Kubernetes Security Management:''
• Role-Based Access Control (RBAC): Defines and enforces access policies.
• Pod Security Policies: Restricts pod behaviors for security compliance.
• Network Policies: Controls communication between pods.
• Secrets Management: Safely handles sensitive information.
• Container Runtime Security: Ensures container runtime security practices.
• Security Contexts: Defines security settings at the pod or container level.
Remember: K8s security layers: RBAC (who), NetworkPolicy (what), PSA (how), encryption (data).
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-security.tsv</code></html>
//Merged from 3 source atoms.//
''Related atoms''
* [[What security best practices do you follow in regards to the Kubernetes cluster?]]
* [[Explain "Security Context"]]
Q: What is PodSecurity and how can it be configured in a Kubernetes cluster?
A: * PodSecurity in Kubernetes: PodSecurity refers to policies and configurations that control the security context of pods.
* It includes settings related to running as a privileged user, allowing privileged containers, and more.
* PodSecurityPolicy was removed in K8s 1.25. Use the built-in Pod Security Admission controller with labels (enforce/audit/warn) and Pod Security Standards (privileged/baseline/restricted).
Remember: K8s security layers: RBAC (who), NetworkPolicy (what), PSA (how), encryption (data).
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-security.tsv</code></html>
//Merged from 2 source atoms.//
''Related atoms''
* [[What happens you create a pod and you DON'T specify a service account?]]
* [[How do you implement encryption for data in transit and at rest in Kubernetes?]]
* [[Kubernetes Secrets: storage, access, and best practices]]
Gatekeeper is a validating (mutating TBA) webhook that enforces CRD-based policies executed by Open Policy Agent (OPA). On every request sent to the Kubernetes API server, Gatekeeper forwards the request and the applicable policies to OPA to evaluate whether any policy is violated. If a violation is found, Gatekeeper returns the policy error message and the request is rejected before reaching the cluster. If no violation is found, the request proceeds normally.
Gatekeeper docs: https://open-policy-agent.github.io/gatekeeper/website/docs
Context — Kubernetes security layers:
* RBAC: who can act
* NetworkPolicy: what can communicate
* PSA (Pod Security Admission): how pods run
* Encryption: data at rest/in transit
* OPA Gatekeeper: policy-as-code enforcement at admission time
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-security.tsv</code></html>
//Merged from 2 source atoms.//
Q: What are some use cases for using Network Policies?
A: - Security: You want to prevent from everyone to communicate with a certain pod for security reasons
** Controlling network traffic: You would like to deny network flow between two specific nodes
Remember: K8s security layers: RBAC (who), NetworkPolicy (what), PSA (how), encryption (data).
Example: Isolate a database pod so only the API pod can reach it: NetworkPolicy with ingress from pods labeled app=api.
Remember: NetworkPolicies are additive — multiple policies on the same pod combine their allowed traffic.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-security.tsv</code></html>
''Related atoms''
* [[True or False? If no network policies are applied to a pod, then no connections to or f…]]
Q: Explain how Service Accounts are different from User Accounts
A: - User accounts are global while Service accounts unique per namespace
** User accounts are meant for humans or client processes while Service accounts are for processes which run in pods
Remember: K8s security layers: RBAC (who), NetworkPolicy (what), PSA (how), encryption (data).
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-security.tsv</code></html>
''Related atoms''
* [[What happens you create a pod and you DON'T specify a service account?]]
* [[How to list Service Accounts?]]
* [[Explain what are "Service Accounts" and in which scenario would use create/use one]]
Q: What is the purpose of an admission controller in Kubernetes, and how can you extend it?
A: * Admission Controller in Kubernetes:
* Admission controllers validate and mutate Kubernetes resources before they are persisted.
They enforce policies and security measures.
* Extending Admission Controllers:
* Custom admission controllers can be created to enforce specific organization or application-specific policies.
* Use the Kubernetes admission webhook mechanism to extend admission control.
Remember: Admission controllers intercept API requests. Types: validating and mutating.
Example: LimitRanger, PodSecurity, OPA/Gatekeeper — common admission controllers.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-security.tsv</code></html>
''Related atoms''
* [[OPA Gatekeeper: policy enforcement webhook in Kubernetes]]
* [[How does Kubernetes manage containerized applications?]]
Q: What is Conftest and how does it validate configuration files?
A: Conftest allows you to write tests against structured files. You can think of it as tests library for Kubernetes resources.
It is mostly used in testing environments such as CI pipelines or local hooks.
Remember: K8s security layers: RBAC (who), NetworkPolicy (what), PSA (how), encryption (data).
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-security.tsv</code></html>
''Related atoms''
* [[What is Datree? How is it different from Conftest?]]
Q: What happens you create a pod and you DON'T specify a service account?
A: The pod is automatically assigned with the default service account (in the namespace where the pod is running).
Remember: K8s security layers: RBAC (who), NetworkPolicy (what), PSA (how), encryption (data).
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-security.tsv</code></html>
''Related atoms''
* [[What is PodSecurity and how can it be configured in a Kubernetes cluster?]]
* [[Explain what are "Service Accounts" and in which scenario would use create/use one]]
* [[Service account tokens are mounted by default even if the pod never calls the Kubernetes API]]
Q: How to list Service Accounts?
A: <html><code>kubectl get serviceaccounts</code></html>
Remember: K8s security layers: RBAC (who), NetworkPolicy (what), PSA (how), encryption (data).
Remember: <html><code>kubectl get sa</code></html> (short for serviceaccounts). Every namespace has a default SA created automatically.
Gotcha: The default SA may have more permissions than you expect — always audit RoleBindings referencing it.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-security.tsv</code></html>
''Related atoms''
* [[What are the security risks of using the default ServiceAccount and how do you audit fo…]]
* [[Explain how Service Accounts are different from User Accounts]]
Q: What is Datree? How is it different from Conftest?
A: Same as Conftest, it is used for policy testing and enforcement. The difference is that it comes with built-in policies.
Remember: K8s security layers: RBAC (who), NetworkPolicy (what), PSA (how), encryption (data).
Remember: Datree = Conftest + built-in policy library. Conftest = bring your own policies (Rego). Datree is easier to start; Conftest is more flexible.
Gotcha: Both tools validate YAML before deployment — they are not runtime enforcers like OPA Gatekeeper.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-security.tsv</code></html>
''Related atoms''
* [[What is Conftest and how does it validate configuration files?]]
Q: How do you implement encryption for data in transit and at rest in Kubernetes?
A: * Encryption in Kubernetes:
* Data in Transit: Use Transport Layer Security (TLS) for encrypting communication between components and pods.
* Data at Rest: Leverage storage providers that support encryption or use tools like dm-crypt for node-level encryption.
* For secrets, use encryption mechanisms provided by Kubernetes, such as sealed secrets.
Remember: K8s security layers: RBAC (who), NetworkPolicy (what), PSA (how), encryption (data).
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-security.tsv</code></html>
''Related atoms''
* [[How does Kubernetes manage security, and what are some best practices?]]
* [[What is PodSecurity and how can it be configured in a Kubernetes cluster?]]
* [[Kubernetes Secrets: storage, access, and best practices]]
Q: Explain what are "Service Accounts" and in which scenario would use create/use one
A: [[Kubernetes.io|https://kubernetes.io/docs/tasks/configure-pod-container/configure-service-account]]: "A service account provides an identity for processes that run in a Pod."
An example of when to use one:
You define a pipeline that needs to build and push an image. In order to have sufficient permissions to build an push an image, that pipeline would require a service account with sufficient permissions.
Remember: K8s security layers: RBAC (who), NetworkPolicy (what), PSA (how), encryption (data).
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-security.tsv</code></html>
''Related atoms''
* [[What happens you create a pod and you DON'T specify a service account?]]
* [[Explain how Service Accounts are different from User Accounts]]
Q: Which Kubernetes concept would you use to control traffic flow at the IP address or port level?
A: Network Policies
Remember: K8s security layers: RBAC (who), NetworkPolicy (what), PSA (how), encryption (data).
Remember: NetworkPolicy = L3/L4 firewall for pods. Controls IP and port-level traffic between pods and external endpoints.
Gotcha: NetworkPolicies require a CNI plugin that supports them (Calico, Cilium). Flannel does NOT enforce them.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-security.tsv</code></html>
''Related atoms''
* [[What are some use cases for using Network Policies?]]
* [[CNI plugin: Kubernetes container networking standard]]
Q: What security best practices do you follow in regards to the Kubernetes cluster?
A: * Secure inter-service communication (one way is to use Istio to provide mutual TLS)
* Isolate different resources into separate namespaces based on some logical groups
* Use supported container runtime (if you use Docker then drop it because it's deprecated. You might want to CRI-O as an engine and podman for CLI)
* Test properly changes to the cluster (e.g. consider using Datree to prevent kubernetes misconfigurations)
* Limit who can do what (by using for example OPA gatekeeper) in the cluster
* Use NetworkPolicy to apply network security
* Consider using tools (e.g.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-security.tsv</code></html>
''Related atoms''
* [[How does Kubernetes manage security, and what are some best practices?]]
* [[What actions or operations you consider as best practices when it comes to Kubernetes?]]
* [[Do you have experience with deploying a Kubernetes cluster? If so, can you describe the…]]
Q: Explain "Security Context"
A: [[kubernetes.io|https://kubernetes.io/docs/tasks/configure-pod-container/security-context]]: "A security context defines privilege and access control settings for a Pod or Container."
Gotcha: PSP removed in K8s 1.25. Use Pod Security Admission (PSA): enforce, audit, warn.
Remember: PSA levels: Privileged, Baseline, Restricted. Mnemonic: "PBR."
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-security.tsv</code></html>
''Related atoms''
* [[How does Kubernetes manage security, and what are some best practices?]]
* [[What happens you create a pod and you DON'T specify a service account?]]
The LoadBalancer service type exposes a Kubernetes service externally by requesting the cloud provider (AWS ELB/NLB, GCP LB, Azure LB) to provision an external load balancer. It builds atop NodePort — Kubernetes creates both the NodePort rules and the cloud load balancer infrastructure, resulting in a single external IP that forwards all traffic to the service. The external IP shows <html><code><pending></code></html> while provisioning; once assigned, external clients connect directly to the load balancer. On bare-metal clusters without a cloud provider, the external IP remains <html><code><pending></code></html> indefinitely; MetalLB or a similar solution is required for on-premises use. Cost scales with the number of load balancers provisioned. This is the standard way to expose services to the internet without running an in-cluster ingress controller.
----
''Sources''
* <html><code>training/library/topics/k8s-services-and-ingress/primer.md</code></html>
* <html><code>training/interactive/knowledge/data/cards/k8s-services.tsv</code></html>
//Merged from 2 source atoms.//
<html><code>kubectl expose</code></html> creates a Service from an existing workload resource (Deployment, ReplicaSet, Pod, etc.).
''Syntax:''
<html><pre><code class="language-plaintext">kubectl expose <resource-type> <name> \
[--name=<service-name>] \
--port=<service-port> \
--target-port=<container-port> \
--type=<ClusterIP|NodePort|LoadBalancer></code></pre></html>
''Key distinctions:''
* <html><code>--port</code></html> — the port the Service listens on (cluster-facing).
* <html><code>--target-port</code></html> — the port the container is actually listening on. These two values are independent and often differ.
* <html><code>--type</code></html> defaults to <html><code>ClusterIP</code></html>; use <html><code>NodePort</code></html> or <html><code>LoadBalancer</code></html> when external access is needed.
* <html><code>--name</code></html> is required when exposing a ReplicaSet (or any resource where the generated name would clash); optional for Deployments.
''Examples:''
<html><pre><code class="language-bash"># Expose a Deployment via LoadBalancer on port 8080
kubectl expose deployment alle --type=LoadBalancer --port=8080
# Expose a Deployment: service port 80, container port 8080
kubectl expose deployment web --port=80 --target-port=8080 --type=ClusterIP
# Expose a ReplicaSet as a named NodePort service
kubectl expose rs my-rs --name=my-svc --target-port=8080 --type=NodePort</code></pre></html>
The target-port value must match the port the application process binds inside the container.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-services.tsv</code></html>
//Merged from 3 source atoms.//
''Related atoms''
* [[True or False? the target port, in the case of running the following command, will be e…]]
Q: How to list Ingress in your namespace?
A: <html><code>kubectl get ingress</code></html> lists all Ingress resources in the current namespace showing hosts, paths, and backends. Add <html><code>-o wide</code></html> for additional details like the ingress class and address.
Gotcha: Ingress requires an Ingress Controller (nginx, traefik, ALB) to be installed — without one, Ingress resources exist but do not route traffic.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-services.tsv</code></html>
''Related atoms''
* [[Ingress 404: controller missing, wrong ingressClassName, or no endpoints]]
* [[Complete the following configuration file to make it Ingress]]
ClusterIP is the default Kubernetes Service type. It exposes the Service on a virtual IP that is reachable only from within the cluster, providing no external access. It is the correct choice for service-to-service communication (e.g., frontend pods calling backend pods).
Under the hood, kube-proxy programs iptables or IPVS rules to load-balance traffic across the matching pods. Within the cluster, the Service is reachable by its DNS name: <html><code>http://<service>.<namespace>.svc.cluster.local</code></html>.
The four Service types are ClusterIP, NodePort, LoadBalancer, and ExternalName (mnemonic: CNLE). To expose a ClusterIP Service externally, promote it to NodePort or LoadBalancer instead.
Key facts:
* Default type when <html><code>spec.type</code></html> is omitted.
* Allocates an internal cluster IP; pods outside the cluster cannot reach it directly.
* Ideal for internal microservice communication while keeping services hidden from external networks.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-services.tsv</code></html>
* <html><code>training/library/topics/k8s-services-and-ingress/primer.md</code></html>
//Merged from 3 source atoms.//
''Related atoms''
* [[Kubernetes Service: stable networking endpoint for pods]]
* [[Pod IPs are ephemeral; use a Service for stable addressing]]
* [[Kubernetes headless services return pod IPs directly via DNS]]
Q: How readiness probe status affect Services when they are combined?
A: Only containers whose state set to Success will be able to receive requests sent to the Service.
Remember: Services use label selectors to find pods. Change labels→change which pods get traffic.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-services.tsv</code></html>
''Related atoms''
* [[How do readiness probes interact with rolling deployments?]]
* [[What happens when a readiness probe fails on a Kubernetes pod?]]
Q: How to configure TLS with Ingress?
A: Add tls and secretName entries.
<html><pre><code class="language-plaintext">spec:
tls:
- hosts:
- some_app.com
secretName: someapp-secret-tls</code></pre></html>
Remember: Services use label selectors to find pods. Change labels→change which pods get traffic.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-services.tsv</code></html>
''Related atoms''
* [[Complete the following configuration file to make it Ingress]]
Q: How make an app accessible on private or external network?
A: Using a Kubernetes Service, which provides a stable virtual IP and DNS name that routes traffic to a set of pods matched by label selectors. Four types: ClusterIP (internal), NodePort (external via node ports), LoadBalancer (cloud LB), ExternalName (DNS alias). Services decouple consumers from individual pod IPs, which change on restart.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-services.tsv</code></html>
''Related atoms''
* [[Kubernetes Service: stable networking endpoint for pods]]
Q: How would you map a service to an external address?
A: Use the ExternalName service type, which maps a service to an external DNS name (e.g., an RDS endpoint or external API).
Example: <html><code>type: ExternalName</code></html> with <html><code>externalName: my.database.example.com</code></html>. This creates a CNAME record in cluster DNS.
Gotcha: ExternalName services do not proxy traffic — they only provide DNS resolution.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-services.tsv</code></html>
Q: Describe what happens when a container tries to connect with its corresponding Service for the first time. Explain who added each of the components you include in your description
A: - The container looks at the nameserver defined in /etc/resolv.conf
** The container queries the nameserver so the address is resolved to the Service IP
** Requests sent to the Service IP are forwarded with iptables rules (or other chosen software) to the endpoint(s).
Explanation as to who added them:
** The nameserver in the container is added by kubelet during the scheduling of the Pod, by using kube-dns
** The DNS record of the service is added by kube-dns during the Service creation
** iptables rules are added by kube-proxy during Endpoint and Service creation
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-services.tsv</code></html>
''Related atoms''
* [[How does Kubernetes handle DNS resolution for services and pods?]]
* [[Kubernetes DNS automatically assigns and resolves names to services and pods]]
* [[Discuss the role of kube-proxy in Kubernetes networking.]]
Q: What are important steps in defining/adding a Service?
A: 1. Making sure that targetPort of the Service is matching the containerPort of the Pod
# Making sure that selector matches at least one of the Pod's labels
Remember: port=Service port, targetPort=container port, nodePort=external. Three distinct values.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-services.tsv</code></html>
''Related atoms''
* [[How to turn the following service into an external one?]]
* [[How to create a pod and a service with one command?]]
* [[kubectl expose: create a Service from a workload resource]]
Q: After creating a service that forwards incoming external traffic to the containerized application, how to make sure it works?
A: You can run <html><code>curl <SERVICE IP>:<SERVICE PORT></code></html> to examine the output.
Remember: Services use label selectors to find pods. Change labels→change which pods get traffic.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-services.tsv</code></html>
''Related atoms''
* [[How to turn the following service into an external one?]]
* [[How to verify that a certain service configured to forward the requests to a given pod]]
* [[Explain what will happen when running apply on the following block]]
A Kubernetes Service is an abstraction that exposes an application running on a set of Pods as a network service, providing a stable IP address and DNS name that persists regardless of individual pod lifecycles. It enables both internal (cluster-scoped) and external connectivity to containerized applications.
Services use label selectors to dynamically determine which pods receive traffic. Modifying a pod's labels changes which Service(s) route traffic to it — making label management a key operational concern.
Without a Service, pods must be addressed by ephemeral IP addresses that change on restart. The Service layer decouples consumers from the volatile pod layer, enabling reliable service discovery and load distribution across all matching pods.
Reference: https://kubernetes.io/docs/concepts/services-networking/service
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-services.tsv</code></html>
//Merged from 2 source atoms.//
''Related atoms''
* [[Service and Ingress divide Layer 4 and Layer 7 routing]]
* [[How does Kubernetes handle DNS resolution for services and pods?]]
* [[ClusterIP: Kubernetes default internal-only Service type]]
Q: How to create a pod and a service with one command?
A: kubectl run nginx --image=nginx --restart=Never --port 80 --expose
Example: <html><code>kubectl expose deployment web --port=80 --target-port=8080 --type=ClusterIP</code></html>
Remember: <html><code>--port</code></html>=Service port, <html><code>--target-port</code></html>=container port. Two distinct values.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-services.tsv</code></html>
''Related atoms''
* [[kubectl expose: create a Service from a workload resource]]
* [[How can you find out information on a Service related to a certain Pod if all you can u…]]
* [[Deploy a pod called "my-pod" using the nginx:alpine image]]
Q: True or False? the target port, in the case of running the following command, will be exposed only on one of the Kubernetes cluster nodes but it will routed to all the pods
A: False. It will be exposed on every node of the cluster and will be routed to one of the Pods (which belong to the ReplicaSet)
Example: <html><code>kubectl expose deployment web --port=80 --target-port=8080 --type=ClusterIP</code></html>
Remember: <html><code>--port</code></html>=Service port, <html><code>--target-port</code></html>=container port. Two distinct values.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-services.tsv</code></html>
''Related atoms''
* [[kubectl expose: create a Service from a workload resource]]
* [[What the following command does?]]
* [[Explain what will happen when running apply on the following block]]
Q: Explain what will happen when running apply on the following block
A: It creates a new Service of the type "NodePort" which means it can be used for internal and external communication with the app.
The port of the application is 8080 and the requests will forwarded to this port. The exposed port is 2017. As a note, this is not a common practice, to specify the nodePort.
The port used TCP (instead of UDP) and this is also the default so you don't have to specify it.
The selector used by the Service to know to which Pods to forward the requests. In this case, Pods with the label "type: backend" and "service: some-app".
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-services.tsv</code></html>
''Related atoms''
* [[True or False? the target port, in the case of running the following command, will be e…]]
* [[kubectl expose: create a Service from a workload resource]]
* [[After creating a service that forwards incoming external traffic to the containerized a…]]
Q: How to get information on a certain service?
A: <html><code>kubectl describe service <SERVICE_NAME></code></html>
It's more common to use <html><code>kubectl describe svc ...</code></html>
Remember: Services use label selectors to find pods. Change labels→change which pods get traffic.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-services.tsv</code></html>
''Related atoms''
* [[How can you find out information on a Service related to a certain Pod if all you can u…]]
* [[How do you find pods that match a particular label selector?]]
* [[List all the pods with the label "env=prod"]]
Q: In case of two pods, if there is an egress policy on the source denining traffic and ingress policy on the destination that allows traffic then, traffic will be allowed or denied?
A: Denied. Both source and destination policies has to allow traffic for it to be allowed.
Remember: Services use label selectors to find pods. Change labels→change which pods get traffic.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-services.tsv</code></html>
''Related atoms''
* [[True or False? If no network policies are applied to a pod, then no connections to or f…]]
* [[What happens when you apply a NetworkPolicy with an empty podSelector and policyTypes: …]]
Q: How to turn the following service into an external one?
A: Adding <html><code>type: LoadBalancer</code></html> and <html><code>nodePort</code></html>
<html><pre><code class="language-plaintext">spec:
selector:
app: some-app
type: LoadBalancer
ports:
- protocol: TCP
port: 8081
targetPort: 8081
nodePort: 32412</code></pre></html>
Remember: Services use label selectors to find pods. Change labels→change which pods get traffic.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-services.tsv</code></html>
''Related atoms''
* [[After creating a service that forwards incoming external traffic to the containerized a…]]
* [[How to configure a default backend?]]
* [[kubectl expose: create a Service from a workload resource]]
Q: What are some use cases for using Ingress?
A: * Multiple sub-domains (multiple host entries, each with its own service)
* One domain with multiple services (multiple paths where each one is mapped to a different service/application)
Remember: Services use label selectors to find pods. Change labels→change which pods get traffic.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-services.tsv</code></html>
''Related atoms''
* [[Service and Ingress divide Layer 4 and Layer 7 routing]]
* [[What problem does Ingress solve that Services alone cannot?]]
* [[What is Ingress Default Backend?]]
Q: How to list the endpoints of a certain app?
A: <html><code>kubectl get ep <name></code></html>
Under the hood: Endpoints object tracks matching pod IPs. No pods = empty Endpoints = no traffic.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-services.tsv</code></html>
''Related atoms''
* [[How to verify that a certain service configured to forward the requests to a given pod]]
* [[List all the pods with the label "env=prod"]]
Q: An internal load balancer in Kubernetes is called ____ and an external load balancer is called ____
A: An internal load balancer in Kubernetes is called Service and an external load balancer is Ingress
Remember: Services use label selectors to find pods. Change labels→change which pods get traffic.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-services.tsv</code></html>
''Related atoms''
* [[What is an "Ingress"?]]
* [[Service and Ingress divide Layer 4 and Layer 7 routing]]
* [[ClusterIP: Kubernetes default internal-only Service type]]
Q: Why using a wildcard in ingress host may lead to issues?
A: The reason you should not wildcard value in a host (like <html><code>- host: *</code></html>) is because you basically tell your Kubernetes cluster to forward all the traffic to the container where you used this ingress. This may cause the entire cluster to go down.
Remember: Services use label selectors to find pods. Change labels→change which pods get traffic.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-services.tsv</code></html>
''Related atoms''
* [[What are some use cases for using Ingress?]]
* [[What is Ingress Default Backend?]]
* [[What happens when you apply a NetworkPolicy with an empty podSelector and policyTypes: …]]
Q: Complete the following configuration file to make it Ingress
A: There are several ways to answer this question.
<html><pre><code class="language-plaintext">apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
name: someapp-ingress
spec:
rules:
- host: my.host
http:
paths:
- path: /
pathType: Prefix
backend:
service:
name: someapp-internal-service
port:
number: 8080</code></pre></html>
Remember: Services use label selectors to find pods. Change labels→change which pods get traffic.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-services.tsv</code></html>
''Related atoms''
* [[How to configure a default backend?]]
* [[Ingress 404: controller missing, wrong ingressClassName, or no endpoints]]
Q: What is Ingress Controller?
A: An implementation for Ingress. It's basically another pod (or set of pods) that does evaluates and processes Ingress rules and this it manages all the redirections.
There are multiple Ingress Controller implementations (the one from Kubernetes is Kubernetes Nginx Ingress Controller).
Remember: Services use label selectors to find pods. Change labels→change which pods get traffic.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-services.tsv</code></html>
''Related atoms''
* [[What is Ingress Default Backend?]]
Q: What is an "Ingress"?
A: Manages external access to services, typically via HTTP.
Remember: Services use label selectors to find pods. Change labels→change which pods get traffic.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-services.tsv</code></html>
''Related atoms''
* [[Service and Ingress divide Layer 4 and Layer 7 routing]]
* [[An internal load balancer in Kubernetes is called ____ and an external load balancer is…]]
Q: How can you get a static IP for a Kubernetes load balancer?
A: A static IP for the Kubernetes load balancer can be achieved by changing DNS records since the Kubernetes Master can assign a new static IP address. You can also specify a loadBalancerIP in your service spec when using cloud providers that support it, or reserve a static IP through your cloud provider and then reference it in your service configuration.
Remember: Services use label selectors to find pods. Change labels→change which pods get traffic.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-services.tsv</code></html>
''Related atoms''
* [[Pod IPs are ephemeral; use a Service for stable addressing]]
* [[Kubernetes headless services return pod IPs directly via DNS]]
Q: When would you use the "LoadBalancer" type
A: Mostly when you would like to combine it with cloud provider's load balancer
Remember: LoadBalancer = NodePort + cloud LB. Auto-provisions on cloud providers.
Gotcha: Bare-metal → stays Pending. Use MetalLB for on-prem.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-services.tsv</code></html>
''Related atoms''
* [[Service type LoadBalancer provisions cloud load balancer]]
* [[Describe in high level what happens when you run kubectl expose deployment remo --type=…]]
Q: Describe in high level what happens when you run kubectl expose deployment remo --type=LoadBalancer --port 8080
A: 1. Kubectl sends a request to Kubernetes API to create a Service object
# Kubernetes asks the cloud provider (e.g. AWS, GCP, Azure) to provision a load balancer
# The newly created load balancer forwards incoming traffic to relevant worker node(s) which forwards the traffic to the relevant containers
Remember: LoadBalancer = NodePort + cloud LB. Auto-provisions on cloud providers.
Gotcha: Bare-metal → stays Pending. Use MetalLB for on-prem.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-services.tsv</code></html>
''Related atoms''
* [[kubectl expose: create a Service from a workload resource]]
* [[What happens when a readiness probe fails on a Kubernetes pod?]]
* [[Service type LoadBalancer provisions cloud load balancer]]
Q: Describe in detail what happens when you create a service
A: 1. Kubectl sends a request to the API server to create a Service
# The controller detects there is a new Service
# Endpoint objects created with the same name as the service, by the controller
# The controller is using the Service selector to identify the endpoints
# kube-proxy detects there is a new endpoint object + new service and adds iptables rules to capture traffic to the Service port and redirect it to endpoints
# kube-dns detects there is a new Service and adds the container record to the dns server
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-services.tsv</code></html>
''Related atoms''
* [[How does Kubernetes handle DNS resolution for services and pods?]]
* [[How does a Kubernetes Service find the right Pods to route traffic to?]]
Q: How to verify that a certain service configured to forward the requests to a given pod
A: Run <html><code>kubectl describe service</code></html> and see if the IPs from "Endpoints" match any IPs from the output of <html><code>kubectl get pod -o wide</code></html>
Remember: Services use label selectors to find pods. Change labels→change which pods get traffic.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-services.tsv</code></html>
''Related atoms''
* [[After creating a service that forwards incoming external traffic to the containerized a…]]
* [[Pod IPs are ephemeral; use a Service for stable addressing]]
Q: What is Ingress Default Backend?
A: It specifies what do with an incoming request to the Kubernetes cluster that isn't mapped to any backend (= no rule to for mapping the request to a service). If the default backend service isn't defined, it's recommended to define so users still see some kind of message instead of nothing or unclear error.
Remember: Services use label selectors to find pods. Change labels→change which pods get traffic.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-services.tsv</code></html>
''Related atoms''
* [[What is Ingress Controller?]]
* [[Service and Ingress divide Layer 4 and Layer 7 routing]]
Q: How to configure a default backend?
A: Create Service resource that specifies the name of the default backend as reflected in <html><code>kubectl describe ingress ...</code></html> and the port under the ports section.
Remember: Services use label selectors to find pods. Change labels→change which pods get traffic.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-services.tsv</code></html>
''Related atoms''
* [[Complete the following configuration file to make it Ingress]]
* [[How to turn the following service into an external one?]]
* [[Explain the meaning of "http", "host" and "backend" directives]]
Q: Explain the meaning of "http", "host" and "backend" directives
A: host is the entry point of the cluster so basically a valid domain address that maps to cluster's node IP address
the http line used for specifying that incoming requests will be forwarded to the internal service using http.
backend is referencing the internal service (serviceName is the name under metadata and servicePort is the port under the ports section).
Remember: Services use label selectors to find pods. Change labels→change which pods get traffic.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-services.tsv</code></html>
''Related atoms''
* [[What is Ingress Default Backend?]]
* [[Explain what will happen when running apply on the following block]]
Q: What would you use to route traffic from outside the Kubernetes cluster to services within a cluster?
A: Ingress. It exposes HTTP/HTTPS routes to services inside the cluster, with support for host/path-based routing, TLS termination, and load balancing — managed by an Ingress Controller (nginx, traefik, etc.).
Remember: Services use label selectors to find pods. Change labels→change which pods get traffic.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-services.tsv</code></html>
''Related atoms''
* [[Service and Ingress divide Layer 4 and Layer 7 routing]]
Q: What is a Kubernetes StatefulSet and when would you use it?
A: StatefulSet is the workload API object used to manage stateful applications. Manages the deployment and scaling of a set of Pods, and provides guarantees about the ordering and uniqueness of these Pods.[[Learn more|https://kubernetes.io/docs/concepts/workloads/controllers/statefulset/]]
Remember: StatefulSet = stable identity: pod-0/pod-1, stable storage, ordered operations.
Gotcha: Deleting StatefulSet does NOT delete PVCs. By design for data safety.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-workloads.tsv</code></html>
//Merged from 2 source atoms.//
''Related atoms''
* [[StatefulSets and volumeClaimTemplates: stable storage per replica]]
* [[Deleting a Deployment (not its pods) is the correct removal path]]
Kubernetes storage is organized into three interdependent layers. PersistentVolumes (PVs) are cluster-level resources representing actual storage. PersistentVolumeClaims (PVCs) are namespace-scoped requests for storage made by developers or workloads. StorageClasses define provisioner, provider-specific parameters, reclaim policy, and volume binding mode.
Pods never reference PVs directly; they reference PVCs, and Kubernetes binds PVCs to suitable PVs based on access mode, storage class, capacity, and label selectors. A bound PV-PVC relationship is exclusive — no other PVC can claim that PV.
Two provisioning models exist. In static provisioning, administrators create PVs ahead of time (analogous to allocating disk in a data center). In dynamic provisioning — introduced in Kubernetes 1.4 via StorageClasses — a PVC referencing a StorageClass triggers automatic PV creation, eliminating the admin bottleneck and enabling a self-service model.
This architecture cleanly separates concerns: storage requests (what a workload needs) from storage provisioning (how storage is created and governed). Administrators control capacity and performance tiers through PVs and StorageClasses; developers simply claim what they need through PVCs. The binding chain is: StorageClass → PV → PVC → Pod mount.
----
''Sources''
* <html><code>training/library/topics/k8s-storage/primer.md</code></html>
* <html><code>training/library/topics/k8s-storage/trivia.md</code></html>
* <html><code>training/interactive/knowledge/data/cards/k8s-storage.tsv</code></html>
//Merged from 3 source atoms.//
''Related atoms''
* [[Kubernetes dynamic provisioning and WaitForFirstConsumer zone awareness]]
* [[StatefulSets and volumeClaimTemplates: stable storage per replica]]
Dynamic provisioning eliminates manual PV creation. When a pod references a PVC, Kubernetes checks for a matching PV; if none exists, it invokes the CSI driver named in the StorageClass to create one automatically. The full flow: (1) Pod references a PVC. (2) PVC references a StorageClass. (3) Kubernetes calls the CSI driver named in the StorageClass. (4) The driver creates the underlying volume. (5) A PV is automatically created and bound to the PVC. (6) The volume is mounted into the pod.
The StorageClass <html><code>volumeBindingMode</code></html> controls provisioning timing. <html><code>Immediate</code></html> provisions as soon as the PVC is created. <html><code>WaitForFirstConsumer</code></html> delays provisioning until a pod is actually scheduled to use the volume. <html><code>WaitForFirstConsumer</code></html> is the safer choice on cloud providers: it ensures the volume is created in the same availability zone as the scheduled node, avoiding cross-zone latency penalties and data-transfer charges. A PVC in Pending state under this mode is normal behavior, not a failure — the volume is intentionally deferred until pod scheduling resolves the target zone.
Storage abstraction hierarchy: StorageClass → PV → PVC → Pod mount.
----
''Sources''
* <html><code>training/library/topics/k8s-storage/primer.md</code></html>
* <html><code>training/interactive/knowledge/data/cards/k8s-storage.tsv</code></html>
//Merged from 2 source atoms.//
''Related atoms''
* [[Kubernetes storage abstraction: PV, PVC, and StorageClass layers]]
* [[A PVC is stuck in Pending state. Walk through your debugging process.]]
The Container Storage Interface (CSI) defines the standard protocol for connecting storage systems to Kubernetes. Each cloud and storage vendor provides a CSI driver implemented as two components: a controller Deployment (handles provisioning, attachment, and snapshots) and a node DaemonSet (handles mount/unmount on every node).
Well-known drivers include <html><code>ebs.csi.aws.com</code></html> (AWS block storage), <html><code>efs.csi.aws.com</code></html> (AWS shared NFS), <html><code>pd.csi.storage.gke.io</code></html> (GCE persistent disk), <html><code>disk.csi.azure.com</code></html>, and <html><code>file.csi.azure.com</code></html>. The driver translates StorageClass parameters into provider-specific API calls — for example, EBS IOPS and throughput settings.
To verify CSI driver health:
* <html><code>kubectl get pods -n kube-system -l app=<csi-controller></code></html> — confirm controller pods are running
* <html><code>kubectl get csinodes</code></html> — shows which drivers are registered per node
* <html><code>kubectl get csidrivers</code></html> — shows cluster-wide driver registrations
* Inspect driver logs for provisioning failures
Unhealthy drivers manifest as PVCs stuck in Pending state or pod mount failures. The broader storage abstraction chain is: StorageClass → PV → PVC → Pod mount; dynamic provisioning automates PV creation via the CSI controller.
----
''Sources''
* <html><code>training/library/topics/k8s-storage/primer.md</code></html>
* <html><code>training/interactive/knowledge/data/cards/k8s-storage.tsv</code></html>
//Merged from 2 source atoms.//
''Related atoms''
* [[Volume Snapshots in Kubernetes]]
StatefulSets use <html><code>volumeClaimTemplates</code></html> to provision one PersistentVolumeClaim per replica. Each PVC follows the naming pattern <html><code><template-name>-<statefulset-name>-<ordinal></code></html>, producing stable identifiers such as <html><code>pgdata-postgres-0</code></html> and <html><code>pgdata-postgres-1</code></html>. When a StatefulSet pod is rescheduled, it reattaches to its existing PVC by name — this is how stateful workloads like databases survive pod failure and rollout.
Scaling down does not delete PVCs. Deleting the StatefulSet itself does not delete PVCs. PVCs must be manually deleted to reclaim storage. Scaling back up reattaches to existing PVCs by ordinal.
The <html><code>WaitForFirstConsumer</code></html> binding mode is almost always correct for StatefulSet storage; it leaves PVCs in Pending state until a pod references them — this is normal, not an error.
Storage abstraction hierarchy: StorageClass → PV → PVC → Pod mount. Dynamic provisioning automates PV creation via StorageClass.
Common operational trap: teams delete StatefulSets expecting automatic PVC cleanup and later find orphaned, expensive volumes accumulating.
----
''Sources''
* <html><code>training/library/topics/k8s-storage/primer.md</code></html>
* <html><code>training/interactive/knowledge/data/cards/k8s-storage.tsv</code></html>
//Merged from 2 source atoms.//
''Related atoms''
* [[PV reclaim policies: Retain preserves data, Delete destroys it]]
* [[Kubernetes storage abstraction: PV, PVC, and StorageClass layers]]
* [[What is a Kubernetes StatefulSet and when would you use it?]]
Q: What is the relationship between a PersistentVolume (PV) and a PersistentVolumeClaim (PVC)?
A: A PV is a cluster-level storage resource. A PVC is a namespace-scoped request for storage. Pods reference PVCs, and Kubernetes binds PVCs to matching PVs based on access mode, storage class, and capacity. The binding is exclusive.
Remember: Storage abstraction: StorageClass→PV→PVC→Pod mount. Dynamic provisioning automates PV.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-storage.tsv</code></html>
* <html><code>training/library/topics/mental-models-core/k8s-pv-pvc.md</code></html>
//Merged from 2 source atoms.//
''Related atoms''
* [[Ephemeral vs. Persistent Volumes in Kubernetes Pods]]
* [[Kubernetes dynamic provisioning and WaitForFirstConsumer zone awareness]]
Q: What is a Persistent Volume (PV) and Persistent Volume Claim (PVC) in Kubernetes?
A: * Persistent Volume (PV):
* Represents a piece of storage in the cluster that has been provisioned by an administrator.
* Can be used to store data independently of any particular pod.
* Provides a way to manage storage resources in a cluster.
* Persistent Volume Claim (PVC):
* Represents a request for storage by a user or pod.
* Binds to a Persistent Volume, making the storage available to the pod.
* Allows for dynamic provisioning of storage resources.
Persistent Volumes and Persistent Volume Claims provide a mechanism for decoupling storage from pod lifecycles.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-storage.tsv</code></html>
''Related atoms''
* [[What is a volume in regards to Kubernetes?]]
Q: What are the three classic PersistentVolume access modes and what does each allow?
A: ReadWriteOnce (RWO): mounted read-write by a single node. ReadOnlyMany (ROX): mounted read-only by many nodes. ReadWriteMany (RWX): mounted read-write by many nodes. RWO is the most common since block storage is single-attach.
Remember: RWO=ReadWriteOnce, ROX=ReadOnlyMany, RWX=ReadWriteMany. O=One, X=Many.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-storage.tsv</code></html>
''Related atoms''
* [[What types of persistent volumes are there?]]
* [[Ephemeral vs. Persistent Volumes in Kubernetes Pods]]
Q: What is the purpose of a StorageClass in Kubernetes?
A: A StorageClass defines how storage is provisioned. It names a provisioner (CSI driver), sets provider-specific parameters (disk type, IOPS), configures the reclaim policy, and controls volume binding mode. It enables dynamic provisioning so admins do not need to pre-create PVs.
Remember: StorageClass = PV factory. Dynamic provisioning — no pre-created PVs needed.
Example: <html><code>kubectl get sc</code></html>. Default SC handles PVCs that omit storageClassName.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-storage.tsv</code></html>
''Related atoms''
* [[Kubernetes dynamic provisioning and WaitForFirstConsumer zone awareness]]
* [[Kubernetes does not provide data persistence by default]]
Q: What is the difference between Immediate and WaitForFirstConsumer volume binding modes?
A: Immediate provisions the volume as soon as the PVC is created. WaitForFirstConsumer delays provisioning until a pod using the PVC is scheduled to a node. WaitForFirstConsumer is critical on cloud providers to ensure the volume is created in the same availability zone as the node.
Remember: Storage abstraction: StorageClass→PV→PVC→Pod mount. Dynamic provisioning automates PV.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-storage.tsv</code></html>
''Related atoms''
* [[What is the relationship between a PersistentVolume (PV) and a PersistentVolumeClaim (P…]]
* [[Kubernetes storage abstraction: PV, PVC, and StorageClass layers]]
A VolumeSnapshot captures a PVC's data at a specific point in time — the Kubernetes equivalent of a cloud disk snapshot.
To create one: define a VolumeSnapshot resource referencing a source PVC and a VolumeSnapshotClass. The VolumeSnapshotClass declares which CSI driver handles the snapshot operation.
To restore: create a new PVC with a <html><code>dataSource</code></html> block pointing to the snapshot (<html><code>kind: VolumeSnapshot</code></html>). The CSI driver provisions a new volume pre-populated from the snapshot.
Prerequisites:
* The CSI driver must advertise the Snapshot capability.
* A VolumeSnapshotClass must exist (<html><code>kubectl get volumesnapshotclass</code></html> — empty output means no snapshot support).
* Not all storage providers support snapshots; verify before depending on this feature.
Storage abstraction reminder: StorageClass → PV → PVC → Pod mount. Dynamic provisioning automates PV creation.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-storage.tsv</code></html>
//Merged from 2 source atoms.//
''Related atoms''
* [[Kubernetes does not provide data persistence by default]]
Q: A pod fails to start with a multi-attach error on an RWO volume. What happened and how do you fix it?
A: An RWO volume is still attached to a previous node, usually because the old node was not properly drained or is in a NotReady state. The volume cannot attach to the new node until it is detached from the old one. Fix: force-detach the volume via the cloud provider API or delete the old VolumeAttachment object. Prevent by using proper node drain procedures and setting appropriate pod disruption budgets.
Remember: Storage abstraction: StorageClass→PV→PVC→Pod mount. Dynamic provisioning automates PV.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-storage.tsv</code></html>
''Related atoms''
* [[Kubernetes PVC expansion: StorageClass gate and two-phase resize]]
Q: How does PVC expansion work and what are the operational risks?
A: PVC expansion requires allowVolumeExpansion: true on the StorageClass. Patch the PVC with a larger storage request. Most CSI drivers support online expansion but some require a pod restart for filesystem resize. Risks: expansion is one-way (cannot shrink), filesystem resize can fail leaving the volume in a resizing condition, and not all storage backends support online expansion. Always snapshot before expanding.
Remember: Storage abstraction: StorageClass→PV→PVC→Pod mount. Dynamic provisioning automates PV.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-storage.tsv</code></html>
''Related atoms''
* [[What volume types are you familiar with?]]
Q: What reclaim policies are there?
A: * Retain
* Delete
Note: Recycle was deprecated in K8s 1.22 and removed in 1.24. Use dynamic provisioning instead.
Remember: Retain(manual), Delete(auto), Recycle(deprecated). Dynamic default=Delete.
Gotcha: Retain→PV becomes Released, not Available. Manual intervention needed.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-storage.tsv</code></html>
''Related atoms''
* [[PV reclaim policies: Retain preserves data, Delete destroys it]]
Q: Explain "Dynamic Provisioning" and "Static Provisioning"
A: The main difference relies on the moment when you want to configure storage. For instance, if you need to pre-populate data in a volume, you choose static provisioning. Whereas, if you need to create volumes on demand, you go for dynamic provisioning.
Remember: Storage abstraction: StorageClass→PV→PVC→Pod mount. Dynamic provisioning automates PV.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-storage.tsv</code></html>
''Related atoms''
* [[Kubernetes storage abstraction: PV, PVC, and StorageClass layers]]
* [[What is the purpose of a StorageClass in Kubernetes?]]
* [[What volume types are you familiar with?]]
Q: Which problems, volumes in Kubernetes solve?
A: 1. Sharing files between containers running in the same Pod
# Storage in containers is ephemeral - it usually doesn't last for long. For example, when a container crashes, you lose all on-disk data. Certain volumes allows to manage such situation by persistent volumes
Remember: Storage abstraction: StorageClass→PV→PVC→Pod mount. Dynamic provisioning automates PV.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-storage.tsv</code></html>
''Related atoms''
* [[What is a Persistent Volume (PV) and Persistent Volume Claim (PVC) in Kubernetes?]]
* [[True or False? A volume defined in Pod can be accessed by all the containers of that Pod]]
* [[What volume types are you familiar with?]]
Q: True or False? Kubernetes provides data persistence out of the box.
A: False. Kubernetes does not provide data persistence by default. Container filesystems are ephemeral — data written inside a container is lost when the Pod restarts or is replaced.
Kubernetes provides the abstraction framework for storage but does not manage physical persistence itself. Storage backends — cloud block devices (EBS, GCE PD), network filesystems (NFS, Ceph), or local disks — handle actual durability.
The storage abstraction stack: StorageClass → PersistentVolume (PV) → PersistentVolumeClaim (PVC) → Pod volume mount.
* PV: cluster-scoped resource representing a piece of physical storage (EBS volume, NFS share, local path).
* PVC: namespace-scoped request that binds to a matching PV based on access mode and capacity.
* StorageClass: enables dynamic provisioning, automatically creating PVs to satisfy new PVCs without manual pre-provisioning.
Without PV/PVC wiring, all pod storage is ephemeral.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-storage.tsv</code></html>
//Merged from 2 source atoms.//
Q: True or False? A volume defined in Pod can be accessed by all the containers of that Pod
A: True. A volume defined in a Pod spec is accessible to all containers in that Pod. Each container mounts it via volumeMounts, and they can share data through it.
Remember: Storage abstraction: StorageClass→PV→PVC→Pod mount. Dynamic provisioning automates PV.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-storage.tsv</code></html>
''Related atoms''
* [[Which problems, volumes in Kubernetes solve?]]
* [[What is a Persistent Volume (PV) and Persistent Volume Claim (PVC) in Kubernetes?]]
Q: What is a volume in regards to Kubernetes?
A: A directory accessible by the containers inside a certain Pod and containers. The mechanism responsible for creating the directory, managing it, ... mainly depends on the volume type.
Remember: Storage abstraction: StorageClass→PV→PVC→Pod mount. Dynamic provisioning automates PV.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-storage.tsv</code></html>
''Related atoms''
* [[What is a Persistent Volume (PV) and Persistent Volume Claim (PVC) in Kubernetes?]]
* [[Kubernetes storage abstraction: PV, PVC, and StorageClass layers]]
Q: What volume types are you familiar with?
A: * emptyDir: created when a Pod assigned to a node and ceases to exist when the Pod is no longer running on that node
* hostPath: mounts a path from the host itself. Usually not used due to security risks but has multiple use-cases where it's needed like access to some internal host paths (<html><code>/sys</code></html>, <html><code>/var/lib</code></html>, etc.)
Remember: Storage abstraction: StorageClass→PV→PVC→Pod mount. Dynamic provisioning automates PV.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-storage.tsv</code></html>
//Merged from 2 source atoms.//
''Related atoms''
* [[Which problems, volumes in Kubernetes solve?]]
Ephemeral volumes have the lifetime of a Pod — they are created when the Pod starts and destroyed when it stops. Persistent Volumes (PVs) exist independently of any Pod's lifecycle, allowing data to survive Pod restarts, rescheduling, or deletion.
Ephemeral volume types include emptyDir, configMap, and secret. Persistent storage is provisioned via a PersistentVolumeClaim (PVC) bound to a PV.
PV lifecycle states: Available → Bound → Released.
Access modes:
* RWO (ReadWriteOnce): one node may mount for writing.
* ROX (ReadOnlyMany): many nodes may mount read-only.
* RWX (ReadWriteMany): many nodes may mount for writing.
Mnemonic: O = One, X = Many.
Analogy: Ephemeral volumes are like RAM — fast and temporary. Persistent volumes are like disk — slower but durable across Pod boundaries.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-storage.tsv</code></html>
//Merged from 2 source atoms.//
Q: What types of persistent volumes are there?
A: * NFS
* iSCSI
* CephFS
* ...
Remember: PV lifecycle: Available→Bound→Released. Modes: RWO, ROX, RWX.
Remember: O=One, X=Many. RWO=one writer, ROX=many readers, RWX=many writers.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-storage.tsv</code></html>
''Related atoms''
* [[What are the three classic PersistentVolume access modes and what does each allow?]]
* [[Ephemeral vs. Persistent Volumes in Kubernetes Pods]]
Q: How do you determine if a pod is Pending due to resource pressure?
A: kubectl describe nodes | grep -A 5 'Allocated resources' shows used vs allocatable. If requests exceed available, the scheduler can't place the pod. kubectl get events --field-selector reason=FailedScheduling confirms.
Remember: Pending = can't schedule. Causes: no resources, taints, unbound PVC, no nodes.
Gotcha: <html><code>kubectl describe pod</code></html> shows FailedScheduling reason.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-troubleshooting.tsv</code></html>
''Related atoms''
* [[Kubernetes node resource pressure: diagnosis, eviction, and prevention]]
Q: How do you use kubectl events for troubleshooting?
A: kubectl get events --sort-by=.lastTimestamp shows recent cluster events. kubectl get events --field-selector involvedObject.name=<pod> filters to one resource. Events expire after 1 hour by default — check quickly after an issue.
Remember: Flow: Get→Describe→Logs→Exec. Mnemonic: "GDLE."
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-troubleshooting.tsv</code></html>
''Related atoms''
* [[What is the standard three-command diagnostic flow when a pod is misbehaving?]]
Q: When do you use kubectl logs vs kubectl describe?
A: Use logs to see application output (stdout/stderr). Use describe to see Kubernetes-level info: scheduling decisions, probe results, image pulls, resource limits, events. Start with describe for cluster issues, logs for app issues.
Example: <html><code>kubectl logs pod --previous --tail=100</code></html> — last 100 lines from crashed container.
Gotcha: Logs lost on pod deletion. Set up log aggregation (Fluentd/Loki) for persistence.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-troubleshooting.tsv</code></html>
''Related atoms''
* [[Why do container logs sometimes bring down Kubernetes nodes?]]
* [[kubectl logs --previous retrieves logs from last crashed container]]
Q: How do you test connectivity to a Service from inside the cluster?
A: kubectl run tmp --image=busybox --rm -it -- wget -qO- http://<svc>.<ns>.svc.cluster.local:<port>. Or use kubectl exec into an existing pod. Check DNS resolution: nslookup <svc>.<ns>.svc.cluster.local.
Remember: K8s troubleshooting: Get→Describe→Logs→Exec. Mnemonic: "GDLE."
Gotcha: Always check Events with <html><code>kubectl describe</code></html> — they tell WHY, not just WHAT failed.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-troubleshooting.tsv</code></html>
''Related atoms''
* [[Users unable to reach an application running on a Pod on Kubernetes. What might be the …]]
Q: A pod with a PVC can't start and shows a multi-attach error. What's wrong?
A: The PV is ReadWriteOnce (RWO) and is already mounted on another node. This happens during rolling updates when old and new pods are on different nodes. Fix: use Recreate strategy instead of RollingUpdate, or switch to ReadWriteMany if the storage supports it.
Remember: K8s troubleshooting: Get→Describe→Logs→Exec. Mnemonic: "GDLE."
Gotcha: Always check Events with <html><code>kubectl describe</code></html> — they tell WHY, not just WHAT failed.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-troubleshooting.tsv</code></html>
Use <html><code>kubectl rollout history deployment/<name></code></html> to list all recorded revisions. Add <html><code>--revision=<n></code></html> to inspect the details of a specific revision. To revert, use <html><code>kubectl rollout undo deployment/<name></code></html> to go back one revision, or <html><code>kubectl rollout undo deployment/<name> --to-revision=<n></code></html> to target a specific version.
Revisions are stored in old ReplicaSets — do not delete them or you lose rollback targets.
For Helm-managed releases, the equivalent is <html><code>helm rollback <RELEASE_NAME> <REVISION_ID></code></html>, which reverts the release to the specified Helm revision.
Diagnostic note: always run <html><code>kubectl describe deployment/<name></code></html> and inspect the Events section — it explains WHY a rollout failed, not just that it did. Troubleshooting order: Get → Describe → Logs → Exec.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-troubleshooting.tsv</code></html>
* <html><code>training/interactive/knowledge/data/cards/k8s-workloads.tsv</code></html>
//Merged from 2 source atoms.//
''Related atoms''
* [[How to view revision history for a certain release?]]
Q: How do you find pods that match a particular label selector?
A: kubectl get pods -l app=myapp,env=prod (comma = AND). kubectl get pods -l 'app in (myapp,otherapp)' (set-based). kubectl get pods --show-labels to see all labels. Mismatched labels are the top cause of Service/Deployment issues.
Remember: K8s troubleshooting: Get→Describe→Logs→Exec. Mnemonic: "GDLE."
Gotcha: Always check Events with <html><code>kubectl describe</code></html> — they tell WHY, not just WHAT failed.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-troubleshooting.tsv</code></html>
''Related atoms''
* [[Kubernetes Service empty endpoints: label-selector mismatch diagnosis]]
* [[How to get information on a certain service?]]
Q: A developer says their app can't reach a service. They're in different namespaces. What's the fix?
A: Use the FQDN: <service>.<namespace>.svc.cluster.local. Short names only resolve within the same namespace. Also check: NetworkPolicy may restrict cross-namespace traffic. kubectl get netpol -n <ns> to verify.
Remember: K8s troubleshooting: Get→Describe→Logs→Exec. Mnemonic: "GDLE."
Gotcha: Always check Events with <html><code>kubectl describe</code></html> — they tell WHY, not just WHAT failed.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-troubleshooting.tsv</code></html>
''Related atoms''
* [[How do you test connectivity to a Service from inside the cluster?]]
* [[What can you find in kube-system namespace?]]
* [[Check how many namespaces are there]]
Q: What is the difference between resources.requests.memory and resources.limits.memory in a Kubernetes pod spec?
A: requests.memory is the scheduling guarantee (the kubelet reserves this amount on the node). limits.memory is the hard ceiling enforced by the Linux cgroup — exceeding it triggers OOMKill.
Remember: K8s troubleshooting: Get→Describe→Logs→Exec. Mnemonic: "GDLE."
Gotcha: Always check Events with <html><code>kubectl describe</code></html> — they tell WHY, not just WHAT failed.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-troubleshooting.tsv</code></html>
''Related atoms''
* [[True or False? Resource limits applied on a Pod level meaning, if limits is 2gb RAM and…]]
* [[Explain why one would specify resource limits in regards to Pods]]
Q: Why does a Java application with -Xmx1g in a container limited to 512Mi get OOMKilled, and how do you fix it?
A: The JVM requests 1GB of heap from the OS, but the cgroup enforces a 512Mi ceiling and kills the process. Fix by using -XX:MaxRAMPercentage=75.0 so the JVM sizes its heap relative to the container's memory limit.
Remember: OOMKilled = exit 137 (128+SIGKILL). Exceeded memory limit. Increase or fix leak.
Gotcha: Check <html><code>kubectl describe pod</code></html> for Reason: OOMKilled in Last State.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-troubleshooting.tsv</code></html>
''Related atoms''
* [[What is a LimitRange and how does it prevent OOMKilled caused by missing resource limits?]]
The Linux kernel assigns each process an OOM kill score equal to its <html><code>oom_score</code></html> (resident set size as a fraction of total system memory) plus its <html><code>oom_score_adj</code></html> (range -1000 to 1000, stored in <html><code>/proc/<pid>/oom_score_adj</code></html>). When memory is exhausted, the process with the highest combined score is killed first. A value of -1000 makes a process immune; +1000 makes it the most likely target.
Kubernetes sets <html><code>oom_score_adj</code></html> automatically based on a pod's QoS class: Guaranteed pods (requests == limits for all containers) receive -997 and are killed last; Burstable pods (requests < limits, or partial) receive 2–999 scaled proportionally by memory request ratio and are killed second; BestEffort pods (no requests or limits) receive 1000 and are killed first.
Under node-level memory pressure, the kernel evaluates all processes across all pods — not just the container that exceeded its own limit. This is why correct QoS class assignment is protective: a Guaranteed pod is the last to be sacrificed even if another pod on the same node causes the exhaustion.
To investigate OOM events: search <html><code>dmesg</code></html> for "invoked oom-killer" to find the triggering allocation and see the killed process's total-vm and anon-rss; check <html><code>/proc/meminfo</code></html> for current pressure; scan <html><code>/proc/<pid>/oom_score</code></html> to predict the next victim. In Kubernetes, <html><code>kubectl describe pod</code></html> shows <html><code>Reason: OOMKilled</code></html> in Last State; exit code 137 (128 + SIGKILL) confirms an OOM kill. Critical non-Kubernetes processes can be protected by writing a negative value directly to their <html><code>/proc/<pid>/oom_score_adj</code></html>.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-troubleshooting.tsv</code></html>
* <html><code>training/library/topics/kernel-troubleshooting/street_ops.md</code></html>
* <html><code>training/library/topics/oomkilled/primer.md</code></html>
* <html><code>training/library/topics/oomkilled/trivia.md</code></html>
//Merged from 4 source atoms.//
Q: Which Prometheus metric should you use to predict OOMKill, and why not container_memory_usage_bytes?
A: Use container_memory_working_set_bytes because it excludes inactive file cache and reflects what the OOM killer actually evaluates. container_memory_usage_bytes includes reclaimable cache and overstates true pressure.
Remember: OOMKilled = exit 137 (128+SIGKILL). Exceeded memory limit. Increase or fix leak.
Gotcha: Check <html><code>kubectl describe pod</code></html> for Reason: OOMKilled in Last State.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-troubleshooting.tsv</code></html>
''Related atoms''
* [[How do you distinguish a container-level OOM from a node-level OOM, and what commands r…]]
Q: What is a LimitRange and how does it prevent OOMKilled caused by missing resource limits?
A: A LimitRange is a namespace-scoped object that sets default memory requests and limits for containers that do not specify their own. It ensures every pod has a cgroup ceiling, preventing unbounded memory consumption that causes node-level OOM.
Remember: OOMKilled = exit 137 (128+SIGKILL). Exceeded memory limit. Increase or fix leak.
Gotcha: Check <html><code>kubectl describe pod</code></html> for Reason: OOMKilled in Last State.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-troubleshooting.tsv</code></html>
''Related atoms''
* [[How do you prevent high memory usage in your Kubernetes cluster and possibly issues lik…]]
Kubernetes exposes two distinct memory-related failure modes: eviction and OOMKill.
''OOMKill'' fires when a container's cgroup exceeds its memory limit. The kernel kills the offending process (exit code 137 = 128 + SIGKILL). Two scopes exist: container-level OOM kills only the affected container in isolation; node-level OOM (kernel global OOM killer) picks victims across all processes on the host, potentially killing multiple pods simultaneously. Diagnose via <html><code>kubectl describe pod</code></html> — look for <html><code>Reason: OOMKilled</code></html> in Last State, and check <html><code>dmesg</code></html> and kubelet logs for node-level events.
''Eviction'' is kubelet's proactive defense against node memory pressure, firing before the kernel OOM killer. Kubelet supports <html><code>--eviction-hard</code></html> (e.g., <html><code>memory.available<100Mi</code></html>) and <html><code>--eviction-soft</code></html> thresholds with grace periods. When available memory crosses the hard threshold — or the soft threshold for its full grace period — kubelet evicts pods starting with BestEffort, then Burstable. Critically, eviction can terminate a pod whose usage is still below its memory limit if node pressure is high enough; the evicted pod receives the message "The node was low on resource: memory," not OOMKilled.
If eviction cannot free memory fast enough and memory is fully exhausted, the kernel OOM killer fires as a last resort.
A common misconfiguration: correct limits set but requests underspecified or omitted, leaving pods vulnerable to eviction before they approach their limits. Both failure modes are disruptive but have different triggers: OOMKill is the cgroup limit enforcer; eviction is the kubelet's node-pressure defense.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-troubleshooting.tsv</code></html>
* <html><code>training/library/topics/oomkilled/primer.md</code></html>
* <html><code>training/library/topics/oomkilled/trivia.md</code></html>
//Merged from 3 source atoms.//
''Related atoms''
* [[Pod deletion: SIGTERM, grace period, and SIGKILL]]
Q: How does the Vertical Pod Autoscaler help prevent OOMKilled, and what are its risks?
A: VPA analyzes historical memory usage and recommends or automatically sets requests and limits. It provides lowerBound, target, and upperBound recommendations. Risks: in UpdateMode it restarts pods to apply new limits (disruption), it can undersize limits if load patterns are spiky, and it conflicts with HPA on the same resource — never use VPA and HPA both scaling on memory.
Remember: OOMKilled = exit 137 (128+SIGKILL). Exceeded memory limit. Increase or fix leak.
Gotcha: Check <html><code>kubectl describe pod</code></html> for Reason: OOMKilled in Last State.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-troubleshooting.tsv</code></html>
Q: What is the difference between exit codes 126 and 127 in a container?
A: Exit code 126 means the entrypoint binary exists but cannot be executed (permission denied). Exit code 127 means the entrypoint binary does not exist in the image (command not found).
Remember: K8s troubleshooting: Get→Describe→Logs→Exec. Mnemonic: "GDLE."
Gotcha: Always check Events with <html><code>kubectl describe</code></html> — they tell WHY, not just WHAT failed.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-troubleshooting.tsv</code></html>
Q: How do init containers help prevent CrashLoopBackOff caused by missing dependencies?
A: Init containers run before the main container and block startup until they succeed. You can use an init container to wait for a dependency (e.g., polling a database port with nc -z) so the main container only starts when its dependencies are actually ready.
Remember: CrashLoopBackOff = start-crash-retry with exponential backoff (10s→20s→40s→5min).
Gotcha: <html><code>kubectl logs pod --previous</code></html> shows crash reason. Common: missing config, OOM, wrong cmd.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-troubleshooting.tsv</code></html>
''Related atoms''
* [[Distinguish CrashLoopBackOff from other pod failure states]]
Q: What is the PID 1 problem in containers and how does it cause exit code 137 on pod termination?
A: In a container, the entrypoint runs as PID 1. If PID 1 does not handle SIGTERM, Kubernetes sends SIGTERM on shutdown, the process ignores it, Kubernetes waits the terminationGracePeriodSeconds (default 30s), then sends SIGKILL — resulting in exit code 137. Fix by using exec form in Dockerfile or a lightweight init system like tini.
Remember: K8s troubleshooting: Get→Describe→Logs→Exec. Mnemonic: "GDLE."
Gotcha: Always check Events with <html><code>kubectl describe</code></html> — they tell WHY, not just WHAT failed.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-troubleshooting.tsv</code></html>
''Related atoms''
* [[Pod deletion: SIGTERM, grace period, and SIGKILL]]
* [[Debugging a failing or non-starting pod in Kubernetes]]
A DaemonSet ensures that all (or some) nodes run a copy of a pod. As nodes are added to the cluster, pods are added to them; as nodes are removed, those pods are garbage collected. Deleting a DaemonSet cleans up the pods it created.
Typical use cases: log collectors, monitoring agents, network plugins — workloads that must run exactly once per host rather than at arbitrary replica counts.
Scheduling mechanism: prior to Kubernetes 1.12, placement was enforced via the <html><code>NodeName</code></html> attribute. From 1.12 onward, DaemonSets use the regular scheduler with node affinity to guarantee one copy per eligible node.
Workload hierarchy: Deployment / StatefulSet / DaemonSet → ReplicaSet → Pod.
Gotcha: Use <html><code>kubectl explain <resource></code></html> for field reference without leaving the CLI. Node affinity rules — not replica counts — are what enforce the one-pod-per-node invariant.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-workloads.tsv</code></html>
//Merged from 3 source atoms.//
''Related atoms''
* [[Kubernetes Control Plane: components and responsibilities]]
* [[True or False? Once a Pod is assisgned to a worker node, it will only run on that node,…]]
* [[Kubernetes Static Pods: definition, use cases, and DaemonSet contrast]]
Kubernetes Deployments handle rolling updates by creating a new ReplicaSet with the updated pod template—triggered by any spec change (new image, resources, env vars)—and gradually scaling it up while scaling the old ReplicaSet down. This ensures zero-downtime transitions.
Rollback is the same mechanism in reverse. The Deployment retains a history of previous ReplicaSets (held at zero replicas). Running <html><code>kubectl rollout undo</code></html> repoints the Deployment to a prior ReplicaSet, scaling it back up while scaling the current one down, ensuring a controlled, availability-preserving revert.
The ownership chain is Deployment → ReplicaSet(s) → Pods. Old ReplicaSets are kept at zero replicas to support rollback; deleting a Deployment cascades to all owned ReplicaSets and Pods. Manually editing a ReplicaSet owned by a Deployment is dangerous: the Deployment's reconciliation loop will overwrite those changes to match its spec.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-workloads.tsv</code></html>
* <html><code>training/library/topics/mental-models-core/k8s-deployment-replicaset-pod.md</code></html>
//Merged from 2 source atoms.//
''Related atoms''
* [[What is the relationship between Deployment, ReplicaSet, and Pod?]]
* [[Scaling a Deployment in Kubernetes]]
* [[What happens after you edit a deployment and change the image?]]
Q: How to check which container image was used as part of replica set called "repli"?
A: <html><code>kubectl describe rs repli | grep -i image</code></html> shows the container image used by the ReplicaSet. You can also use <html><code>kubectl get rs repli -o jsonpath='{.spec.template.spec.containers[*].image}'</code></html> for a clean output.
Gotcha: always use specific image tags (nginx:1.25.3) in production — the 'latest' tag is mutable and can change unexpectedly.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-workloads.tsv</code></html>
''Related atoms''
* [[How to modify a replica set called "rori" to use a different image?]]
Q: Explain what is CronJob and what is it used for
A: A CronJob creates Jobs on a repeating schedule. One CronJob object is like one line of a crontab (cron table) file. It runs a job periodically on a given schedule, written in Cron format.
Remember: CronJob = Job on schedule. Fields: minute hour day month weekday. "MHDMW."
Gotcha: <html><code>concurrencyPolicy: Forbid</code></html> prevents overlap. Default allows multiple Jobs.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-workloads.tsv</code></html>
''Related atoms''
* [[What issue might arise from using the following CronJob and how to fix it?]]
Q: Create a deployment called "pluck" using the image "redis" and make sure it runs 5 replicas
A: <html><code>kubectl create deployment pluck --image=redis</code></html>
<html><code>kubectl scale deployment pluck --replicas=5</code></html>
Example: <html><code>kubectl create deployment web --image=nginx --replicas=3 --port=80</code></html>
Remember: Deployments manage ReplicaSets→Pods. <html><code>kubectl rollout</code></html> for updates.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-workloads.tsv</code></html>
''Related atoms''
* [[Create a deployment with the following properties:]]
* [[What rollout/deployment strategies are you familiar with?]]
* [[Fix the following deployment manifest]]
Q: Create a file definition/manifest of a deployment called "dep", with 3 replicas that uses the image 'redis'
A: <html><code>k create deploy dep -o yaml --image=redis --dry-run=client --replicas 3 > deployment.yaml </code></html>
Example: <html><code>kubectl create deployment web --image=nginx --replicas=3 --port=80</code></html>
Remember: Deployments manage ReplicaSets→Pods. <html><code>kubectl rollout</code></html> for updates.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-workloads.tsv</code></html>
''Related atoms''
* [[How to create a deployment with the image "nginx:alpine"?]]
* [[How does rolling deployment work in Kubernetes?]]
* [[What is a "Deployment" in Kubernetes?]]
Q: What possible issue can arise from using the following spec and how to fix it?
A: If the cron job fails, the next job will not replace the previous one due to the "concurrencyPolicy" value which is "Allow". It will keep spawning new jobs and so eventually the system will be filled with failed cron jobs.
To avoid such problem, the "concurrencyPolicy" value should be either "Replace" or "Forbid".
Remember: Workload hierarchy: Deployment/StatefulSet/DaemonSet→ReplicaSet→Pod.
Gotcha: Use <html><code>kubectl explain <resource></code></html> for field reference without leaving the CLI.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-workloads.tsv</code></html>
''Related atoms''
* [[How does Kubernetes ensure high availability of applications?]]
Q: How to delete a replica set called "rori"?
A: <html><code>kubectl delete rs rori</code></html> removes the ReplicaSet and all pods it manages. The workload hierarchy is: Deployment -> ReplicaSet -> Pod.
Gotcha: if the ReplicaSet was created by a Deployment, the Deployment will immediately recreate it. Delete the parent Deployment instead to fully remove the workload.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-workloads.tsv</code></html>
''Related atoms''
* [[Deleting a Deployment (not its pods) is the correct removal path]]
* [[How to modify a replica set called "rori" to use a different image?]]
* [[True or False? Deleting a ReplicaSet will delete the Pods it created]]
Q: Is it possible to delete ReplicaSet without deleting the Pods it created?
A: Yes, with <html><code>kubectl delete rs rori --cascade=orphan</code></html> (or the older <html><code>--cascade=false</code></html>). This removes the ReplicaSet object but leaves its pods running as orphans without a controller. Useful when you want to adopt pods under a new controller.
Gotcha: orphaned pods will not be recreated if they crash — they lose their self-healing capability.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-workloads.tsv</code></html>
''Related atoms''
* [[What will happen when a Pod, created by ReplicaSet, is deleted directly with kubectl de…]]
Q: What is the difference between a Job and a CronJob in Kubernetes?
A: * Job: A Job in Kubernetes runs a pod to completion and then terminates.
* It is used for short-lived, batch-style processes, ensuring that the task is executed once successfully.
* CronJob: A CronJob is a higher-level abstraction that schedules jobs at specified intervals using cron expressions.
* It is suitable for recurring, automated tasks that need to run periodically.
* While Jobs are designed for one-time execution of tasks, CronJobs provide a scheduling mechanism for recurring tasks.
Remember: CronJob = Job on schedule. Fields: minute hour day month weekday. "MHDMW."
Gotcha: <html><code>concurrencyPolicy: Forbid</code></html> prevents overlap. Default allows multiple Jobs.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-workloads.tsv</code></html>
''Related atoms''
* [[What is the job of the kube-scheduler?]]
Q: Deploy a pod called "my-pod" using the nginx:alpine image
A: <html><code>kubectl run my-pod --image=nginx:alpine</code></html>
If you are a Kubernetes beginner you should know that this is not a common way to run Pods. The common way is to run a Deployment which in turn runs Pod(s).
In addition, Pods and/or Deployments are usually defined in files rather than executed directly using only the CLI arguments.
Remember: Use specific tags (:1.25.3), never :latest in prod. Latest is mutable.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-workloads.tsv</code></html>
''Related atoms''
* [[Fix the following deployment manifest]]
* [[Create a deployment with the following properties:]]
* [[How to create a pod and a service with one command?]]
Q: How to create a deployment with the image "nginx:alpine"?
A: <html><code>kubectl create deployment my-first-deployment --image=nginx:alpine</code></html>
OR
<html><pre><code class="language-plaintext">cat << EOF | kubectl create -f -
apiVersion: apps/v1
kind: Deployment
metadata:
name: nginx
spec:
replicas: 1
selector:
matchLabels:
app: nginx
template:
metadata:
labels:
app: nginx
spec:
containers:
- name: nginx
image: nginx:alpine</code></pre></html>
Example: <html><code>kubectl create deployment web --image=nginx --replicas=3 --port=80</code></html>
Remember: Deployments manage ReplicaSets→Pods. <html><code>kubectl rollout</code></html> for updates.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-workloads.tsv</code></html>
''Related atoms''
* [[Fix the following deployment manifest]]
* [[Create a file definition/manifest of a deployment called "dep", with 3 replicas that us…]]
* [[Create a deployment with the following properties:]]
Q: How to scale a deployment to 8 replicas?
A: <html><code>kubectl scale deploy <DEPLOYMENT_NAME> --replicas=8</code></html> sets the desired replica count to
# Kubernetes gradually creates new pods to match the target. For auto-scaling, use HorizontalPodAutoscaler: <html><code>kubectl autoscale deployment <name> --min=2 --max=10 --cpu-percent=80</code></html>.
Gotcha: manual scaling is overridden if HPA is active on the same deployment.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-workloads.tsv</code></html>
''Related atoms''
* [[What happens when you set replicas in a Deployment manifest and also use HPA?]]
* [[How can PodDisruptionBudget conflict with HPA scale-down, and what is the best practice…]]
Q: In case of a ReplicaSet, Which field is mandatory in the spec section?
A: The field <html><code>template</code></html> in spec section is mandatory. It's used by the ReplicaSet to create new Pods when needed.
Remember: Don't create ReplicaSets directly. Use Deployments for rolling updates+rollbacks.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-workloads.tsv</code></html>
''Related atoms''
* [[True or False? Pods specified by the selector field of ReplicaSet must be created by th…]]
* [[What fields are mandatory with any Kubernetes object?]]
* [[What is the relationship between Deployment, ReplicaSet, and Pod?]]
Q: What the following block of lines does?
A: It defines a replicaset for Pods whose type is set to "backend" so at any given point of time there will be 2 concurrent Pods running.
Remember: Workload hierarchy: Deployment/StatefulSet/DaemonSet→ReplicaSet→Pod.
Gotcha: Use <html><code>kubectl explain <resource></code></html> for field reference without leaving the CLI.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-workloads.tsv</code></html>
''Related atoms''
* [[How does Kubernetes ensure high availability of applications?]]
* [[What the following command does?]]
* [[What Kubernetes objects are there?]]
Q: What is the difference between a Pod and a Deployment in Kubernetes?
A: A Pod is the smallest deployable unit in Kubernetes - it represents one or more
containers that share storage and network resources.
A Deployment is a higher-level controller that manages Pods:
* Ensures desired number of Pod replicas are running
* Handles rolling updates and rollbacks
* Provides declarative updates for Pods and ReplicaSets
* Self-healing: recreates Pods if they fail
Key differences:
* Pod: Single instance, no auto-recovery if deleted
* Deployment: Manages multiple Pod replicas, auto-recovers failures
You rarely create Pods directly in production.
Remember: Deployment > ReplicaSet > Pod. Three layers for declarative management.
Example: <html><code>kubectl create deployment web --image=nginx:1.25 --replicas=3 --port=80</code></html>
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-workloads.tsv</code></html>
//Merged from 2 source atoms.//
''Related atoms''
* [[What happens when you delete a deployment?]]
* [[How does rolling deployment work in Kubernetes?]]
* [[How Service and Deployment are connected?]]
Q: What is the difference between a StatefulSet and a Deployment in Kubernetes?
A: ''Deployment:''
• Used for stateless applications.
• Provides a way to manage and scale replica sets.
• Pods created by a deployment are not uniquely identified.
• Suitable for applications that can be easily replicated and scaled horizontally.
''StatefulSet:''
• Designed for stateful applications with unique network identities and stable hostnames.
• Maintains a unique identifier for each pod, allowing for ordered scaling and predictable naming.
• Suitable for applications that require stable network identities and persistent storage.
Remember: Deployment > ReplicaSet > Pod. Three layers for declarative management.
Example: <html><code>kubectl create deployment web --image=nginx:1.25 --replicas=3 --port=80</code></html>
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-workloads.tsv</code></html>
//Merged from 2 source atoms.//
''Related atoms''
* [[How do you define a Kubernetes Deployment?]]
* [[Scaling a Deployment in Kubernetes]]
* [[How does rolling deployment work in Kubernetes?]]
Q: Fix the following ReplicaSet definition
A: The selector doesn't match the label (cache vs cachy). To solve it, fix cachy so it's cache instead.
Remember: Don't create ReplicaSets directly. Use Deployments for rolling updates+rollbacks.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-workloads.tsv</code></html>
''Related atoms''
* [[You've created a ReplicaSet, how to check whether the ReplicaSet found matching Pods or…]]
* [[True or False? Pods specified by the selector field of ReplicaSet must be created by th…]]
* [[True or False? Removing the label from a Pod that is tracked by a ReplicaSet, will caus…]]
Q: Discuss the implications of pod sprawl and how to manage it effectively in Kubernetes.
A: Implications of Pod Sprawl:
• Increased resource consumption.
• More complex cluster management.
• Potential impact on network performance.
Management Strategies:
• Implement resource quotas and limits.
• Use Horizontal Pod Autoscaling to adjust pod counts dynamically.
• Regularly review and decommission unused or redundant pods. projects/knowledge/interview/kubernetes/368-discuss-the-implications-of-pod-sprawl-and-how-to-.txt
Remember: Workload hierarchy: Deployment/StatefulSet/DaemonSet→ReplicaSet→Pod.
Gotcha: Use <html><code>kubectl explain <resource></code></html> for field reference without leaving the CLI.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-workloads.tsv</code></html>
''Related atoms''
* [[What is the purpose of Horizontal Pod Autoscaling in Kubernetes?]]
* [[How does Kubernetes ensure high availability of applications?]]
* [[What are the challenges in managing stateful applications in Kubernetes?]]
A ReplicaSet maintains a stable set of identical Pod replicas running at any given time, guaranteeing availability of a specified number of Pods. The ReplicaSet controller continuously reconciles actual state with desired state in both directions: if there are more Pods than defined, it removes the excess; if there are fewer, it creates replacements. This makes ReplicaSets the primary mechanism for self-healing at the pod level.
Best practice: do not create ReplicaSets directly. Use Deployments instead — a Deployment manages ReplicaSets automatically and adds rolling update and rollback capabilities on top of the replica-count guarantee.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-workloads.tsv</code></html>
//Merged from 2 source atoms.//
''Related atoms''
* [[DaemonSet vs Deployment vs ReplicaSet in Kubernetes]]
* [[Scale a ReplicaSet with kubectl scale rs]]
* [[Describe the sequence of events in case of creating a ReplicaSet]]
Q: Fix the following deployment manifest
A: Change <html><code>kind: Deploy</code></html> to <html><code>kind: Deployment</code></html>
Example: <html><code>kubectl create deployment web --image=nginx --replicas=3 --port=80</code></html>
Remember: Deployments manage ReplicaSets→Pods. <html><code>kubectl rollout</code></html> for updates.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-workloads.tsv</code></html>
''Related atoms''
* [[How to create a deployment with the image "nginx:alpine"?]]
* [[Create a deployment with the following properties:]]
* [[Deploy a pod called "my-pod" using the nginx:alpine image]]
Q: How does Kubernetes handle rolling updates with zero downtime?
A: Rolling Updates with Zero Downtime:
* Kubernetes updates a Deployment by creating a new ReplicaSet alongside the existing one.
* New pods are gradually created and added to the new ReplicaSet, while old pods are gracefully terminated.
* This ensures a controlled transition without disrupting the availability of the application.
* During a rolling update, Kubernetes progressively replaces old pods with new ones, ensuring that there is always a sufficient number of healthy pods.
Remember: Workload hierarchy: Deployment/StatefulSet/DaemonSet→ReplicaSet→Pod.
Gotcha: Use <html><code>kubectl explain <resource></code></html> for field reference without leaving the CLI.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-workloads.tsv</code></html>
''Related atoms''
* [[How do readiness probes interact with rolling deployments?]]
* [[How does Kubernetes ensure high availability of applications?]]
* [[Kubernetes ReplicaSet: maintaining stable pod replica count]]
Q: ReplicaSets are running the moment the user executed the command to create them (like kubectl create -f rs.yaml)
A: False. It can take some time, depends on what exactly you are running. To see if they are up and running, run <html><code>kubectl get rs</code></html> and watch the 'READY' column.
Remember: Don't create ReplicaSets directly. Use Deployments for rolling updates+rollbacks.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-workloads.tsv</code></html>
''Related atoms''
* [[True or False? Pods specified by the selector field of ReplicaSet must be created by th…]]
* [[Describe the sequence of events in case of creating a ReplicaSet]]
* [[What the following output of kubectl get rs means?]]
To adjust the number of running pod replicas for a Deployment, use <html><code>kubectl scale</code></html>:
<html><pre><code class="language-plaintext">kubectl scale deployment/<name> --replicas=N</code></pre></html>
Alternatively, update the <html><code>replicas</code></html> field in the Deployment manifest and apply it with <html><code>kubectl apply -f</code></html>. Both methods instruct the Deployment's controller to reconcile the underlying ReplicaSet, which adds or removes Pods until the actual count matches the desired count.
Example — create a Deployment already at three replicas:
<html><pre><code class="language-plaintext">kubectl create deployment web --image=nginx --replicas=3 --port=80</code></pre></html>
Key relationships to keep in mind:
* Deployments manage ReplicaSets, which manage Pods.
* <html><code>kubectl scale</code></html> is an imperative shortcut; the manifest approach is declarative and version-controlled.
* <html><code>kubectl rollout</code></html> handles rolling updates and rollbacks, which is distinct from replica scaling.
* The application itself must be designed to run multiple instances correctly (stateless or with shared-state handling) for horizontal scaling to be safe.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-workloads.tsv</code></html>
//Merged from 2 source atoms.//
''Related atoms''
* [[What is a "Deployment" in Kubernetes?]]
* [[Create a deployment with the following properties:]]
* [[Kubernetes Deployment rolling updates and rollbacks]]
Q: How do you define a Kubernetes Deployment?
A: A Kubernetes Deployment is a resource object used to declare the desired state for a set of pods. It provides declarative updates to applications, allowing users to describe how an application should run and scale over time. Deployments enable rolling updates, rollbacks, and scaling of applications without manual intervention.
Key elements of a Deployment include:
* Pod Template: Defines the desired state of the pods.
* Replica Count: Specifies the desired number of pod replicas.
* Update Strategy: Defines how updates should be applied (e.g., rolling updates).\
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-workloads.tsv</code></html>
''Related atoms''
* [[Describe the Kubernetes API versioning strategy.]]
* [[You have a microservices-based application. How would you deploy and manage it in Kuber…]]
* [[What ways are you familiar with to implement deployment strategies (like canary, blue/g…]]
Q: What the following in a Deployment configuration file means?
A: USER_PASSWORD environment variable will store the value from password key in the secret called "some-secret"
In other words, you reference a value from a Kubernetes Secret.
Remember: Deployment > ReplicaSet > Pod. Three layers for declarative management.
Example: <html><code>kubectl create deployment web --image=nginx:1.25 --replicas=3 --port=80</code></html>
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-workloads.tsv</code></html>
''Related atoms''
* [[What is a "Deployment" in Kubernetes?]]
* [[What happens when you delete a deployment?]]
* [[What is the difference between a StatefulSet and a Deployment in Kubernetes?]]
Q: How to delete a deployment?
A: One way is by specifying the deployment name: <html><code>kubectl delete deployment [deployment_name]</code></html>
Another way is using the deployment configuration file: <html><code>kubectl delete -f deployment.yaml</code></html>
Example: <html><code>kubectl create deployment web --image=nginx --replicas=3 --port=80</code></html>
Remember: Deployments manage ReplicaSets→Pods. <html><code>kubectl rollout</code></html> for updates.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-workloads.tsv</code></html>
''Related atoms''
* [[Deleting a Deployment (not its pods) is the correct removal path]]
* [[Create a file definition/manifest of a deployment called "dep", with 3 replicas that us…]]
* [[What is a "Deployment" in Kubernetes?]]
Q: True or False? Deleting a ReplicaSet will delete the Pods it created
A: True (and not only the Pods but anything else it created).
Remember: Don't create ReplicaSets directly. Use Deployments for rolling updates+rollbacks.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-workloads.tsv</code></html>
''Related atoms''
* [[Deleting a Deployment (not its pods) is the correct removal path]]
* [[True or False? In case of a ReplicaSet, if Pods specified in the selector field don't e…]]
* [[What happens when you delete a deployment?]]
Q: True or False? If a ReplicaSet defines 2 replicas but there 3 Pods running matching the ReplicaSet selector, it will do nothing
A: False. It will terminate one of the Pods to reach the desired state of 2 replicas.
Remember: Don't create ReplicaSets directly. Use Deployments for rolling updates+rollbacks.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-workloads.tsv</code></html>
''Related atoms''
* [[What will happen when a Pod, created by ReplicaSet, is deleted directly with kubectl de…]]
* [[True or False? Once a Pod is assisgned to a worker node, it will only run on that node,…]]
* [[You've created a ReplicaSet, how to check whether the ReplicaSet found matching Pods or…]]
Q: True or False? Removing the label from a Pod that is tracked by a ReplicaSet, will cause the ReplicaSet to create a new Pod
A: True. When the label, used by a ReplicaSet in the selector field, removed from a Pod, that Pod no longer controlled by the ReplicaSet and the ReplicaSet will create a new Pod to compensate for the one it "lost".
Remember: Don't create ReplicaSets directly. Use Deployments for rolling updates+rollbacks.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-workloads.tsv</code></html>
''Related atoms''
* [[What will happen when a Pod, created by ReplicaSet, is deleted directly with kubectl de…]]
* [[Is it possible to delete ReplicaSet without deleting the Pods it created?]]
Q: True or False? Pods specified by the selector field of ReplicaSet must be created by the ReplicaSet itself
A: False. The Pods can be already running and initially they can be created by any object. It doesn't matter for the ReplicaSet and not a requirement for it to acquire and monitor them.
Remember: Don't create ReplicaSets directly. Use Deployments for rolling updates+rollbacks.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-workloads.tsv</code></html>
''Related atoms''
* [[Describe the sequence of events in case of creating a ReplicaSet]]
* [[What is the relationship between Deployment, ReplicaSet, and Pod?]]
* [[ReplicaSets are running the moment the user executed the command to create them (like k…]]
Q: True or False? In case of a ReplicaSet, if Pods specified in the selector field don't exists, the ReplicaSet will wait for them to run before doing anything
A: False. A ReplicaSet actively creates pods to match its desired replica count — it does not wait for pods to appear on their own. If the selector matches existing pods, they count toward the desired number. If not enough exist, the ReplicaSet creates new ones immediately. This is the core reconciliation loop of Kubernetes controllers.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-workloads.tsv</code></html>
''Related atoms''
* [[True or False? Deleting a ReplicaSet will delete the Pods it created]]
* [[Kubernetes ReplicaSet: maintaining stable pod replica count]]
* [[Describe the sequence of events in case of creating a ReplicaSet]]
Q: What issue might arise from using the following CronJob and how to fix it?
A: The following lines placed under the template:
<html><pre><code class="language-plaintext">concurrencyPolicy: Forbid
successfulJobsHistoryLimit: 1
failedJobsHistoryLimit: 1</code></pre></html>
As a result this configuration isn't part of the cron job spec hence the cron job has no limits which can cause issues like OOM and potentially lead to API server being down.
To fix it, these lines should placed in the spec of the cron job, above or under the "schedule" directive in the above example.
Remember: CronJob = Job on schedule. Fields: minute hour day month weekday. "MHDMW."
Gotcha: <html><code>concurrencyPolicy: Forbid</code></html> prevents overlap. Default allows multiple Jobs.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-workloads.tsv</code></html>
''Related atoms''
* [[Explain what is CronJob and what is it used for]]
* [[What is the difference between a Job and a CronJob in Kubernetes?]]
''DaemonSet'' ensures that exactly one copy of a Pod runs on all (or a selected subset of) nodes. It is used for cluster-level services and agents that must be present on every node — log collectors, monitoring agents, network plugins, etc.
''Deployment'' manages stateless applications that can be replicated and scaled horizontally across arbitrary nodes. It targets a desired replica count rather than node coverage, and supports rolling updates and rollbacks.
''ReplicaSet'' maintains a stable count of replica Pods running at any given time. It is the lower-level primitive that a Deployment wraps; do not create ReplicaSets directly — use Deployments instead to gain rolling-update and rollback semantics on top.
Key distinctions:
* DaemonSet targets //every node// (one Pod per node); Deployment and ReplicaSet target a //replica count// spread across available nodes.
* DaemonSet is for node-scoped infrastructure concerns; Deployment is for horizontally scalable application workloads.
* ReplicaSet and Deployment are interchangeable at the scheduling level, but Deployment adds lifecycle management; DaemonSet is a fundamentally different scheduling model.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-workloads.tsv</code></html>
//Merged from 2 source atoms.//
''Related atoms''
* [[Kubernetes ReplicaSet: maintaining stable pod replica count]]
* [[What is a "Deployment" in Kubernetes?]]
Q: How Service and Deployment are connected?
A: The truth is they aren't connected. Service points to Pod(s) directly, without connecting to the Deployment in any way.
Example: <html><code>kubectl create deployment web --image=nginx --replicas=3 --port=80</code></html>
Remember: Deployments manage ReplicaSets→Pods. <html><code>kubectl rollout</code></html> for updates.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-workloads.tsv</code></html>
''Related atoms''
* [[What is the difference between a Pod and a Deployment in Kubernetes?]]
* [[What happens behind the scenes when you create a Deployment object?]]
* [[How does rolling deployment work in Kubernetes?]]
Q: What happens when you delete a deployment?
A: The pod related to the deployment will terminate and the replicaset will be removed.
Remember: Deployment > ReplicaSet > Pod. Three layers for declarative management.
Example: <html><code>kubectl create deployment web --image=nginx:1.25 --replicas=3 --port=80</code></html>
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-workloads.tsv</code></html>
''Related atoms''
* [[What is the difference between a Pod and a Deployment in Kubernetes?]]
* [[How does rolling deployment work in Kubernetes?]]
* [[True or False? Deleting a ReplicaSet will delete the Pods it created]]
Q: What happens after you edit a deployment and change the image?
A: The pod will terminate and another, new pod, will be created.
Also, when looking at the replicaset, you'll see the old replica doesn't have any pods and a new replicaset is created.
Remember: Deployment > ReplicaSet > Pod. Three layers for declarative management.
Example: <html><code>kubectl create deployment web --image=nginx:1.25 --replicas=3 --port=80</code></html>
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-workloads.tsv</code></html>
''Related atoms''
* [[Kubernetes Deployment rolling updates and rollbacks]]
* [[What will happen when a Pod, created by ReplicaSet, is deleted directly with kubectl de…]]
* [[What is the relationship between Deployment, ReplicaSet, and Pod?]]
Q: What happens behind the scenes when you create a Deployment object?
A: The following occurs when you run <html><code>kubectl create deployment some_deployment --image=nginx</code></html>
# HTTP request sent to kubernetes API server on the cluster to create a new deployment
# A new Pod object is created and scheduled to one of the workers nodes
# Kublet on the worker node notices the new Pod and instructs the Container runtime engine to pull the image from the registry
# A new container is created using the image that was just pulled
Remember: Deployment > ReplicaSet > Pod. Three layers for declarative management.
Example: <html><code>kubectl create deployment web --image=nginx:1.25 --replicas=3 --port=80</code></html>
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-workloads.tsv</code></html>
''Related atoms''
* [[How Service and Deployment are connected?]]
* [[What is the difference between a StatefulSet and a Deployment in Kubernetes?]]
* [[What rollout/deployment strategies are you familiar with?]]
Q: Explain Canary deployments/rollouts in detail
A: Canary deployment steps:
# Traffic coming from users through a load balancer to the application which is currently version 1
Users -> Load Balancer -> App Version 1
# A new application version 2 is deployed (while version 1 still running) and part of the traffic is redirected to the new version
Users -> Load Balancer ->(95% of the traffic) App Version 1
->(5% of the traffic) App Version 2
3.
Remember: Deployment > ReplicaSet > Pod. Three layers for declarative management.
Example: <html><code>kubectl create deployment web --image=nginx:1.25 --replicas=3 --port=80</code></html>
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-workloads.tsv</code></html>
''Related atoms''
* [[What is a "Deployment" in Kubernetes?]]
* [[What happens behind the scenes when you create a Deployment object?]]
Q: What ways are you familiar with to implement deployment strategies (like canary, blue/green) in Kubernetes?
A: There are multiple ways. One of them is Argo Rollouts.
Remember: Deployment > ReplicaSet > Pod. Three layers for declarative management.
Example: <html><code>kubectl create deployment web --image=nginx:1.25 --replicas=3 --port=80</code></html>
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-workloads.tsv</code></html>
''Related atoms''
* [[How do you define a Kubernetes Deployment?]]
* [[What is a "Deployment" in Kubernetes?]]
* [[How does rolling deployment work in Kubernetes?]]
Q: What rollout/deployment strategies are you familiar with?
A: * Blue/Green Deployments: You deploy a new version of your app, while old version still running, and you start redirecting traffic to the new version of the app
* Canary Deployments: You deploy a new version of your app and start redirecting ''portion'' of your users/traffic to the new version. So you the migration to the new version is much more gradual
Remember: Deployment > ReplicaSet > Pod. Three layers for declarative management.
Example: <html><code>kubectl create deployment web --image=nginx:1.25 --replicas=3 --port=80</code></html>
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-workloads.tsv</code></html>
''Related atoms''
* [[What is a "Deployment" in Kubernetes?]]
* [[What happens when you delete a deployment?]]
* [[Create a deployment called "pluck" using the image "redis" and make sure it runs 5 repl…]]
Q: Explain Blue/Green deployments/rollouts in detail
A: Blue/Green deployment steps:
# Traffic coming from users through a load balancer to the application which is currently version 1
Users -> Load Balancer -> App Version 1
# A new application version 2 is deployed (while version 1 still running)
Users -> Load Balancer -> App Version 1
App Version 2
# If version 2 runs properly, traffic switched to it instead of version 1
User -> Load Balancer App version 1
-> App Version 2
4.
Remember: Deployment > ReplicaSet > Pod. Three layers for declarative management.
Example: <html><code>kubectl create deployment web --image=nginx:1.25 --replicas=3 --port=80</code></html>
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-workloads.tsv</code></html>
''Related atoms''
* [[What is a "Deployment" in Kubernetes?]]
* [[Create a file definition/manifest of a deployment called "dep", with 3 replicas that us…]]
* [[What happens behind the scenes when you create a Deployment object?]]
Q: What is the job of the kube-scheduler?
A: The kube-scheduler assigns nodes to newly created pods. It watches for newly created pods that have no node assigned, and selects a node for them to run on based on resource requirements, hardware/software/policy constraints, affinity and anti-affinity specifications, data locality, and workload interference.
Remember: Job = run-to-completion. Tracks successful completions. "One-time task."
Example: <html><code>kubectl create job db-migrate --image=myapp -- python manage.py migrate</code></html>
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-workloads.tsv</code></html>
''Related atoms''
* [[True or False? The scheduler is responsible for both deciding where a Pod will run and …]]
* [[Kubernetes custom schedulers: deploy and use]]
* [[Kubernetes DaemonSet: one pod per node]]
Q: How to list all daemonsets in the current namespace?
A: <html><code>kubectl get ds</code></html>
Remember: DaemonSet = one pod per node. For: log collectors, monitoring, network plugins.
Gotcha: DaemonSets use node affinity to ensure one copy per eligible node.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-workloads.tsv</code></html>
''Related atoms''
* [[You are looking for a Pod called "atreus". How to check in which namespace it runs?]]
* [[How to list ReplicaSets in the current namespace?]]
* [[Which command lists all Pods in a Kubernetes cluster?]]
Q: How to verify a deployment was created?
A: <html><code>kubectl get deployments</code></html> or <html><code>kubectl get deploy</code></html>
This command lists all the Deployment objects created and exist in the cluster. It doesn't mean the deployments are ready and running. This can be checked with the "READY" and "AVAILABLE" columns.
Example: <html><code>kubectl create deployment web --image=nginx --replicas=3 --port=80</code></html>
Remember: Deployments manage ReplicaSets→Pods. <html><code>kubectl rollout</code></html> for updates.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-workloads.tsv</code></html>
''Related atoms''
* [[Create a file definition/manifest of a deployment called "dep", with 3 replicas that us…]]
* [[ReplicaSets are running the moment the user executed the command to create them (like k…]]
* [[How Service and Deployment are connected?]]
Q: How to list ReplicaSets in the current namespace?
A: <html><code>kubectl get rs</code></html> lists all ReplicaSets in the current namespace showing NAME, DESIRED, CURRENT, and READY counts. Add <html><code>-o wide</code></html> to see container images and selectors. In practice, you rarely interact with ReplicaSets directly — Deployments create and manage them automatically during rollouts and rollbacks.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-workloads.tsv</code></html>
''Related atoms''
* [[How to list all daemonsets in the current namespace?]]
* [[How to check how many Pods are ready as part of a replica set called "repli"?]]
* [[You are looking for a Pod called "atreus". How to check in which namespace it runs?]]
Q: What is a "Deployment" in Kubernetes?
A: A Kubernetes Deployment is used to tell Kubernetes how to create or modify instances of the pods that hold a containerized application.
Deployments can scale the number of replica pods, enable rollout of updated code in a controlled manner, or roll back to an earlier deployment version if necessary.
A Deployment is a declarative statement for the desired state for Pods and Replica Sets.
Remember: Deployment > ReplicaSet > Pod. Three layers for declarative management.
Example: <html><code>kubectl create deployment web --image=nginx:1.25 --replicas=3 --port=80</code></html>
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-workloads.tsv</code></html>
//Merged from 2 source atoms.//
''Related atoms''
* [[Scaling a Deployment in Kubernetes]]
* [[What the following in a Deployment configuration file means?]]
* [[Create a file definition/manifest of a deployment called "dep", with 3 replicas that us…]]
Q: How does rolling deployment work in Kubernetes?
A: ''Rolling Deployment:''
* Strategy for updating an application without downtime.
* Involves gradually replacing old pods with new ones.
* Ensures a smooth transition, as each new pod is ready before an old one is terminated.
''Deployment Update:'' A new version of the application is deployed using a Deployment resource.
* ReplicaSet Transition: Kubernetes creates a new ReplicaSet for the updated version while maintaining the old one.
* Pod Replacement: Pods are gradually replaced by new ones, ensuring a controlled rollout.\
Example: <html><code>kubectl create deployment web --image=nginx --replicas=3 --port=80</code></html>
Remember: Deployments manage ReplicaSets→Pods. <html><code>kubectl rollout</code></html> for updates.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-workloads.tsv</code></html>
''Related atoms''
* [[What is the difference between a Pod and a Deployment in Kubernetes?]]
* [[What happens when you delete a deployment?]]
* [[How do you define a Kubernetes Deployment?]]
Q: How does Kubernetes ensure high availability of applications?
A: Kubernetes ensures high availability through several mechanisms:
* ReplicaSets: Deployments in Kubernetes are often managed by ReplicaSets, which ensure a specified number of replicas (pod instances) are running at all times. If a pod or node fails, ReplicaSets automatically create replacements on healthy nodes.
* Pod Distribution: Kubernetes spreads pods across multiple nodes to avoid a single point of failure.
Remember: Workload hierarchy: Deployment/StatefulSet/DaemonSet→ReplicaSet→Pod.
Gotcha: Use <html><code>kubectl explain <resource></code></html> for field reference without leaving the CLI.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-workloads.tsv</code></html>
''Related atoms''
* [[Discuss the implications of pod sprawl and how to manage it effectively in Kubernetes.]]
* [[Kubernetes ReplicaSet: maintaining stable pod replica count]]
* [[How does Kubernetes handle rolling updates with zero downtime?]]
Q: What are the challenges in managing stateful applications in Kubernetes?
A: * Challenges in Managing Stateful Applications: Persistent Storage: Ensuring data persistence and managing stateful data storage.
* Unique Network Identities: Assigning stable network identities to stateful pods.
* Orderly Scaling: Scaling stateful applications in a specific order to maintain relationships.
* Backup and Restore: Implementing robust backup and restore mechanisms for data.
* Stateful applications often have unique challenges compared to stateless ones, especially related to data persistence and maintaining stable identities.
Remember: Workload hierarchy: Deployment/StatefulSet/DaemonSet→ReplicaSet→Pod.
Gotcha: Use <html><code>kubectl explain <resource></code></html> for field reference without leaving the CLI.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-workloads.tsv</code></html>
//Merged from 2 source atoms.//
''Related atoms''
* [[How does Kubernetes manage containerized applications?]]
* [[What is the difference between a StatefulSet and a Deployment in Kubernetes?]]
Q: How to check how many Pods are ready as part of a replica set called "repli"?
A: <html><code>k describe rs repli | grep -i "Pods Status"</code></html>
Remember: Workload hierarchy: Deployment/StatefulSet/DaemonSet→ReplicaSet→Pod.
Gotcha: Use <html><code>kubectl explain <resource></code></html> for field reference without leaving the CLI.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-workloads.tsv</code></html>
''Related atoms''
* [[You've created a ReplicaSet, how to check whether the ReplicaSet found matching Pods or…]]
* [[How to list ReplicaSets in the current namespace?]]
* [[List all the pods with the label "env=prod"]]
<html><code>kubectl scale rs <name> --replicas=<n></code></html> sets the target replica count for a ReplicaSet. Kubernetes reconciles the current state toward the target:
* ''Scale up'': new pods are created immediately to reach the target count.
* ''Scale down'': excess pods are terminated gracefully, respecting <html><code>terminationGracePeriodSeconds</code></html> (default 30 s). If pods have <html><code>preStop</code></html> hooks or need to drain active connections, ensure the grace period is long enough to avoid dropped requests.
Example — scale up to 5: <html><code>kubectl scale rs rori --replicas=5</code></html>
Example — scale down to 1: <html><code>kubectl scale rs rori --replicas=1</code></html>
''Gotcha'': scaling a ReplicaSet directly may conflict with the desired state of its parent Deployment, which manages its own replica count. For production workloads, prefer scaling through the Deployment (<html><code>kubectl scale deploy</code></html>) or use a HorizontalPodAutoscaler (HPA) for automatic scaling based on CPU/memory metrics.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-workloads.tsv</code></html>
//Merged from 2 source atoms.//
''Related atoms''
* [[Kubernetes ReplicaSet: maintaining stable pod replica count]]
* [[What is the purpose of Horizontal Pod Autoscaling in Kubernetes?]]
Q: Describe the sequence of events in case of creating a ReplicaSet
A: * The client (e.g. kubectl) sends a request to the API server to create a ReplicaSet
* The Controller detects there is a new event requesting for a ReplicaSet
* The controller creates new Pod definitions (the exact number depends on what is defined in the ReplicaSet definition)
* The scheduler detects unassigned Pods and decides to which nodes to assign the Pods.
Remember: Don't create ReplicaSets directly. Use Deployments for rolling updates+rollbacks.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-workloads.tsv</code></html>
''Related atoms''
* [[Kubernetes ReplicaSet: maintaining stable pod replica count]]
* [[True or False? Pods specified by the selector field of ReplicaSet must be created by th…]]
* [[You've created a ReplicaSet, how to check whether the ReplicaSet found matching Pods or…]]
Q: What are some use cases for using a DaemonSet?
A: * Monitoring: You would like to perform monitoring on every node part of cluster. For example datadog pod runs on every node using a daemonset
* Logging: You would like to having logging set up on every node part of your cluster
* Networking: there is networking component you need on every node for all nodes to communicate between them
Remember: DaemonSet = one pod per node. For: log collectors, monitoring, network plugins.
Gotcha: DaemonSets use node affinity to ensure one copy per eligible node.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-workloads.tsv</code></html>
''Related atoms''
* [[DaemonSet vs Deployment vs ReplicaSet in Kubernetes]]
* [[Kubernetes DaemonSet: one pod per node]]
Q: What will happen when a Pod, created by ReplicaSet, is deleted directly with kubectl delete po ...?
A: The ReplicaSet will create a new Pod in order to reach the desired number of replicas.
Remember: Don't create ReplicaSets directly. Use Deployments for rolling updates+rollbacks.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-workloads.tsv</code></html>
''Related atoms''
* [[What happens after you edit a deployment and change the image?]]
* [[Kubernetes ReplicaSet: maintaining stable pod replica count]]
Q: You've created a ReplicaSet, how to check whether the ReplicaSet found matching Pods or it created new Pods?
A: <html><code>kubectl describe rs <ReplicaSet Name></code></html>
It will be visible under <html><code>Events</code></html> (the very last lines)
Remember: Don't create ReplicaSets directly. Use Deployments for rolling updates+rollbacks.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-workloads.tsv</code></html>
''Related atoms''
* [[Describe the sequence of events in case of creating a ReplicaSet]]
* [[How to check how many Pods are ready as part of a replica set called "repli"?]]
* [[How to list ReplicaSets in the current namespace?]]
Q: How to modify a replica set called "rori" to use a different image?
A: <html><code>kubectl edit rs rori</code></html> opens the ReplicaSet manifest in your default editor. Change the container image under spec.template.spec.containers[].image.
Gotcha: editing the ReplicaSet template does not automatically restart existing pods — you must delete old pods manually for the new image to take effect. Use Deployments for automatic rolling updates.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-workloads.tsv</code></html>
''Related atoms''
* [[How to delete a replica set called "rori"?]]
* [[What happens after you edit a deployment and change the image?]]
* [[How to check which container image was used as part of replica set called "repli"?]]
Q: What is the difference between a replica set and a replication controller?
A: A Replication Controller (RC) is a wrapper on a pod. This provides additional functionality to the pods, which offers replicas. It monitors the pods and automatically restarts them if they fail. If the node fails, this controller will respawn all the pods of that node on another node. If the pods die, they won't be spawned again unless wrapped around a replica set.
Replica Set (RS) is the next-generation replication controller. This kind of support has some selector types and supports both equality-based and set-based selectors. It allows filtering by label values and keys.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-workloads.tsv</code></html>
''Related atoms''
* [[True or False? In case of a ReplicaSet, if Pods specified in the selector field don't e…]]
* [[DaemonSet vs Deployment vs ReplicaSet in Kubernetes]]
* [[Kubernetes controllers and kube-controller-manager]]
Q: Create a deployment with the following properties:
A: <html><code>kubectl create deployment blufer --image=python --replicas=3 -o yaml --dry-run=client > deployment.yaml</code></html>
Add the following section (<html><code>vi deployment.yaml</code></html>):
<html><pre><code class="language-plaintext">spec:
affinity:
nodeAffinity:
requiredDuringSchedulingIgnoredDuringExecution:
nodeSelectorTerms:
- matchExpressions:
- key: blufer
operator: Exists</code></pre></html>
<html><code>kubectl apply -f deployment.yaml</code></html>
Example: <html><code>kubectl create deployment web --image=nginx --replicas=3 --port=80</code></html>
Remember: Deployments manage ReplicaSets→Pods. <html><code>kubectl rollout</code></html> for updates.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-workloads.tsv</code></html>
''Related atoms''
* [[Create a deployment called "pluck" using the image "redis" and make sure it runs 5 repl…]]
* [[Scaling a Deployment in Kubernetes]]
* [[How to create a deployment with the image "nginx:alpine"?]]
Q: How to edit a deployment?
A: <html><code>kubectl edit deployment <DEPLOYMENT_NAME></code></html> opens the manifest in your editor for live editing. Changes to the pod template (image, env, resources) trigger an automatic rolling update. Alternative: <html><code>kubectl set image deployment/<name> container=image:tag</code></html> for quick image updates without opening an editor. Use <html><code>kubectl rollout status</code></html> to monitor the update progress.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-workloads.tsv</code></html>
''Related atoms''
* [[What happens after you edit a deployment and change the image?]]
* [[How does rolling deployment work in Kubernetes?]]
* [[How do you define a Kubernetes Deployment?]]
Q: True or False? The same as there are "Static Pods" there are other static resources like "deployments" and "replicasets"
A: False. Static Pods are unique — they are managed directly by the kubelet on a specific node, not by the API server. Deployments, ReplicaSets, and other workloads are API-managed resources with no 'static' equivalent.
Remember: Don't create ReplicaSets directly. Use Deployments for rolling updates+rollbacks.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-workloads.tsv</code></html>
''Related atoms''
* [[Kubernetes Static Pods: definition, use cases, and DaemonSet contrast]]
* [[How to identify which Pods are Static Pods?]]
* [[How does Kubernetes ensure high availability of applications?]]
A Service provides a stable Layer 4 (transport) endpoint for pod communication — it assigns a cluster-internal IP and DNS name, handles load balancing across pod replicas, and abstracts away pod churn. An Ingress sits on top, adding Layer 7 (application/HTTP) routing decisions. It can route based on hostnames, paths, and other HTTP attributes, directing traffic to Services rather than directly to pods. This layering means you always define a Service first, then point your Ingress rules at that Service. Services come in types — ClusterIP for internal use, NodePort for node-level exposure, LoadBalancer for cloud-provisioned external endpoints — while Ingress is a single, standardized way to express L7 routing intent, delegating to whatever controller implementation is installed.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-services.tsv</code></html>
* <html><code>training/library/topics/mental-models-core/k8s-service-vs-ingress.md</code></html>
//Merged from 5 source atoms.//
''Related atoms''
* [[Kubernetes Service: stable networking endpoint for pods]]
When the kernel's OOM killer terminates a process, it sends SIGKILL, which cannot be caught or handled gracefully. The process exits with code 137, which equals 128 + 9 (SIGKILL signal number). This is the clearest diagnostic marker for an OOMKill event. Unlike exit code 143 (128 + 15, SIGTERM), which indicates a graceful shutdown, an exit code of 137 means the process had no opportunity to clean up, flush buffers, or close connections. In Kubernetes, you can identify OOMKilled pods by checking <html><code>lastState.terminated.reason</code></html> in <html><code>kubectl describe pod</code></html> output or by querying the pod's status JSON. Because the kernel kills immediately without notification, there is no final log line — the container simply stops. Monitoring for exit code 137 is the most reliable way to catch OOM events in practice.
----
''Sources''
* <html><code>training/interactive/knowledge/data/cards/k8s-troubleshooting.tsv</code></html>
* <html><code>training/library/topics/oomkilled/trivia.md</code></html>
* <html><code>training/library/topics/oomkilled/street_ops.md</code></html>
* <html><code>training/library/topics/oomkilled/primer.md</code></html>
//Merged from 4 source atoms.//
''Related atoms''
* [[Pod deletion: SIGTERM, grace period, and SIGKILL]]
* [[Which Prometheus metric should you use to predict OOMKill, and why not container_memory…]]
* [[A pod is OOMKilled but the app's memory usage looks normal. What happened?]]
!! MOC — flashcards
Every atom extracted from a ''flashcard'' source (483 total).
* [[A PVC is stuck in Pending state. Walk through your debugging process.]]
* [[A container is OOMKilled but the application memory profiler shows usage well below the…]]
* [[A container keeps restarting with exit code 137. Describe your troubleshooting steps.]]
* [[A developer reports they cannot exec into pods despite having pod access. Walk through …]]
* [[A developer says their app can't reach a service. They're in different namespaces. What…]]
* [[A pod can reach the internet but not other pods. What do you check?]]
* [[A pod fails to start with a multi-attach error on an RWO volume. What happened and how …]]
* [[A pod has a main container limited to 512Mi and an Istio sidecar limited to 256Mi. The …]]
* [[A pod is OOMKilled but the app's memory usage looks normal. What happened?]]
* [[A pod is stuck in ImagePullBackOff. What are the common causes?]]
* [[A pod is stuck in Init:0/2 status. How do you debug it?]]
* [[A pod with a PVC can't start and shows a multi-attach error. What's wrong?]]
* [[A rolling update is stuck because the PDB minAvailable equals the replica count. Why is…]]
* [[After creating a service that forwards incoming external traffic to the containerized a…]]
* [[After creating a service, how to check it was created?]]
* [[After running kubectl run database --image mongo you see the status is "CrashLoopBackOf…]]
* [[Aggregated ClusterRoles compose permissions via label selectors]]
* [[An app takes 120 seconds to initialize. Liveness probe kills it before startup complete…]]
* [[An engineer form your organization asked whether there is a way to prevent from Pods (w…]]
* [[An engineer form your organization told you he is interested only in seeing his team re…]]
* [[An internal load balancer in Kubernetes is called ____ and an external load balancer is…]]
* [[Are there any tools, projects you are using for building Operators?]]
* [[Assuming you have multiple schedulers, how to know which scheduler was used for a given…]]
* [[Bare pods have no controller and stay dead on crash]]
* [[Binding type determines scope; ClusterRole can be namespace-scoped]]
* [[CNI plugin: Kubernetes container networking standard]]
* [[CNI plugins: Calico, Cilium, and Flannel compared]]
* [[CSI architecture in Kubernetes and driver health verification]]
* [[Check Deployment rollout history and roll back to a specific revision]]
* [[Check how many namespaces are there]]
* [[Check if there are any limits on one of the pods in your cluster]]
* [[Check if there are taints on node "master"]]
* [[Check what labels one of your nodes in the cluster has]]
* [[ClusterIP: Kubernetes default internal-only Service type]]
* [[Complete the following configuration file to make it Ingress]]
* [[ConfigMap updates do not automatically restart Pods]]
* [[Conntrack table exhaustion silently drops Kubernetes NAT connections]]
* [[Create a deployment called "pluck" using the image "redis" and make sure it runs 5 repl…]]
* [[Create a deployment with the following properties:]]
* [[Create a file definition/manifest of a deployment called "dep", with 3 replicas that us…]]
* [[Create a list of all nodes in JSON format and store it in a file called "some_nodes.json"]]
* [[Create a pod called "kartos" in the namespace dev. The pod should be using the "redis" …]]
* [[Create a static pod with the image python that runs the command sleep 2017]]
* [[Create a taint on one of the nodes in your cluster with key of "app" and value of "web"…]]
* [[DaemonSet vs Deployment vs ReplicaSet in Kubernetes]]
* [[Debugging CrashLoopBackOff pods: logs, debug containers, node inspection]]
* [[Debugging a failing or non-starting pod in Kubernetes]]
* [[Default-deny egress NetworkPolicies must explicitly allow DNS]]
* [[Deleting a Deployment (not its pods) is the correct removal path]]
* [[Dependency checks in liveness probes cause cascading restarts]]
* [[Deploy a pod called "my-pod" using the nginx:alpine image]]
* [[Describe a least-privilege RBAC pattern for a CI/CD deployer service account.]]
* [[Describe how would you delete a static Pod]]
* [[Describe in detail what happens when you create a service]]
* [[Describe in detail what is the Operator Lifecycle Manager]]
* [[Describe in high level what happens when you run kubectl expose deployment remo --type=…]]
* [[Describe in high-level how Kustomize works]]
* [[Describe shortly and in high-level, what happens when you run kubectl get nodes]]
* [[Describe the Kubernetes API versioning strategy.]]
* [[Describe the role of etcd in a Kubernetes cluster.]]
* [[Describe the sequence of events in case of creating a ReplicaSet]]
* [[Describe what happens when a container tries to connect with its corresponding Service …]]
* [[Diagnose and fix failing Kubernetes liveness/readiness probes]]
* [[Diagnosing Kubernetes pods stuck in Pending state]]
* [[Diagnosing a stuck Kubernetes Deployment rollout]]
* [[Diagnosing pods that won't schedule on a node due to taints]]
* [[Discuss the considerations for migrating an application from a monolithic architecture …]]
* [[Discuss the differences between OpenShift and vanilla Kubernetes.]]
* [[Discuss the implications of pod sprawl and how to manage it effectively in Kubernetes.]]
* [[Discuss the relationship between Kubernetes and container runtimes like Docker and cont…]]
* [[Discuss the role of kube-proxy in Kubernetes networking.]]
* [[Distinguish CrashLoopBackOff from other pod failure states]]
* [[Do you have experience with deploying a Kubernetes cluster? If so, can you describe the…]]
* [[Ephemeral vs. Persistent Volumes in Kubernetes Pods]]
* [[Exit code 137 in containers means OOM kill, not app error]]
* [[Explain "Dynamic Provisioning" and "Static Provisioning"]]
* [[Explain "Security Context"]]
* [[Explain Blue/Green deployments/rollouts in detail]]
* [[Explain Canary deployments/rollouts in detail]]
* [[Explain how ConfigMap and Secret updates are handled in Kubernetes.]]
* [[Explain how Service Accounts are different from User Accounts]]
* [[Explain how you would manage configuration drift in a Kubernetes environment.]]
* [[Explain how you would monitor and scale a critical production application in Kubernetes.]]
* [[Explain the Helm Chart Directory Structure]]
* [[Explain the concept of Namespaces in Kubernetes.]]
* [[Explain the concept of PodDisruptionBudget in Kubernetes.]]
* [[Explain the differences between ClusterIP, NodePort, and LoadBalancer service types.]]
* [[Explain the escalate and bind verbs. Why are they dangerous?]]
* [[Explain the main components of Kubernetes architecture.]]
* [[Explain the meaning of "http", "host" and "backend" directives]]
* [[Explain the need for Kustomize by describing actual use cases]]
* [[Explain the purpose of kubelet in the Kubernetes cluster.]]
* [[Explain the purpose of the following lines]]
* [[Explain the selectPolicy field in HPA behavior and how Max vs Min affect scaling aggres…]]
* [[Explain the working of the master node in Kubernetes?]]
* [[Explain what are "Service Accounts" and in which scenario would use create/use one]]
* [[Explain what is CronJob and what is it used for]]
* [[Explain what will happen when running apply on the following block]]
* [[Explain why one would specify resource limits in regards to Pods]]
* [[Fix the following ReplicaSet definition]]
* [[Fix the following deployment manifest]]
* [[HPA and VPA on the same metric create conflicting feedback loops]]
* [[HPA keeps scaling to max replicas even when average CPU is low. What could cause this?]]
* [[HPA requires resource requests to calculate utilization percentage]]
* [[HPA shows <unknown>/80% for CPU target and won't scale. What is wrong?]]
* [[Helm: Kubernetes package manager for bundling and distributing YAML]]
* [[How Helm supports release management?]]
* [[How Service and Deployment are connected?]]
* [[How are Kubernetes and Docker related?]]
* [[How can PodDisruptionBudget conflict with HPA scale-down, and what is the best practice…]]
* [[How can you find out information on a Service related to a certain Pod if all you can u…]]
* [[How can you get a static IP for a Kubernetes load balancer?]]
* [[How do init containers help prevent CrashLoopBackOff caused by missing dependencies?]]
* [[How do readiness probes interact with rolling deployments?]]
* [[How do you capture network traffic inside a running pod without modifying its image?]]
* [[How do you define a Kubernetes Deployment?]]
* [[How do you determine if a pod is Pending due to resource pressure?]]
* [[How do you distinguish a container-level OOM from a node-level OOM, and what commands r…]]
* [[How do you find pods that match a particular label selector?]]
* [[How do you implement encryption for data in transit and at rest in Kubernetes?]]
* [[How do you list deployed releases?]]
* [[How do you prevent high memory usage in your Kubernetes cluster and possibly issues lik…]]
* [[How do you right-size memory limits for a Kubernetes deployment?]]
* [[How do you search for charts?]]
* [[How do you test connectivity to a Service from inside the cluster?]]
* [[How do you use kubectl auth can-i to debug RBAC permissions?]]
* [[How do you use kubectl events for troubleshooting?]]
* [[How do you verify that an ImagePullSecret is correct?]]
* [[How do you verify that metrics-server is running and providing data?]]
* [[How does Kubernetes ensure high availability of applications?]]
* [[How does Kubernetes handle DNS resolution for services and pods?]]
* [[How does Kubernetes handle rolling updates with zero downtime?]]
* [[How does Kubernetes integrate with cloud providers like AWS, Azure, and GCP?]]
* [[How does Kubernetes manage containerized applications?]]
* [[How does Kubernetes manage security, and what are some best practices?]]
* [[How does PVC expansion work and what are the operational risks?]]
* [[How does a Kubernetes Service find the right Pods to route traffic to?]]
* [[How does rolling deployment work in Kubernetes?]]
* [[How does the Vertical Pod Autoscaler help prevent OOMKilled, and what are its risks?]]
* [[How make an app accessible on private or external network?]]
* [[How many containers can a pod contain?]]
* [[How readiness probe status affect Services when they are combined?]]
* [[How should liveness and readiness probe endpoints differ in what they check, and why?]]
* [[How to check how many Pods are ready as part of a replica set called "repli"?]]
* [[How to check to which worker node the pods were scheduled to? In other words, how to ch…]]
* [[How to check which container image was used as part of replica set called "repli"?]]
* [[How to commit secrets to Git and in general how to use encrypted secrets?]]
* [[How to configure TLS with Ingress?]]
* [[How to configure a default backend?]]
* [[How to confirm a container is running after running the command kubectl run web --image…]]
* [[How to create a Secret from a key and value?]]
* [[How to create a deployment with the image "nginx:alpine"?]]
* [[How to create a pod and a service with one command?]]
* [[How to create components in a namespace?]]
* [[How to delete a deployment?]]
* [[How to delete a replica set called "rori"?]]
* [[How to delete all pods whose status is not "Running"?]]
* [[How to display the resources usages of pods?]]
* [[How to edit a deployment?]]
* [[How to execute the command "ls" in an existing pod?]]
* [[How to get information on a certain service?]]
* [[How to get list of resources which are not bound to a specific namespace?]]
* [[How to get the name of the current namespace?]]
* [[How to identify which Pods are Static Pods?]]
* [[How to list Ingress in your namespace?]]
* [[How to list ReplicaSets in the current namespace?]]
* [[How to list Service Accounts?]]
* [[How to list all daemonsets in the current namespace?]]
* [[How to list the endpoints of a certain app?]]
* [[How to modify a replica set called "rori" to use a different image?]]
* [[How to scale a deployment to 8 replicas?]]
* [[How to schedule a pod on a node called "node1"?]]
* [[How to switch to another namespace? In other words how to change active namespace?]]
* [[How to turn the following service into an external one?]]
* [[How to upgrade a release?]]
* [[How to use ConfigMaps?]]
* [[How to verify a deployment was created?]]
* [[How to verify that a certain service configured to forward the requests to a given pod]]
* [[How to view revision history for a certain release?]]
* [[How view all the pods running in all the namespaces?]]
* [[How would you approach version upgrades of Kubernetes in a production environment?]]
* [[How would you map a service to an external address?]]
* [[In a Go-based operator using Kubebuilder, what does returning ctrl.Result{RequeueAfter:…]]
* [[In case of a ReplicaSet, Which field is mandatory in the spec section?]]
* [[In case of two pods, if there is an egress policy on the source denining traffic and in…]]
* [[Ingress 404: controller missing, wrong ingressClassName, or no endpoints]]
* [[Investigate and recover a Kubernetes NotReady node]]
* [[Is it possible to delete ReplicaSet without deleting the Pods it created?]]
* [[Is it possible to override values in values.yaml file when installing a chart?]]
* [[It is said that Helm is also Templating Engine. What does it mean?]]
* [[JVM GC pauses cause liveness probe failures if timeoutSeconds is too low]]
* [[KEDA extends HPA to scale on 60+ event sources and to zero]]
* [[Karpenter vs Cluster Autoscaler: key differences]]
* [[Kubernetes Annotations vs Labels]]
* [[Kubernetes Control Plane: components and responsibilities]]
* [[Kubernetes DNS automatically assigns and resolves names to services and pods]]
* [[Kubernetes DaemonSet: one pod per node]]
* [[Kubernetes Deployment rolling updates and rollbacks]]
* [[Kubernetes Labels and Selectors]]
* [[Kubernetes NetworkPolicy: whitelist semantics and operational traps]]
* [[Kubernetes Operator: components and control loop pattern]]
* [[Kubernetes Operator: purpose and usage]]
* [[Kubernetes PVC expansion: StorageClass gate and two-phase resize]]
* [[Kubernetes Pod: the smallest deployable unit]]
* [[Kubernetes QoS classes: criteria and eviction order]]
* [[Kubernetes ReplicaSet: maintaining stable pod replica count]]
* [[Kubernetes Secrets are base64-encoded, not encrypted, in etcd by default]]
* [[Kubernetes Secrets: storage, access, and best practices]]
* [[Kubernetes Service empty endpoints: label-selector mismatch diagnosis]]
* [[Kubernetes Service: stable networking endpoint for pods]]
* [[Kubernetes Static Pods: definition, use cases, and DaemonSet contrast]]
* [[Kubernetes controllers and kube-controller-manager]]
* [[Kubernetes custom schedulers: deploy and use]]
* [[Kubernetes does not provide data persistence by default]]
* [[Kubernetes dynamic provisioning and WaitForFirstConsumer zone awareness]]
* [[Kubernetes headless services return pod IPs directly via DNS]]
* [[Kubernetes ingress controllers: edge routing and policy enforcement]]
* [[Kubernetes memory failure modes: eviction vs OOMKill]]
* [[Kubernetes monitoring solutions overview]]
* [[Kubernetes ndots:5 causes excessive external DNS queries]]
* [[Kubernetes node resource pressure: diagnosis, eviction, and prevention]]
* [[Kubernetes node upgrade procedure]]
* [[Kubernetes originally used ABAC authorization; RBAC became the default in version 1.8]]
* [[Kubernetes storage abstraction: PV, PVC, and StorageClass layers]]
* [[Kubernetes taints and tolerations: effects, syntax, and use cases]]
* [[Kubernetes: definition, origin, and core features]]
* [[List all the pods with the label "env=prod"]]
* [[Manual node scaling vs Cluster Autoscaler in Kubernetes]]
* [[Memory is a poor primary HPA metric because apps hold allocations]]
* [[Name the initial namespaces from which Kubernetes starts?]]
* [[Name three operator-building frameworks and when you would choose each.]]
* [[OOMKill sends SIGKILL (exit code 137)]]
* [[OPA Gatekeeper: policy enforcement webhook in Kubernetes]]
* [[PV reclaim policies: Retain preserves data, Delete destroys it]]
* [[Perhaps a general question but, you suspect one of the pods is having issues, you don't…]]
* [[Pod IPs are ephemeral; use a Service for stable addressing]]
* [[Pod deletion: SIGTERM, grace period, and SIGKILL]]
* [[Pod reaches Service by ClusterIP but internal DNS fails: diagnosis]]
* [[PodDisruptionBudgets block kubectl drain indefinitely]]
* [[Pods are Pending cluster-wide after a control plane upgrade. All worker nodes show a No…]]
* [[Pods are being evicted with the message 'The node was low on resource: ephemeral-storag…]]
* [[Pods from the same Deployment keep landing on the same node, causing a single point of …]]
* [[Pods without memory limits cause node-wide evictions and OOM kills]]
* [[RBAC subresources require explicit rules; parent grants do not cascade]]
* [[ReplicaSets are running the moment the user executed the command to create them (like k…]]
* [[Role is namespace-scoped; ClusterRole is cluster-wide]]
* [[Role of kube-apiserver in Kubernetes]]
* [[Run a command to view all nodes of the cluster]]
* [[Run a pod called "yay2" with the image "python". Make sure it has resources request of …]]
* [[Scale a ReplicaSet with kubectl scale rs]]
* [[Scaling a Deployment in Kubernetes]]
* [[Service account tokens are mounted by default even if the pod never calls the Kubernetes API]]
* [[Service and Ingress divide Layer 4 and Layer 7 routing]]
* [[Service type LoadBalancer provisions cloud load balancer]]
* [[Startup probes replaced the initialDelaySeconds compromise in Kubernetes 1.18]]
* [[StatefulSets and volumeClaimTemplates: stable storage per replica]]
* [[Tell me about your Kubernetes experience.]]
* [[True or False? A single Pod can be split across multiple nodes]]
* [[True or False? A volume defined in Pod can be accessed by all the containers of that Pod]]
* [[True or False? By default there is no communication between two Pods in two different n…]]
* [[True or False? By default, pods are isolated. This means they are unable to receive tra…]]
* [[True or False? Deleting a ReplicaSet will delete the Pods it created]]
* [[True or False? Each Pod, when created, gets its own public IP address]]
* [[True or False? Every cluster must have 0 or more master nodes and at least 1 worker]]
* [[True or False? If a ReplicaSet defines 2 replicas but there 3 Pods running matching the…]]
* [[True or False? If no network policies are applied to a pod, then no connections to or f…]]
* [[True or False? In case of a ReplicaSet, if Pods specified in the selector field don't e…]]
* [[True or False? Memory is a compressible resource, meaning that when a container reach t…]]
* [[True or False? Once a Pod is assisgned to a worker node, it will only run on that node,…]]
* [[True or False? Pods specified by the selector field of ReplicaSet must be created by th…]]
* [[True or False? Removing the label from a Pod that is tracked by a ReplicaSet, will caus…]]
* [[True or False? Resource limits applied on a Pod level meaning, if limits is 2gb RAM and…]]
* [[True or False? Sensitive data, like credentials, should be stored in a ConfigMap]]
* [[True or False? The "Pending" phase means the Pod was not yet accepted by the Kubernetes…]]
* [[True or False? The same as there are "Static Pods" there are other static resources lik…]]
* [[True or False? The scheduler is responsible for both deciding where a Pod will run and …]]
* [[True or False? Using the node affinity type "preferredDuringSchedulingIgnoredDuringExec…]]
* [[True or False? When a namespace is deleted all resources in that namespace are not dele…]]
* [[True or False? With namespaces you can limit the resources consumed by the users/teams]]
* [[True or False? storing data in a Secret component makes it automatically secured]]
* [[True or False? the target port, in the case of running the following command, will be e…]]
* [[Users unable to reach an application running on a Pod on Kubernetes. What might be the …]]
* [[Using node affinity, set a Pod to schedule on a node where the key is "region" and valu…]]
* [[Volume Snapshots in Kubernetes]]
* [[What DNS record format does Kubernetes create for a ClusterIP Service?]]
* [[What Kubernetes objects are there?]]
* [[What Kubernetes objects do you usually use when deploying applications in Kubernetes?]]
* [[What QoS classes are there?]]
* [[What actions or operations you consider as best practices when it comes to Kubernetes?]]
* [[What are Custom Resource Definitions (CRDs) in Kubernetes?]]
* [[What are Kubernetes selectors and how do they match resources?]]
* [[What are all the phases/steps of a control loop?]]
* [[What are common causes of kubelet failures and how do they manifest?]]
* [[What are federated clusters?]]
* [[What are finalizers in the context of Kubernetes operators, and why are they needed?]]
* [[What are important steps in defining/adding a Service?]]
* [[What are some of Kubernetes features?]]
* [[What are some use cases for using Helm template file?]]
* [[What are some use cases for using Ingress?]]
* [[What are some use cases for using Network Policies?]]
* [[What are some use cases for using a DaemonSet?]]
* [[What are the challenges in managing stateful applications in Kubernetes?]]
* [[What are the components of a worker node (aka data plane)?]]
* [[What are the five levels of the operator maturity model?]]
* [[What are the four metric types supported by HPA v2 and when would you use each?]]
* [[What are the four probe mechanisms Kubernetes supports?]]
* [[What are the main differences between Docker Swarm and Kubernetes?]]
* [[What are the possible Pod phases?]]
* [[What are the prerequisites for HPA to work?]]
* [[What are the security risks of using the default ServiceAccount and how do you audit fo…]]
* [[What are the standard RBAC verbs in Kubernetes and what API operations do they map to?]]
* [[What are the three classic PersistentVolume access modes and what does each allow?]]
* [[What are the top 3 causes of CrashLoopBackOff?]]
* [[What are the types of controller managers?]]
* [[What are your thoughts on "Pods are not meant to be created directly"?]]
* [[What can you find in kube-system namespace?]]
* [[What causes a node drain to get stuck and how do you troubleshoot it?]]
* [[What challenges do you anticipate when managing large-scale Kubernetes clusters, and ho…]]
* [[What combination of mechanisms ensures zero-downtime during node maintenance?]]
* [[What components the Operator Framework consists of?]]
* [[What do you understand by Cloud controller manager?]]
* [[What does '<unknown>/50%' mean in HPA status?]]
* [[What does HPA do and what problem remains after you enable it?]]
* [[What does a Kubernetes liveness probe determine?]]
* [[What does being cloud-native mean?]]
* [[What does failureThreshold control, and how does it interact with periodSeconds?]]
* [[What does the "ErrImagePull" status of a Pod means?]]
* [[What does the node status contain?]]
* [[What fields are mandatory with any Kubernetes object?]]
* [[What formula does the HPA use to compute the desired replica count?]]
* [[What happens after you edit a deployment and change the image?]]
* [[What happens behind the scenes when you create a Deployment object?]]
* [[What happens to running pods if if you stop Kubelet on the worker nodes?]]
* [[What happens when a readiness probe fails on a Kubernetes pod?]]
* [[What happens when you apply a NetworkPolicy with an empty podSelector and policyTypes: …]]
* [[What happens when you delete a deployment?]]
* [[What happens when you run a Pod with kubectl?]]
* [[What happens when you set replicas in a Deployment manifest and also use HPA?]]
* [[What happens you create a pod and you DON'T specify a service account?]]
* [[What is Conftest and how does it validate configuration files?]]
* [[What is Container resource monitoring?]]
* [[What is Datree? How is it different from Conftest?]]
* [[What is Heapster in Kubernetes?]]
* [[What is Helm, and how is it used in Kubernetes?]]
* [[What is Ingress Controller?]]
* [[What is Ingress Default Backend?]]
* [[What is Istio? What is it used for?]]
* [[What is Minikube and when would you use it?]]
* [[What is NodePort service type in Kubernetes?]]
* [[What is PodSecurity and how can it be configured in a Kubernetes cluster?]]
* [[What is Resource Quota?]]
* [[What is a "Deployment" in Kubernetes?]]
* [[What is a ConfigMap, and how is it used in Kubernetes?]]
* [[What is a Kubernetes Cluster?]]
* [[What is a Kubernetes Secret and how does it store sensitive data?]]
* [[What is a Kubernetes StatefulSet and when would you use it?]]
* [[What is a Kubernetes cluster and what are its components?]]
* [[What is a LimitRange and how does it prevent OOMKilled caused by missing resource limits?]]
* [[What is a Persistent Volume (PV) and Persistent Volume Claim (PVC) in Kubernetes?]]
* [[What is a PodDisruptionBudget (PDB) and why does it create tension with node drains?]]
* [[What is a cluster of containers in Kubernetes?]]
* [[What is a node in Kubernetes?]]
* [[What is a volume in regards to Kubernetes?]]
* [[What is an "Ingress"?]]
* [[What is kubconfig? What do you use it for?]]
* [[What is kubectl and how is it used to manage Kubernetes clusters?]]
* [[What is orchestration when it comes to software and DevOps?]]
* [[What is the Google Container Engine?]]
* [[What is the Kubernetes concept chain?]]
* [[What is the Operator Framework?]]
* [[What is the PID 1 problem in containers and how does it cause exit code 137 on pod term…]]
* [[What is the constraint on successThreshold for liveness and startup probes, and why doe…]]
* [[What is the default HPA scale-down stabilization window?]]
* [[What is the difference between Immediate and WaitForFirstConsumer volume binding modes?]]
* [[What is the difference between a Job and a CronJob in Kubernetes?]]
* [[What is the difference between a Pod and a Deployment in Kubernetes?]]
* [[What is the difference between a StatefulSet and a Deployment in Kubernetes?]]
* [[What is the difference between a replica set and a replication controller?]]
* [[What is the difference between cordoning and draining a node?]]
* [[What is the difference between deploying applications on hosts and containers?]]
* [[What is the difference between exit codes 126 and 127 in a container?]]
* [[What is the difference between liveness and readiness probes in Kubernetes?]]
* [[What is the difference between resource requests and limits, and which one affects sche…]]
* [[What is the difference between resources.requests.memory and resources.limits.memory in…]]
* [[What is the fundamental rule of the Kubernetes pod networking model?]]
* [[What is the job of the kube-scheduler?]]
* [[What is the key difference between kube-proxy iptables mode and IPVS mode?]]
* [[What is the lifecycle of a Kubernetes node?]]
* [[What is the problem with the following Secret file:]]
* [[What is the purpose of Horizontal Pod Autoscaling in Kubernetes?]]
* [[What is the purpose of a StorageClass in Kubernetes?]]
* [[What is the purpose of a startup probe?]]
* [[What is the purpose of an admission controller in Kubernetes, and how can you extend it?]]
* [[What is the reconciliation loop in a Kubernetes operator?]]
* [[What is the relationship between Deployment, ReplicaSet, and Pod?]]
* [[What is the relationship between a PersistentVolume (PV) and a PersistentVolumeClaim (P…]]
* [[What is the role of kube-apiserver aggregation layer in Kubernetes?]]
* [[What is the standard three-command diagnostic flow when a pod is misbehaving?]]
* [[What issue might arise from using the following CronJob and how to fix it?]]
* [[What kube-node-lease contains?]]
* [[What kube-public contains?]]
* [[What kubectl command shows whether a pod was OOMKilled, and what fields do you look for?]]
* [[What kubectl describe pod [pod name] does? command does?]]
* [[What kubectl get componentstatus does?]]
* [[What node conditions does Kubernetes monitor for health, and what happens when a node g…]]
* [[What openshift-operator-lifecycle-manager namespace includes?]]
* [[What possible issue can arise from using the following spec and how to fix it?]]
* [[What problem does Ingress solve that Services alone cannot?]]
* [[What problem does a ConfigMap solve?]]
* [[What problem does a Deployment solve that bare Pods cannot?]]
* [[What problems does Kubernetes actually solve?]]
* [[What process is responsible for running and installing the different controllers?]]
* [[What process runs on Kubernetes Master Node?]]
* [[What reclaim policies are there?]]
* [[What rollout/deployment strategies are you familiar with?]]
* [[What security best practices do you follow in regards to the Kubernetes cluster?]]
* [[What special namespaces are there by default when creating a Kubernetes cluster?]]
* [[What the following block of lines does?]]
* [[What the following command does?]]
* [[What the following in a Deployment configuration file means?]]
* [[What the following output of kubectl get rs means?]]
* [[What the master node is responsible for?]]
* [[What three kubectl commands form the basic CrashLoopBackOff diagnostic workflow?]]
* [[What type: Opaque in a secret file means? What other types are there?]]
* [[What types of persistent volumes are there?]]
* [[What use cases exist for running multiple containers in a single pod?]]
* [[What volume types are you familiar with?]]
* [[What ways are you familiar with to implement deployment strategies (like canary, blue/g…]]
* [[What will happen when a Pod, created by ReplicaSet, is deleted directly with kubectl de…]]
* [[What would you use to route traffic from outside the Kubernetes cluster to services wit…]]
* [[What's the difference between ImagePullBackOff and ErrImagePull?]]
* [[When HPA is configured with multiple metrics, how does it decide the replica count?]]
* [[When do you use kubectl logs vs kubectl describe?]]
* [[When or why NOT to use Kubernetes?]]
* [[When would you use the "LoadBalancer" type]]
* [[Where static Pods manifests are located?]]
* [[Which Kubernetes concept would you use to control traffic flow at the IP address or por…]]
* [[Which Prometheus metric should you use to predict OOMKill, and why not container_memory…]]
* [[Which command lists all Pods in a Kubernetes cluster?]]
* [[Which command will list all the object types in a cluster?]]
* [[Which components can't be created within a namespace?]]
* [[Which problems, volumes in Kubernetes solve?]]
* [[Which resources are accessible from different namespaces?]]
* [[Which service and in which namespace the following file is referencing?]]
* [[While namespaces do provide scope for resources, they are not isolating them]]
* [[Why do Ingress resources do nothing by themselves?]]
* [[Why do Kubernetes clusters fail at scale?]]
* [[Why do container logs sometimes bring down Kubernetes nodes?]]
* [[Why do some teams set CPU requests but no CPU limits?]]
* [[Why do we need Operators?]]
* [[Why does Kubernetes stress Linux more than VMs?]]
* [[Why does a Java application with -Xmx1g in a container limited to 512Mi get OOMKilled, …]]
* [[Why etcd? Why not some SQL or NoSQL database?]]
* [[Why is a PodDisruptionBudget important when using node autoscaling?]]
* [[Why is etcd performance tied to disk latency more than CPU?]]
* [[Why it's common to have only one container per Pod in most cases?]]
* [[Why must the Reconcile function in a Kubernetes operator be idempotent?]]
* [[Why should operators use the status subresource for status updates instead of updating …]]
* [[Why should you use a Secret instead of a ConfigMap for passwords?]]
* [[Why there is no such command in Kubernetes? kubectl get containers]]
* [[Why to use namespaces? What is the problem with using one default namespace?]]
* [[Why using a wildcard in ingress host may lead to issues?]]
* [[Wildcard RBAC rules grant unrestricted access and enable privilege escalation]]
* [[Would you use Helm, Go or something else for creating an Operator?]]
* [[You applied a taint with k taint node minikube app=web:NoSchedule on the only node in y…]]
* [[You are looking for a Pod called "atreus". How to check in which namespace it runs?]]
* [[You are managing multiple Kubernetes clusters. How do you quickly change between the cl…]]
* [[You create a resource but cannot find it with kubectl get. What namespace-related mista…]]
* [[You deploy an Envoy sidecar with your app container. After a rolling update, requests f…]]
* [[You encounter a performance issue in a Kubernetes cluster. How do you diagnose and reso…]]
* [[You have a microservices-based application. How would you deploy and manage it in Kuber…]]
* [[You have one Kubernetes cluster and multiple teams that would like to use it. You would…]]
* [[You would like to limit the number of resources being used in your cluster. For example…]]
* [[You've created a ReplicaSet, how to check whether the ReplicaSet found matching Pods or…]]
* [[kube-scheduler: pod scheduling in the Kubernetes control plane]]
* [[kubeconfig: composable context and credential store for kubectl]]
* [[kubectl expose: create a Service from a workload resource]]
* [[kubectl logs --previous retrieves logs from last crashed container]]
* [[kubectl logs: view container logs in a Pod]]
* [[kubectl rollout undo targets previous revision, not known-good]]
* [[oom_score_adj and Kubernetes QoS determine OOM kill priority]]
* [[ownerReferences in operators: semantics, use, and footgun]]
* [[startupProbe prevents false crash-loops during slow initialization]]
!! MOC — confidence high
Atoms with confidence in the ''high'' band (469 total).
* [[A PVC is stuck in Pending state. Walk through your debugging process.]]
* [[A container is OOMKilled but the application memory profiler shows usage well below the…]]
* [[A container keeps restarting with exit code 137. Describe your troubleshooting steps.]]
* [[A developer reports they cannot exec into pods despite having pod access. Walk through …]]
* [[A developer says their app can't reach a service. They're in different namespaces. What…]]
* [[A pod can reach the internet but not other pods. What do you check?]]
* [[A pod fails to start with a multi-attach error on an RWO volume. What happened and how …]]
* [[A pod has a main container limited to 512Mi and an Istio sidecar limited to 256Mi. The …]]
* [[A pod is OOMKilled but the app's memory usage looks normal. What happened?]]
* [[A pod is stuck in ImagePullBackOff. What are the common causes?]]
* [[A pod is stuck in Init:0/2 status. How do you debug it?]]
* [[A pod with a PVC can't start and shows a multi-attach error. What's wrong?]]
* [[A rolling update is stuck because the PDB minAvailable equals the replica count. Why is…]]
* [[After creating a service that forwards incoming external traffic to the containerized a…]]
* [[After creating a service, how to check it was created?]]
* [[After running kubectl run database --image mongo you see the status is "CrashLoopBackOf…]]
* [[Aggregated ClusterRoles compose permissions via label selectors]]
* [[An app takes 120 seconds to initialize. Liveness probe kills it before startup complete…]]
* [[An engineer form your organization asked whether there is a way to prevent from Pods (w…]]
* [[An engineer form your organization told you he is interested only in seeing his team re…]]
* [[An internal load balancer in Kubernetes is called ____ and an external load balancer is…]]
* [[Are there any tools, projects you are using for building Operators?]]
* [[Assuming you have multiple schedulers, how to know which scheduler was used for a given…]]
* [[Bare pods have no controller and stay dead on crash]]
* [[Binding type determines scope; ClusterRole can be namespace-scoped]]
* [[CNI plugins: Calico, Cilium, and Flannel compared]]
* [[CSI architecture in Kubernetes and driver health verification]]
* [[Check how many namespaces are there]]
* [[Check if there are any limits on one of the pods in your cluster]]
* [[Check if there are taints on node "master"]]
* [[Check what labels one of your nodes in the cluster has]]
* [[ClusterIP: Kubernetes default internal-only Service type]]
* [[Complete the following configuration file to make it Ingress]]
* [[ConfigMap updates do not automatically restart Pods]]
* [[Conntrack table exhaustion silently drops Kubernetes NAT connections]]
* [[Create a deployment called "pluck" using the image "redis" and make sure it runs 5 repl…]]
* [[Create a deployment with the following properties:]]
* [[Create a file definition/manifest of a deployment called "dep", with 3 replicas that us…]]
* [[Create a list of all nodes in JSON format and store it in a file called "some_nodes.json"]]
* [[Create a pod called "kartos" in the namespace dev. The pod should be using the "redis" …]]
* [[Create a static pod with the image python that runs the command sleep 2017]]
* [[Create a taint on one of the nodes in your cluster with key of "app" and value of "web"…]]
* [[Debugging CrashLoopBackOff pods: logs, debug containers, node inspection]]
* [[Debugging a failing or non-starting pod in Kubernetes]]
* [[Default-deny egress NetworkPolicies must explicitly allow DNS]]
* [[Deleting a Deployment (not its pods) is the correct removal path]]
* [[Dependency checks in liveness probes cause cascading restarts]]
* [[Deploy a pod called "my-pod" using the nginx:alpine image]]
* [[Describe a least-privilege RBAC pattern for a CI/CD deployer service account.]]
* [[Describe how would you delete a static Pod]]
* [[Describe in detail what happens when you create a service]]
* [[Describe in detail what is the Operator Lifecycle Manager]]
* [[Describe in high level what happens when you run kubectl expose deployment remo --type=…]]
* [[Describe in high-level how Kustomize works]]
* [[Describe shortly and in high-level, what happens when you run kubectl get nodes]]
* [[Describe the Kubernetes API versioning strategy.]]
* [[Describe the role of etcd in a Kubernetes cluster.]]
* [[Describe the sequence of events in case of creating a ReplicaSet]]
* [[Describe what happens when a container tries to connect with its corresponding Service …]]
* [[Diagnose and fix failing Kubernetes liveness/readiness probes]]
* [[Diagnosing Kubernetes pods stuck in Pending state]]
* [[Diagnosing a stuck Kubernetes Deployment rollout]]
* [[Diagnosing pods that won't schedule on a node due to taints]]
* [[Discuss the considerations for migrating an application from a monolithic architecture …]]
* [[Discuss the differences between OpenShift and vanilla Kubernetes.]]
* [[Discuss the implications of pod sprawl and how to manage it effectively in Kubernetes.]]
* [[Discuss the relationship between Kubernetes and container runtimes like Docker and cont…]]
* [[Discuss the role of kube-proxy in Kubernetes networking.]]
* [[Distinguish CrashLoopBackOff from other pod failure states]]
* [[Do you have experience with deploying a Kubernetes cluster? If so, can you describe the…]]
* [[Ephemeral vs. Persistent Volumes in Kubernetes Pods]]
* [[Exit code 137 in containers means OOM kill, not app error]]
* [[Explain "Dynamic Provisioning" and "Static Provisioning"]]
* [[Explain "Security Context"]]
* [[Explain Blue/Green deployments/rollouts in detail]]
* [[Explain Canary deployments/rollouts in detail]]
* [[Explain how ConfigMap and Secret updates are handled in Kubernetes.]]
* [[Explain how Service Accounts are different from User Accounts]]
* [[Explain how you would manage configuration drift in a Kubernetes environment.]]
* [[Explain how you would monitor and scale a critical production application in Kubernetes.]]
* [[Explain the Helm Chart Directory Structure]]
* [[Explain the concept of Namespaces in Kubernetes.]]
* [[Explain the concept of PodDisruptionBudget in Kubernetes.]]
* [[Explain the differences between ClusterIP, NodePort, and LoadBalancer service types.]]
* [[Explain the escalate and bind verbs. Why are they dangerous?]]
* [[Explain the main components of Kubernetes architecture.]]
* [[Explain the meaning of "http", "host" and "backend" directives]]
* [[Explain the need for Kustomize by describing actual use cases]]
* [[Explain the purpose of kubelet in the Kubernetes cluster.]]
* [[Explain the purpose of the following lines]]
* [[Explain the selectPolicy field in HPA behavior and how Max vs Min affect scaling aggres…]]
* [[Explain the working of the master node in Kubernetes?]]
* [[Explain what are "Service Accounts" and in which scenario would use create/use one]]
* [[Explain what is CronJob and what is it used for]]
* [[Explain what will happen when running apply on the following block]]
* [[Explain why one would specify resource limits in regards to Pods]]
* [[Fix the following ReplicaSet definition]]
* [[Fix the following deployment manifest]]
* [[HPA and VPA on the same metric create conflicting feedback loops]]
* [[HPA keeps scaling to max replicas even when average CPU is low. What could cause this?]]
* [[HPA requires resource requests to calculate utilization percentage]]
* [[HPA shows <unknown>/80% for CPU target and won't scale. What is wrong?]]
* [[Helm: Kubernetes package manager for bundling and distributing YAML]]
* [[How Helm supports release management?]]
* [[How Service and Deployment are connected?]]
* [[How are Kubernetes and Docker related?]]
* [[How can PodDisruptionBudget conflict with HPA scale-down, and what is the best practice…]]
* [[How can you find out information on a Service related to a certain Pod if all you can u…]]
* [[How can you get a static IP for a Kubernetes load balancer?]]
* [[How do init containers help prevent CrashLoopBackOff caused by missing dependencies?]]
* [[How do readiness probes interact with rolling deployments?]]
* [[How do you capture network traffic inside a running pod without modifying its image?]]
* [[How do you define a Kubernetes Deployment?]]
* [[How do you determine if a pod is Pending due to resource pressure?]]
* [[How do you distinguish a container-level OOM from a node-level OOM, and what commands r…]]
* [[How do you find pods that match a particular label selector?]]
* [[How do you implement encryption for data in transit and at rest in Kubernetes?]]
* [[How do you list deployed releases?]]
* [[How do you prevent high memory usage in your Kubernetes cluster and possibly issues lik…]]
* [[How do you right-size memory limits for a Kubernetes deployment?]]
* [[How do you search for charts?]]
* [[How do you test connectivity to a Service from inside the cluster?]]
* [[How do you use kubectl auth can-i to debug RBAC permissions?]]
* [[How do you use kubectl events for troubleshooting?]]
* [[How do you verify that an ImagePullSecret is correct?]]
* [[How do you verify that metrics-server is running and providing data?]]
* [[How does Kubernetes ensure high availability of applications?]]
* [[How does Kubernetes handle DNS resolution for services and pods?]]
* [[How does Kubernetes handle rolling updates with zero downtime?]]
* [[How does Kubernetes integrate with cloud providers like AWS, Azure, and GCP?]]
* [[How does Kubernetes manage containerized applications?]]
* [[How does Kubernetes manage security, and what are some best practices?]]
* [[How does PVC expansion work and what are the operational risks?]]
* [[How does a Kubernetes Service find the right Pods to route traffic to?]]
* [[How does rolling deployment work in Kubernetes?]]
* [[How does the Vertical Pod Autoscaler help prevent OOMKilled, and what are its risks?]]
* [[How make an app accessible on private or external network?]]
* [[How many containers can a pod contain?]]
* [[How readiness probe status affect Services when they are combined?]]
* [[How should liveness and readiness probe endpoints differ in what they check, and why?]]
* [[How to check how many Pods are ready as part of a replica set called "repli"?]]
* [[How to check to which worker node the pods were scheduled to? In other words, how to ch…]]
* [[How to check which container image was used as part of replica set called "repli"?]]
* [[How to commit secrets to Git and in general how to use encrypted secrets?]]
* [[How to configure TLS with Ingress?]]
* [[How to configure a default backend?]]
* [[How to confirm a container is running after running the command kubectl run web --image…]]
* [[How to create a Secret from a key and value?]]
* [[How to create a deployment with the image "nginx:alpine"?]]
* [[How to create a pod and a service with one command?]]
* [[How to create components in a namespace?]]
* [[How to delete a deployment?]]
* [[How to delete a replica set called "rori"?]]
* [[How to delete all pods whose status is not "Running"?]]
* [[How to display the resources usages of pods?]]
* [[How to edit a deployment?]]
* [[How to execute the command "ls" in an existing pod?]]
* [[How to get information on a certain service?]]
* [[How to get list of resources which are not bound to a specific namespace?]]
* [[How to get the name of the current namespace?]]
* [[How to identify which Pods are Static Pods?]]
* [[How to list Ingress in your namespace?]]
* [[How to list ReplicaSets in the current namespace?]]
* [[How to list Service Accounts?]]
* [[How to list all daemonsets in the current namespace?]]
* [[How to list the endpoints of a certain app?]]
* [[How to modify a replica set called "rori" to use a different image?]]
* [[How to scale a deployment to 8 replicas?]]
* [[How to schedule a pod on a node called "node1"?]]
* [[How to switch to another namespace? In other words how to change active namespace?]]
* [[How to turn the following service into an external one?]]
* [[How to upgrade a release?]]
* [[How to use ConfigMaps?]]
* [[How to verify a deployment was created?]]
* [[How to verify that a certain service configured to forward the requests to a given pod]]
* [[How to view revision history for a certain release?]]
* [[How view all the pods running in all the namespaces?]]
* [[How would you approach version upgrades of Kubernetes in a production environment?]]
* [[How would you map a service to an external address?]]
* [[In a Go-based operator using Kubebuilder, what does returning ctrl.Result{RequeueAfter:…]]
* [[In case of a ReplicaSet, Which field is mandatory in the spec section?]]
* [[In case of two pods, if there is an egress policy on the source denining traffic and in…]]
* [[Ingress 404: controller missing, wrong ingressClassName, or no endpoints]]
* [[Is it possible to delete ReplicaSet without deleting the Pods it created?]]
* [[Is it possible to override values in values.yaml file when installing a chart?]]
* [[It is said that Helm is also Templating Engine. What does it mean?]]
* [[JVM GC pauses cause liveness probe failures if timeoutSeconds is too low]]
* [[KEDA extends HPA to scale on 60+ event sources and to zero]]
* [[Karpenter vs Cluster Autoscaler: key differences]]
* [[Kubernetes Annotations vs Labels]]
* [[Kubernetes Control Plane: components and responsibilities]]
* [[Kubernetes DNS automatically assigns and resolves names to services and pods]]
* [[Kubernetes DaemonSet: one pod per node]]
* [[Kubernetes Labels and Selectors]]
* [[Kubernetes NetworkPolicy: whitelist semantics and operational traps]]
* [[Kubernetes Operator: components and control loop pattern]]
* [[Kubernetes Operator: purpose and usage]]
* [[Kubernetes PVC expansion: StorageClass gate and two-phase resize]]
* [[Kubernetes Pod: the smallest deployable unit]]
* [[Kubernetes QoS classes: criteria and eviction order]]
* [[Kubernetes ReplicaSet: maintaining stable pod replica count]]
* [[Kubernetes Secrets are base64-encoded, not encrypted, in etcd by default]]
* [[Kubernetes Secrets: storage, access, and best practices]]
* [[Kubernetes Service empty endpoints: label-selector mismatch diagnosis]]
* [[Kubernetes Service: stable networking endpoint for pods]]
* [[Kubernetes controllers and kube-controller-manager]]
* [[Kubernetes does not provide data persistence by default]]
* [[Kubernetes dynamic provisioning and WaitForFirstConsumer zone awareness]]
* [[Kubernetes headless services return pod IPs directly via DNS]]
* [[Kubernetes monitoring solutions overview]]
* [[Kubernetes ndots:5 causes excessive external DNS queries]]
* [[Kubernetes node upgrade procedure]]
* [[Kubernetes originally used ABAC authorization; RBAC became the default in version 1.8]]
* [[Kubernetes storage abstraction: PV, PVC, and StorageClass layers]]
* [[Kubernetes taints and tolerations: effects, syntax, and use cases]]
* [[List all the pods with the label "env=prod"]]
* [[Manual node scaling vs Cluster Autoscaler in Kubernetes]]
* [[Memory is a poor primary HPA metric because apps hold allocations]]
* [[Name the initial namespaces from which Kubernetes starts?]]
* [[Name three operator-building frameworks and when you would choose each.]]
* [[OOMKill sends SIGKILL (exit code 137)]]
* [[OPA Gatekeeper: policy enforcement webhook in Kubernetes]]
* [[Perhaps a general question but, you suspect one of the pods is having issues, you don't…]]
* [[Pod IPs are ephemeral; use a Service for stable addressing]]
* [[Pod deletion: SIGTERM, grace period, and SIGKILL]]
* [[Pod reaches Service by ClusterIP but internal DNS fails: diagnosis]]
* [[PodDisruptionBudgets block kubectl drain indefinitely]]
* [[Pods are Pending cluster-wide after a control plane upgrade. All worker nodes show a No…]]
* [[Pods are being evicted with the message 'The node was low on resource: ephemeral-storag…]]
* [[Pods from the same Deployment keep landing on the same node, causing a single point of …]]
* [[RBAC subresources require explicit rules; parent grants do not cascade]]
* [[ReplicaSets are running the moment the user executed the command to create them (like k…]]
* [[Role is namespace-scoped; ClusterRole is cluster-wide]]
* [[Role of kube-apiserver in Kubernetes]]
* [[Run a command to view all nodes of the cluster]]
* [[Run a pod called "yay2" with the image "python". Make sure it has resources request of …]]
* [[Scale a ReplicaSet with kubectl scale rs]]
* [[Scaling a Deployment in Kubernetes]]
* [[Service account tokens are mounted by default even if the pod never calls the Kubernetes API]]
* [[Service and Ingress divide Layer 4 and Layer 7 routing]]
* [[Service type LoadBalancer provisions cloud load balancer]]
* [[Startup probes replaced the initialDelaySeconds compromise in Kubernetes 1.18]]
* [[StatefulSets and volumeClaimTemplates: stable storage per replica]]
* [[Tell me about your Kubernetes experience.]]
* [[True or False? A single Pod can be split across multiple nodes]]
* [[True or False? A volume defined in Pod can be accessed by all the containers of that Pod]]
* [[True or False? By default there is no communication between two Pods in two different n…]]
* [[True or False? By default, pods are isolated. This means they are unable to receive tra…]]
* [[True or False? Deleting a ReplicaSet will delete the Pods it created]]
* [[True or False? Each Pod, when created, gets its own public IP address]]
* [[True or False? Every cluster must have 0 or more master nodes and at least 1 worker]]
* [[True or False? If a ReplicaSet defines 2 replicas but there 3 Pods running matching the…]]
* [[True or False? If no network policies are applied to a pod, then no connections to or f…]]
* [[True or False? In case of a ReplicaSet, if Pods specified in the selector field don't e…]]
* [[True or False? Memory is a compressible resource, meaning that when a container reach t…]]
* [[True or False? Once a Pod is assisgned to a worker node, it will only run on that node,…]]
* [[True or False? Pods specified by the selector field of ReplicaSet must be created by th…]]
* [[True or False? Removing the label from a Pod that is tracked by a ReplicaSet, will caus…]]
* [[True or False? Resource limits applied on a Pod level meaning, if limits is 2gb RAM and…]]
* [[True or False? Sensitive data, like credentials, should be stored in a ConfigMap]]
* [[True or False? The "Pending" phase means the Pod was not yet accepted by the Kubernetes…]]
* [[True or False? The same as there are "Static Pods" there are other static resources lik…]]
* [[True or False? The scheduler is responsible for both deciding where a Pod will run and …]]
* [[True or False? Using the node affinity type "preferredDuringSchedulingIgnoredDuringExec…]]
* [[True or False? When a namespace is deleted all resources in that namespace are not dele…]]
* [[True or False? With namespaces you can limit the resources consumed by the users/teams]]
* [[True or False? storing data in a Secret component makes it automatically secured]]
* [[True or False? the target port, in the case of running the following command, will be e…]]
* [[Users unable to reach an application running on a Pod on Kubernetes. What might be the …]]
* [[Using node affinity, set a Pod to schedule on a node where the key is "region" and valu…]]
* [[Volume Snapshots in Kubernetes]]
* [[What DNS record format does Kubernetes create for a ClusterIP Service?]]
* [[What Kubernetes objects are there?]]
* [[What Kubernetes objects do you usually use when deploying applications in Kubernetes?]]
* [[What QoS classes are there?]]
* [[What actions or operations you consider as best practices when it comes to Kubernetes?]]
* [[What are Custom Resource Definitions (CRDs) in Kubernetes?]]
* [[What are Kubernetes selectors and how do they match resources?]]
* [[What are all the phases/steps of a control loop?]]
* [[What are common causes of kubelet failures and how do they manifest?]]
* [[What are federated clusters?]]
* [[What are finalizers in the context of Kubernetes operators, and why are they needed?]]
* [[What are important steps in defining/adding a Service?]]
* [[What are some of Kubernetes features?]]
* [[What are some use cases for using Helm template file?]]
* [[What are some use cases for using Ingress?]]
* [[What are some use cases for using Network Policies?]]
* [[What are some use cases for using a DaemonSet?]]
* [[What are the challenges in managing stateful applications in Kubernetes?]]
* [[What are the components of a worker node (aka data plane)?]]
* [[What are the five levels of the operator maturity model?]]
* [[What are the four metric types supported by HPA v2 and when would you use each?]]
* [[What are the four probe mechanisms Kubernetes supports?]]
* [[What are the main differences between Docker Swarm and Kubernetes?]]
* [[What are the possible Pod phases?]]
* [[What are the prerequisites for HPA to work?]]
* [[What are the security risks of using the default ServiceAccount and how do you audit fo…]]
* [[What are the standard RBAC verbs in Kubernetes and what API operations do they map to?]]
* [[What are the three classic PersistentVolume access modes and what does each allow?]]
* [[What are the top 3 causes of CrashLoopBackOff?]]
* [[What are the types of controller managers?]]
* [[What are your thoughts on "Pods are not meant to be created directly"?]]
* [[What can you find in kube-system namespace?]]
* [[What causes a node drain to get stuck and how do you troubleshoot it?]]
* [[What challenges do you anticipate when managing large-scale Kubernetes clusters, and ho…]]
* [[What combination of mechanisms ensures zero-downtime during node maintenance?]]
* [[What components the Operator Framework consists of?]]
* [[What do you understand by Cloud controller manager?]]
* [[What does '<unknown>/50%' mean in HPA status?]]
* [[What does HPA do and what problem remains after you enable it?]]
* [[What does a Kubernetes liveness probe determine?]]
* [[What does being cloud-native mean?]]
* [[What does failureThreshold control, and how does it interact with periodSeconds?]]
* [[What does the "ErrImagePull" status of a Pod means?]]
* [[What does the node status contain?]]
* [[What fields are mandatory with any Kubernetes object?]]
* [[What formula does the HPA use to compute the desired replica count?]]
* [[What happens after you edit a deployment and change the image?]]
* [[What happens behind the scenes when you create a Deployment object?]]
* [[What happens to running pods if if you stop Kubelet on the worker nodes?]]
* [[What happens when a readiness probe fails on a Kubernetes pod?]]
* [[What happens when you apply a NetworkPolicy with an empty podSelector and policyTypes: …]]
* [[What happens when you delete a deployment?]]
* [[What happens when you run a Pod with kubectl?]]
* [[What happens when you set replicas in a Deployment manifest and also use HPA?]]
* [[What happens you create a pod and you DON'T specify a service account?]]
* [[What is Conftest and how does it validate configuration files?]]
* [[What is Container resource monitoring?]]
* [[What is Datree? How is it different from Conftest?]]
* [[What is Heapster in Kubernetes?]]
* [[What is Helm, and how is it used in Kubernetes?]]
* [[What is Ingress Controller?]]
* [[What is Ingress Default Backend?]]
* [[What is Istio? What is it used for?]]
* [[What is Minikube and when would you use it?]]
* [[What is NodePort service type in Kubernetes?]]
* [[What is PodSecurity and how can it be configured in a Kubernetes cluster?]]
* [[What is Resource Quota?]]
* [[What is a "Deployment" in Kubernetes?]]
* [[What is a ConfigMap, and how is it used in Kubernetes?]]
* [[What is a Kubernetes Cluster?]]
* [[What is a Kubernetes Secret and how does it store sensitive data?]]
* [[What is a Kubernetes StatefulSet and when would you use it?]]
* [[What is a Kubernetes cluster and what are its components?]]
* [[What is a LimitRange and how does it prevent OOMKilled caused by missing resource limits?]]
* [[What is a Persistent Volume (PV) and Persistent Volume Claim (PVC) in Kubernetes?]]
* [[What is a PodDisruptionBudget (PDB) and why does it create tension with node drains?]]
* [[What is a cluster of containers in Kubernetes?]]
* [[What is a node in Kubernetes?]]
* [[What is a volume in regards to Kubernetes?]]
* [[What is an "Ingress"?]]
* [[What is kubconfig? What do you use it for?]]
* [[What is kubectl and how is it used to manage Kubernetes clusters?]]
* [[What is orchestration when it comes to software and DevOps?]]
* [[What is the Google Container Engine?]]
* [[What is the Kubernetes concept chain?]]
* [[What is the Operator Framework?]]
* [[What is the PID 1 problem in containers and how does it cause exit code 137 on pod term…]]
* [[What is the constraint on successThreshold for liveness and startup probes, and why doe…]]
* [[What is the default HPA scale-down stabilization window?]]
* [[What is the difference between Immediate and WaitForFirstConsumer volume binding modes?]]
* [[What is the difference between a Job and a CronJob in Kubernetes?]]
* [[What is the difference between a Pod and a Deployment in Kubernetes?]]
* [[What is the difference between a StatefulSet and a Deployment in Kubernetes?]]
* [[What is the difference between a replica set and a replication controller?]]
* [[What is the difference between cordoning and draining a node?]]
* [[What is the difference between deploying applications on hosts and containers?]]
* [[What is the difference between exit codes 126 and 127 in a container?]]
* [[What is the difference between liveness and readiness probes in Kubernetes?]]
* [[What is the difference between resource requests and limits, and which one affects sche…]]
* [[What is the difference between resources.requests.memory and resources.limits.memory in…]]
* [[What is the fundamental rule of the Kubernetes pod networking model?]]
* [[What is the job of the kube-scheduler?]]
* [[What is the key difference between kube-proxy iptables mode and IPVS mode?]]
* [[What is the lifecycle of a Kubernetes node?]]
* [[What is the problem with the following Secret file:]]
* [[What is the purpose of Horizontal Pod Autoscaling in Kubernetes?]]
* [[What is the purpose of a StorageClass in Kubernetes?]]
* [[What is the purpose of a startup probe?]]
* [[What is the purpose of an admission controller in Kubernetes, and how can you extend it?]]
* [[What is the reconciliation loop in a Kubernetes operator?]]
* [[What is the relationship between Deployment, ReplicaSet, and Pod?]]
* [[What is the relationship between a PersistentVolume (PV) and a PersistentVolumeClaim (P…]]
* [[What is the role of kube-apiserver aggregation layer in Kubernetes?]]
* [[What is the standard three-command diagnostic flow when a pod is misbehaving?]]
* [[What issue might arise from using the following CronJob and how to fix it?]]
* [[What kube-node-lease contains?]]
* [[What kube-public contains?]]
* [[What kubectl command shows whether a pod was OOMKilled, and what fields do you look for?]]
* [[What kubectl describe pod [pod name] does? command does?]]
* [[What kubectl get componentstatus does?]]
* [[What node conditions does Kubernetes monitor for health, and what happens when a node g…]]
* [[What openshift-operator-lifecycle-manager namespace includes?]]
* [[What possible issue can arise from using the following spec and how to fix it?]]
* [[What problem does Ingress solve that Services alone cannot?]]
* [[What problem does a ConfigMap solve?]]
* [[What problem does a Deployment solve that bare Pods cannot?]]
* [[What problems does Kubernetes actually solve?]]
* [[What process is responsible for running and installing the different controllers?]]
* [[What process runs on Kubernetes Master Node?]]
* [[What reclaim policies are there?]]
* [[What rollout/deployment strategies are you familiar with?]]
* [[What security best practices do you follow in regards to the Kubernetes cluster?]]
* [[What special namespaces are there by default when creating a Kubernetes cluster?]]
* [[What the following block of lines does?]]
* [[What the following command does?]]
* [[What the following in a Deployment configuration file means?]]
* [[What the following output of kubectl get rs means?]]
* [[What the master node is responsible for?]]
* [[What three kubectl commands form the basic CrashLoopBackOff diagnostic workflow?]]
* [[What type: Opaque in a secret file means? What other types are there?]]
* [[What types of persistent volumes are there?]]
* [[What use cases exist for running multiple containers in a single pod?]]
* [[What volume types are you familiar with?]]
* [[What ways are you familiar with to implement deployment strategies (like canary, blue/g…]]
* [[What will happen when a Pod, created by ReplicaSet, is deleted directly with kubectl de…]]
* [[What would you use to route traffic from outside the Kubernetes cluster to services wit…]]
* [[What's the difference between ImagePullBackOff and ErrImagePull?]]
* [[When HPA is configured with multiple metrics, how does it decide the replica count?]]
* [[When do you use kubectl logs vs kubectl describe?]]
* [[When or why NOT to use Kubernetes?]]
* [[When would you use the "LoadBalancer" type]]
* [[Where static Pods manifests are located?]]
* [[Which Kubernetes concept would you use to control traffic flow at the IP address or por…]]
* [[Which Prometheus metric should you use to predict OOMKill, and why not container_memory…]]
* [[Which command lists all Pods in a Kubernetes cluster?]]
* [[Which command will list all the object types in a cluster?]]
* [[Which components can't be created within a namespace?]]
* [[Which problems, volumes in Kubernetes solve?]]
* [[Which resources are accessible from different namespaces?]]
* [[Which service and in which namespace the following file is referencing?]]
* [[While namespaces do provide scope for resources, they are not isolating them]]
* [[Why do Ingress resources do nothing by themselves?]]
* [[Why do Kubernetes clusters fail at scale?]]
* [[Why do container logs sometimes bring down Kubernetes nodes?]]
* [[Why do some teams set CPU requests but no CPU limits?]]
* [[Why do we need Operators?]]
* [[Why does Kubernetes stress Linux more than VMs?]]
* [[Why does a Java application with -Xmx1g in a container limited to 512Mi get OOMKilled, …]]
* [[Why etcd? Why not some SQL or NoSQL database?]]
* [[Why is a PodDisruptionBudget important when using node autoscaling?]]
* [[Why is etcd performance tied to disk latency more than CPU?]]
* [[Why it's common to have only one container per Pod in most cases?]]
* [[Why must the Reconcile function in a Kubernetes operator be idempotent?]]
* [[Why should operators use the status subresource for status updates instead of updating …]]
* [[Why should you use a Secret instead of a ConfigMap for passwords?]]
* [[Why there is no such command in Kubernetes? kubectl get containers]]
* [[Why to use namespaces? What is the problem with using one default namespace?]]
* [[Why using a wildcard in ingress host may lead to issues?]]
* [[Wildcard RBAC rules grant unrestricted access and enable privilege escalation]]
* [[Would you use Helm, Go or something else for creating an Operator?]]
* [[You applied a taint with k taint node minikube app=web:NoSchedule on the only node in y…]]
* [[You are looking for a Pod called "atreus". How to check in which namespace it runs?]]
* [[You are managing multiple Kubernetes clusters. How do you quickly change between the cl…]]
* [[You create a resource but cannot find it with kubectl get. What namespace-related mista…]]
* [[You deploy an Envoy sidecar with your app container. After a rolling update, requests f…]]
* [[You encounter a performance issue in a Kubernetes cluster. How do you diagnose and reso…]]
* [[You have a microservices-based application. How would you deploy and manage it in Kuber…]]
* [[You have one Kubernetes cluster and multiple teams that would like to use it. You would…]]
* [[You would like to limit the number of resources being used in your cluster. For example…]]
* [[You've created a ReplicaSet, how to check whether the ReplicaSet found matching Pods or…]]
* [[kube-scheduler: pod scheduling in the Kubernetes control plane]]
* [[kubeconfig: composable context and credential store for kubectl]]
* [[kubectl expose: create a Service from a workload resource]]
* [[kubectl logs --previous retrieves logs from last crashed container]]
* [[kubectl logs: view container logs in a Pod]]
* [[kubectl rollout undo targets previous revision, not known-good]]
* [[oom_score_adj and Kubernetes QoS determine OOM kill priority]]
* [[startupProbe prevents false crash-loops during slow initialization]]
!! MOC — confidence mid
Atoms with confidence in the ''mid'' band (14 total).
* [[CNI plugin: Kubernetes container networking standard]]
* [[Check Deployment rollout history and roll back to a specific revision]]
* [[DaemonSet vs Deployment vs ReplicaSet in Kubernetes]]
* [[Investigate and recover a Kubernetes NotReady node]]
* [[Kubernetes Deployment rolling updates and rollbacks]]
* [[Kubernetes Static Pods: definition, use cases, and DaemonSet contrast]]
* [[Kubernetes custom schedulers: deploy and use]]
* [[Kubernetes ingress controllers: edge routing and policy enforcement]]
* [[Kubernetes memory failure modes: eviction vs OOMKill]]
* [[Kubernetes node resource pressure: diagnosis, eviction, and prevention]]
* [[Kubernetes: definition, origin, and core features]]
* [[PV reclaim policies: Retain preserves data, Delete destroys it]]
* [[Pods without memory limits cause node-wide evictions and OOM kills]]
* [[ownerReferences in operators: semantics, use, and footgun]]
!! MOC — merged atoms
Atoms that were consolidated from 2+ source concepts during cross-dedup (128 total). Reading these is a cheap way to see where the pipeline found duplication worth collapsing.
* [[A PVC is stuck in Pending state. Walk through your debugging process.]]
* [[Aggregated ClusterRoles compose permissions via label selectors]]
* [[Bare pods have no controller and stay dead on crash]]
* [[Binding type determines scope; ClusterRole can be namespace-scoped]]
* [[CNI plugin: Kubernetes container networking standard]]
* [[CNI plugins: Calico, Cilium, and Flannel compared]]
* [[CSI architecture in Kubernetes and driver health verification]]
* [[Check Deployment rollout history and roll back to a specific revision]]
* [[Check how many namespaces are there]]
* [[ClusterIP: Kubernetes default internal-only Service type]]
* [[ConfigMap updates do not automatically restart Pods]]
* [[Conntrack table exhaustion silently drops Kubernetes NAT connections]]
* [[Create a taint on one of the nodes in your cluster with key of "app" and value of "web"…]]
* [[DaemonSet vs Deployment vs ReplicaSet in Kubernetes]]
* [[Debugging CrashLoopBackOff pods: logs, debug containers, node inspection]]
* [[Debugging a failing or non-starting pod in Kubernetes]]
* [[Default-deny egress NetworkPolicies must explicitly allow DNS]]
* [[Deleting a Deployment (not its pods) is the correct removal path]]
* [[Dependency checks in liveness probes cause cascading restarts]]
* [[Describe the role of etcd in a Kubernetes cluster.]]
* [[Diagnose and fix failing Kubernetes liveness/readiness probes]]
* [[Diagnosing Kubernetes pods stuck in Pending state]]
* [[Diagnosing a stuck Kubernetes Deployment rollout]]
* [[Diagnosing pods that won't schedule on a node due to taints]]
* [[Discuss the role of kube-proxy in Kubernetes networking.]]
* [[Distinguish CrashLoopBackOff from other pod failure states]]
* [[Ephemeral vs. Persistent Volumes in Kubernetes Pods]]
* [[Exit code 137 in containers means OOM kill, not app error]]
* [[Explain the Helm Chart Directory Structure]]
* [[Explain the concept of Namespaces in Kubernetes.]]
* [[Explain the differences between ClusterIP, NodePort, and LoadBalancer service types.]]
* [[Explain the main components of Kubernetes architecture.]]
* [[Explain the purpose of kubelet in the Kubernetes cluster.]]
* [[HPA and VPA on the same metric create conflicting feedback loops]]
* [[HPA requires resource requests to calculate utilization percentage]]
* [[Helm: Kubernetes package manager for bundling and distributing YAML]]
* [[How does Kubernetes manage security, and what are some best practices?]]
* [[How to create a Secret from a key and value?]]
* [[How to get list of resources which are not bound to a specific namespace?]]
* [[Ingress 404: controller missing, wrong ingressClassName, or no endpoints]]
* [[Investigate and recover a Kubernetes NotReady node]]
* [[JVM GC pauses cause liveness probe failures if timeoutSeconds is too low]]
* [[KEDA extends HPA to scale on 60+ event sources and to zero]]
* [[Karpenter vs Cluster Autoscaler: key differences]]
* [[Kubernetes Annotations vs Labels]]
* [[Kubernetes Control Plane: components and responsibilities]]
* [[Kubernetes DNS automatically assigns and resolves names to services and pods]]
* [[Kubernetes DaemonSet: one pod per node]]
* [[Kubernetes Deployment rolling updates and rollbacks]]
* [[Kubernetes Labels and Selectors]]
* [[Kubernetes NetworkPolicy: whitelist semantics and operational traps]]
* [[Kubernetes Operator: components and control loop pattern]]
* [[Kubernetes Operator: purpose and usage]]
* [[Kubernetes PVC expansion: StorageClass gate and two-phase resize]]
* [[Kubernetes Pod: the smallest deployable unit]]
* [[Kubernetes QoS classes: criteria and eviction order]]
* [[Kubernetes ReplicaSet: maintaining stable pod replica count]]
* [[Kubernetes Secrets are base64-encoded, not encrypted, in etcd by default]]
* [[Kubernetes Secrets: storage, access, and best practices]]
* [[Kubernetes Service empty endpoints: label-selector mismatch diagnosis]]
* [[Kubernetes Service: stable networking endpoint for pods]]
* [[Kubernetes Static Pods: definition, use cases, and DaemonSet contrast]]
* [[Kubernetes controllers and kube-controller-manager]]
* [[Kubernetes custom schedulers: deploy and use]]
* [[Kubernetes does not provide data persistence by default]]
* [[Kubernetes dynamic provisioning and WaitForFirstConsumer zone awareness]]
* [[Kubernetes headless services return pod IPs directly via DNS]]
* [[Kubernetes ingress controllers: edge routing and policy enforcement]]
* [[Kubernetes memory failure modes: eviction vs OOMKill]]
* [[Kubernetes monitoring solutions overview]]
* [[Kubernetes ndots:5 causes excessive external DNS queries]]
* [[Kubernetes node resource pressure: diagnosis, eviction, and prevention]]
* [[Kubernetes node upgrade procedure]]
* [[Kubernetes originally used ABAC authorization; RBAC became the default in version 1.8]]
* [[Kubernetes storage abstraction: PV, PVC, and StorageClass layers]]
* [[Kubernetes taints and tolerations: effects, syntax, and use cases]]
* [[Kubernetes: definition, origin, and core features]]
* [[Manual node scaling vs Cluster Autoscaler in Kubernetes]]
* [[Memory is a poor primary HPA metric because apps hold allocations]]
* [[OOMKill sends SIGKILL (exit code 137)]]
* [[OPA Gatekeeper: policy enforcement webhook in Kubernetes]]
* [[PV reclaim policies: Retain preserves data, Delete destroys it]]
* [[Pod IPs are ephemeral; use a Service for stable addressing]]
* [[Pod deletion: SIGTERM, grace period, and SIGKILL]]
* [[Pod reaches Service by ClusterIP but internal DNS fails: diagnosis]]
* [[PodDisruptionBudgets block kubectl drain indefinitely]]
* [[Pods without memory limits cause node-wide evictions and OOM kills]]
* [[RBAC subresources require explicit rules; parent grants do not cascade]]
* [[Role is namespace-scoped; ClusterRole is cluster-wide]]
* [[Role of kube-apiserver in Kubernetes]]
* [[Run a pod called "yay2" with the image "python". Make sure it has resources request of …]]
* [[Scale a ReplicaSet with kubectl scale rs]]
* [[Scaling a Deployment in Kubernetes]]
* [[Service account tokens are mounted by default even if the pod never calls the Kubernetes API]]
* [[Service and Ingress divide Layer 4 and Layer 7 routing]]
* [[Service type LoadBalancer provisions cloud load balancer]]
* [[Startup probes replaced the initialDelaySeconds compromise in Kubernetes 1.18]]
* [[StatefulSets and volumeClaimTemplates: stable storage per replica]]
* [[Using node affinity, set a Pod to schedule on a node where the key is "region" and valu…]]
* [[Volume Snapshots in Kubernetes]]
* [[What are Custom Resource Definitions (CRDs) in Kubernetes?]]
* [[What are the challenges in managing stateful applications in Kubernetes?]]
* [[What happens when a readiness probe fails on a Kubernetes pod?]]
* [[What is Helm, and how is it used in Kubernetes?]]
* [[What is PodSecurity and how can it be configured in a Kubernetes cluster?]]
* [[What is a "Deployment" in Kubernetes?]]
* [[What is a ConfigMap, and how is it used in Kubernetes?]]
* [[What is a Kubernetes StatefulSet and when would you use it?]]
* [[What is a node in Kubernetes?]]
* [[What is the difference between a Pod and a Deployment in Kubernetes?]]
* [[What is the difference between a StatefulSet and a Deployment in Kubernetes?]]
* [[What is the difference between liveness and readiness probes in Kubernetes?]]
* [[What is the difference between resource requests and limits, and which one affects sche…]]
* [[What is the purpose of Horizontal Pod Autoscaling in Kubernetes?]]
* [[What is the relationship between a PersistentVolume (PV) and a PersistentVolumeClaim (P…]]
* [[What volume types are you familiar with?]]
* [[Wildcard RBAC rules grant unrestricted access and enable privilege escalation]]
* [[You applied a taint with k taint node minikube app=web:NoSchedule on the only node in y…]]
* [[You would like to limit the number of resources being used in your cluster. For example…]]
* [[kube-scheduler: pod scheduling in the Kubernetes control plane]]
* [[kubeconfig: composable context and credential store for kubectl]]
* [[kubectl expose: create a Service from a workload resource]]
* [[kubectl logs --previous retrieves logs from last crashed container]]
* [[kubectl logs: view container logs in a Pod]]
* [[kubectl rollout undo targets previous revision, not known-good]]
* [[oom_score_adj and Kubernetes QoS determine OOM kill priority]]
* [[ownerReferences in operators: semantics, use, and footgun]]
* [[startupProbe prevents false crash-loops during slow initialization]]
!! Maps of Content
! By kind
* [[MOC: flashcards]]
! By confidence
* [[MOC: confidence high]]
* [[MOC: confidence mid]]
! Quality
* [[MOC: merged atoms]]
canonical atomic concepts from k8s-ops and related sources
GettingStarted
[[MOC: index]]
[[GettingStarted]]
[[MOC: index]]
----
''By kind''
* [[MOC: footguns]]
* [[MOC: trivia]]
* [[MOC: flashcards]]
* [[MOC: compendium q&a]]
* [[MOC: primer]]
* [[MOC: anti-primer]]
* [[MOC: street ops]]
* [[MOC: cheatsheet]]
----
''By confidence''
* [[MOC: confidence high]]
* [[MOC: confidence mid]]
----
''Quality''
* [[MOC: merged atoms]]
!! Kubernetes — atoms deck
''483 atoms'' — canonical atomic concepts from k8s-ops and related sources
//Each tiddler is one atomic concept. Sources and related atoms are listed in the footer. Cross-dedup has already merged duplicates, so every concept should appear exactly once.//
! Start here
* [[MOC: index]] — master index of MOCs
* [[MOC: merged atoms]] — concepts consolidated from multiple sources
! Breakdown — by kind
|!Kind|!Count|!Jump|h
|Flashcards|483|[[MOC: flashcards]]|
! Breakdown — by confidence
|!Band|!Count|!Jump|h
|high|469|[[MOC: confidence high]]|
|mid|14|[[MOC: confidence mid]]|
! Pipeline provenance
* Raw extracted atoms (whole corpus): 17641
* After intra-source dedup (whole corpus): 7753
* In this deck (domain = k8s, post cross-dedup): 483
* Atoms in this deck merged from multiple sources: 128
See ''FULL-CORPUS-REPORT.md'' in the grokzett repo for full metrics.
/* grokzett TWC deck theme overrides */
body, #contentWrapper { font-family: -apple-system, "Segoe UI", Roboto, "Helvetica Neue", sans-serif; }
.tiddler { margin-bottom: 1.5em; }
.tiddler .title { font-size: 1.4em; letter-spacing: -0.01em; }
.viewer { line-height: 1.55; }
.viewer h1, .viewer h2, .viewer h3 { border-bottom: none; margin-top: 1.2em; }
.viewer h1 { font-size: 1.35em; }
.viewer h2 { font-size: 1.20em; }
.viewer h3 { font-size: 1.05em; color: #444; }
.viewer blockquote { border-left: 3px solid #c9d6df; margin: 1em 0; padding: 0.2em 1em; background: #f6f9fb; color: #333; }
.viewer code { background: #f2f2f2; padding: 1px 5px; border-radius: 3px; font-size: 0.92em; }
.viewer pre { background: #282c34; color: #abb2bf; padding: 0.8em 1em; border-radius: 6px; overflow-x: auto; font-size: 0.88em; line-height: 1.4; }
.viewer pre code { background: transparent; padding: 0; color: inherit; border-radius: 0; }
.viewer pre.literal { background: #282c34; color: #abb2bf; }
.viewer table { border-collapse: collapse; margin: 1em 0; font-size: 0.92em; width: auto; }
.viewer th, .viewer td { border: 1px solid #d8dee4; padding: 6px 10px; }
.viewer th { background: #eef2f6; text-align: left; }
.viewer tr:nth-child(even) td { background: #fbfcfd; }
.viewer hr { border: none; border-top: 1px solid #dde3e9; margin: 1.5em 0; }
.tagged { background: #eef5fb; padding: 4px 8px; border-radius: 4px; margin-right: 4px; }
/* grokzett status controls (injected by features.js per tiddler) */
.grokzett-status {
font-size: 0.85em; margin: 0.3em 0 1em; color: #555;
display: flex; align-items: center; gap: 6px;
}
.grokzett-status .gz-label { color: #888; letter-spacing: 0.02em; }
.grokzett-status button {
border: 1px solid #c9d6df; background: #f7f9fb; padding: 2px 10px;
font-size: 0.92em; cursor: pointer; border-radius: 12px; color: #333;
}
.grokzett-status button:hover { background: #e6edf3; }
.grokzett-status button.active { background: #2b6cb0; color: white; border-color: #2b6cb0; }
/* Fixed filter bar (top-right) */
.grokzett-filter-bar {
position: fixed; top: 10px; right: 10px; z-index: 10000;
background: #2d3748; color: #e2e8f0;
padding: 8px 14px; border-radius: 8px;
font-size: 0.85em; font-family: -apple-system, "Segoe UI", Roboto, sans-serif;
display: flex; align-items: center; gap: 12px;
box-shadow: 0 3px 12px rgba(0,0,0,0.25);
}
.grokzett-filter-bar .gz-title { font-weight: 600; letter-spacing: 0.02em; }
.grokzett-filter-bar select, .grokzett-filter-bar button {
background: #4a5568; color: #e2e8f0; border: 1px solid #718096;
border-radius: 4px; padding: 2px 8px; font: inherit;
}
.grokzett-filter-bar button { cursor: pointer; }
.grokzett-filter-bar button:hover { background: #5a6678; }
.grokzett-filter-bar label { display: flex; align-items: center; gap: 4px; }
.grokzett-filter-bar #gz-counts { opacity: 0.75; font-size: 0.95em; }
/* Subtle status badge on each tiddler (color bar at left) */
.tiddler[data-gz-status="learning"] { border-left: 3px solid #ed8936; padding-left: 8px; }
.tiddler[data-gz-status="known"] { border-left: 3px solid #48bb78; padding-left: 8px; opacity: 0.7; }
/* Cloze spans — trivia mode */
.gz-cloze {
background: #f6e05e; color: #f6e05e; border-radius: 3px; padding: 0 4px;
cursor: pointer; user-select: none; transition: color 0.15s;
}
.gz-cloze.revealed { color: #533f03; background: #fef6ad; }
/* Flashcard Q/A layout */
.gz-flash-a {
margin-top: 1em; padding: 0.8em 1em; background: #f7fafc;
border-radius: 6px; border: 1px solid #e2e8f0; position: relative;
}
.gz-flash-a.hidden > * { filter: blur(5px); user-select: none; }
.gz-flash-a.hidden { cursor: pointer; }
.gz-flash-a.hidden::after {
content: "click to reveal answer";
position: absolute; left: 0; right: 0; top: 50%; transform: translateY(-50%);
text-align: center; color: #718096; font-size: 0.9em; letter-spacing: 0.05em;
pointer-events: none;
}
.gz-flash-a:not(.hidden) { border-left: 4px solid #48bb78; }