---
tags:
- k8s
- l1
- flashcard-deck
- k8s-ops
---
<!-- wiki:breadcrumb:start -->
[Portal](../../../../library/portal/index.md) | **Level:** [L1: Foundations](../../../../library/portal/levels.md) | **Topics:** [Kubernetes Core](../../../../library/portal/topics.md) | **Domain:** Kubernetes
<!-- wiki:breadcrumb:end -->

id	category	difficulty	tags	question	answer	source_path
k8s-ops/0020ca6b4d48	k8s-ops	hard	kubernetes, etcd-ops, cert-rotation, kubectl	Do you have experience with deploying a Kubernetes cluster? If so, can you describe the process in high-level?	1. Create multiple instances you will use as Kubernetes nodes/workers. Create also an instance to act as the Master. The instances can be provisioned in a cloud or they can be virtual machines on bare metal hosts.\n2. Provision a certificate authority that will be used to generate TLS certificates for the different components of a Kubernetes cluster (kubelet, etcd, ...)\n  1. Generate a certificate and private key for the different components\n3. Generate kubeconfigs so the different clients of Kubernetes can locate the API servers and authenticate.\n4. Generate encryption key that will be used for encrypting the cluster data\n5. Create an etcd cluster	projects/knowledge/interview/kubernetes/020-do-you-have-experience-with-deploying-a-kubernetes.txt
k8s-ops/006e085c205c	k8s-ops	hard	kubernetes, logging, troubleshooting, kubectl	After running kubectl run database --image mongo you see the status is "CrashLoopBackOff". What could possibly went wrong and what do you do to confirm?	CrashLoopBackOff means the Pod is starting, crashing, starting...and so it repeats itself. \nThere are many different reasons to get this error - lack of permissions, init-container misconfiguration, persistent volume connection issue, etc.\n\nOne of the ways to check why it happened is to run `kubectl describe po <POD_NAME>` and having a look at the exit code\n\n```\n Last State: Terminated\n Reason: Error\n Exit Code: 100\n```\n\nAnother way to check what's going on, is to run `kubectl logs <POD_NAME>`. This will provide us with the logs from the containers running in that Pod.	projects/knowledge/interview/kubernetes/039-after-running-kubectl-run-database-image-mongo-you.txt
k8s-ops/01ba29ccb332	k8s-ops	medium	kubernetes, upgrade, helm, operator	Would you use Helm, Go or something else for creating an Operator?	Depends on the scope and maturity of the Operator. If it mainly covers installation and upgrades, Helm might be enough. If you want to go for Lifecycle management, insights and auto-pilot, this is where you'd probably use Go.\n\nRemember: Helm = K8s package manager. Chart=package, Release=instance, Repo=store.\n\nExample: `helm install my-app bitnami/nginx --set service.type=LoadBalancer`	projects/knowledge/interview/kubernetes/211-would-you-use-helm-go-or-something-else-for-creati.txt
k8s-ops/05670bcdc5a9	k8s-ops	medium	kubernetes, kubectl	Why there is no such command in Kubernetes? kubectl get containers	Because a container is not a Kubernetes object. The smallest object unit in Kubernetes is a Pod. In a single Pod you can find one or more containers.\n\nRemember: `kubectl explain <resource>` for fields. `kubectl api-resources` for all resource types.	projects/knowledge/interview/kubernetes/008-why-there-is-no-such-command-in-kubernetes-kubectl.txt
k8s-ops/080422f633e8	k8s-ops	hard	kubernetes, upgrade, backup, helm	How would you approach version upgrades of Kubernetes in a production environment?	**Version Upgrades in Production:* • \n • Conduct thorough testing in a staging environment before production.\n • Follow Kubernetes documentation and release notes for upgrade procedures.\n • Use tools like kubeadm for streamlined upgrade processes.\n • Ensure backups and have a rollback plan in case of issues. projects/knowledge/interview/kubernetes/366-how-would-you-approach-version-upgrades-of-kuberne.txt\n\nRemember: Version skew: kubelet ≤1 minor behind API server. Upgrade control plane first.	
k8s-ops/086fb1946fe1	k8s-ops	easy	kubernetes, helm	What is Helm and how does it help manage Kubernetes applications?	Package manager for Kubernetes. Basically the ability to package YAML files and distribute them to other users and apply them in the cluster(s).\n\nAs a concept it's quite common and can be found in many platforms and services. Think for example on package managers in operating systems. If you use Fedora/RHEL that would be dnf. If you use Ubuntu then, apt. If you don't use Linux, then a different question should be asked and it's why? but that's another topic :)	projects/knowledge/interview/kubernetes/252-what-is-helm.txt
k8s-ops/0a77b6d78cd5	k8s-ops	medium	kubernetes, operator	What components the Operator Framework consists of?	1. Operator SDK - allows developers to build operators\n2. Operator Lifecycle Manager - helps to install, update and generally manage the lifecycle of all operators\n3. Operator Metering - Enables usage reporting for operators that provide specialized services\n4.\n\nRemember: `kubectl explain <resource>` for fields. `kubectl api-resources` for all resource types.	projects/knowledge/interview/kubernetes/207-what-components-the-operator-framework-consists-of.txt
k8s-ops/0c305ca54b3b	k8s-ops	easy	kubernetes, upgrade, helm	What is the role of Helm and how does it simplify Kubernetes deployments?	* Role of Helm: Helm is a package manager for Kubernetes applications, simplifying deployment and management.\n* It uses charts (packages of pre-configured Kubernetes resources) to define, install, and upgrade applications.\n* Helm streamlines the release process and promotes reusability.\n\nRemember: Helm = K8s package manager. Chart=package, Release=instance, Repo=store.\n\nExample: `helm install my-app bitnami/nginx --set service.type=LoadBalancer`	projects/knowledge/interview/kubernetes/360-what-is-the-role-of-helm-and-how-does-it-simplify-.txt
k8s-ops/1fa3ce1e8acd	k8s-ops	hard	kubernetes, kubectl	You are managing multiple Kubernetes clusters. How do you quickly change between the clusters using kubectl?	`kubectl config use-context <context-name>` switches between clusters by changing the active context in your kubeconfig. List available contexts with `kubectl config get-contexts`. Each context combines a cluster, user, and namespace. Consider using `kubectx` for faster switching in multi-cluster environments.	projects/knowledge/interview/kubernetes/018-you-are-managing-multiple-kubernetes-clusters-how-.txt
k8s-ops/24ae8d82a401	k8s-ops	easy	kubernetes, helm	What are some use cases for using Helm template file?	* Deploy the same application across multiple different environments\n* CI/CD\n\nRemember: Helm = K8s package manager. Chart=package, Release=instance, Repo=store.\n\nExample: `helm install my-app bitnami/nginx --set service.type=LoadBalancer`	projects/knowledge/interview/kubernetes/256-what-are-some-use-cases-for-using-helm-template-fi.txt
k8s-ops/3436b599d045	k8s-ops	medium	kubernetes, helm	It is said that Helm is also Templating Engine. What does it mean?	It is useful for scenarios where you have multiple applications and all are similar, so there are minor differences in their configuration files and most values are the same. With Helm you can define a common blueprint for all of them and the values that are not fixed and change can be placeholders. This is called a template file and it looks similar to the following\n\n```\napiVersion: v1\nkind: Pod\nmetadata:\n  name: {[ .Values.name ]}\nspec:\n  containers:\n  - name: {{ .Values.container.name }}\n  image: {{ .Values.container.image }}\n  port: {{ .Values.container.port }}\n```\n\nThe values themselves will in separate file:\n\n```\nname: some-app\ncontainer:\n  name: some-app-container\n  image: some-app-image\n  port: 1991\n```	projects/knowledge/interview/kubernetes/255-it-is-said-that-helm-is-also-templating-engine-wha.txt
k8s-ops/3d6ac78c701e	k8s-ops	medium	kubernetes, operator	What components the Operator consists of?	1. CRD (Custom Resource Definition) - You are fanmiliar with Kubernetes resources like Deployment, Pod, Service, etc. CRD is also a resource, but one that you or the developer the operator defines.\n2. Controller - Custom control loop which runs against the CRD\n\nRemember: `kubectl explain <resource>` for fields. `kubectl api-resources` for all resource types.	projects/knowledge/interview/kubernetes/203-what-components-the-operator-consists-of.txt
k8s-ops/3de047cd0cbf	k8s-ops	medium	kubernetes, upgrade, backup, helm	How Helm supports release management?	Helm allows you to upgrade, remove and rollback to previous versions of charts. In version 2 of Helm it was with what is known as "Tiller". In version 3, it was removed due to security concerns.\n\nRemember: Helm = K8s package manager. Chart=package, Release=instance, Repo=store.\n\nExample: `helm install my-app bitnami/nginx --set service.type=LoadBalancer`	projects/knowledge/interview/kubernetes/258-how-helm-supports-release-management.txt
k8s-ops/3e305694b333	k8s-ops	medium	kubernetes, operator	How does a Kubernetes Operator work using the control loop pattern?	It uses the control loop used by Kubernetes in general. It watches for changes in the application state. The difference is that is uses a custom control loop.\n\nIn addition, it also makes use of CRD's (Custom Resources Definitions) so basically it extends Kubernetes API.\n\nRemember: `kubectl explain <resource>` for fields. `kubectl api-resources` for all resource types.	projects/knowledge/interview/kubernetes/205-how-operator-works.txt
k8s-ops/4ad4163b6ad8	k8s-ops	medium	kubernetes, kubectl	Run a command to view all nodes of the cluster	`kubectl get nodes`\n\nNote: You might want to create an alias (`alias k=kubectl`) and get used to `k get no`\n\nRemember: `kubectl explain <resource>` for fields. `kubectl api-resources` for all resource types.	projects/knowledge/interview/kubernetes/024-run-a-command-to-view-all-nodes-of-the-cluster.txt
k8s-ops/4fe9c4b7d3a7	k8s-ops	hard	kubernetes, troubleshooting, kubectl	You try to run a Pod but it's in "Pending" state. What might be the reason?	One possible reason is that the scheduler which supposed to schedule Pods on nodes, is not running. To verify it, you can run `kubectl get po -A | grep scheduler` or check directly in `kube-system` namespace.\n\nRemember: `kubectl explain <resource>` for fields. `kubectl api-resources` for all resource types.	projects/knowledge/interview/kubernetes/062-you-try-to-run-a-pod-but-its-in-pending-state-what.txt
k8s-ops/52860e63ae92	k8s-ops	hard	kubernetes, monitoring, logging, etcd-ops	Describe how the monitoring solution you are working with monitors Kubernetes	Common Kubernetes monitoring solutions:\n\n- **Prometheus + Grafana**: Most popular open-source stack. Prometheus scrapes metrics via ServiceMonitors; Grafana provides dashboards. Alertmanager handles alerts\n- **metrics-server**: Lightweight, in-memory. Powers `kubectl top` but no persistence or alerting\n- **Datadog/New Relic/Dynatrace**: Commercial SaaS platforms with auto-discovery, APM, and built-in dashboards\n- **ELK/Loki**: Log aggregation (Elasticsearch or Loki + Grafana for unified metrics/logs)\n\nKey things to monitor: node resources, pod status, API server latency, etcd health, PV usage, network policies.\n\nRemember: `kubectl explain <resource>` for fields. `kubectl api-resources` for all resource types.	projects/knowledge/interview/kubernetes/294-describe-how-the-monitoring-solution-you-are-worki.txt
k8s-ops/5ad515931d7d	k8s-ops	medium	kubernetes, helm	How do you list deployed releases?	`helm ls` or `helm list` shows all deployed Helm releases in the current namespace. Add `--all-namespaces` or `-A` to see releases across all namespaces. The output shows NAME, NAMESPACE, REVISION, STATUS, CHART, and APP VERSION. Use `--filter <regex>` to search for specific releases.	projects/knowledge/interview/kubernetes/261-how-do-you-list-deployed-releases.txt
k8s-ops/5c030cc1261a	k8s-ops	medium	kubernetes, kubectl	How to display the resources usages of pods?	`kubectl top pod` shows real-time CPU and memory usage per pod. Requires metrics-server to be installed in the cluster. Add `--containers` to see per-container metrics within pods. Use `--sort-by=cpu` or `--sort-by=memory` to find resource-hungry pods quickly. \nGotcha: if metrics-server is not installed, this command returns an error.	projects/knowledge/interview/kubernetes/197-how-to-display-the-resources-usages-of-pods.txt
k8s-ops/641d0218b41f	k8s-ops	medium	kubernetes, kubectl	Explain the role of kubeconfig in connecting to a Kubernetes cluster.	Role of kubeconfig:\n* kubeconfig is a file that stores cluster information, user credentials, and context.\n* It is used by the kubectl command-line tool to interact with a Kubernetes cluster.\n* kubeconfig allows users to switch between different clusters and contexts.\n\nRemember: Default: `~/.kube/config`. Override: KUBECONFIG env var or --kubeconfig flag.\n\nGotcha: Merge: `KUBECONFIG=f1:f2 kubectl config view --merge --flatten > merged`	projects/knowledge/interview/kubernetes/370-explain-the-role-of-kubeconfig-in-connecting-to-a-.txt
k8s-ops/6e84a39b833d	k8s-ops	medium	kubernetes, upgrade, scaling, operator	Why do we need Operators?	The process of managing stateful applications in Kubernetes isn't as straightforward as managing stateless applications where reaching the desired status and upgrades are both handled the same way for every replica. In stateful applications, upgrading each replica might require different handling due to the stateful nature of the app, each replica might be in a different status. As a result, we often need a human operator to manage stateful applications. Kubernetes Operator is suppose to assist with this.\n\nThis also help with automating a standard process on multiple Kubernetes clusters	projects/knowledge/interview/kubernetes/202-why-do-we-need-operators.txt
k8s-ops/72497b8323f3	k8s-ops	medium	kubernetes, helm, kustomize	Explain the need for Kustomize by describing actual use cases	* You have an helm chart of an application used by multiple teams in your organization and there is a requirement to add annotation to the app specifying the name of the of team owning the app\n  * Without Kustomize you would need to copy the files (chart template in this case) and modify it to include the specific annotations we need\n  * With Kustomize you don't need to copy the entire repo or files\n* You are asked to apply a change/patch to some app without modifying the original files of the app\n  * With Kustomize you can define kustomization.yml file that defines these customizations so you don't need to touch the original app files	projects/knowledge/interview/kubernetes/295-explain-the-need-for-kustomize-by-describing-actua.txt
k8s-ops/72e16a4e5a3c	k8s-ops	easy	kubernetes, monitoring	What is Heapster in Kubernetes?	Heapster is a performance monitoring and metrics collection system for data collected by the Kubelet. This aggregator is natively supported and runs like any other pod within a Kubernetes cluster, which allows it to discover and query usage data from all nodes within the cluster. Note: Heapster has been deprecated and replaced by the Metrics Server.\n\nRemember: `kubectl explain <resource>` for fields. `kubectl api-resources` for all resource types.	projects/knowledge/interview/kubernetes/392-what-is-heapster-in-kubernetes.txt
k8s-ops/7305f11be662	k8s-ops	hard	kubernetes, helm, kustomize	Explain how you would manage configuration drift in a Kubernetes environment.	Managing Configuration Drift:\n* Regularly audit configurations using tools like kube-score.\n* Use version control for configuration files to track changes.\n* Implement GitOps practices for declarative cluster configuration.\n* Leverage Helm or Kustomize for consistent application configuration.\n\nRemember: `kubectl explain <resource>` for fields. `kubectl api-resources` for all resource types.	projects/knowledge/interview/kubernetes/373-explain-how-you-would-manage-configuration-drift-i.txt
k8s-ops/73369665bfb1	k8s-ops	medium	kubernetes, kubectl	What is kubconfig? What do you use it for?	A kubeconfig file is a file used to configure access to Kubernetes when used in conjunction with the kubectl commandline tool (or other clients).\nUse kubeconfig files to organize information about clusters, users, namespaces, and authentication mechanisms.\n\nRemember: `kubectl explain <resource>` for fields. `kubectl api-resources` for all resource types.	projects/knowledge/interview/kubernetes/210-what-is-kubconfig-what-do-you-use-it-for.txt
k8s-ops/76fdf1aeb1bd	k8s-ops	medium	kubernetes, monitoring	What monitoring solutions are you familiar with in regards to Kubernetes?	There are many types of monitoring solutions for Kubernetes. Some open-source, some are in-memory, some of them cost money, ... here is a short list:\n\n* metrics-server: in-memory open source monitoring\n* datadog: $$$\n* prometheus: open source monitoring solution\n\nRemember: `kubectl explain <resource>` for fields. `kubectl api-resources` for all resource types.	projects/knowledge/interview/kubernetes/293-what-monitoring-solutions-are-you-familiar-with-in.txt
k8s-ops/78e1eb5649e6	k8s-ops	medium	kubernetes, cert-rotation, helm	Why do we need Helm? What would be the use case for using it?	Sometimes when you would like to deploy a certain application to your cluster, you need to create multiple YAML files/components like: Secret, Service, ConfigMap, etc. This can be tedious task. So it would make sense to ease the process by introducing something that will allow us to share these bundle of YAMLs every time we would like to add an application to our cluster. This something is called Helm.\n\nA common scenario is having multiple Kubernetes clusters (prod, dev, staging). Instead of individually applying different YAMLs in each cluster, it makes more sense to create one Chart and install it in every cluster.\n\nAnother scenario is, you would like to share what you've created with the community. For people and companies to easily deploy your application in their cluster.	projects/knowledge/interview/kubernetes/253-why-do-we-need-helm-what-would-be-the-use-case-for.txt
k8s-ops/7a0090ad8a16	k8s-ops	medium	kubernetes, kubectl	Create a list of all nodes in JSON format and store it in a file called "some_nodes.json"	`kubectl get nodes -o json > some_nodes.json` exports all node objects in JSON format. Use `-o jsonpath='{.items[*].metadata.name}'` for just the names. Other output formats: `-o yaml`, `-o wide` (extra columns), `-o name` (just resource names). Pipe to `jq` for filtering: `kubectl get nodes -o json | jq '.items[].status.conditions'`.	projects/knowledge/interview/kubernetes/025-create-a-list-of-all-nodes-in-json-format-and-stor.txt
k8s-ops/7a4e3b4c97ca	k8s-ops	medium	kubernetes, kubectl	After creating a service, how to check it was created?	`kubectl get svc` lists all services in the current namespace showing NAME, TYPE, CLUSTER-IP, EXTERNAL-IP, and PORT(S). Add `-o wide` for selector details or `--all-namespaces` for cluster-wide view. Use `kubectl describe svc <name>` for endpoint details and to verify pods are being targeted correctly.	projects/knowledge/interview/kubernetes/091-after-creating-a-service-how-to-check-it-was-creat.txt
k8s-ops/7b7c1ec4e6a8	k8s-ops	medium	kubernetes, helm, kubectl	Is it possible to override values in values.yaml file when installing a chart?	Yes. You can pass another values file:\n`helm install --values=override-values.yaml [CHART_NAME]`\n\nOr directly on the command line: `helm install --set some_key=some_value`\n\nRemember: `kubectl explain <resource>` for fields. `kubectl api-resources` for all resource types.	projects/knowledge/interview/kubernetes/260-is-it-possible-to-override-values-in-valuesyaml-fi.txt
k8s-ops/7bdaaf7cc17f	k8s-ops	medium	kubernetes, upgrade, operator, kubectl	Describe in detail what is the Operator Lifecycle Manager	It's part of the Operator Framework, used for managing the lifecycle of operators. It basically extends Kubernetes so a user can use a declarative way to manage operators (installation, upgrade, ...).\n\nRemember: `kubectl explain <resource>` for fields. `kubectl api-resources` for all resource types.	projects/knowledge/interview/kubernetes/208-describe-in-detail-what-is-the-operator-lifecycle-.txt
k8s-ops/7c7081a8ee4f	k8s-ops	medium	kubernetes, operator	Explain the purpose and usage of Operators in the Kubernetes ecosystem.	* Operators in Kubernetes: Operators are custom controllers that extend Kubernetes functionality.\n* They automate complex operational tasks for managing applications.\n* Operators use custom resources to define application-specific behaviors.\n\nRemember: `kubectl explain <resource>` for fields. `kubectl api-resources` for all resource types.	projects/knowledge/interview/kubernetes/361-explain-the-purpose-and-usage-of-operators-in-the-.txt
k8s-ops/7f8f59e60c6d	k8s-ops	medium	kubernetes, helm	Discuss the use of Helm charts for application packaging in Kubernetes.	Helm Charts: Helm Charts are packages of pre-configured Kubernetes resources.\nThey include templates for deployments, services, ConfigMaps, and other resources needed for an application.\nHelm Charts simplify the deployment and versioning of Kubernetes applications.\nHelm Charts encapsulate the complexity of deploying applications in Kubernetes, making it easy to share and reproduce application deployments. They provide a standardized and reusable way to package, deploy, and manage applications across different environments.\n\nRemember: Helm = K8s package manager. Chart=package, Release=instance, Repo=store.\n\nExample: `helm install my-app bitnami/nginx --set service.type=LoadBalancer`	projects/knowledge/interview/kubernetes/335-discuss-the-use-of-helm-charts-for-application-pac.txt
k8s-ops/80f2157ea80e	k8s-ops	medium	kubernetes, kubectl	Check if there are any limits on one of the pods in your cluster	`kubectl describe pod <POD_NAME> | grep -i limits` shows resource limits (CPU, memory) set on the pod's containers. You can also use `kubectl get pod <name> -o jsonpath='{.spec.containers[*].resources}'` for structured output. \nGotcha: pods without limits can consume unbounded resources and affect other workloads on the same node.	projects/knowledge/interview/kubernetes/290-check-if-there-are-any-limits-on-one-of-the-pods-i.txt
k8s-ops/86ab244573bd	k8s-ops	easy	kubernetes, kubectl, commands	Which command lists all Pods in a Kubernetes cluster?	kubectl get pods (for current namespace) or kubectl get pods -A (for all namespaces) will list all pods and their status.\n\nRemember: `kubectl explain <resource>` for fields. `kubectl api-resources` for all resource types.	projects/knowledge/interview/kubernetes/426-list-all-pods.txt
k8s-ops/90c47048e4a1	k8s-ops	medium	kubernetes, logging, scaling, kubectl	What the following output of kubectl get rs means?	The replicaset `web` has 2 replicas. It seems that the containers inside the Pod(s) are not yet running since the value of READY is 0. It might be normal since it takes time for some containers to start running and it might be due to an error. Running `kubectl describe po POD_NAME` or `kubectl logs POD_NAME` can give us more information.\n\nRemember: `kubectl explain <resource>` for fields. `kubectl api-resources` for all resource types.	projects/knowledge/interview/kubernetes/132-what-the-following-output-of-kubectl-get-rs-means-.txt
k8s-ops/96009dfd98b6	k8s-ops	hard	kubernetes, logging, troubleshooting, kubectl	Describe the steps involved in troubleshooting a pod that is not starting in Kubernetes.	**Troubleshooting Steps:**\n • Check Pod Status: Use kubectl get pods to check the pod's current status.\n • Pod Events: Use kubectl describe pod <pod-name> to view events and identify issues.\n • Logs: Examine container logs using kubectl logs <pod-name>.\n • Resource Constraints: Verify resource requests and limits in the pod spec.\n • Pod Configuration: Review the pod's configuration, including ConfigMaps and Secrets.\n • Network Issues: Check network policies and connectivity.\n • Image Availability: Ensure the container image is accessible and correct.\n • Health Probes: Inspect readiness	
k8s-ops/961ee52fad02	k8s-ops	hard	kubernetes, logging, troubleshooting, kubectl	How do you debug a crash-looping pod?	Systematic approach:\n\n**1. Logs first**:\n```bash\nkubectl logs pod-name\nkubectl logs pod-name --previous  # Previous crash\n```\n\n**2. Events/describe**:\n```bash\nkubectl describe pod pod-name\n```\nLook for: OOMKilled, image pull errors, failed mounts, probe failures.\n\n**3. Container command/env**:\n* Wrong entrypoint?\n* Missing environment variables?\n* Bad config mounted?\n\n**4. Resource limits**:\n* Memory limit too low → OOMKilled\n* CPU throttling causing timeouts\n\n**5. Debug container** (if logs don't help):\n```bash\nkubectl debug pod-name -it --image=busybox\n```\n\nMost crashes are: missing config, wrong image tag, resource limits, or dependency not ready.\n\nRemember: Debug flow: Get→Describe→Logs→Exec. Mnemonic: "GDLE."	projects/knowledge/interview/kubernetes/377-how-do-you-debug-a-crash-looping-pod.txt
k8s-ops/970d6e4f8d62	k8s-ops	hard	kubernetes, networking, conntrack, troubleshooting, performance	Kubernetes cluster randomly drops traffic under load. Nodes look healthy. What's the root cause?	Most likely conntrack table exhaustion from NAT/service explosion.\n\nThe problem:\n- Every Service + Pod combination creates conntrack entries\n- kube-proxy uses iptables NAT for service routing\n- Each connection = conntrack entry\n- Default nf_conntrack_max often too low (65536)\n\nWhy traffic drops silently:\n- New connections can't create conntrack entries\n- Packets dropped in kernel, no error to application\n- No log by default (must enable)\n- Appears as random timeouts\n\nRemember: `kubectl explain <resource>` for fields. `kubectl api-resources` for all resource types.	projects/knowledge/interview/kubernetes/400-k8s-random-traffic-drop.txt
k8s-ops/98d95b87bd77	k8s-ops	medium	kubernetes, kubectl, kustomize	Describe in high-level how Kustomize works	1. You add kustomization.yml file in the folder of the app you would like to customize.\n   1. You define the customizations you would like to perform\n2. You run `kustomize build APP_PATH` where your kustomization.yml also resides\n\nRemember: `kubectl explain <resource>` for fields. `kubectl api-resources` for all resource types.	projects/knowledge/interview/kubernetes/296-describe-in-high-level-how-kustomize-works.txt
k8s-ops/9d5fadee7c90	k8s-ops	medium	kubernetes, helm	Explain "Helm Charts"	Helm Charts is a bundle of YAML files. A bundle that you can consume from repositories or create your own and publish it to the repositories.\n\nRemember: Helm = K8s package manager. Chart=package, Release=instance, Repo=store.\n\nExample: `helm install my-app bitnami/nginx --set service.type=LoadBalancer`	projects/knowledge/interview/kubernetes/254-explain-helm-charts.txt
k8s-ops/a1d421909746	k8s-ops	easy	kubernetes, kubectl	What is kubectl and how is it used to manage Kubernetes clusters?	Kubectl is a CLI (command-line interface) that is used to run commands against Kubernetes clusters. As such, it controls the Kubernetes cluster manager through different create and manage commands on the Kubernetes components. It is the primary tool for interacting with Kubernetes clusters.\n\nRemember: `kubectl explain <resource>` for fields. `kubectl api-resources` for all resource types.	projects/knowledge/interview/kubernetes/410-what-is-kubectl.txt
k8s-ops/a2f5a9ae1d03	k8s-ops	hard	kubernetes, logging, troubleshooting, operations	Why do container logs sometimes bring down Kubernetes nodes?	Uncontrolled log growth exhausts disk space or inodes, causing kubelet failure.\n\nThe cascade:\n\n1. Application logs to stdout/stderr\n2. Container runtime captures to JSON files\n3. Files grow without bound (misconfigured rotation)\n4. Disk fills OR inodes exhausted\n5. kubelet can't function (needs disk for pods)\n6. Node goes NotReady\n7. Pods rescheduled, bring their logs elsewhere\n8. Repeat on next node\n\nSpecific failure modes:\n\nExample: `kubectl logs -f pod --previous` — follow live or show crashed container logs.\n\nRemember: Multi-container: `--container=name`. Errors without it on multi-container pods.	projects/knowledge/interview/kubernetes/401-container-logs-kill-nodes.txt
k8s-ops/a98e96e3e68e	k8s-ops	medium	kubernetes, logging, scaling, troubleshooting	Perhaps a general question but, you suspect one of the pods is having issues, you don't know what exactly. What do you do?	Start by inspecting the pods status. we can use the command `kubectl get pods` (--all-namespaces for pods in system namespace) \n\nIf we see "Error" status, we can keep debugging by running the command `kubectl describe pod [name]`. In case we still don't see anything useful we can try stern for log tailing. \n\nIn case we find out there was a temporary issue with the pod or the system, we can try restarting the pod with the following `kubectl scale deployment [name] --replicas=0` \n\nSetting the replicas to 0 will shut down the process. Now start it with `kubectl scale deployment [name] --replicas=1`	projects/knowledge/interview/kubernetes/198-perhaps-a-general-question-but-you-suspect-one-of-.txt
k8s-ops/aa2d758c1a7a	k8s-ops	hard	kubernetes, logging, troubleshooting, kubectl	An engineer in your team runs a Pod but the status he sees is "CrashLoopBackOff". What does it means? How to identify the issue?	The container failed to run (due to different reasons) and Kubernetes tries to run the Pod again after some delay (= BackOff time).\n\nSome reasons for it to fail:\n  - Misconfiguration - misspelling, non supported value, etc.\n  - Resource not available - nodes are down, PV not mounted, etc.\n\nSome ways to debug:\n\n1. `kubectl describe pod POD_NAME`\n   1. Focus on `State` (which should be Waiting, CrashLoopBackOff) and `Last State` which should tell what happened before (as in why it failed)\n2. Run `kubectl logs mypod`\n   1. This should provide an accurate output of \n   2. For specific container, you can add `-c CONTAINER_NAME`	projects/knowledge/interview/kubernetes/302-an-engineer-in-your-team-runs-a-pod-but-the-status.txt
k8s-ops/ae4f88e0f61d	k8s-ops	medium	kubernetes, kubectl	How to execute the command "ls" in an existing pod?	`kubectl exec some-pod -it -- ls` runs the `ls` command inside a running pod's container. For a shell: `kubectl exec -it pod-name -- /bin/sh`. In multi-container pods, specify the container: `kubectl exec -it pod-name --container=sidecar -- /bin/sh`. \nGotcha: exec requires the container to have the binary installed — distroless images may lack common tools.	projects/knowledge/interview/kubernetes/192-how-to-execute-the-command-ls-in-an-existing-pod.txt
k8s-ops/ae5702f1bed5	k8s-ops	medium	kubernetes, logging, operator	What openshift-operator-lifecycle-manager namespace includes?	It includes:\n\n  * catalog-operator - Resolving and installing ClusterServiceVersions the resource they specify.\n  * olm-operator - Deploys applications defined by ClusterServiceVersion resource\n\nRemember: `kubectl explain <resource>` for fields. `kubectl api-resources` for all resource types.	projects/knowledge/interview/kubernetes/209-what-openshift-operator-lifecycle-manager-namespac.txt
k8s-ops/b270bbb6b1e9	k8s-ops	medium	kubernetes, cert-rotation, helm	How to view revision history for a certain release?	`helm history RELEASE_NAME` shows the revision history including REVISION number, STATUS, CHART version, and DESCRIPTION. This lets you see what changed between deployments and identify which revision to rollback to. \nExample: `helm history my-app` then `helm rollback my-app 3` to revert to revision 3.	projects/knowledge/interview/kubernetes/263-how-to-view-revision-history-for-a-certain-release.txt
k8s-ops/b92d4d36d480	k8s-ops	medium	kubernetes, monitoring, scaling	Explain how you would monitor and scale a critical production application in Kubernetes.	**Monitoring:* • \n • Use monitoring tools like Prometheus, Grafana, or Kubernetes-native solutions.\n • Set up alerts based on key metrics, including resource utilization and application health.\n • Monitor pod and node status, and track events.\n**Scaling:* • \n • Utilize Horizontal Pod Autoscaling (HPA) based on metrics like CPU or custom metrics.\n • Consider Vertical Pod Autoscaling for adjusting resource limits dynamically.\n • Implement Cluster Autoscaler for scaling the node pool based on demand.\n • Effective monitoring involves selecting appropriate tools and setting up alerts.	
k8s-ops/b9829690c775	k8s-ops	easy	kubernetes, operator	What is the Operator Framework?	open source toolkit used to manage k8s native applications, called operators, in an automated and efficient way.\n\nRemember: `kubectl explain <resource>` for fields. `kubectl api-resources` for all resource types.	projects/knowledge/interview/kubernetes/206-what-is-the-operator-framework.txt
k8s-ops/ba9af627f0cd	k8s-ops	easy	kubernetes, etcd-ops, kubectl	Describe shortly and in high-level, what happens when you run kubectl get nodes	1. Your user is getting authenticated\n2. Request is validated by the kube-apiserver\n3. Data is retrieved from etcd\n\nRemember: get=summary list, describe=detailed+events. "get=glance, describe=deep dive."	projects/knowledge/interview/kubernetes/013-describe-shortly-and-in-high-level-what-happens-wh.txt
k8s-ops/be08a8e9b776	k8s-ops	easy	kubernetes, monitoring	What is Container resource monitoring?	Container resource monitoring refers to the activity that collects the metrics and tracks the health of containerized applications and microservices environments. It helps to improve health and performance and also makes sure that they operate smoothly. Common tools include Prometheus, Grafana, and the Kubernetes Metrics Server.\n\nRemember: `kubectl explain <resource>` for fields. `kubectl api-resources` for all resource types.	projects/knowledge/interview/kubernetes/405-what-is-container-resource-monitoring.txt
k8s-ops/c95315153999	k8s-ops	easy	kubernetes, kubectl	How do you delete a pod in Kubernetes and what happens after?	`kubectl delete pod pod_name`\n\nGotcha: delete sends SIGTERM, 30s grace, then SIGKILL. `--grace-period=0 --force` skips.\n\nRemember: Deleting a Deployment also removes its ReplicaSets and Pods.	projects/knowledge/interview/kubernetes/058-how-to-delete-a-pod.txt
k8s-ops/cb51ec5765ef	k8s-ops	medium	kubernetes, scaling, kubectl	What the following command does?	It exposes a ReplicaSet by creating a service called 'replicaset-svc'. The exposed port is 2017 (this is the port used by the application) and the service type is NodePort which means it will be reachable externally.\n\nRemember: `kubectl explain <resource>` for fields. `kubectl api-resources` for all resource types.	projects/knowledge/interview/kubernetes/098-what-the-following-command-does-kubectl-expose-rs-.txt
k8s-ops/cf1451a380ff	k8s-ops	medium	kubernetes, upgrade, helm	How to upgrade a release?	`helm upgrade RELEASE_NAME CHART_NAME` updates a deployed release with new chart values or a new chart version. Add `--set key=value` for inline overrides or `-f values.yaml` for file-based config. Use `--dry-run` to preview changes before applying. \nGotcha: always pin chart versions in production to avoid unexpected upgrades.	projects/knowledge/interview/kubernetes/264-how-to-upgrade-a-release.txt
k8s-ops/e0efb16006ca	k8s-ops	medium	kubernetes, upgrade, helm, kubectl	What is Helm, and how is it used in Kubernetes?	* Helm: Helm is a package manager for Kubernetes applications.\n* It simplifies the deployment and management of Kubernetes applications by packaging them into charts.\n* Charts are pre-configured Kubernetes resource definitions that can be easily deployed and versioned.\n* Helm allows users to define, install, and upgrade even the most complex Kubernetes applications with a single command. \n* Charts encapsulate all the required Kubernetes resources and configurations, making it easier to share and reproduce application deployments across different environments.\n\nRemember: Helm = K8s package manager. Chart=package, Release=instance, Repo=store.\n\nExample: `helm install my-app bitnami/nginx --set service.type=LoadBalancer`	projects/knowledge/interview/kubernetes/327-what-is-helm-and-how-is-it-used-in-kubernetes.txt
k8s-ops/e26af8ee2ebf	k8s-ops	medium	kubernetes, helm	Explain the Helm Chart Directory Structure	someChart/     -> the name of the chart\n  Chart.yaml   -> meta information on the chart\n  values.yaml  -> values for template files\n  charts/      -> chart dependencies\n  templates/   -> templates files :)\n\nRemember: Helm = K8s package manager. Chart=package, Release=instance, Repo=store.\n\nExample: `helm install my-app bitnami/nginx --set service.type=LoadBalancer`	projects/knowledge/interview/kubernetes/257-explain-the-helm-chart-directory-structure.txt
k8s-ops/e8d9e1aef418	k8s-ops	hard	kubernetes, logging, scaling, troubleshooting	You encounter a performance issue in a Kubernetes cluster. How do you diagnose and resolve it?	**Diagnosing and Resolving Performance Issues:* • \n • Resource Utilization: Check CPU, memory, and storage usage for nodes and pods.\n • Logs and Events: Analyze container logs and Kubernetes events for anomalies.\n • Network: Examine network policies, traffic, and potential bottlenecks.\n • Pod Placement: Review node placement and resource allocation for pods.\n • Kubernetes Components: Inspect the health and performance of Kubernetes control plane components.\n • Application Code: Review application code for performance bottlenecks.\n • Scaling: Consider scaling resources based on demand.\n	
k8s-ops/eabcb3172b4b	k8s-ops	medium	kubernetes, operator	Are there any tools, projects you are using for building Operators?	This one is based more on a personal experience and taste...\n\n* Operator Framework\n* Kubebuilder\n* Controller Runtime\n...\n\nRemember: `kubectl explain <resource>` for fields. `kubectl api-resources` for all resource types.	projects/knowledge/interview/kubernetes/212-are-there-any-tools-projects-you-are-using-for-bui.txt
k8s-ops/f0a155493d5f	k8s-ops	medium	kubernetes, troubleshooting, kubectl	What does the "ErrImagePull" status of a Pod means?	It wasn't able to pull the image specified for running the container(s). This can happen if the client didn't authenticated for example. \nMore details can be obtained with `kubectl describe po <POD_NAME>`.\n\nRemember: `kubectl explain <resource>` for fields. `kubectl api-resources` for all resource types.	projects/knowledge/interview/kubernetes/042-what-does-the-errimagepull-status-of-a-pod-means.txt
k8s-ops/f1a2b4c6e6b5	k8s-ops	medium	kubernetes, helm	How do you search for charts?	`helm search hub <keyword>` searches Artifact Hub for charts across all repositories. `helm search repo <keyword>` searches only locally-added repositories. \nExample: `helm search hub prometheus` finds Prometheus charts from multiple publishers. Add `--max-col-width 80` for readable descriptions and `--version <semver>` to filter by chart version.	projects/knowledge/interview/kubernetes/259-how-do-you-search-for-charts.txt
k8s-ops/f1e56e6f0c6e	k8s-ops	hard	kubernetes, logging, troubleshooting, cert-rotation	Users unable to reach an application running on a Pod on Kubernetes. What might be the issue and how to check?	Troubleshoot Kubernetes application connectivity layer by layer:\n\n1. **Pod status**: `kubectl get pods` / `kubectl describe pod` — check for CrashLoopBackOff, Pending, ImagePullBackOff\n2. **Pod health**: `kubectl exec -it pod -- curl localhost:PORT` — verify app responds inside container\n3. **Service & endpoints**: `kubectl get svc` / `kubectl get endpoints` — confirm selector matches pod labels\n4. **Network policies**: Check if ingress/egress rules are blocking traffic\n5. **DNS**: `kubectl exec -- nslookup service-name` — verify CoreDNS resolution\n6. **Ingress/LB**: Check ingress rules, TLS config, and controller logs\n7. **External access**: Verify NodePort, LoadBalancer provisioning, or DNS pointing to cluster	projects/knowledge/interview/kubernetes/267-users-unable-to-reach-an-application-running-on-a-.txt
k8s-ops/f690bab317ff	k8s-ops	easy	kubernetes, kubectl	Which command will list all the object types in a cluster?	`kubectl api-resources` lists all resource types (pods, services, deployments, etc.) available in the cluster. Shows NAME, SHORTNAMES, APIVERSION, NAMESPACED, and KIND columns. Use `--namespaced=true` to filter namespaced resources only. Combine with `kubectl explain <resource>` to explore field schemas for any listed resource.	projects/knowledge/interview/kubernetes/021-which-command-will-list-all-the-object-types-in-a-.txt
k8s-ops/ffb889bcab28	k8s-ops	medium	kubernetes, troubleshooting, pods	How do you debug a failing pod?	Check the events, logs, container status, resource limits, readiness/liveness probes, and images. If needed: kubectl describe, kubectl logs, and verify networking, secrets, configmaps, and node health.\n\nRemember: Debug flow: Get→Describe→Logs→Exec. Mnemonic: "GDLE."	projects/knowledge/interview/kubernetes/375-how-do-you-debug-a-failing-pod.txt
k8s-ops/train-lab01-a	k8s-ops	hard	kubernetes, probes, training	What happens when a readiness probe fails on a Kubernetes pod?	The pod is removed from service endpoints, so it stops receiving traffic. Unlike a liveness probe failure (which restarts the container), a readiness failure just takes the pod out of the load balancer. The pod keeps running. See: training/interactive/runtime-labs/lab-runtime-01-rollout-probe-failure/\n\nRemember: `kubectl explain <resource>` for fields. `kubectl api-resources` for all resource types.	training/interactive/runtime-labs/lab-runtime-01-rollout-probe-failure/
k8s-ops/train-lab01-b	k8s-ops	medium	kubernetes, deployments, training	Why does a deployment get stuck in 'Progressing' status?	A deployment gets stuck when new pods fail to become Ready within the progressDeadlineSeconds (default 600s). Common causes: readiness probe failure, ImagePullBackOff, OOMKilled on startup, or missing dependencies. Check: kubectl rollout status, describe pod events.\n\nRemember: `kubectl explain <resource>` for fields. `kubectl api-resources` for all resource types.	training/interactive/runtime-labs/lab-runtime-01-rollout-probe-failure/
k8s-ops/train-lab02-a	k8s-ops	easy	kubernetes, hpa, training	What does '<unknown>/50%' mean in HPA status?	It means the HPA cannot read CPU metrics. Usually because metrics-server is not installed, not healthy, or the deployment doesn't have CPU resource requests defined. HPA needs requests to calculate percentage-based utilization.\n\nRemember: `kubectl explain <resource>` for fields. `kubectl api-resources` for all resource types.	training/interactive/runtime-labs/lab-runtime-02-hpa-live-scaling/
k8s-ops/train-lab02-b	k8s-ops	easy	kubernetes, hpa, training	What are the prerequisites for HPA to work?	1) metrics-server must be installed and healthy; 2) The target deployment must have resource requests (at minimum CPU requests for CPU-based scaling); 3) The metrics API must be accessible (/apis/metrics.k8s.io/v1beta1).\n\nRemember: `kubectl explain <resource>` for fields. `kubectl api-resources` for all resource types.	training/interactive/runtime-labs/lab-runtime-02-hpa-live-scaling/
k8s-ops/train-lab08-a	k8s-ops	easy	kubernetes, oom, training	What does exit code 137 mean for a Kubernetes container?	Exit code 137 = 128 + 9 (SIGKILL). The container was killed by the kernel OOM killer because it exceeded its memory limit. Check with: kubectl describe pod | grep 'Last State' — look for 'Reason: OOMKilled'.\n\nRemember: `kubectl explain <resource>` for fields. `kubectl api-resources` for all resource types.	training/interactive/runtime-labs/lab-runtime-08-resource-limits-oom/
k8s-ops/train-lab08-b	k8s-ops	medium	kubernetes, resources, training	How do you right-size memory limits for a Kubernetes deployment?	1) Set requests based on steady-state usage (observe via kubectl top or Prometheus); 2) Set limits 1.5-2x requests to allow for spikes; 3) Monitor with container_memory_working_set_bytes metric; 4) Set alerts at 80% of limit. Never set limits = requests for variable workloads.\n\nRemember: `kubectl explain <resource>` for fields. `kubectl api-resources` for all resource types.	training/interactive/runtime-labs/lab-runtime-08-resource-limits-oom/
k8s-ops/train-runbook-a	k8s-ops	hard	kubernetes, crashloop, training	What are the top 3 causes of CrashLoopBackOff?	1) Application error on startup (bad config, missing env var, import error); 2) OOMKilled (memory limit too low); 3) Liveness probe failing too aggressively (app healthy but probe times out). Always check 'kubectl logs --previous' to see the crash reason.\n\nRemember: `kubectl explain <resource>` for fields. `kubectl api-resources` for all resource types.	training/library/runbooks/crashloopbackoff.md
k8s-ops/train-runbook-b	k8s-ops	medium	kubernetes, images, training	What's the difference between ImagePullBackOff and ErrImagePull?	ErrImagePull is the first failure to pull an image. ImagePullBackOff means Kubernetes tried, failed, and is now waiting with exponential backoff before retrying. Common causes: wrong tag, private registry without credentials, or image not imported into k3s.\n\nRemember: `kubectl explain <resource>` for fields. `kubectl api-resources` for all resource types.	training/library/runbooks/imagepullbackoff.md
k8s-hpa/a3f7b1c9d2e4	k8s-ops	easy	k8s, hpa, autoscaling, metrics	What must be set on pods for HPA CPU-based scaling to work?	Resource requests (resources.requests.cpu) must be defined. HPA computes utilization as currentUsage / request, so without requests the metric is undefined and HPA cannot function.\n\nRemember: `kubectl explain <resource>` for fields. `kubectl api-resources` for all resource types.	training/library/topics/k8s-ops/primer.md
k8s-hpa/b8e2d4f6a1c3	k8s-ops	easy	k8s, hpa, autoscaling	What is the default HPA scale-down stabilization window?	300 seconds (5 minutes). The controller looks back over this window and picks the highest (most conservative) replica count recommendation to prevent flapping.\n\nExample: `kubectl scale deployment web --replicas=5`. HPA for automatic scaling.	training/library/topics/k8s-ops/primer.md
k8s-hpa/c1d5e7f9b3a2	k8s-ops	easy	k8s, hpa, metrics-server	How do you verify that metrics-server is running and providing data?	Run kubectl top nodes and kubectl top pods. If they return metrics, the server is working. Also check kubectl get apiservices | grep metrics and kubectl -n kube-system get pods -l k8s-app=metrics-server.\n\nRemember: `kubectl explain <resource>` for fields. `kubectl api-resources` for all resource types.	training/library/topics/k8s-ops/primer.md
k8s-hpa/d4a6c8e0f2b1	k8s-ops	medium	k8s, hpa, autoscaling, formula	What formula does the HPA use to compute the desired replica count?	desiredReplicas = ceil(currentReplicas * (currentMetricValue / desiredMetricValue)). For example, if current CPU is 90% and target is 70%, the scale factor is 90/70 = 1.28, so replicas increase by roughly 28%.\n\nRemember: `kubectl explain <resource>` for fields. `kubectl api-resources` for all resource types.	training/library/topics/k8s-ops/primer.md
k8s-hpa/e7b9d1f3a5c4	k8s-ops	medium	k8s, hpa, custom-metrics	What are the four metric types supported by HPA v2 and when would you use each?	Resource (CPU/memory from metrics-server), Pods (per-pod app metrics like RPS via custom metrics adapter), External (cloud service metrics like SQS queue depth via external adapter), and Object (metrics from a specific Kubernetes object like Ingress RPS via custom adapter).\n\nRemember: `kubectl explain <resource>` for fields. `kubectl api-resources` for all resource types.	training/library/topics/k8s-ops/primer.md
k8s-hpa/f2a4c6e8d0b3	k8s-ops	medium	k8s, hpa, vpa, autoscaling	Why should VPA and HPA not target the same metric?	They will conflict. HPA adjusts replica count based on per-pod utilization, while VPA adjusts resource requests on individual pods. If both act on CPU, HPA might scale out while VPA simultaneously changes the request denominator, causing an unstable feedback loop.\n\nRemember: `kubectl explain <resource>` for fields. `kubectl api-resources` for all resource types.	training/library/topics/k8s-ops/primer.md
k8s-hpa/a1c3e5d7f9b2	k8s-ops	medium	k8s, hpa, memory, scaling	Why is memory generally a poor primary metric for HPA scaling?	Many applications (JVM, Python) allocate memory and never release it even after load drops. Memory utilization stays high regardless of current demand, so HPA never scales down. CPU is preferred because it correlates better with active request load.\n\nRemember: `kubectl explain <resource>` for fields. `kubectl api-resources` for all resource types.	training/library/topics/k8s-ops/primer.md
k8s-hpa/b5d7f9a1c3e4	k8s-ops	medium	k8s, hpa, multiple-metrics	When HPA is configured with multiple metrics, how does it decide the replica count?	It evaluates each metric independently and takes the maximum desired replica count across all metrics. The most demanding metric wins. This means combining metrics with very different response characteristics can lead to unexpected scaling behavior.\n\nRemember: `kubectl explain <resource>` for fields. `kubectl api-resources` for all resource types.	training/library/topics/k8s-ops/primer.md
k8s-hpa/c8e0a2d4f6b1	k8s-ops	hard	k8s, hpa, behavior, scaling-policies	Explain the selectPolicy field in HPA behavior and how Max vs Min affect scaling aggressiveness.	selectPolicy determines which policy to apply when multiple policies are defined. Max picks whichever policy allows the largest change (most aggressive scaling). Min picks the smallest change (most conservative). Disabled prevents scaling in that direction entirely. For example, with both a Percent(100%) and Pods(5) policy on scaleUp with selectPolicy: Max, the HPA uses whichever allows adding more pods.\n\nRemember: `kubectl explain <resource>` for fields. `kubectl api-resources` for all resource types.	training/library/topics/k8s-ops/primer.md
k8s-hpa/d9f1b3a5c7e2	k8s-ops	hard	k8s, hpa, pdb, disruption	How can PodDisruptionBudget conflict with HPA scale-down, and what is the best practice to avoid it?	PDB enforces minAvailable during voluntary disruptions. HPA sets the desired replica count, but if scaling down would violate PDB constraints during node drains or spot terminations, evictions are blocked. The HPA controller itself does not check PDB. Best practice: set HPA minReplicas to at least what PDB requires as minimum available.\n\nExample: `kubectl scale deployment web --replicas=5`. HPA for automatic scaling.	training/library/topics/k8s-ops/primer.md
k8s-hpa/e2a4c6b8d0f3	k8s-ops	hard	k8s, hpa, keda, scale-to-zero	Why can HPA not scale to zero, and what are the alternatives?	HPA requires minReplicas >= 1. It cannot scale to zero because with zero pods there are no metrics to evaluate for scale-up decisions. For scale-to-zero capability (cost savings on idle workloads), use KEDA (Kubernetes Event-Driven Autoscaling), Knative Serving, or a custom controller that can scale from zero based on external signals like queue depth or incoming HTTP requests.\n\nExample: `kubectl scale deployment web --replicas=5`. HPA for automatic scaling.	training/library/topics/k8s-ops/primer.md
k8s-probes/a1b2c3d4e5f6	k8s-ops	easy	k8s, probes, liveness	What does a Kubernetes liveness probe determine?	Whether the container is still alive. If the liveness probe fails, Kubernetes kills and restarts the container.\n\nRemember: `kubectl explain <resource>` for fields. `kubectl api-resources` for all resource types.	training/library/topics/k8s-ops/primer.md
k8s-probes/b2c3d4e5f6a7	k8s-ops	easy	k8s, probes, readiness	What happens when a readiness probe fails?	The pod is removed from Service endpoints so it stops receiving traffic, but it is NOT restarted. Traffic is routed to other healthy pods.\n\nRemember: `kubectl explain <resource>` for fields. `kubectl api-resources` for all resource types.	training/library/topics/k8s-ops/primer.md
k8s-probes/c3d4e5f6a7b8	k8s-ops	easy	k8s, probes, mechanisms	What are the four probe mechanisms Kubernetes supports?	httpGet (HTTP GET returning 2xx/3xx), tcpSocket (TCP port is open), exec (command exits 0), and grpc (gRPC health check, Kubernetes 1.24+).\n\nRemember: `kubectl explain <resource>` for fields. `kubectl api-resources` for all resource types.	training/library/topics/k8s-ops/primer.md
k8s-probes/d4e5f6a7b8c9	k8s-ops	easy	k8s, probes, startup	What is the purpose of a startup probe?	It tells Kubernetes the container is still booting. While the startup probe is running, liveness and readiness probes are disabled. Once it succeeds, it never runs again and the other probes take over.\n\nRemember: `kubectl explain <resource>` for fields. `kubectl api-resources` for all resource types.	training/library/topics/k8s-ops/primer.md
k8s-probes/e5f6a7b8c9d0	k8s-ops	medium	k8s, probes, liveness, anti-pattern	Why is checking database connectivity in a liveness probe dangerous?	If the database goes down, all pods fail liveness simultaneously, Kubernetes restarts them all, they thundering-herd the database on reconnection, and the cycle repeats — a cascading restart storm.\n\nRemember: `kubectl explain <resource>` for fields. `kubectl api-resources` for all resource types.	training/library/topics/k8s-ops/primer.md
k8s-probes/f6a7b8c9d0e1	k8s-ops	medium	k8s, probes, parameters	What does failureThreshold control, and how does it interact with periodSeconds?	failureThreshold is the number of consecutive probe failures before Kubernetes takes action. Combined with periodSeconds, it sets the detection window: e.g., failureThreshold=3 and periodSeconds=10 means 30 seconds of failures before restart or endpoint removal.\n\nRemember: `kubectl explain <resource>` for fields. `kubectl api-resources` for all resource types.	training/library/topics/k8s-ops/primer.md
k8s-probes/a7b8c9d0e1f2	k8s-ops	medium	k8s, probes, deployment, rollout	How do readiness probes interact with rolling deployments?	New pods must pass their readiness probe before receiving traffic and before old pods are terminated. If new pods never become ready (e.g., broken config), the rollout stalls and old pods continue serving — preventing a bad deploy from causing downtime.\n\nRemember: `kubectl explain <resource>` for fields. `kubectl api-resources` for all resource types.	training/library/topics/k8s-ops/primer.md
k8s-probes/b8c9d0e1f2a3	k8s-ops	medium	k8s, probes, startup, jvm	Before startup probes existed, how did operators handle slow-starting containers, and why was that approach fragile?	They set a large initialDelaySeconds on the liveness probe. This was fragile because if the app started faster, detection of a truly dead container was delayed; if it started slower, the liveness probe would kill it during boot.\n\nRemember: `kubectl explain <resource>` for fields. `kubectl api-resources` for all resource types.	training/library/topics/k8s-ops/primer.md
k8s-probes/c9d0e1f2a3b4	k8s-ops	hard	k8s, probes, design, architecture	How should liveness and readiness probe endpoints differ in what they check, and why?	Liveness should only verify the process is alive and responsive (no dependency checks) — it answers "should I restart?" Readiness should check dependencies, cache warmth, and load — it answers "can I serve traffic?" Using the same endpoint for both causes dependency failures to trigger unnecessary restarts.\n\nRemember: `kubectl explain <resource>` for fields. `kubectl api-resources` for all resource types.	training/library/topics/k8s-ops/primer.md
k8s-probes/d0e1f2a3b4c5	k8s-ops	hard	k8s, probes, debugging	A pod shows RESTARTS=5 and readiness probe failures in kubectl describe. Walk through how you would debug this.	1) kubectl describe pod to read probe failure events and identify which probe is failing. 2) Check lastState for exit codes (e.g., 137=OOMKilled). 3) kubectl exec into the pod and curl the probe endpoint manually to see the actual response. 4) kubectl logs --previous to read logs from the crashed container. 5) Verify probe port matches the container port and the endpoint returns the expected status code.\n\nRemember: Debug flow: Get→Describe→Logs→Exec. Mnemonic: "GDLE."	training/library/topics/k8s-ops/primer.md
k8s-probes/e1f2a3b4c5d6	k8s-ops	hard	k8s, probes, timeout, gc	Why can JVM garbage collection cause liveness probe failures, and how do you mitigate it?	Full GC pauses can stop the JVM for several seconds. If timeoutSeconds is set too low (e.g., 1s) and a GC pause exceeds that, the probe times out and counts as a failure. Mitigation: increase timeoutSeconds to exceed worst-case GC pause duration, use a startup probe for slow JVM boot, and tune GC to reduce pause times.\n\nRemember: `kubectl explain <resource>` for fields. `kubectl api-resources` for all resource types.	training/library/topics/k8s-ops/primer.md
k8s-probes/f2a3b4c5d6e7	k8s-ops	hard	k8s, probes, successThreshold, parameters	What is the constraint on successThreshold for liveness and startup probes, and why does it matter for readiness?	For liveness and startup probes, successThreshold must be 1 (Kubernetes ignores other values). Only readiness probes can require multiple consecutive successes before the pod is added back to endpoints. This matters because you may want a readiness probe to confirm stability (e.g., successThreshold=2) before resuming traffic after a transient failure.\n\nRemember: `kubectl explain <resource>` for fields. `kubectl api-resources` for all resource types.	training/library/topics/k8s-ops/primer.md

<!-- wiki:related:start -->
---

## Wiki Navigation

### Related Content

- [Adversarial Interview Gauntlet (30 sequences)](../../../../library/interview-scenarios/gauntlet/README.md) (Scenario, L2) — Kubernetes Core
- [Case Study: Alert Storm — Flapping Health Checks](../../../../library/case-studies/cross-domain/alert-storm-flapping-healthchecks/README.md) (Case Study, L2) — Kubernetes Core
- [Case Study: Canary Deploy Routing to Wrong Backend — Ingress Misconfigured](../../../../library/case-studies/cross-domain/canary-deploy-wrong-backend-ingress/README.md) (Case Study, L2) — Kubernetes Core
- [Case Study: CrashLoopBackOff No Logs](../../../../library/case-studies/kubernetes_ops/crashloopbackoff-no-logs/README.md) (Case Study, L1) — Kubernetes Core
- [Case Study: DNS Looks Broken — TLS Expired, Fix Is Cert-Manager](../../../../library/case-studies/cross-domain/dns-tls-certmanager/README.md) (Case Study, L2) — Kubernetes Core
- [Case Study: DaemonSet Blocks Eviction](../../../../library/case-studies/kubernetes_ops/daemonset-blocks-eviction/README.md) (Case Study, L2) — Kubernetes Core
- [Case Study: Deployment Stuck — ImagePull Auth Failure, Vault Secret Rotation](../../../../library/case-studies/cross-domain/deployment-stuck-imagepull-vault/README.md) (Case Study, L2) — Kubernetes Core
- [Case Study: Drain Blocked by PDB](../../../../library/case-studies/kubernetes_ops/drain-blocked-by-pdb/README.md) (Case Study, L2) — Kubernetes Core
- [Case Study: HPA Flapping — Metrics Server Clock Skew, Fix Is NTP](../../../../library/case-studies/cross-domain/hpa-flapping-clock-skew-ntp/README.md) (Case Study, L2) — Kubernetes Core
- [Case Study: ImagePullBackOff Registry Auth](../../../../library/case-studies/kubernetes_ops/imagepullbackoff-registry-auth/README.md) (Case Study, L1) — Kubernetes Core

<!-- wiki:related:end -->
