---
tags:
- devops
- l1
- flashcard-deck
- grokdevops-training
---
<!-- wiki:breadcrumb:start -->
[Portal](../../../../library/portal/index.md) | **Level:** [L1: Foundations](../../../../library/portal/levels.md) | **Topics:** [Career Engineering](../../../../library/portal/topics.md) | **Domain:** DevOps & Tooling
<!-- wiki:breadcrumb:end -->

id	category	difficulty	tags	question	answer	source_path
grokdevops-training/001-crashloop-triage	grokdevops-training	medium	crashloopbackoff,debugging	A pod is in CrashLoopBackOff. What are your first 3 commands?	1. `kubectl get pods -n grokdevops` -- check status and restart count\n2. `kubectl logs -n grokdevops deploy/grokdevops --previous` -- get crash logs\n3. `kubectl describe pod -n grokdevops -l app.kubernetes.io/name=grokdevops` -- check events and exit code	training/library/runbooks/crashloopbackoff.md
grokdevops-training/002-exit-code-137	grokdevops-training	medium	oomkilled,exit-codes	What does exit code 137 mean in a Kubernetes pod?	Exit code 137 = 128 + 9 (SIGKILL). The container was killed by the Linux kernel's OOM killer because it exceeded its cgroup memory limit. Check: `kubectl describe pod | grep OOMKilled`	training/library/runbooks/oomkilled.md
grokdevops-training/003-readiness-vs-liveness	grokdevops-training	medium	probes,readiness,liveness	What is the difference between readiness and liveness probes in terms of what K8s does when they fail?	Readiness probe failure: pod is removed from Service endpoints (no traffic routed to it), but pod keeps running. Liveness probe failure: kubelet restarts the container. Key insight: readiness gates traffic, liveness gates lifecycle.	training/library/runbooks/readiness_probe_failed.md
grokdevops-training/004-hpa-unknown	grokdevops-training	hard	hpa,metrics,scaling	HPA shows '<unknown>/50%' for CPU. What are the two most likely causes?	1. metrics-server is not installed (no CPU metrics available). Check: `kubectl get --raw /apis/metrics.k8s.io/v1beta1/nodes`\n2. Deployment has no CPU resource requests defined. HPA calculates percentage as actual/request * 100 -- no request = undefined.	training/library/runbooks/hpa_not_scaling.md
grokdevops-training/005-imagepull-no-logs	grokdevops-training	easy	imagepullbackoff,debugging	A pod is in ImagePullBackOff. Can you check its logs? Why or why not?	No. The container never started, so there are no container logs. Use `kubectl describe pod` to see the pull error in Events. Check the image name: `kubectl get deploy -o jsonpath='{.spec.template.spec.containers[0].image}'`	training/library/runbooks/imagepullbackoff.md
grokdevops-training/006-rollout-stuck	grokdevops-training	medium	deployment,rollout,probes	A deployment shows 'Progressing' for 20 minutes. What's happening to old pods?	Old ReplicaSet pods continue serving traffic. During a RollingUpdate, K8s won't terminate old pods until new ones pass readiness probes. If new pods never become ready, old pods stay running indefinitely (up to progressDeadlineSeconds, default 600s).	training/library/interview-scenarios/01-deployment-stuck-progressing.md
grokdevops-training/007-service-no-endpoints	grokdevops-training	medium	service,endpoints,selector	A Service shows 0 endpoints. What do you check?	1. `kubectl get endpoints <svc> -n grokdevops` -- confirm empty\n2. Compare service selector to pod labels: `kubectl get svc <svc> -o yaml | grep selector` vs `kubectl get pods --show-labels`\n3. Check if pods are Ready (unready pods are excluded from endpoints)	training/library/runbooks/ingress_404.md
grokdevops-training/008-helm-rollback	grokdevops-training	medium	helm,rollback,revision	How does `helm rollback` work under the hood?	Helm reads the manifests from a previous revision (stored as a Secret in the namespace), and re-applies them via a 3-way merge. Rollback creates a NEW revision -- it doesn't delete the failed one. The revision history is append-only.	training/library/runbooks/helm_upgrade_failed.md
grokdevops-training/009-helm-pending-upgrade	grokdevops-training	hard	helm,failed,stuck	A Helm release is stuck in 'pending-upgrade' state. How do you recover?	Run: `helm rollback <release> <last-good-revision> -n <namespace>`. The pending-upgrade state means the previous upgrade never completed. Rolling back to a known-good revision resets the state.\n\nGotcha: If rollback also fails, you may need to manually delete the broken Helm secret: `kubectl delete secret sh.helm.release.v1.<name>.v<N>`.\n\nRemember: Helm stores release state in Kubernetes Secrets. Stuck states mean a corrupted release secret.	training/library/runbooks/helm_upgrade_failed.md
grokdevops-training/010-servicemonitor-labels	grokdevops-training	hard	prometheus,servicemonitor,labels	Prometheus shows no data for your app. ServiceMonitor exists. What's the most likely issue?	Label selector mismatch. Two things must match: (1) Prometheus's serviceMonitorSelector must find the ServiceMonitor, (2) ServiceMonitor's spec.selector.matchLabels must match the Service's labels. Check both: `kubectl get servicemonitor -o yaml` and `kubectl get svc --show-labels`	training/library/runbooks/prometheus_target_down.md
grokdevops-training/011-promtail-missing	grokdevops-training	medium	loki,promtail,daemonset	Logs stopped appearing in Grafana Loki. What's the first thing to check?	Check if Promtail pods are running: `kubectl get pods -n monitoring -l app.kubernetes.io/name=promtail`. Promtail is a DaemonSet that collects logs from nodes. If pods are missing, check the DaemonSet for nodeSelector/toleration issues.	training/library/runbooks/loki_no_logs.md
grokdevops-training/012-networkpolicy-dns	grokdevops-training	hard	networkpolicy,dns,egress	You applied a NetworkPolicy and now DNS doesn't work. Why?	Adding a NetworkPolicy enables default-deny for the specified direction (ingress/egress). If you created an egress policy without allowing port 53 UDP to kube-system, DNS queries are blocked. Fix: add an egress rule allowing UDP port 53 to any namespace.	training/library/runbooks/networkpolicy_block.md
grokdevops-training/013-networkpolicy-default	grokdevops-training	medium	networkpolicy,default-deny	In a namespace with no NetworkPolicies, what traffic is allowed?	All traffic (ingress and egress) is allowed by default. NetworkPolicies are additive: adding the first policy for a direction implicitly denies everything not explicitly allowed by any policy.\n\nRemember: No NetworkPolicy = allow all. First policy = default deny for that direction. Most counter-intuitive K8s networking concept.\n\nGotcha: NetworkPolicies require a CNI that supports them (Calico, Cilium). Flannel does NOT enforce NetworkPolicies.	training/library/runbooks/networkpolicy_block.md
grokdevops-training/014-rbac-deny-default	grokdevops-training	medium	rbac,authorization	What is the default RBAC behavior in Kubernetes?	Deny-by-default. Without a Role/ClusterRole + RoleBinding/ClusterRoleBinding granting access, all API requests are denied. Use `kubectl auth can-i <verb> <resource> --as=<sa>` to test permissions.\n\nRemember: RBAC = deny by default. NetworkPolicy = allow by default. Opposite defaults — a common source of confusion.\n\nGotcha: system:anonymous and system:unauthenticated groups have some discovery permissions by default.	training/library/runbooks/rbac_forbidden.md
grokdevops-training/015-dns-fqdn	grokdevops-training	medium	dns,service,fqdn	What is the full DNS name (FQDN) for a service called 'grokdevops' in namespace 'grokdevops'?	grokdevops.grokdevops.svc.cluster.local. Format: <service>.<namespace>.svc.cluster.local. Short names work within the same namespace due to search domains in /etc/resolv.conf.\n\nRemember: DNS format: <svc>.<ns>.svc.cluster.local. Within the same namespace, just the service name works.\n\nGotcha: Cross-namespace calls need at least <svc>.<ns>. The full FQDN with trailing dot bypasses search domain expansion.	training/library/runbooks/dns_resolution.md
grokdevops-training/016-trivy-base-image	grokdevops-training	medium	trivy,security,docker	Most CRITICAL CVEs in a Trivy scan come from what layer?	The OS base image layer (apt/apk packages). The most impactful fix is usually updating the base image (e.g., python:3.12-slim-bookworm instead of python:3.9-slim-buster), not patching individual packages.\n\nRemember: Fix base image first, then application dependencies. 80% of CVEs come from the OS layer.\n\nGotcha: Distroless and Alpine images have far fewer CVEs than Debian/Ubuntu base images. Consider switching for production.	training/interactive/runtime-labs/lab-runtime-06-trivy-fail-to-green/
grokdevops-training/017-gitops-drift	grokdevops-training	medium	gitops,drift,argocd	What causes 'configuration drift' in a GitOps-managed cluster?	Manual changes via kubectl (scale, edit, set env) that bypass the Git-managed desired state. In a GitOps setup (ArgoCD/Flux), these changes will be reverted on the next reconciliation loop. Fix: always go through Git, never kubectl directly in production.	training/interactive/runtime-labs/lab-runtime-07-gitops-sync-and-drift/
grokdevops-training/018-oomkilled-vs-eviction	grokdevops-training	hard	oomkilled,eviction,resources	What's the difference between OOMKilled and pod eviction?	OOMKilled: container exceeded its cgroup memory limit (set by resources.limits.memory). Kernel kills the process (exit 137). Eviction: kubelet removes pods when NODE memory is under pressure. OOMKilled is per-container; eviction is per-node.	training/library/runbooks/oomkilled.md
grokdevops-training/019-helm-atomic	grokdevops-training	medium	helm,atomic,upgrade	What does the --atomic flag do in `helm upgrade`?	If the upgrade fails (pods don't become ready within --timeout), Helm automatically rolls back to the previous revision. Without --atomic, a failed upgrade leaves the release in 'failed' state and you must manually rollback.	training/library/runbooks/helm_upgrade_failed.md
grokdevops-training/020-kubectl-previous	grokdevops-training	easy	kubectl,logs,debugging	When should you use `kubectl logs --previous`?	When a container has crashed and restarted. Current logs may be empty (new container just started). --previous shows logs from the LAST terminated container instance. Only works if there was a previous instance.	training/library/runbooks/crashloopbackoff.md
grokdevops-training/021-probe-timing	grokdevops-training	medium	probes,startup,timing	A slow-starting application keeps getting killed by the liveness probe. What setting do you adjust?	Increase initialDelaySeconds on the liveness probe, or better: use a startupProbe. startupProbe disables liveness and readiness probes until the app signals it has started. This is preferred for slow-starting apps (e.g., Java).	training/library/runbooks/readiness_probe_failed.md
grokdevops-training/022-ingress-pathtype	grokdevops-training	medium	ingress,path,routing	What's the difference between pathType: Prefix and pathType: Exact in an Ingress rule?	Prefix: matches the URL path prefix (e.g., /api matches /api, /api/v1, /api/users). Exact: matches only the exact path (e.g., /api matches only /api, not /api/v1). Most apps should use Prefix.	training/library/runbooks/ingress_404.md
grokdevops-training/023-helm-secrets	grokdevops-training	hard	helm,revision,secrets	Where does Helm store release information?	As Secrets in the release's namespace (default driver). Each revision is a separate Secret named sh.helm.release.v1.<name>.v<revision>. This is why `helm list` works without a local state file -- it reads from the cluster.	training/library/runbooks/helm_upgrade_failed.md
grokdevops-training/024-metrics-server	grokdevops-training	medium	metrics-server,hpa,kubectl-top	`kubectl top pods` returns 'error: Metrics API not available'. What's missing?	metrics-server is not installed. Install: `kubectl apply -f https://github.com/kubernetes-sigs/metrics-server/releases/latest/download/components.yaml`. For k3s, also add `--kubelet-insecure-tls` arg.\n\nGotcha: metrics-server requires TLS access to kubelets. In k3s/minikube, add --kubelet-insecure-tls to bypass self-signed cert issues.\n\nRemember: metrics-server provides the Metrics API for `kubectl top` and HPA. Without it, neither works.	training/library/runbooks/hpa_not_scaling.md
grokdevops-training/025-ndots	grokdevops-training	hard	dns,ndots,resolv.conf	What is the ndots setting in /etc/resolv.conf and why does it matter?	Default ndots:5 in K8s means any name with fewer than 5 dots is first tried with search domain suffixes before absolute lookup. This means 'google.com' (1 dot) gets tried as google.com.grokdevops.svc.cluster.local first. Can cause slow DNS if external names are used frequently.	training/library/runbooks/dns_resolution.md
grokdevops-training/026-break-fix-cycle	grokdevops-training	easy	methodology,debugging	What is the correct order for a break/fix debugging cycle?	1. Observe symptoms (kubectl get pods, logs, events)\n2. Form hypothesis (ranked by likelihood)\n3. Test hypothesis (targeted command)\n4. Fix (minimal change)\n5. Verify (confirm symptom is gone)\n6. Teardown (clean up any debug artifacts)	training/START_HERE.md
grokdevops-training/027-chaos-safety	grokdevops-training	easy	chaos,safety	What two flags should chaos scripts support for safety?	--dry-run (preview what would change without applying) and --yes (explicit confirmation required before destructive action). This prevents accidental chaos in production.\n\nRemember: --dry-run = preview, --yes = confirm. Both flags together mean show me what would happen, then do it.\n\nAnalogy: Like a pilot pre-flight checklist — never skip the safety checks before introducing controlled failure.	training/interactive/chaos/README.md
grokdevops-training/028-forensics-bundle	grokdevops-training	medium	incident,forensics,evidence	What should an incident forensics bundle contain?	Pod status, events, logs (current + previous), describe output, resource usage (top), HPA status, service endpoints, recent events. Capture BEFORE fixing so evidence isn't lost. Use: `make incident-forensics`\n\nRemember: Capture BEFORE fixing — evidence disappears after restart. The forensics bundle is your incident black box recorder.\n\nGotcha: `kubectl logs --previous` only works if the container has restarted. Capture current logs too.	training/interactive/incidents/forensics.sh
grokdevops-training/029-cgroup-memory	grokdevops-training	hard	cgroups,memory,oomkill	How does Kubernetes enforce memory limits at the Linux level?	K8s sets cgroup memory.limit_in_bytes for each container. When the process's RSS exceeds this limit, the kernel's OOM killer sends SIGKILL (signal 9). Exit code = 128+9 = 137. This is a kernel-level enforcement, not a K8s decision.	training/library/runbooks/oomkilled.md
grokdevops-training/030-request-vs-limit	grokdevops-training	medium	resources,request,limit	What's the difference between resource requests and limits?	Request: guaranteed minimum. Used for scheduling (K8s places pod on node with enough capacity). Limit: maximum allowed. Enforced by cgroups (CPU throttled, memory OOMKilled). A pod can use more CPU than requested (up to limit) but will be killed if it exceeds memory limit.	training/library/runbooks/oomkilled.md
grokdevops-training/031-coreDNS-check	grokdevops-training	medium	dns,coredns,debugging	How do you verify CoreDNS is working?	1. Check pods: `kubectl get pods -n kube-system -l k8s-app=kube-dns`\n2. Test resolution: `kubectl run dns-test --rm -it --restart=Never --image=busybox:1.36 -- nslookup kubernetes.default.svc.cluster.local`\n3. Check logs: `kubectl logs -n kube-system -l k8s-app=kube-dns`	training/library/runbooks/dns_resolution.md
grokdevops-training/032-hpa-cooldown	grokdevops-training	hard	hpa,scaling,cooldown	How long does HPA wait before scaling down after load decreases?	Default stabilization window for scale-down is 5 minutes (300 seconds). This prevents flapping. Scale-up is faster (15 seconds default). Both are configurable via HPA behavior spec.\n\nRemember: Scale-up = 15s default, scale-down = 5min default. Asymmetric by design to prevent flapping.\n\nGotcha: HPA behavior spec (K8s 1.18+) allows customizing scale-up/down policies, stabilization windows, and rate limits.	training/library/runbooks/hpa_not_scaling.md
grokdevops-training/033-helm-dry-run	grokdevops-training	easy	helm,dry-run,validation	How do you test a Helm upgrade before applying it?	`helm upgrade <release> <chart> -f values.yaml --dry-run` renders templates and validates against the K8s API without creating any resources. Add `--debug` for verbose template output.\n\nGotcha: --dry-run validates against the cluster API but does NOT create resources. Use `helm template` for offline rendering without cluster access.\n\nRemember: --dry-run + --debug = full rendered output with values. Essential for debugging template issues.	training/library/runbooks/helm_upgrade_failed.md
grokdevops-training/034-progressDeadline	grokdevops-training	medium	deployment,rollout,deadline	What is progressDeadlineSeconds and what's the default?	Default: 600 seconds (10 minutes). If a deployment's rollout doesn't make progress for this duration, K8s marks it as Failed. 'Progress' means at least one new pod became Ready. This doesn't auto-rollback -- it just updates the status condition.	training/library/interview-scenarios/01-deployment-stuck-progressing.md
grokdevops-training/035-service-selector	grokdevops-training	easy	service,selector,labels	How does a Kubernetes Service know which pods to route traffic to?	Via label selectors. The Service's spec.selector must match labels on the target pods. Only pods that are Ready (readiness probe passing) are included in the Service's endpoints.\n\nRemember: Service selector to Pod labels to Endpoints. A mismatch at any point = no traffic routing.\n\nDebug clue: `kubectl get endpoints <svc>` shows which pod IPs are in the pool. Empty = selector mismatch or no ready pods.	training/library/runbooks/ingress_404.md
grokdevops-training/036-prom-scrape-interval	grokdevops-training	medium	prometheus,scrape,interval	How often does Prometheus scrape targets by default?	Default scrape interval is 30 seconds. After changing ServiceMonitor labels, wait at least one scrape interval for Prometheus to detect the new config and scrape the target.\n\nGotcha: Changing scrape interval affects PromQL functions like rate(). Use $__rate_interval in Grafana to auto-adjust.\n\nRemember: 15s interval = more granularity but more storage. 60s = less storage but may miss short spikes.	training/library/runbooks/prometheus_target_down.md
grokdevops-training/037-tempo-vs-promtail	grokdevops-training	medium	tempo,promtail,traces,logs	What's the key architectural difference between how Prometheus/Promtail collect data vs Tempo?	Prometheus scrapes (pulls) metrics from /metrics endpoints. Promtail tails (pushes) log files to Loki. Tempo is receive-only: the application must push traces via OTLP protocol. Tempo doesn't collect -- it waits.	training/library/runbooks/tempo_no_traces.md
grokdevops-training/038-rolling-update	grokdevops-training	medium	deployment,rolling-update,strategy	During a RollingUpdate, when does K8s terminate old pods?	Only after new pods pass readiness probes and are added to Service endpoints. maxUnavailable controls how many old pods can be down simultaneously. maxSurge controls how many extra pods can exist. Default: 25% each.	training/library/interview-scenarios/01-deployment-stuck-progressing.md
grokdevops-training/039-k3s-image-import	grokdevops-training	easy	k3s,docker,image	How do you make a locally-built Docker image available in k3s?	`docker save <image>:<tag> | sudo k3s ctr images import -`. k3s uses containerd (not Docker), so Docker images must be explicitly imported. Alternative: push to a registry and pull.\n\nGotcha: k3s uses containerd, not Docker daemon. `docker images` and `k3s crictl images` are separate image stores.\n\nRemember: In production, always use a registry (Docker Hub, GHCR, ECR). Local import is for dev/testing only.	training/library/runbooks/imagepullbackoff.md
grokdevops-training/040-investigation-loop	grokdevops-training	medium	methodology,investigation	What is the recommended investigation loop in this training system?	1. `make incident YES=1` -- inject a random failure\n2. `make investigate` -- see step-by-step investigation plan\n3. Use kubectl/helm to gather evidence\n4. `make hint` if stuck (progressive hints 1-4)\n5. Fix the issue\n6. `make incident-resolve` -- mark resolved and record time	training/interactive/investigation/README.md
grokdevops-training/041-daemonset-node	grokdevops-training	medium	daemonset,nodeSelector,scheduling	A DaemonSet shows DESIRED=0. What's wrong?	No nodes match the DaemonSet's nodeSelector or tolerations. Check: `kubectl get daemonset -o yaml | grep -A5 nodeSelector`. Remove the impossible selector or add matching labels to nodes.\n\nDebug clue: DESIRED=0 means the scheduler found zero qualifying nodes. Check nodeSelector, tolerations, and node labels.\n\nRemember: DaemonSets run one pod per matching node. 0 matching nodes = 0 desired pods.	training/library/runbooks/loki_no_logs.md
grokdevops-training/042-rbac-can-i	grokdevops-training	easy	rbac,authorization,testing	How do you test if a service account has a specific permission?	`kubectl auth can-i <verb> <resource> -n <namespace> --as=system:serviceaccount:<namespace>:<sa-name>`. Returns 'yes' or 'no'. Example: `kubectl auth can-i list pods -n grokdevops --as=system:serviceaccount:grokdevops:default`\n\nRemember: `kubectl auth can-i --list` shows ALL permissions for the current user. Add -n <ns> for namespace-scoped.\n\nGotcha: Service account format is system:serviceaccount:<namespace>:<name>. Missing the prefix = wrong identity.	training/library/runbooks/rbac_forbidden.md
grokdevops-training/043-pod-readiness-gate	grokdevops-training	hard	readiness,endpoints,traffic	A pod is Running but not Ready. Does it receive traffic from the Service?	No. Kubernetes removes non-Ready pods from Service endpoints. The kube-proxy/iptables rules won't route traffic to pods that haven't passed their readiness probe. The pod stays running but is effectively invisible to the Service.	training/library/runbooks/readiness_probe_failed.md
grokdevops-training/044-helm-template	grokdevops-training	easy	helm,template,debug	How do you see what Kubernetes manifests a Helm chart would generate without deploying?	`helm template <release> <chart> -f values.yaml` renders all templates to stdout. No cluster interaction needed. Useful for debugging template issues before upgrade.\n\nRemember: `helm template` = offline rendering (no cluster needed). `helm install --dry-run` = server-side rendering (validates against cluster API).\n\nGotcha: `helm template` does not evaluate lookup functions — they always return empty without a cluster connection.	training/library/runbooks/helm_upgrade_failed.md
grokdevops-training/045-kubectl-jsonpath	grokdevops-training	medium	kubectl,jsonpath,output	How do you extract a specific field from a Kubernetes resource using kubectl?	`kubectl get <resource> -o jsonpath='{.spec.template.spec.containers[0].image}'`. Use python3 -m json.tool to format JSON output. For multiple fields, use custom-columns: `-o custom-columns='NAME:.metadata.name,IMAGE:.spec.template.spec.containers[0].image'`\n\nRemember: JSONPath starts with . for the root. Array access: [0]. Wildcard: [*]. Recursive: ..\n\nGotcha: JSONPath in kubectl uses single quotes around the expression. Escape carefully in shell scripts.	training/kubectl-debugging-cheatsheet.md
grokdevops-training/046-configmap-pod-restart	grokdevops-training	medium	configmap,restart,deployment	You updated a ConfigMap. Why don't pods see the new values?	Pods using ConfigMap via envFrom or env don't auto-restart when the ConfigMap changes. You must restart the deployment: `kubectl rollout restart deployment/<name>`. ConfigMaps mounted as volumes DO auto-update (after kubelet sync period, ~1 minute), but the app must re-read the file.	training/library/runbooks/crashloopbackoff.md
grokdevops-training/047-loki-query	grokdevops-training	easy	loki,grafana,query	What is the Loki query to see logs from the grokdevops namespace?	`{namespace=\grokdevops\"}`. Add filters: `{namespace=\"grokdevops\"} |= \"error\"` for lines containing 'error'. Use `|~` for regex matching."\n\nRemember: LogQL syntax: {label=value} for stream selection, |= for contains, != for exclude, |~ for regex. Pipe operators chain left to right.\n\nGotcha: Label values must be quoted in LogQL. Unquoted values cause parse errors.	training/library/runbooks/loki_no_logs.md
grokdevops-training/048-ingress-controller	grokdevops-training	medium	ingress,controller,traefik	An Ingress resource exists but returns 404. First thing to check after the Ingress spec?	Check if the ingress controller is running: `kubectl get pods -n kube-system -l app.kubernetes.io/name=traefik` (k3s default) or `-l app.kubernetes.io/name=ingress-nginx`. No controller = Ingress resources are ignored.\n\nRemember: Ingress resource = routing rules. Ingress controller = software that implements them. Without a controller, rules are ignored.\n\nGotcha: k3s ships with Traefik by default. If you disabled it at install (--disable traefik), install your own.	training/library/runbooks/ingress_404.md
grokdevops-training/049-events-sorted	grokdevops-training	easy	events,debugging,kubectl	How do you see recent events in a namespace, sorted by time?	`kubectl get events -n grokdevops --sort-by='.lastTimestamp' | tail -20`. Events are ephemeral (default TTL: 1 hour). Check early in your investigation or evidence may be gone.\n\nGotcha: K8s events have a default TTL of 1 hour. If you investigate too late, evidence is gone. Capture events early in triage.\n\nRemember: For persistent event storage, send events to a log aggregator (Loki, Elasticsearch) via an events exporter.	training/kubectl-debugging-cheatsheet.md
grokdevops-training/050-helm-revision-0	grokdevops-training	medium	helm,rollback,revision	In `helm rollback <release> 0`, what does revision 0 mean?	Revision 0 means 'roll back to the previous revision' (one before current). It's a shortcut. To roll back to a specific revision, use: `helm rollback <release> <revision-number>`\n\nGotcha: After rollback, always verify with `helm status` and `kubectl get pods`. Rollback creates a NEW revision — it does not delete the failed one.\n\nRemember: `helm history <release>` shows all revisions including rollbacks.	training/library/runbooks/helm_upgrade_failed.md

<!-- wiki:related:start -->
---

## Wiki Navigation

### Related Content

- [Career Engineering for Ops People](../../../../library/topics/career-engineering/index.md) (Topic Pack, L0) — Career Engineering
- [Corporate IT Fluency for Engineers](../../../../library/topics/corporate-it-fluency/index.md) (Topic Pack, L0) — Career Engineering

<!-- wiki:related:end -->
