GrokDevOps Wiki β€” All Pages

Every page on the wiki, by title. Press Ctrl+F (or Cmd+F) to search titles, descriptions, and topics β€” then click through to the real page.

Looking for a word or phrase *inside* a page instead of a page itself? Try the full-text version (much bigger file, same idea).

3615 pages indexed.

Home (1)

Welcome · Central hub for 300+ DevOps training exercises, labs, runbooks, and case studies covering Linux, Kubernetes, networking, and more.

Library / Portal (689)

/proc Filesystem · Portal | All Topics | Domain: Linux | Tier: 2
Ai devops tools · 1. What are three things you should never paste into an AI tool like ChatGPT or Claude when asking for DevOps help?
AI Tools for DevOps · Portal | All Topics | Domain: DevOps & Tooling | Tier: 4
AI/ML Infrastructure Ops · Portal | All Topics | Domain: DevOps & Tooling | Tier: 4
Alerting · 32 cards β€” 🟒 5 easy | 🟑 9 medium | πŸ”΄ 6 hard
Alerting · 1. How does Alertmanager route and group alerts?
Alerting Rules · Portal | All Topics | Domain: Observability | Tier: 2
All Pages (Ctrl+F Index) · Every page on the wiki, by title/description/topic, for when search doesn't help
Ansible · Portal | Tag Cloud | 35 assets across 4 topics
Ansible · Portal | All Topics | Domain: DevOps & Tooling | Tier: 1
Ansible core · 56 cards β€” 🟒 9 easy | 🟑 28 medium | πŸ”΄ 4 hard
Ansible deep dive · 1. You define app_version in inventory group_vars, playbook group_vars, and host_vars. All three conflict. Which value wins, and how would you override all of them for a single run?
Ansible Deep Dive · Portal | All Topics | Domain: DevOps & Tooling | Tier: 2
Ansible Hub · Ansible is a tier-1 core skill β€” one of the three most important technologies in this training system (Linux > Ansible > Python). This page collects every piece of Ansible content in one place.
Ansible ops · 55 cards β€” 🟒 10 easy | 🟑 18 medium | πŸ”΄ 12 hard
Ansible playbooks · 3. What language are Ansible playbooks written in?
Api gateway · 16 cards β€” 🟒 3 easy | 🟑 4 medium | πŸ”΄ 3 hard
Api gateway · 1. What is the role of an ingress controller in Kubernetes, and how does it differ from a LoadBalancer service per app?
API Gateways & Ingress · Portal | All Topics | Domain: Kubernetes | Tier: 3
Argo · 65 cards β€” 🟒 12 easy | 🟑 26 medium | πŸ”΄ 12 hard
Argo Workflows · Portal | All Topics | Domain: Kubernetes | Tier: 2
ArgoCD & GitOps · Portal | All Topics | Domain: DevOps & Tooling | Tier: 1
Arp · 18 cards β€” 🟒 3 easy | 🟑 4 medium | πŸ”΄ 3 hard
Arp · 1. What does ARP do and when is it used?
ARP · Portal | All Topics | Domain: Networking | Tier: 4
Audit logging · 30 cards β€” 🟒 4 easy | 🟑 9 medium | πŸ”΄ 7 hard
Audit logging · 1. What is auditd and when do you use it?
Audit Logging · Portal | All Topics | Domain: Security | Tier: 3
awk · Portal | All Topics | Domain: CLI Tools | Tier: 1
Aws advanced · 15 cards β€” 🟒 1 easy | 🟑 13 medium | πŸ”΄ 1 hard
AWS CloudWatch · Portal | All Topics | Domain: Cloud | Tier: 2
Aws compute · 105 cards β€” 🟒 23 easy | 🟑 53 medium | πŸ”΄ 23 hard
Aws database · 66 cards β€” 🟒 8 easy | 🟑 31 medium | πŸ”΄ 13 hard
Aws devops · 41 cards β€” 🟒 10 easy | 🟑 10 medium | πŸ”΄ 6 hard
Aws ec2 · 1. Your t3.micro runs at 15% CPU consistently. Initially responsive, after a week it becomes extremely slow. CPU metrics show 100% utilization in spikes. What is happening?
AWS EC2 · Portal | All Topics | Domain: Cloud | Tier: 1
AWS ECS · Portal | All Topics | Domain: Cloud | Tier: 2
Aws general · 70 cards β€” 🟒 19 easy | 🟑 25 medium | πŸ”΄ 11 hard
Aws iam · 1. A bucket policy allows s3:GetObject for Principal: '*'. An IAM policy attached to user Alice explicitly denies s3:GetObject. Can Alice read objects from the bucket?
AWS IAM · Portal | All Topics | Domain: Cloud | Tier: 1
AWS Lambda · Portal | All Topics | Domain: Cloud | Tier: 2
Aws networking · 130 cards β€” 🟒 20 easy | 🟑 77 medium | πŸ”΄ 27 hard
Aws networking · 1. You create a /24 subnet and plan to run 250 instances. At 251, instance launch fails with 'InsufficientFreeAddressesInSubnet'. A /24 has 256 addresses. Where did the other 5 go?
AWS Networking · Portal | All Topics | Domain: Cloud | Tier: 1
AWS Route 53 · Portal | All Topics | Domain: Cloud | Tier: 2
Aws s3 deep dive · 1. You set a bucket policy denying all public access, but an object was uploaded with a public-read ACL. Can the public read that object?
AWS S3 Deep Dive · Portal | All Topics | Domain: Cloud | Tier: 1
Aws security · 63 cards β€” 🟒 18 easy | 🟑 25 medium | πŸ”΄ 14 hard
Aws storage · 85 cards β€” 🟒 14 easy | 🟑 45 medium | πŸ”΄ 20 hard
Aws troubleshooting · 31 cards β€” 🟒 10 easy | 🟑 14 medium | πŸ”΄ 1 hard
Aws troubleshooting · 1. An EC2 instance can't reach the internet. What do you check?
AWS Troubleshooting · Portal | All Topics | Domain: Cloud | Tier: 1
Azure · 108 cards β€” 🟒 23 easy | 🟑 47 medium | πŸ”΄ 23 hard
Azure Blob Storage · Portal | All Topics | Domain: Cloud | Tier: 1
Azure troubleshooting · 31 cards β€” 🟒 7 easy | 🟑 12 medium | πŸ”΄ 6 hard
Azure troubleshooting · 1. An Azure VM can't reach the internet. What do you check?
Azure Troubleshooting · Portal | All Topics | Domain: Cloud | Tier: 1
Azure Virtual Machines · Portal | All Topics | Domain: Cloud | Tier: 1
Azure Virtual Network · Portal | All Topics | Domain: Cloud | Tier: 1
Backstage & Developer Portals · Portal | All Topics | Domain: DevOps & Tooling | Tier: 3
Backup & Restore · Portal | All Topics | Domain: Security | Tier: 2
Backup restore · 27 cards β€” 🟒 4 easy | 🟑 10 medium | πŸ”΄ 6 hard
Backup restore · 1. How do you validate that your backups actually work?
Bash · 86 cards β€” 🟒 23 easy | 🟑 47 medium | πŸ”΄ 1 hard
Bash / Shell Scripting · Portal | All Topics | Domain: Linux | Tier: 1
Bash advanced · 1. How do you safely handle filenames with spaces and special characters?
Bash scripting · 1. What does set -euo pipefail do in a bash script?
BGP EVPN / VXLAN · Portal | All Topics | Domain: Networking | Tier: 2
Binary · 26 cards β€” 🟒 5 easy | 🟑 12 medium | πŸ”΄ 3 hard
Binary · 1. A log shows a file size of 0x400 bytes. How many bytes is that in decimal?
Binary & Number Representation · Portal | All Topics | Domain: Linux | Tier: 3
BMC (Baseboard Management Controller) · Portal | All Topics | Domain: Datacenter & Hardware | Tier: 1
By Level · Browse all learning content by difficulty level, across all domains. Level definitions: L0 (Entry), L1 (Foundations), L2 (Operations), L3 (Advanced).
By Topic · Find content by topic regardless of domain. Search for what you need.
Capacity planning · 26 cards β€” 🟒 5 easy | 🟑 9 medium | πŸ”΄ 6 hard
Capacity planning · 1. What is the difference between utilization and saturation, and why does it matter for capacity planning?
Capacity Planning · Portal | All Topics | Domain: DevOps & Tooling | Tier: 3
Career · 1. What is the 'impact formula' for writing strong resume bullet points as an ops engineer?
Career Engineering · Portal | All Topics | Domain: DevOps & Tooling | Tier: 4
Case Studies · Incident narratives with symptoms, diagnostic questions, solutions, and self-grading.
Ceph Storage · Portal | All Topics | Domain: DevOps & Tooling | Tier: 2
cert-manager · Portal | All Topics | Domain: Kubernetes | Tier: 1
Certificates · Portal | All Topics | Domain: DevOps & Tooling | Tier: 2
cgroups & Linux Namespaces · Portal | All Topics | Domain: Linux | Tier: 2
Change management · 16 cards β€” 🟒 3 easy | 🟑 4 medium | πŸ”΄ 3 hard
Change management · 1. What are the three categories of changes and how do they differ in approval process?
Change Management · Portal | All Topics | Domain: DevOps & Tooling | Tier: 1
Chaos engineering · 26 cards β€” 🟒 5 easy | 🟑 10 medium | πŸ”΄ 5 hard
Chaos engineering · 1. What is the difference between chaos engineering and just breaking things?
Chaos Engineering · Portal | All Topics | Domain: DevOps & Tooling | Tier: 3
Cheatsheets · Quick-reference command and pattern sheets. 34 cheatsheets.
Chef · 110 cards β€” 🟒 31 easy | 🟑 34 medium | πŸ”΄ 30 hard
CI/CD · Portal | Tag Cloud | 20 assets across 6 topics
CI/CD · Portal | All Topics | Domain: DevOps & Tooling | Tier: 1
CI/CD Patterns · Portal | All Topics | Domain: DevOps & Tooling | Tier: 2
CI/CD Pipelines Realities · Portal | All Topics | Domain: DevOps & Tooling | Tier: 2
Cicd · 65 cards β€” 🟒 10 easy | 🟑 27 medium | πŸ”΄ 13 hard
Cicd · 1. What are the typical stages in a CI/CD pipeline?
Cilium · 16 cards β€” 🟒 3 easy | 🟑 4 medium | πŸ”΄ 3 hard
Cilium & eBPF Networking · Portal | All Topics | Domain: Kubernetes | Tier: 2
Circleci · 31 cards β€” 🟒 5 easy | 🟑 8 medium | πŸ”΄ 3 hard
Cisco · 18 cards β€” 🟒 3 easy | 🟑 4 medium | πŸ”΄ 3 hard
Cisco · 1. What do the two parts of 'show interface' status mean: 'is up, line protocol is down' vs 'is up, line protocol is up'?
Cisco CLI · Portal | All Topics | Domain: Networking | Tier: 1
Claude code · 31 cards β€” 🟒 9 easy | 🟑 11 medium | πŸ”΄ 5 hard
Claude code · 1. What is the difference between Claude Code's Read tool and using cat in the Bash tool?
Claude Code · Portal | All Topics | Domain: DevOps & Tooling | Tier: 4
CLI Tools · Portal | Tag Cloud | 26 assets across 17 topics
Cli tools · 41 cards β€” 🟒 12 easy | 🟑 13 medium | πŸ”΄ 7 hard
Cloud · Portal | Tag Cloud | 42 assets across 17 topics
Cloud · 33 cards β€” 🟒 7 easy | 🟑 5 medium | πŸ”΄ 6 hard
Cloud deep dive · 1. What is the difference between a Security Group and a Network ACL (NACL) in AWS?
Cloud Deep Dive · Portal | All Topics | Domain: Cloud | Tier: 1
Compliance · 27 cards β€” 🟒 4 easy | 🟑 10 medium | πŸ”΄ 6 hard
Compliance · 1. What are the five levels of compliance maturity, and what level should a production environment target?
Compliance & Audit · Portal | All Topics | Domain: Security | Tier: 3
CompTIA Security+ (SY0-701) · Portal | All Topics | Domain: Security | Tier: 2
Configuration Management · Portal | All Topics | Domain: DevOps & Tooling | Tier: 1
Consul · 36 cards β€” 🟒 9 easy | 🟑 15 medium | πŸ”΄ 6 hard
Consul · 1. What is HashiCorp Consul and what are its primary use cases?
Container base images · 26 cards β€” 🟒 6 easy | 🟑 9 medium | πŸ”΄ 5 hard
Container base images · 1. What are the main container base image options and their key tradeoffs?
Container Base Images · Portal | All Topics | Domain: Kubernetes | Tier: 1
Container Image Optimization (alias β†’ container_images) · Portal | All Topics | Domain: Kubernetes | Tier: 2
Container Image Scanning (alias β†’ container_images) · Portal | All Topics | Domain: Security | Tier: 1
Container runtime · 36 cards β€” 🟒 7 easy | 🟑 10 medium | πŸ”΄ 8 hard
Container runtime · 1. What is the relationship between containerd, runc, and the shim?
Container Runtimes · Portal | All Topics | Domain: Kubernetes | Tier: 3
Containers Deep Dive · Portal | All Topics | Domain: DevOps & Tooling | Tier: 1
Content Hub · Browse all DevOps learning content by type, domain, or tier.
Continuous Profiling · Portal | All Topics | Domain: Observability | Tier: 2
Corporate it · 1. Your company uses Active Directory for centralized auth. A new developer cannot access the internal wiki. Where do you check first?
Corporate it fluency · 42 cards β€” 🟒 23 easy | 🟑 3 medium | πŸ”΄ 10 hard
Corporate IT Fluency · Portal | All Topics | Domain: DevOps & Tooling | Tier: 4
Crashloop · 1. What does the CrashLoopBackOff status mean in Kubernetes?
Crashloopbackoff · 16 cards β€” 🟒 4 easy | 🟑 4 medium | πŸ”΄ 2 hard
CrashLoopBackOff · Portal | All Topics | Domain: Kubernetes | Tier: 2
CrashLoopBackOff (alias) · Portal | All Topics | Domain: Kubernetes | Tier: 2
Cron · 17 cards β€” 🟒 3 easy | 🟑 5 medium | πŸ”΄ 3 hard
Cron · 1. What are the five fields in a cron schedule expression and what does '/5 * * * ' mean?
Cron & Job Scheduling · Portal | All Topics | Domain: Linux | Tier: 1
Cross-Domain Content · Content that spans multiple technology domains β€” the crown jewels of the wiki
Crossplane · Portal | All Topics | Domain: DevOps & Tooling | Tier: 3
Css · 20 cards β€” 🟒 5 easy | 🟑 7 medium | πŸ”΄ 2 hard
CSS · Portal | All Topics | Domain: DevOps & Tooling | Tier: 4
curl & wget · Portal | All Topics | Domain: CLI Tools | Tier: 1
Dagger / CI as Code · Portal | All Topics | Domain: DevOps & Tooling | Tier: 4
Data Modeling · Portal | All Topics | Domain: DevOps & Tooling | Tier: 2
Database internals · 1. What are the ACID properties of a database transaction?
Database Internals · Portal | All Topics | Domain: DevOps & Tooling | Tier: 3
Database Locking · Portal | All Topics | Domain: Kubernetes | Tier: 2
Database Operations · Portal | All Topics | Domain: Kubernetes | Tier: 2
Database ops · 26 cards β€” 🟒 3 easy | 🟑 6 medium | πŸ”΄ 3 hard
Database Replication · Portal | All Topics | Domain: Kubernetes | Tier: 2
Databases · Portal | Tag Cloud | 26 assets across 17 topics
Databases · 32 cards β€” 🟒 9 easy | 🟑 9 medium | πŸ”΄ 7 hard
Databases · Portal | All Topics | Domain: DevOps & Tooling | Tier: 2
Datacenter · Portal | Tag Cloud | 53 assets across 19 topics
Datacenter · 134 cards β€” 🟒 30 easy | 🟑 55 medium | πŸ”΄ 34 hard
Datacenter RAID & Disk Failures (alias β†’ disk_and_storage_ops) · Portal | All Topics | Domain: Datacenter & Hardware | Tier: 1
Datadog · 33 cards β€” 🟒 5 easy | 🟑 11 medium | πŸ”΄ 2 hard
Datascience · 22 cards β€” 🟒 4 easy | 🟑 7 medium | πŸ”΄ 4 hard
Db internals · 96 cards β€” 🟒 22 easy | 🟑 40 medium | πŸ”΄ 28 hard
Debian & Ubuntu Ecosystem · Portal | All Topics | Domain: Linux | Tier: 1
Debian ubuntu · 27 cards β€” 🟒 6 easy | 🟑 12 medium | πŸ”΄ 2 hard
Debian ubuntu · 1. What is the difference between 'apt update' and 'apt upgrade'?
Debugging methodology · 38 cards β€” 🟒 10 easy | 🟑 15 medium | πŸ”΄ 7 hard
Debugging methodology · 1. What are the five steps of the scientific method applied to debugging?
Debugging Methodology · Portal | All Topics | Domain: DevOps & Tooling | Tier: 1
Deep Dives · Extended reference documents on advanced topics. 23 deep dives.
Deep Thinking · Extended analysis, reflection, and thinking-out-loud notes. The 'why behind the why' for complex topics.
Deployments · Portal | All Topics | Domain: DevOps & Tooling | Tier: 1
Devops · 80 cards β€” 🟒 14 easy | 🟑 39 medium | πŸ”΄ 17 hard
Dhcp · 16 cards β€” 🟒 3 easy | 🟑 4 medium | πŸ”΄ 3 hard
Dhcp · 1. What are the four steps of the DHCP process (DORA) and what transport protocol does it use?
DHCP & IP Address Management · Portal | All Topics | Domain: Networking | Tier: 1
Disaster recovery · 26 cards β€” 🟒 4 easy | 🟑 9 medium | πŸ”΄ 7 hard
Disaster recovery · 1. Explain the 3-2-1 backup rule and give a concrete example of how to implement it for a production database.
Disaster Recovery · Portal | All Topics | Domain: Security | Tier: 1
Disk & Storage Ops · Portal | All Topics | Domain: Linux | Tier: 1
Distributed systems · 1. What is the CAP theorem and how does it apply to choosing a database for a microservices architecture?
Distributed Systems Fundamentals · Portal | All Topics | Domain: DevOps & Tooling | Tier: 1
Dnf · 37 cards β€” 🟒 9 easy | 🟑 15 medium | πŸ”΄ 7 hard
Dnf · 1. On RHEL 8+, what happens when you run the yum command?
DNF Package Manager · Portal | All Topics | Domain: Linux | Tier: 1
Dns · 40 cards β€” 🟒 8 easy | 🟑 10 medium | πŸ”΄ 7 hard
Dns · 1. dig shows the correct IP but the app can't connect. What do you check?
DNS · Portal | All Topics | Domain: Networking | Tier: 1
DNS Deep Dive · Portal | All Topics | Domain: Networking | Tier: 1
Dnssec · 16 cards β€” 🟒 3 easy | 🟑 4 medium | πŸ”΄ 3 hard
DNSSEC & DNS Security · Portal | All Topics | Domain: Networking | Tier: 2
Docker · Portal | Tag Cloud | 30 assets across 6 topics
Docker · 1. Why should you use multi-stage builds?
Docker / Containers · Portal | All Topics | Domain: Kubernetes | Tier: 1
Docker basics · 105 cards β€” 🟒 20 easy | 🟑 47 medium | πŸ”΄ 23 hard
Docker networking · 35 cards β€” 🟒 5 easy | 🟑 10 medium | πŸ”΄ 5 hard
Docker ops · 91 cards β€” 🟒 19 easy | 🟑 38 medium | πŸ”΄ 19 hard
Docker security · 33 cards β€” 🟒 6 easy | 🟑 14 medium | πŸ”΄ 6 hard
Docker storage · 24 cards β€” 🟒 4 easy | 🟑 9 medium | πŸ”΄ 4 hard
DORA Metrics & DevEx · Portal | All Topics | Domain: DevOps & Tooling | Tier: 2
Drills · Pre-configured study sessions. Run these with the session runner:
Drills Index · Muscle-memory CLI exercises. 30 drill sets.
Ebpf · 29 cards β€” 🟒 5 easy | 🟑 8 medium | πŸ”΄ 7 hard
Ebpf · 1. What is eBPF and why is it safer than loading a kernel module for observability?
eBPF · Portal | All Topics | Domain: Linux | Tier: 1
Edge & IoT Infrastructure · Portal | All Topics | Domain: DevOps & Tooling | Tier: 4
Edge iot · 16 cards β€” 🟒 3 easy | 🟑 4 medium | πŸ”΄ 3 hard
Edge iot · 1. What is the A/B partition scheme for OTA updates, and why is it essential for edge devices?
Elasticsearch · 84 cards β€” 🟒 20 easy | 🟑 38 medium | πŸ”΄ 20 hard
Elasticsearch · 1. What is the difference between shards and replicas in Elasticsearch?
Elasticsearch · Portal | All Topics | Domain: DevOps & Tooling | Tier: 4
Email infrastructure · 1. What are SPF, DKIM, and DMARC, and why do you need all three for reliable email delivery?
Email Infrastructure · Portal | All Topics | Domain: Networking | Tier: 1
Environment Variables · Portal | All Topics | Domain: Linux | Tier: 1
Envoy · 36 cards β€” 🟒 10 easy | 🟑 14 medium | πŸ”΄ 6 hard
Envoy · 1. What is Envoy proxy and why is it used as the data plane in most service meshes?
Envoy Proxy · Portal | All Topics | Domain: Kubernetes | Tier: 2
Etcd · 23 cards β€” 🟒 3 easy | 🟑 5 medium | πŸ”΄ 3 hard
Etcd · 1. What does etcd store in a Kubernetes cluster?
etcd · Portal | All Topics | Domain: Kubernetes | Tier: 2
Everything (Ctrl+F Fallback) · One page with the full text of every wiki page, for when search doesn't help
Exercise Map · How the break/fix exercises align with the rest of the training system. Use this to know which exercises to do before or after a lab, runbook, or interview scenario.
Fd · 28 cards β€” 🟒 5 easy | 🟑 10 medium | πŸ”΄ 5 hard
Fd · 1. How does fd differ from find in handling .gitignore files?
fd · Portal | All Topics | Domain: CLI Tools | Tier: 4
Feature flags · 1. What is the difference between a release flag, an experiment flag, and an operational flag?
Feature Flags · Portal | All Topics | Domain: DevOps & Tooling | Tier: 1
Filesystems & Storage · Portal | All Topics | Domain: Linux | Tier: 1
Filesystems storage · 1. What are the key differences between ext4 and XFS?
find · Portal | All Topics | Domain: CLI Tools | Tier: 1
Finops · 19 cards β€” 🟒 3 easy | 🟑 4 medium | πŸ”΄ 3 hard
Finops · 1. In Kubernetes, what is the single biggest cost optimization opportunity, and why?
FinOps · Portal | All Topics | Domain: DevOps & Tooling | Tier: 3
Firewalls · 29 cards β€” 🟒 5 easy | 🟑 9 medium | πŸ”΄ 6 hard
Firewalls · 1. What is the difference between iptables DROP and REJECT?
Firewalls · Portal | All Topics | Domain: Networking | Tier: 1
Firmware · 25 cards β€” 🟒 3 easy | 🟑 4 medium | πŸ”΄ 3 hard
Firmware · 1. Why is firmware management critical in a datacenter?
Firmware / BIOS / UEFI · Portal | All Topics | Domain: Datacenter & Hardware | Tier: 1
Flashcards · Browse all flashcard decks. Each deck contains questions with collapsible answers.
Fleet Operations · Portal | All Topics | Domain: DevOps & Tooling | Tier: 1
Fleet ops · 16 cards β€” 🟒 3 easy | 🟑 4 medium | πŸ”΄ 3 hard
Fleet ops · 1. What is the difference between treating servers as 'pets' vs 'cattle'?
Footguns · Every footgun, gotcha, and common mistake across all topics. Learn from others' pain before it becomes yours.
Forensics · 1. As a first responder to a suspected security incident, what should you do and what should you absolutely NOT do?
Fuzzy Search · Typo-tolerant search across all GrokDevOps training content
Fzf · 29 cards β€” 🟒 4 easy | 🟑 10 medium | πŸ”΄ 6 hard
Fzf · 1. How do you use fzf to interactively select and kill a process?
fzf · Portal | All Topics | Domain: CLI Tools | Tier: 4
Gcp compute · 51 cards β€” 🟒 9 easy | 🟑 18 medium | πŸ”΄ 13 hard
Gcp general · 45 cards β€” 🟒 15 easy | 🟑 14 medium | πŸ”΄ 5 hard
Gcp kubernetes · 48 cards β€” 🟒 13 easy | 🟑 14 medium | πŸ”΄ 10 hard
Gcp networking · 28 cards β€” 🟒 3 easy | 🟑 11 medium | πŸ”΄ 3 hard
Gcp security · 43 cards β€” 🟒 1 easy | 🟑 21 medium | πŸ”΄ 10 hard
Gcp troubleshooting · 36 cards β€” 🟒 7 easy | 🟑 11 medium | πŸ”΄ 7 hard
Gcp troubleshooting · 1. A GCE instance can't reach the internet. What do you check?
GCP Troubleshooting · Portal | All Topics | Domain: Cloud | Tier: 4
Generativeai · 23 cards β€” 🟒 4 easy | 🟑 7 medium | πŸ”΄ 4 hard
Git · Portal | Tag Cloud | 22 assets across 5 topics
Git · 140 cards β€” 🟒 40 easy | 🟑 83 medium | πŸ”΄ 3 hard
Git · 1. What are the three local areas in Git where changes live?
Git · Portal | All Topics | Domain: DevOps & Tooling | Tier: 1
Git advanced · 15 cards β€” 🟒 3 easy | 🟑 10 medium | πŸ”΄ 2 hard
Git Advanced · Portal | All Topics | Domain: DevOps & Tooling | Tier: 2
Git Save Your Ass · Portal | All Topics | Domain: DevOps & Tooling | Tier: 1
Git workflows · 38 cards β€” 🟒 10 easy | 🟑 15 medium | πŸ”΄ 7 hard
Git workflows · 1. Which Git workflow uses environment branches (staging, production) to model the deployment pipeline?
Git Workflows & Branching Strategies · Portal | All Topics | Domain: DevOps & Tooling | Tier: 1
GitHub Actions · Portal | All Topics | Domain: DevOps & Tooling | Tier: 2
Gitops · 48 cards β€” 🟒 8 easy | 🟑 16 medium | πŸ”΄ 9 hard
Gitops · 1. What is GitOps and how does it differ from traditional CI/CD?
GitOps · Portal | All Topics | Domain: DevOps & Tooling | Tier: 2
Grafana · 33 cards β€” 🟒 5 easy | 🟑 9 medium | πŸ”΄ 6 hard
Grafana · 1. How do you use Grafana template variables to make dashboards dynamic?
Grafana · Portal | All Topics | Domain: Observability | Tier: 1
Graphql · 36 cards β€” 🟒 9 easy | 🟑 15 medium | πŸ”΄ 6 hard
Graphql · 1. What is the N+1 query problem in GraphQL and how do you solve it?
GraphQL · Portal | All Topics | Domain: DevOps & Tooling | Tier: 2
grep & Regular Expressions · Portal | All Topics | Domain: CLI Tools | Tier: 1
Grokdevops training · 57 cards β€” 🟒 11 easy | 🟑 29 medium | πŸ”΄ 10 hard
Grpc · 1. What are the four types of gRPC service methods and when would you use each?
gRPC & Protocol Buffers · Portal | All Topics | Domain: Networking | Tier: 3
Hardware Security · Portal | All Topics | Domain: Datacenter & Hardware | Tier: 1
HashiCorp Consul · Portal | All Topics | Domain: Kubernetes | Tier: 2
HashiCorp Vault · Portal | All Topics | Domain: Security | Tier: 2
Helm · Portal | Tag Cloud | 10 assets across 1 topics
Helm · 51 cards β€” 🟒 10 easy | 🟑 15 medium | πŸ”΄ 12 hard
Helm · 1. Helm upgrade failed and the app is broken. How do you recover?
Helm · Portal | All Topics | Domain: DevOps & Tooling | Tier: 1
Homelab · 16 cards β€” 🟒 3 easy | 🟑 4 medium | πŸ”΄ 3 hard
Homelab · 1. What is Proxmox VE and what does it replace in a homelab environment?
Homelab · Portal | All Topics | Domain: DevOps & Tooling | Tier: 4
HPA / Autoscaling · Portal | All Topics | Domain: Kubernetes | Tier: 2
Http protocol · 22 cards β€” 🟒 5 easy | 🟑 7 medium | πŸ”΄ 3 hard
HTTP Protocol · Portal | All Topics | Domain: Networking | Tier: 1
Incident psychology · 16 cards β€” 🟒 3 easy | 🟑 4 medium | πŸ”΄ 3 hard
Incident psychology · 1. What is anchoring bias during an incident, and how do you counter it?
Incident Psychology · Portal | All Topics | Domain: DevOps & Tooling | Tier: 3
Incident response · 27 cards β€” 🟒 4 easy | 🟑 5 medium | πŸ”΄ 3 hard
Incident response · 1. You suspect a server is compromised. What are your first three steps?
Incident Response · Portal | All Topics | Domain: DevOps & Tooling | Tier: 2
Incident triage · 28 cards β€” 🟒 4 easy | 🟑 10 medium | πŸ”΄ 6 hard
Incident Triage · Portal | All Topics | Domain: Security | Tier: 2
Infrastructure Forensics · Portal | All Topics | Domain: Security | Tier: 3
Infrastructure Testing · Portal | All Topics | Domain: DevOps & Tooling | Tier: 2
Inodes · 25 cards β€” 🟒 3 easy | 🟑 5 medium | πŸ”΄ 2 hard
IPMI / ipmitool · Portal | All Topics | Domain: Datacenter & Hardware | Tier: 1
iptables & nftables · Portal | All Topics | Domain: Linux | Tier: 1
Istio · 36 cards β€” 🟒 9 easy | 🟑 15 medium | πŸ”΄ 6 hard
Istio · 1. What are the core components of Istio and what does each one do?
Istio Service Mesh · Portal | All Topics | Domain: Kubernetes | Tier: 2
Jenkins · 25 cards β€” 🟒 2 easy | 🟑 5 medium | πŸ”΄ 3 hard
Jq · 30 cards β€” 🟒 4 easy | 🟑 9 medium | πŸ”΄ 7 hard
Jq · 1. How do you extract all pod names from kubectl output using jq?
jq / JSON Processing · Portal | All Topics | Domain: CLI Tools | Tier: 4
K8s advanced ops · 32 cards β€” 🟒 3 easy | 🟑 13 medium | πŸ”΄ 9 hard
K8s concept chain · 30 cards β€” 🟒 9 easy | 🟑 10 medium | πŸ”΄ 5 hard
K8s config · 44 cards β€” 🟒 7 easy | 🟑 16 medium | πŸ”΄ 6 hard
K8s core · 157 cards β€” 🟒 37 easy | 🟑 74 medium | πŸ”΄ 31 hard
K8s core · 1. A pod is in CrashLoopBackOff. What are your first three commands?
K8s debugging · 1. What are the first three kubectl commands you run when a pod is not working?
K8s Ecosystem · Portal | All Topics | Domain: Kubernetes | Tier: 2
K8s general · 43 cards β€” 🟒 7 easy | 🟑 15 medium | πŸ”΄ 6 hard
K8s HPA (alias β†’ k8s_ops) · Portal | All Topics | Domain: Kubernetes | Tier: 2
K8s networking · 27 cards β€” 🟒 4 easy | 🟑 4 medium | πŸ”΄ 4 hard
K8s node lifecycle · 21 cards β€” 🟒 3 easy | 🟑 4 medium | πŸ”΄ 3 hard
K8s node lifecycle · 1. What is the difference between cordoning and draining a Kubernetes node?
K8s operators · 16 cards β€” 🟒 3 easy | 🟑 4 medium | πŸ”΄ 3 hard
K8s operators · 1. What is a CRD (Custom Resource Definition) and what does it enable?
K8s ops · 110 cards β€” 🟒 21 easy | 🟑 51 medium | πŸ”΄ 23 hard
K8s pods scheduling · 2. kubectl describe pod shows Pending with event 'Insufficient cpu'. Node CPU usage averages 40%. Why won't the scheduler place the pod?
K8s Probes (alias β†’ k8s_ops) · Portal | All Topics | Domain: Kubernetes | Tier: 2
K8s rbac · 25 cards β€” 🟒 4 easy | 🟑 4 medium | πŸ”΄ 4 hard
K8s rbac · 1. What are the four RBAC object types in Kubernetes and what scope does each have?
K8s security · 39 cards β€” 🟒 6 easy | 🟑 14 medium | πŸ”΄ 6 hard
K8s services · 58 cards β€” 🟒 13 easy | 🟑 25 medium | πŸ”΄ 13 hard
K8s services ingress · 1. How does a Kubernetes Service route traffic to pods?
K8s storage · 39 cards β€” 🟒 8 easy | 🟑 12 medium | πŸ”΄ 8 hard
K8s storage · 1. What is the relationship between a PersistentVolume (PV), PersistentVolumeClaim (PVC), and StorageClass in Kubernetes?
K8s troubleshooting · 56 cards β€” 🟒 13 easy | 🟑 25 medium | πŸ”΄ 11 hard
K8s workloads · 84 cards β€” 🟒 19 easy | 🟑 39 medium | πŸ”΄ 20 hard
Kafka · 91 cards β€” 🟒 21 easy | 🟑 42 medium | πŸ”΄ 21 hard
Kafka · 1. What is a Kafka partition and why does partition count matter?
Kafka · Portal | All Topics | Domain: DevOps & Tooling | Tier: 4
Kernel troubleshooting · 1. What command shows kernel error messages with human-readable timestamps, and what patterns should you grep for?
Kernel Troubleshooting · Portal | All Topics | Domain: Linux | Tier: 1
Knowledge Compendiums · 2491 Q&A pairs across Ansible, Linux, and Python β€” the ultimate study resource.
Knowledge Graph · Interactive map of all training assets. Nodes are colored by area, sized by level (L0 small, L3 large). Lines show prerequisite chains and shared topics.
Kubernetes · Portal | Tag Cloud | 138 assets across 31 topics
Kubernetes Concept Chain · Portal | All Topics | Domain: Kubernetes | Tier: 1
Kubernetes Core · Portal | All Topics | Domain: Kubernetes | Tier: 1
Kubernetes Debugging · Portal | All Topics | Domain: Kubernetes | Tier: 2
Kubernetes Networking · Portal | All Topics | Domain: Kubernetes | Tier: 2
Kubernetes Operators · Portal | All Topics | Domain: Kubernetes | Tier: 3
Kubernetes Pods & Scheduling · Portal | All Topics | Domain: Kubernetes | Tier: 1
Kubernetes Services & Ingress · Portal | All Topics | Domain: Kubernetes | Tier: 1
Kubernetes Storage · Portal | All Topics | Domain: Kubernetes | Tier: 2
Kustomize · Portal | All Topics | Domain: Kubernetes | Tier: 2
Lacp · 21 cards β€” 🟒 3 easy | 🟑 4 medium | πŸ”΄ 3 hard
Lacp · 1. What are the two LACP modes and what happens if both sides are set to passive?
LACP / Link Aggregation · Portal | All Topics | Domain: Networking | Tier: 4
Ldap · 16 cards β€” 🟒 3 easy | 🟑 4 medium | πŸ”΄ 3 hard
Ldap · 1. What is a Distinguished Name (DN) in LDAP, and what are its components?
LDAP & Identity Management · Portal | All Topics | Domain: Security | Tier: 3
Learning Paths · The guided-learning surface. Pick a structured journey based on your time and goal β€” all paths are here, from a quick crash course to the full 40-week curriculum.
Least privilege · 1. What is the principle of least privilege and give one concrete Linux example?
Least Privilege · Portal | All Topics | Domain: Security | Tier: 2
Legacy System Archaeology · Portal | All Topics | Domain: DevOps & Tooling | Tier: 4
Legacy systems · 16 cards β€” 🟒 3 easy | 🟑 4 medium | πŸ”΄ 3 hard
Legacy systems · 1. When you inherit a system on your first day, what is the very first survey you should run before changing anything?
Linux · Portal | Tag Cloud | 118 assets across 45 topics
Linux Boot Process · Portal | All Topics | Domain: Linux | Tier: 1
Linux data hoarding · 39 cards β€” 🟒 12 easy | 🟑 14 medium | πŸ”΄ 7 hard
Linux data hoarding · 1. What is the core architectural difference between traditional RAID and the JBOD+mergerfs+SnapRAID approach used in data hoarding?
Linux Data Hoarding · Portal | All Topics | Domain: Linux | Tier: 2
Linux Distribution Comparison · Portal | All Topics | Domain: Linux | Tier: 1
Linux distros · 21 cards β€” 🟒 4 easy | 🟑 8 medium | πŸ”΄ 3 hard
Linux distros · 1. What are the key differences between the Red Hat and Debian Linux families?
Linux filesystem · 141 cards β€” 🟒 35 easy | 🟑 59 medium | πŸ”΄ 32 hard
Linux fundamentals · 122 cards β€” 🟒 50 easy | 🟑 47 medium | πŸ”΄ 10 hard
Linux fundamentals · 2. How do you find which process is using port 8080?
Linux Fundamentals · Portal | All Topics | Domain: Linux | Tier: 1
Linux hardening · 1. What are the first three things you harden on a fresh Linux server?
Linux Hardening · Portal | All Topics | Domain: Linux | Tier: 1
Linux kernel · 69 cards β€” 🟒 18 easy | 🟑 30 medium | πŸ”΄ 15 hard
Linux kernel tuning · 1. A production server is dropping incoming TCP connections under load. dmesg shows 'TCP: request_sock_TCP: Possible SYN flooding on port 443'. What sysctl parameters should you check first and why?
Linux Kernel Tuning · Portal | All Topics | Domain: Linux | Tier: 2
Linux Logging · Portal | All Topics | Domain: Linux | Tier: 1
Linux memory · 36 cards β€” 🟒 9 easy | 🟑 9 medium | πŸ”΄ 11 hard
Linux memory management · 1. What does the OOM killer do and how does it choose which process to kill?
Linux Memory Management · Portal | All Topics | Domain: Linux | Tier: 1
Linux networking · 103 cards β€” 🟒 27 easy | 🟑 44 medium | πŸ”΄ 17 hard
Linux networking · 1. What is the difference between an access port and a trunk port?
Linux Networking Tools · Portal | All Topics | Domain: Networking | Tier: 1
Linux Networking Tools (alias β†’ linux_ops) · Portal | All Topics | Domain: Networking | Tier: 1
Linux Ops Performance Triage · Portal | All Topics | Domain: Linux | Tier: 1
Linux Ops Storage · Portal | All Topics | Domain: Linux | Tier: 1
Linux Ops systemd · Portal | All Topics | Domain: Linux | Tier: 1
Linux performance · 53 cards β€” 🟒 7 easy | 🟑 22 medium | πŸ”΄ 14 hard
Linux performance · 1. What is the USE method for performance analysis?
Linux Performance Tuning · Portal | All Topics | Domain: Linux | Tier: 1
Linux processes · 108 cards β€” 🟒 22 easy | 🟑 62 medium | πŸ”΄ 16 hard
Linux recovery · 1. What is the difference between /dev/random and /dev/urandom?
Linux security · 91 cards β€” 🟒 25 easy | 🟑 35 medium | πŸ”΄ 18 hard
Linux Signals & Process Control · Portal | All Topics | Domain: Linux | Tier: 1
Linux systemd · 42 cards β€” 🟒 8 easy | 🟑 14 medium | πŸ”΄ 5 hard
Linux Text Processing · Portal | All Topics | Domain: Linux | Tier: 1
Linux Users & Permissions · Portal | All Topics | Domain: Linux | Tier: 1
Load balancing · 29 cards β€” 🟒 5 easy | 🟑 9 medium | πŸ”΄ 6 hard
Load balancing · 1. What is the difference between L4 and L7 load balancing?
Load Balancing · Portal | All Topics | Domain: Networking | Tier: 1
Load Testing · Portal | All Topics | Domain: DevOps & Tooling | Tier: 1
Log pipelines · 17 cards β€” 🟒 3 easy | 🟑 4 medium | πŸ”΄ 3 hard
Log pipelines · 1. What is the difference between structured and unstructured logging, and why does it matter for log pipelines?
Log Pipelines · Portal | All Topics | Domain: Observability | Tier: 2
Logging · Portal | All Topics | Domain: Observability | Tier: 2
Loki · 25 cards β€” 🟒 3 easy | 🟑 4 medium | πŸ”΄ 3 hard
Loki · 1. What are LogQL pipeline stages and how do they work?
Loki · Portal | All Topics | Domain: Observability | Tier: 2
LPIC / LFCS Exam · Portal | All Topics | Domain: Linux | Tier: 2
Lpic lfcs · 26 cards β€” 🟒 6 easy | 🟑 12 medium | πŸ”΄ 2 hard
Lpic lfcs · 1. What are setuid, setgid, and sticky bit, and how do you set them?
Make & Build Systems · Portal | All Topics | Domain: DevOps & Tooling | Tier: 2
Mellanox switches · 39 cards β€” 🟒 11 easy | 🟑 15 medium | πŸ”΄ 7 hard
Mellanox switches · 1. What is the Onyx CLI command to save the running configuration to persistent storage?
Mellanox Switches · Portal | All Topics | Domain: Networking | Tier: 2
Mental Models (Core) · Portal | All Topics | Domain: DevOps & Tooling | Tier: 1
Mental models core · 34 cards β€” 🟒 11 easy | 🟑 15 medium | πŸ”΄ 1 hard
Mergerfs · 41 cards β€” 🟒 12 easy | 🟑 16 medium | πŸ”΄ 7 hard
Mergerfs · 1. What are the three branch modes in mergerfs?
mergerfs · Portal | All Topics | Domain: other | Tier: 4
Message queues · 36 cards β€” 🟒 9 easy | 🟑 15 medium | πŸ”΄ 6 hard
Message queues · 1. What is the difference between a message queue (RabbitMQ) and a log-based broker (Kafka), and when do you choose each?
Message Queues · Portal | All Topics | Domain: DevOps & Tooling | Tier: 2
Ml ops · 17 cards β€” 🟒 3 easy | 🟑 4 medium | πŸ”΄ 3 hard
Ml ops · 1. Why is mounting /dev/shm as an emptyDir with medium: Memory critical for PyTorch training jobs in Kubernetes?
Modern cli · 34 cards β€” 🟒 11 easy | 🟑 7 medium | πŸ”΄ 7 hard
Modern cli · 1. What modern CLI tool replaces grep -r for recursive code search, and what is its key speed advantage?
Modern CLI Tools · Portal | All Topics | Domain: CLI Tools | Tier: 4
Modern CLI Workflows · Portal | All Topics | Domain: CLI Tools | Tier: 4
MongoDB Operations · Portal | All Topics | Domain: DevOps & Tooling | Tier: 1
Mongodb ops · 1. What is a MongoDB replica set and what happens when the primary goes down?
Monitoring · 75 cards β€” 🟒 10 easy | 🟑 42 medium | πŸ”΄ 8 hard
Monitoring fundamentals · 1. What are the Four Golden Signals from Google SRE and what does each measure?
Monitoring Fundamentals · Portal | All Topics | Domain: Observability | Tier: 1
Monitoring migration · 1. What are three key differences between legacy monitoring (Nagios/Zabbix) and modern monitoring (Prometheus)?
Monitoring Migration · Portal | All Topics | Domain: Observability | Tier: 1
Mounts & Filesystems (alias) · Portal | All Topics | Domain: Linux | Tier: 1
Mtu · 22 cards β€” 🟒 3 easy | 🟑 4 medium | πŸ”΄ 3 hard
Mtu · 2. What is MTU and what happens when a packet exceeds it?
MTU · Portal | All Topics | Domain: Networking | Tier: 4
Multi tenancy · 16 cards β€” 🟒 3 easy | 🟑 4 medium | πŸ”΄ 3 hard
Multi tenancy · 1. Why must you create a default-deny NetworkPolicy before any allow rules in a multi-tenant cluster?
Multi-Tenancy Patterns · Portal | All Topics | Domain: Kubernetes | Tier: 3
MySQL / MariaDB Operations · Portal | All Topics | Domain: DevOps & Tooling | Tier: 1
Mysql ops · 1. What is the difference between InnoDB and MyISAM, and why is InnoDB the default in modern MySQL?
Nat · 16 cards β€” 🟒 3 easy | 🟑 4 medium | πŸ”΄ 3 hard
Nat · 1. An application behind NAT can reach the internet but inbound connections fail. Why?
NAT · Portal | All Topics | Domain: Networking | Tier: 4
Network Automation · Portal | All Topics | Domain: Networking | Tier: 2
Networking · Portal | Tag Cloud | 108 assets across 31 topics
Networking · 138 cards β€” 🟒 43 easy | 🟑 56 medium | πŸ”΄ 24 hard
Networking Troubleshooting · Portal | All Topics | Domain: Networking | Tier: 1
Networking Troubleshooting Tools · Portal | All Topics | Domain: Networking | Tier: 1
Nginx · 26 cards β€” 🟒 5 easy | 🟑 10 medium | πŸ”΄ 5 hard
Nginx · 1. Why must you always run 'nginx -t' before 'nginx -s reload'?
Nginx & Web Servers · Portal | All Topics | Domain: DevOps & Tooling | Tier: 1
Nix / NixOS · Portal | All Topics | Domain: DevOps & Tooling | Tier: 3
Node Lifecycle & Maintenance · Portal | All Topics | Domain: Kubernetes | Tier: 3
Observability · Portal | Tag Cloud | 53 assets across 16 topics
Observability · 31 cards β€” 🟒 5 easy | 🟑 15 medium | πŸ”΄ 5 hard
Observability Deep Dive · Portal | All Topics | Domain: Observability | Tier: 2
Offensive Security Basics · Portal | All Topics | Domain: Security | Tier: 3
On call · 30 cards β€” 🟒 6 easy | 🟑 8 medium | πŸ”΄ 6 hard
On call · 1. What is the role of an Incident Commander (IC) during a production outage?
On-Call & Incident Command · Portal | All Topics | Domain: DevOps & Tooling | Tier: 2
Oob management · 1. Server is unresponsive to SSH. How do you access it?
Oom · 1. What does OOMKilled mean and what exit code does it produce?
Oomkilled · 16 cards β€” 🟒 3 easy | 🟑 5 medium | πŸ”΄ 2 hard
OOMKilled · Portal | All Topics | Domain: Kubernetes | Tier: 2
OOMKilled (alias) · Portal | All Topics | Domain: Kubernetes | Tier: 2
Open policy agent · 36 cards β€” 🟒 9 easy | 🟑 15 medium | πŸ”΄ 6 hard
Open policy agent · 1. What is Open Policy Agent (OPA) and how does it decouple policy from application code?
Open Policy Agent · Portal | All Topics | Domain: Security | Tier: 2
Openshift · 49 cards β€” 🟒 8 easy | 🟑 18 medium | πŸ”΄ 8 hard
Opentelemetry · 1. What problem does OpenTelemetry solve, and what are the three telemetry signals it unifies?
OpenTelemetry · Portal | All Topics | Domain: Observability | Tier: 3
OpenTofu & Terraform Ecosystem · Portal | All Topics | Domain: DevOps & Tooling | Tier: 3
Ops war stories · 16 cards β€” 🟒 3 easy | 🟑 4 medium | πŸ”΄ 3 hard
Ops war stories · 1. What single question resolves approximately 40% of production incidents within minutes?
Ops War Stories · Portal | All Topics | Domain: DevOps & Tooling | Tier: 4
OpSec Mistakes · Portal | All Topics | Domain: Security | Tier: 3
Out-of-Band Management · Portal | All Topics | Domain: Datacenter & Hardware | Tier: 1
Overview · Training portal β€” navigate the DevOps learning content collection.
Overview · Browse the interactive training content β€” flashcards, quiz questions, incident scenarios, and study drills.
Package management · 21 cards β€” 🟒 4 easy | 🟑 4 medium | πŸ”΄ 4 hard
Package management · 1. What is the difference between dpkg/rpm and apt/dnf?
Package Management · Portal | All Topics | Domain: Linux | Tier: 1
Packer · Portal | All Topics | Domain: DevOps & Tooling | Tier: 3
Packet path · 16 cards β€” 🟒 3 easy | 🟑 4 medium | πŸ”΄ 3 hard
Packet path · 1. A curl request times out. How do you systematically narrow down where the problem is?
Packet Path · Portal | All Topics | Domain: Networking | Tier: 4
Perf · 20 cards β€” 🟒 4 easy | 🟑 7 medium | πŸ”΄ 3 hard
perf Profiling · Portal | All Topics | Domain: Linux | Tier: 2
Performance · Portal | All Topics | Domain: DevOps & Tooling | Tier: 2
Perl · 46 cards β€” 🟒 8 easy | 🟑 15 medium | πŸ”΄ 8 hard
Pipes & Redirection · Portal | All Topics | Domain: Linux | Tier: 1
Platform Engineering · Portal | Tag Cloud | 20 assets across 10 topics
Platform engineering · 17 cards β€” 🟒 3 easy | 🟑 4 medium | πŸ”΄ 3 hard
Platform engineering · 1. What is the fundamental difference between a platform and an ops ticket queue?
Platform Engineering · Portal | All Topics | Domain: DevOps & Tooling | Tier: 3
Policy engines · 21 cards β€” 🟒 3 easy | 🟑 4 medium | πŸ”΄ 3 hard
Policy engines · 1. What problem do policy engines solve that RBAC alone cannot?
Policy Engines · Portal | All Topics | Domain: Kubernetes | Tier: 3
PostgreSQL Operations · Portal | All Topics | Domain: DevOps & Tooling | Tier: 2
Postmortem slo · 20 cards β€” 🟒 3 easy | 🟑 4 medium | πŸ”΄ 3 hard
Postmortem slo · 1. What is an error budget, and what should happen when it is exhausted?
Postmortems & SLOs · Portal | All Topics | Domain: DevOps & Tooling | Tier: 3
Power · 1. What is PUE and what is a good value?
Power & UPS · Portal | All Topics | Domain: Datacenter & Hardware | Tier: 1
PowerShell · Portal | All Topics | Domain: DevOps & Tooling | Tier: 4
Primers · Introduction and orientation for every topic. Start here if you're new to a subject.
Probes (Liveness/Readiness) · Portal | All Topics | Domain: Kubernetes | Tier: 2
Process management · 1. What is the difference between SIGTERM and SIGKILL, and why should you always send SIGTERM first?
Process Management · Portal | All Topics | Domain: Linux | Tier: 1
Progress · Track your learning progress across domains and levels.
Progressive Delivery · Portal | All Topics | Domain: Kubernetes | Tier: 1
Prometheus · 1. What is the difference between a counter and a gauge in Prometheus?
Prometheus · Portal | All Topics | Domain: Observability | Tier: 1
Prometheus deep dive · 1. You add a request_id label (unique per request) to your http_requests_total counter. Within hours, Prometheus memory spikes and queries slow to a crawl. What happened and how do you fix it?
Prometheus Deep Dive · Portal | All Topics | Domain: Observability | Tier: 2
Prometheus stack · 90 cards β€” 🟒 21 easy | 🟑 42 medium | πŸ”΄ 21 hard
Pulumi · Portal | All Topics | Domain: DevOps & Tooling | Tier: 4
Puppet · 70 cards β€” 🟒 17 easy | 🟑 20 medium | πŸ”΄ 18 hard
Pxe · 1. A server won't PXE boot. What do you check?
PXE / Provisioning · Portal | All Topics | Domain: Datacenter & Hardware | Tier: 1
Python · Portal | Tag Cloud | 13 assets across 4 topics
Python · 71 cards β€” 🟒 30 easy | 🟑 17 medium | πŸ”΄ 6 hard
Python algorithms · 37 cards β€” 🟒 6 easy | 🟑 11 medium | πŸ”΄ 3 hard
Python Async & Concurrency · Portal | All Topics | Domain: DevOps & Tooling | Tier: 2
Python automation · 1. Why would you choose Python over Bash for an infrastructure automation script?
Python Automation · Portal | All Topics | Domain: DevOps & Tooling | Tier: 1
Python concepts · 109 cards β€” 🟒 55 easy | 🟑 37 medium | πŸ”΄ 2 hard
Python Debugging · Portal | All Topics | Domain: DevOps & Tooling | Tier: 1
Python oop · 32 cards β€” 🟒 3 easy | 🟑 7 medium | πŸ”΄ 5 hard
Python Packaging · Portal | All Topics | Domain: DevOps & Tooling | Tier: 1
Quiz · 1. What does idempotent mean in Ansible and why does it matter?
Quiz Bank · Browse quiz questions by topic. Each question has a hidden answer.
RabbitMQ & Message Queues · Portal | All Topics | Domain: DevOps & Tooling | Tier: 4
Rack & Stack · Portal | All Topics | Domain: Datacenter & Hardware | Tier: 1
Rack ops · 1. You're racking a new server. What's the cabling checklist?
Raid · 1. RAID 5 is degraded and another disk has a predictive failure. What do you do?
RAID · Portal | All Topics | Domain: Datacenter & Hardware | Tier: 1
Random Discovery · Click any button to get 5 random picks. Click again for a fresh set.
RBAC · Portal | All Topics | Domain: Kubernetes | Tier: 2
Redfish API · Portal | All Topics | Domain: Datacenter & Hardware | Tier: 1
Redis Operations · Portal | All Topics | Domain: DevOps & Tooling | Tier: 3
Regex · 26 cards β€” 🟒 5 easy | 🟑 9 medium | πŸ”΄ 6 hard
Regex · 1. What is the difference between BRE (Basic Regular Expressions) and ERE (Extended Regular Expressions) when using grep?
Regex & Text Wrangling · Portal | All Topics | Domain: CLI Tools | Tier: 2
Reliability Patterns · Portal | All Topics | Domain: DevOps & Tooling | Tier: 2
Replication · 1. What is the difference between synchronous and asynchronous replication?
Repo · 38 cards β€” 🟒 6 easy | 🟑 12 medium | πŸ”΄ 6 hard
Rhce · 56 cards β€” 🟒 19 easy | 🟑 24 medium | πŸ”΄ 7 hard
Rhce · 1. What is the precedence order for Ansible configuration files?
RHCE (EX294) Exam · Portal | All Topics | Domain: DevOps & Tooling | Tier: 1
Ripgrep · 24 cards β€” 🟒 5 easy | 🟑 7 medium | πŸ”΄ 3 hard
Ripgrep · 1. How do you search for a pattern in only Python files, excluding tests?
ripgrep (rg) · Portal | All Topics | Domain: CLI Tools | Tier: 4
Routing · 25 cards β€” 🟒 3 easy | 🟑 4 medium | πŸ”΄ 3 hard
Routing · 1. How does a Linux host decide where to send a packet?
Routing · Portal | All Topics | Domain: Networking | Tier: 1
rsync · Portal | All Topics | Domain: Linux | Tier: 1
Runbook craft · 16 cards β€” 🟒 3 easy | 🟑 4 medium | πŸ”΄ 3 hard
Runbook craft · 1. What are the five sections every effective runbook should have?
Runbook Craft · Portal | All Topics | Domain: DevOps & Tooling | Tier: 1
Runbooks · Incident response procedures and operational playbooks. 56 runbooks.
Runtime Security with Falco · Portal | All Topics | Domain: Security | Tier: 2
S3-Compatible Object Storage · Portal | All Topics | Domain: DevOps & Tooling | Tier: 1
Saltstack · 70 cards β€” 🟒 17 easy | 🟑 22 medium | πŸ”΄ 16 hard
Scenarios · Practice triage, containment, and resolution with these scenario drills.
Search · Flat list of every learning asset. Use Ctrl+F / Cmd+F to find what you need. Topic aliases are included so synonym searches work.
Search · Full-text search across all GrokDevOps wiki content
Secrets management · 40 cards β€” 🟒 8 easy | 🟑 11 medium | πŸ”΄ 8 hard
Secrets management · 1. Why are Kubernetes Secrets not truly secure by default?
Secrets Management · Portal | All Topics | Domain: Security | Tier: 2
Security · Portal | Tag Cloud | 59 assets across 22 topics
Security · 209 cards β€” 🟒 52 easy | 🟑 118 medium | πŸ”΄ 27 hard
Security scanning · 1. What does a container image vulnerability scanner like Trivy check for?
Security Scanning · Portal | All Topics | Domain: Security | Tier: 2
Security+ Hub · Everything for CompTIA Security+ (SY0-701) exam prep lives here β€” one place, all study material.
sed · Portal | All Topics | Domain: CLI Tools | Tier: 1
Selinux · 27 cards β€” 🟒 4 easy | 🟑 10 medium | πŸ”΄ 6 hard
Selinux · 1. What are the three SELinux modes and which one should be used in production?
SELinux & AppArmor · Portal | All Topics | Domain: Security | Tier: 1
Server hardware · 25 cards β€” 🟒 3 easy | 🟑 4 medium | πŸ”΄ 3 hard
Server hardware · 1. What do amber and blue LED indicators mean on most server hardware?
Server Hardware · Portal | All Topics | Domain: Datacenter & Hardware | Tier: 1
Serverless Computing · Portal | All Topics | Domain: Cloud | Tier: 2
Service mesh · 20 cards β€” 🟒 3 easy | 🟑 4 medium | πŸ”΄ 3 hard
Service mesh · 1. What are the two planes of a service mesh, and what does each do?
Service Mesh · Portal | All Topics | Domain: Kubernetes | Tier: 3
Shuffled Trivia Compendium · 2491 questions from Ansible, Linux, and Python β€” shuffled together for cross-topic study. Each question is tagged with its source topic.
SLO Tooling · Portal | All Topics | Domain: Observability | Tier: 2
Software development · 41 cards β€” 🟒 6 easy | 🟑 14 medium | πŸ”΄ 6 hard
Sql · 22 cards β€” 🟒 6 easy | 🟑 4 medium | πŸ”΄ 5 hard
Sql · 1. A query that was fast last week is now slow. EXPLAIN shows a sequential scan instead of an index scan. What are the most common causes?
SQL · Portal | All Topics | Domain: DevOps & Tooling | Tier: 2
SQLite Operations & Internals · Portal | All Topics | Domain: DevOps & Tooling | Tier: 4
Sre · 1. What is toil in the SRE context, and what is the target threshold?
SRE & Incidents · Portal | Tag Cloud | 38 assets across 18 topics
SRE Practices · Portal | All Topics | Domain: DevOps & Tooling | Tier: 2
Ssh deep dive · 1. Why should you disable SSH password authentication and what do you use instead?
SSH Deep Dive · Portal | All Topics | Domain: Linux | Tier: 1
Ssh hygiene · 16 cards β€” 🟒 3 easy | 🟑 4 medium | πŸ”΄ 3 hard
Storage · 43 cards β€” 🟒 8 easy | 🟑 17 medium | πŸ”΄ 8 hard
Storage · 1. Why should you use UUIDs instead of /dev/sdX names in /etc/fstab?
Storage (SAN/NAS/DAS) · Portal | All Topics | Domain: Datacenter & Hardware | Tier: 1
Stp · 17 cards β€” 🟒 3 easy | 🟑 4 medium | πŸ”΄ 3 hard
Stp · 1. Why is Spanning Tree Protocol necessary on a network with redundant Layer 2 links?
STP / Spanning Tree · Portal | All Topics | Domain: Networking | Tier: 4
Strace · 20 cards β€” 🟒 5 easy | 🟑 7 medium | πŸ”΄ 2 hard
strace · Portal | All Topics | Domain: Linux | Tier: 2
Street Ops · Practical operations guides β€” the stuff you actually do in production. Commands, workflows, and real-world patterns.
Subnetting & IP Addressing · Portal | All Topics | Domain: Networking | Tier: 1
Supply Chain Security · Portal | All Topics | Domain: Security | Tier: 2
Synthetic Monitoring · Portal | All Topics | Domain: Observability | Tier: 1
systemctl & journalctl Deep Dive · Portal | All Topics | Domain: Linux | Tier: 1
Systemd · 1. How do you check why a service failed to start?
systemd · Portal | All Topics | Domain: Linux | Tier: 1
Systems thinking · 17 cards β€” 🟒 3 easy | 🟑 4 medium | πŸ”΄ 3 hard
Systems thinking · 1. What is a positive (reinforcing) feedback loop in infrastructure? Give an example.
Systems Thinking · Portal | All Topics | Domain: DevOps & Tooling | Tier: 3
Tags · Browse all wiki content by tag
Tailscale & Zero Trust Networking · Portal | All Topics | Domain: Networking | Tier: 3
tar & Compression · Portal | All Topics | Domain: Linux | Tier: 1
Tcp ip · 1. How do you troubleshoot a connection timeout vs connection refused?
TCP/IP · Portal | All Topics | Domain: Networking | Tier: 1
TCP/IP Deep Dive · Portal | All Topics | Domain: Networking | Tier: 2
Tcpdump · 21 cards β€” 🟒 5 easy | 🟑 7 medium | πŸ”΄ 3 hard
Tempo · 20 cards β€” 🟒 3 easy | 🟑 4 medium | πŸ”΄ 3 hard
Tempo · 1. What is Tempo's role in an observability stack, and how does it receive trace data?
Tempo · Portal | All Topics | Domain: Observability | Tier: 3
Terminal · 21 cards β€” 🟒 5 easy | 🟑 8 medium | πŸ”΄ 2 hard
Terminal Internals · Portal | All Topics | Domain: Linux | Tier: 3
Terraform · Portal | Tag Cloud | 22 assets across 7 topics
Terraform · 1. What is Terraform state and why is it critical?
Terraform · Portal | All Topics | Domain: DevOps & Tooling | Tier: 2
Terraform basics · 142 cards β€” 🟒 27 easy | 🟑 72 medium | πŸ”΄ 28 hard
Terraform deep dive · 1. You have 3 EC2 instances created with count. You need to rename index 1 without destroying and recreating the other two. What happens if you switch from count to for_each, and how do you migrate safely?
Terraform Deep Dive · Portal | All Topics | Domain: DevOps & Tooling | Tier: 2
Terraform modules · 38 cards β€” 🟒 6 easy | 🟑 16 medium | πŸ”΄ 6 hard
Terraform providers · 28 cards β€” 🟒 5 easy | 🟑 9 medium | πŸ”΄ 4 hard
Terraform state · 79 cards β€” 🟒 14 easy | 🟑 40 medium | πŸ”΄ 15 hard
Terraform workflow · 47 cards β€” 🟒 8 easy | 🟑 21 medium | πŸ”΄ 8 hard
Tls · 43 cards β€” 🟒 8 easy | 🟑 11 medium | πŸ”΄ 9 hard
TLS & Certificates Ops · Portal | All Topics | Domain: Security | Tier: 1
TLS & PKI · Portal | All Topics | Domain: Security | Tier: 2
Tls pki · 1. What is a certificate chain and why does order matter?
tmux & screen · Portal | All Topics | Domain: Linux | Tier: 1
Toil Reduction · Portal | All Topics | Domain: DevOps & Tooling | Tier: 2
Topic Coverage Grid · Coverage: 133/134 primer, 133/134 street ops, 133/134 footguns, 26/134 deep dive, 134/134 cards, 134/134 quiz, 36/134 cases, 11/134 scenarios, 26/134 runbooks, 8/134 labs
Tracing · 29 cards β€” 🟒 4 easy | 🟑 9 medium | πŸ”΄ 7 hard
Tracing · 1. When should you use metrics, logs, or traces for debugging?
Tracing · Portal | All Topics | Domain: Observability | Tier: 3
Training Content Graph β€” grokdevops
Trivia · Surprising, historical, and little-known facts across all topics. 209 topics, 2197 facts.
Vendor management · 16 cards β€” 🟒 3 easy | 🟑 4 medium | πŸ”΄ 3 hard
Vendor management · 1. What are the four tiers of vendor support and what can each tier typically do?
Vendor Management & Escalation · Portal | All Topics | Domain: DevOps & Tooling | Tier: 1
Virtualization · 17 cards β€” 🟒 3 easy | 🟑 4 medium | πŸ”΄ 3 hard
Virtualization · 1. What are the three components of the standard Linux virtualization stack, and what does each do?
Virtualization · Portal | All Topics | Domain: Datacenter & Hardware | Tier: 4
Vlans · 25 cards β€” 🟒 3 easy | 🟑 4 medium | πŸ”΄ 3 hard
VLANs · Portal | All Topics | Domain: Networking | Tier: 1
Vmware · 31 cards β€” 🟒 8 easy | 🟑 12 medium | πŸ”΄ 5 hard
VMware · Portal | All Topics | Domain: Datacenter & Hardware | Tier: 2
Vpn · 17 cards β€” 🟒 3 easy | 🟑 4 medium | πŸ”΄ 3 hard
Vpn · 1. What is the difference between split tunneling and full tunneling in a VPN, and when would you choose each?
VPN & Tunneling · Portal | All Topics | Domain: Networking | Tier: 4
VS Code · Portal | All Topics | Domain: DevOps & Tooling | Tier: 4
Vscode · 16 cards β€” 🟒 3 easy | 🟑 4 medium | πŸ”΄ 3 hard
Vscode · 1. What VS Code feature lets you edit files on a remote server as if they were local, and how do you connect?
WebAssembly for Infrastructure · Portal | All Topics | Domain: DevOps & Tooling | Tier: 4
What's New · Recently added and updated wiki content
Wireshark & Packet Analysis · Portal | All Topics | Domain: Networking | Tier: 3
xargs · Portal | All Topics | Domain: CLI Tools | Tier: 1
YAML, JSON & Config Formats · Portal | All Topics | Domain: CLI Tools | Tier: 2
Zuul · 25 cards β€” 🟒 4 easy | 🟑 4 medium | πŸ”΄ 2 hard

Library / Topics (1492)

/proc Filesystem · Guide to the Linux /proc filesystem for process inspection and kernel tuning
/proc Filesystem - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Proc Filesystem.
/proc Filesystem Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Proc Filesystem.
Advanced Bash Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Advanced Bash.
Advanced Bash for Ops · Advanced Bash scripting techniques for system administration and automation
Advanced Bash for Ops - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Advanced Bash.
Advanced Bash β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Advanced Bash.
AI DevOps Tools β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Ai Devops Tools.
AI Tools for DevOps · Overview of AI-powered tools for DevOps automation, monitoring, and incident response
AI Tools for DevOps - Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Ai Devops Tools.
AI Tools for DevOps - Street Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Ai Devops Tools.
AI/ML Ops Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Ai Ml Ops.
AI/ML Ops β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Ai Ml Ops.
Alerting Rules · Guide to designing effective alerting rules that reduce noise and catch real incidents
Alerting Rules Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Alerting Rules.
Alerting Rules β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Alerting Rules.
Ansible Deep Dive - Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Ansible Deep Dive.
Ansible Deep Dive - Street Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Ansible Deep Dive.
Ansible Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Ansible.
Ansible for Infrastructure Automation - Street Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Ansible.
Ansible β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Ansible.
Ansible: idempotence + modules vs plugins vs collections · Explains the differences and relationships in Ansible: idempotence + modules vs plugins vs collections.
Ansible: inventory β€” hosts, groups, vars, targeting · Technical explainer covering Ansible: inventory β€” hosts, groups, vars, targeting concepts and practical usage.
Ansible: playbook vs play vs task vs role vs handler · Explains the differences and relationships in Ansible: playbook vs play vs task vs role vs handler.
Ansible: variable precedence · Explains evaluation order and priority rules for Ansible: variable precedence.
Anti-Primer: Advanced Bash · A narrative walkthrough of common Advanced Bash mistakes and cascading failures in production.
Anti-Primer: AI Devops Tools · A narrative walkthrough of common Ai Devops Tools mistakes and cascading failures in production.
Anti-Primer: AI ML Ops · A narrative walkthrough of common Ai Ml Ops mistakes and cascading failures in production.
Anti-Primer: Alerting Rules · A narrative walkthrough of common Alerting Rules mistakes and cascading failures in production.
Anti-Primer: Ansible · A narrative walkthrough of common Ansible mistakes and cascading failures in production.
Anti-Primer: Ansible Deep Dive · A narrative walkthrough of common Ansible Deep Dive mistakes and cascading failures in production.
Anti-Primer: API Gateways · A narrative walkthrough of common Api Gateways mistakes and cascading failures in production.
Anti-Primer: Argo Workflows · A narrative walkthrough of common Argo Workflows mistakes and cascading failures in production.
Anti-Primer: Argocd Gitops · A narrative walkthrough of common Argocd Gitops mistakes and cascading failures in production.
Anti-Primer: ARP · A narrative walkthrough of common Arp mistakes and cascading failures in production.
Anti-Primer: Audit Logging · A narrative walkthrough of common Audit Logging mistakes and cascading failures in production.
Anti-Primer: awk · A narrative walkthrough of common Awk mistakes and cascading failures in production.
Anti-Primer: AWS Cloudwatch · A narrative walkthrough of common Aws Cloudwatch mistakes and cascading failures in production.
Anti-Primer: AWS EC2 · A narrative walkthrough of common Aws Ec2 mistakes and cascading failures in production.
Anti-Primer: AWS ECS · A narrative walkthrough of common Aws Ecs mistakes and cascading failures in production.
Anti-Primer: AWS IAM · A narrative walkthrough of common Aws Iam mistakes and cascading failures in production.
Anti-Primer: AWS Lambda · A narrative walkthrough of common Aws Lambda mistakes and cascading failures in production.
Anti-Primer: AWS Networking · A narrative walkthrough of common Aws Networking mistakes and cascading failures in production.
Anti-Primer: AWS Route53 · A narrative walkthrough of common Aws Route53 mistakes and cascading failures in production.
Anti-Primer: AWS S3 Deep Dive · A narrative walkthrough of common Aws S3 Deep Dive mistakes and cascading failures in production.
Anti-Primer: AWS Troubleshooting · A narrative walkthrough of common Aws Troubleshooting mistakes and cascading failures in production.
Anti-Primer: Azure Blob Storage · A narrative walkthrough of common Azure Blob Storage mistakes and cascading failures in production.
Anti-Primer: Azure Troubleshooting · A narrative walkthrough of common Azure Troubleshooting mistakes and cascading failures in production.
Anti-Primer: Azure Virtual Machines · A narrative walkthrough of common Azure Virtual Machines mistakes and cascading failures in production.
Anti-Primer: Backstage · A narrative walkthrough of common Backstage mistakes and cascading failures in production.
Anti-Primer: Backup Restore · A narrative walkthrough of common Backup Restore mistakes and cascading failures in production.
Anti-Primer: Bare Metal Provisioning · A narrative walkthrough of common Bare Metal Provisioning mistakes and cascading failures in production.
Anti-Primer: BGP EVPN VXLAN · A narrative walkthrough of common Bgp Evpn Vxlan mistakes and cascading failures in production.
Anti-Primer: Binary And Floats · A narrative walkthrough of common Binary And Floats mistakes and cascading failures in production.
Anti-Primer: Capacity Planning · A narrative walkthrough of common Capacity Planning mistakes and cascading failures in production.
Anti-Primer: Career Engineering · A narrative walkthrough of common Career Engineering mistakes and cascading failures in production.
Anti-Primer: Ceph · A narrative walkthrough of common Ceph mistakes and cascading failures in production.
Anti-Primer: Cert Manager · A narrative walkthrough of common Cert Manager mistakes and cascading failures in production.
Anti-Primer: Cgroups Namespaces · A narrative walkthrough of common Cgroups Namespaces mistakes and cascading failures in production.
Anti-Primer: Change Management · A narrative walkthrough of common Change Management mistakes and cascading failures in production.
Anti-Primer: Chaos Engineering · A narrative walkthrough of common Chaos Engineering mistakes and cascading failures in production.
Anti-Primer: CI/CD · A narrative walkthrough of common Cicd mistakes and cascading failures in production.
Anti-Primer: Cilium · A narrative walkthrough of common Cilium mistakes and cascading failures in production.
Anti-Primer: Cisco Fundamentals For Devops · A narrative walkthrough of common Cisco Fundamentals For Devops mistakes and cascading failures in production.
Anti-Primer: Claude Code · A narrative walkthrough of common Claude Code mistakes and cascading failures in production.
Anti-Primer: Cloud Deep Dive · A narrative walkthrough of common Cloud Deep Dive mistakes and cascading failures in production.
Anti-Primer: Cloud Ops Basics · A narrative walkthrough of common Cloud Ops Basics mistakes and cascading failures in production.
Anti-Primer: Compliance Automation · A narrative walkthrough of common Compliance Automation mistakes and cascading failures in production.
Anti-Primer: Consul · A narrative walkthrough of common Consul mistakes and cascading failures in production.
Anti-Primer: Container Base Images · A narrative walkthrough of common Container Images mistakes and cascading failures in production.
Anti-Primer: Containers Deep Dive · A narrative walkthrough of common Containers Deep Dive mistakes and cascading failures in production.
Anti-Primer: Continuous Profiling · A narrative walkthrough of common Continuous Profiling mistakes and cascading failures in production.
Anti-Primer: Corporate It Fluency · A narrative walkthrough of common Corporate It Fluency mistakes and cascading failures in production.
Anti-Primer: Crashloopbackoff · A narrative walkthrough of common Crashloopbackoff mistakes and cascading failures in production.
Anti-Primer: Cron Scheduling · A narrative walkthrough of common Cron Scheduling mistakes and cascading failures in production.
Anti-Primer: Crossplane · A narrative walkthrough of common Crossplane mistakes and cascading failures in production.
Anti-Primer: CSS Fundamentals · A narrative walkthrough of common Css Fundamentals mistakes and cascading failures in production.
Anti-Primer: Curl And Wget · A narrative walkthrough of common Curl And Wget mistakes and cascading failures in production.
Anti-Primer: Dagger · A narrative walkthrough of common Dagger mistakes and cascading failures in production.
Anti-Primer: Database Internals · A narrative walkthrough of common Database Internals mistakes and cascading failures in production.
Anti-Primer: Database Ops · A narrative walkthrough of common Database Ops mistakes and cascading failures in production.
Anti-Primer: Datacenter · A narrative walkthrough of common Datacenter mistakes and cascading failures in production.
Anti-Primer: Debian Ubuntu · A narrative walkthrough of common Debian Ubuntu mistakes and cascading failures in production.
Anti-Primer: Debugging Methodology · A narrative walkthrough of common Debugging Methodology mistakes and cascading failures in production.
Anti-Primer: Dell Poweredge · A narrative walkthrough of common Dell Poweredge mistakes and cascading failures in production.
Anti-Primer: DHCP IPAM · A narrative walkthrough of common Dhcp Ipam mistakes and cascading failures in production.
Anti-Primer: Disaster Recovery · A narrative walkthrough of common Disaster Recovery mistakes and cascading failures in production.
Anti-Primer: Disk And Storage Ops · A narrative walkthrough of common Disk And Storage Ops mistakes and cascading failures in production.
Anti-Primer: Distributed Systems · A narrative walkthrough of common Distributed Systems mistakes and cascading failures in production.
Anti-Primer: DNS Deep Dive · A narrative walkthrough of common Dns Deep Dive mistakes and cascading failures in production.
Anti-Primer: DNS Ops · A narrative walkthrough of common Dns Ops mistakes and cascading failures in production.
Anti-Primer: DNSSEC · A narrative walkthrough of common Dnssec mistakes and cascading failures in production.
Anti-Primer: Docker · A narrative walkthrough of common Docker mistakes and cascading failures in production.
Anti-Primer: DORA Metrics · A narrative walkthrough of common Dora Metrics mistakes and cascading failures in production.
Anti-Primer: eBPF Observability · A narrative walkthrough of common Ebpf Observability mistakes and cascading failures in production.
Anti-Primer: Edge IoT · A narrative walkthrough of common Edge Iot mistakes and cascading failures in production.
Anti-Primer: Elasticsearch · A narrative walkthrough of common Elasticsearch mistakes and cascading failures in production.
Anti-Primer: Email Infrastructure · A narrative walkthrough of common Email Infrastructure mistakes and cascading failures in production.
Anti-Primer: Environment Variables · A narrative walkthrough of common Environment Variables mistakes and cascading failures in production.
Anti-Primer: Envoy · A narrative walkthrough of common Envoy mistakes and cascading failures in production.
Anti-Primer: Etcd · A narrative walkthrough of common Etcd mistakes and cascading failures in production.
Anti-Primer: Falco · A narrative walkthrough of common Falco mistakes and cascading failures in production.
Anti-Primer: fd · A narrative walkthrough of common Fd mistakes and cascading failures in production.
Anti-Primer: Feature Flags · A narrative walkthrough of common Feature Flags mistakes and cascading failures in production.
Anti-Primer: Find · A narrative walkthrough of common Find mistakes and cascading failures in production.
Anti-Primer: Finops · A narrative walkthrough of common Finops mistakes and cascading failures in production.
Anti-Primer: Firewalls · A narrative walkthrough of common Firewalls mistakes and cascading failures in production.
Anti-Primer: Firmware · A narrative walkthrough of common Firmware mistakes and cascading failures in production.
Anti-Primer: Fleet Ops · A narrative walkthrough of common Fleet Ops mistakes and cascading failures in production.
Anti-Primer: fzf · A narrative walkthrough of common Fzf mistakes and cascading failures in production.
Anti-Primer: GCP Troubleshooting · A narrative walkthrough of common Gcp Troubleshooting mistakes and cascading failures in production.
Anti-Primer: Git · A narrative walkthrough of common Git mistakes and cascading failures in production.
Anti-Primer: Git Advanced · A narrative walkthrough of common Git Advanced mistakes and cascading failures in production.
Anti-Primer: Github Actions · A narrative walkthrough of common Github Actions mistakes and cascading failures in production.
Anti-Primer: Gitops · A narrative walkthrough of common Gitops mistakes and cascading failures in production.
Anti-Primer: Graphql · A narrative walkthrough of common Graphql mistakes and cascading failures in production.
Anti-Primer: Grep And Regex · A narrative walkthrough of common Grep And Regex mistakes and cascading failures in production.
Anti-Primer: gRPC · A narrative walkthrough of common Grpc mistakes and cascading failures in production.
Anti-Primer: Hashicorp Vault · A narrative walkthrough of common Hashicorp Vault mistakes and cascading failures in production.
Anti-Primer: Helm · A narrative walkthrough of common Helm mistakes and cascading failures in production.
Anti-Primer: Homelab · A narrative walkthrough of common Homelab mistakes and cascading failures in production.
Anti-Primer: HTTP Protocol · A narrative walkthrough of common Http Protocol mistakes and cascading failures in production.
Anti-Primer: Incident Command · A narrative walkthrough of common Incident Command mistakes and cascading failures in production.
Anti-Primer: Incident Psychology · A narrative walkthrough of common Incident Psychology mistakes and cascading failures in production.
Anti-Primer: Incident Triage · A narrative walkthrough of common Incident Triage mistakes and cascading failures in production.
Anti-Primer: Infra Forensics · A narrative walkthrough of common Infra Forensics mistakes and cascading failures in production.
Anti-Primer: Infra Testing · A narrative walkthrough of common Infra Testing mistakes and cascading failures in production.
Anti-Primer: Inodes · A narrative walkthrough of common Inodes mistakes and cascading failures in production.
Anti-Primer: IPMI And Ipmitool · A narrative walkthrough of common Ipmi And Ipmitool mistakes and cascading failures in production.
Anti-Primer: Iptables Nftables · A narrative walkthrough of common Iptables Nftables mistakes and cascading failures in production.
Anti-Primer: Istio · A narrative walkthrough of common Istio mistakes and cascading failures in production.
Anti-Primer: jq · A narrative walkthrough of common Jq mistakes and cascading failures in production.
Anti-Primer: Kafka · A narrative walkthrough of common Kafka mistakes and cascading failures in production.
Anti-Primer: Kernel Troubleshooting · A narrative walkthrough of common Kernel Troubleshooting mistakes and cascading failures in production.
Anti-Primer: Kubernetes Debugging Playbook · A narrative walkthrough of common K8S Debugging Playbook mistakes and cascading failures in production.
Anti-Primer: Kubernetes Ecosystem · A narrative walkthrough of common K8S Ecosystem mistakes and cascading failures in production.
Anti-Primer: Kubernetes Networking · A narrative walkthrough of common K8S Networking mistakes and cascading failures in production.
Anti-Primer: Kubernetes Node Lifecycle · A narrative walkthrough of common K8S Node Lifecycle mistakes and cascading failures in production.
Anti-Primer: Kubernetes Ops · A narrative walkthrough of common K8S Ops mistakes and cascading failures in production.
Anti-Primer: Kubernetes Pods And Scheduling · A narrative walkthrough of common K8S Pods And Scheduling mistakes and cascading failures in production.
Anti-Primer: Kubernetes RBAC · A narrative walkthrough of common K8S Rbac mistakes and cascading failures in production.
Anti-Primer: Kubernetes Services And Ingress · A narrative walkthrough of common K8S Services And Ingress mistakes and cascading failures in production.
Anti-Primer: Kubernetes Storage · A narrative walkthrough of common K8S Storage mistakes and cascading failures in production.
Anti-Primer: Kustomize · A narrative walkthrough of common Kustomize mistakes and cascading failures in production.
Anti-Primer: LACP · A narrative walkthrough of common Lacp mistakes and cascading failures in production.
Anti-Primer: LDAP Identity · A narrative walkthrough of common Ldap Identity mistakes and cascading failures in production.
Anti-Primer: Legacy Archaeology · A narrative walkthrough of common Legacy Archaeology mistakes and cascading failures in production.
Anti-Primer: Linux Boot Process · A narrative walkthrough of common Linux Boot Process mistakes and cascading failures in production.
Anti-Primer: Linux Distro Comparison · A narrative walkthrough of common Linux Distro Comparison mistakes and cascading failures in production.
Anti-Primer: Linux Hardening · A narrative walkthrough of common Linux Hardening mistakes and cascading failures in production.
Anti-Primer: Linux Kernel Tuning · A narrative walkthrough of common Linux Kernel Tuning mistakes and cascading failures in production.
Anti-Primer: Linux Logging · A narrative walkthrough of common Linux Logging mistakes and cascading failures in production.
Anti-Primer: Linux Memory Management · A narrative walkthrough of common Linux Memory Management mistakes and cascading failures in production.
Anti-Primer: Linux Ops · A narrative walkthrough of common Linux Ops mistakes and cascading failures in production.
Anti-Primer: Linux Ops Storage · A narrative walkthrough of common Linux Ops Storage mistakes and cascading failures in production.
Anti-Primer: Linux Ops Systemd · A narrative walkthrough of common Linux Ops Systemd mistakes and cascading failures in production.
Anti-Primer: Linux Performance · A narrative walkthrough of common Linux Performance mistakes and cascading failures in production.
Anti-Primer: Linux Signals And Process Control · A narrative walkthrough of common Linux Signals And Process Control mistakes and cascading failures in production.
Anti-Primer: Linux Text Processing · A narrative walkthrough of common Linux Text Processing mistakes and cascading failures in production.
Anti-Primer: Linux Users And Permissions · A narrative walkthrough of common Linux Users And Permissions mistakes and cascading failures in production.
Anti-Primer: Load Balancing · A narrative walkthrough of common Load Balancing mistakes and cascading failures in production.
Anti-Primer: Load Testing · A narrative walkthrough of common Load Testing mistakes and cascading failures in production.
Anti-Primer: Log Pipelines · A narrative walkthrough of common Log Pipelines mistakes and cascading failures in production.
Anti-Primer: LPIC LFCS · A narrative walkthrough of common Lpic Lfcs mistakes and cascading failures in production.
Anti-Primer: Make And Build Systems · A narrative walkthrough of common Make And Build Systems mistakes and cascading failures in production.
Anti-Primer: Message Queues · A narrative walkthrough of common Message Queues mistakes and cascading failures in production.
Anti-Primer: Modern CLI · A narrative walkthrough of common Modern Cli mistakes and cascading failures in production.
Anti-Primer: Modern CLI Workflows · A narrative walkthrough of common Modern Cli Workflows mistakes and cascading failures in production.
Anti-Primer: Mongodb Ops · A narrative walkthrough of common Mongodb Ops mistakes and cascading failures in production.
Anti-Primer: Monitoring Fundamentals · A narrative walkthrough of common Monitoring Fundamentals mistakes and cascading failures in production.
Anti-Primer: Monitoring Migration · A narrative walkthrough of common Monitoring Migration mistakes and cascading failures in production.
Anti-Primer: Mounts Filesystems · A narrative walkthrough of common Mounts Filesystems mistakes and cascading failures in production.
Anti-Primer: MTU · A narrative walkthrough of common Mtu mistakes and cascading failures in production.
Anti-Primer: Multi Tenancy · A narrative walkthrough of common Multi Tenancy mistakes and cascading failures in production.
Anti-Primer: Mysql Ops · A narrative walkthrough of common Mysql Ops mistakes and cascading failures in production.
Anti-Primer: NAT · A narrative walkthrough of common Nat mistakes and cascading failures in production.
Anti-Primer: Network Automation · A narrative walkthrough of common Network Automation mistakes and cascading failures in production.
Anti-Primer: Networking · A narrative walkthrough of common Networking mistakes and cascading failures in production.
Anti-Primer: Networking Troubleshooting · A narrative walkthrough of common Networking Troubleshooting mistakes and cascading failures in production.
Anti-Primer: Nginx Web Servers · A narrative walkthrough of common Nginx Web Servers mistakes and cascading failures in production.
Anti-Primer: Nix · A narrative walkthrough of common Nix mistakes and cascading failures in production.
Anti-Primer: Node Maintenance · A narrative walkthrough of common Node Maintenance mistakes and cascading failures in production.
Anti-Primer: Observability Deep Dive · A narrative walkthrough of common Observability Deep Dive mistakes and cascading failures in production.
Anti-Primer: Offensive Security Basics · A narrative walkthrough of common Offensive Security Basics mistakes and cascading failures in production.
Anti-Primer: Oomkilled · A narrative walkthrough of common Oomkilled mistakes and cascading failures in production.
Anti-Primer: Open Policy Agent · A narrative walkthrough of common Open Policy Agent mistakes and cascading failures in production.
Anti-Primer: Opentelemetry · A narrative walkthrough of common Opentelemetry mistakes and cascading failures in production.
Anti-Primer: Opentofu · A narrative walkthrough of common Opentofu mistakes and cascading failures in production.
Anti-Primer: Ops War Stories · A narrative walkthrough of common Ops War Stories mistakes and cascading failures in production.
Anti-Primer: Opsec Mistakes · A narrative walkthrough of common Opsec Mistakes mistakes and cascading failures in production.
Anti-Primer: Package Management · A narrative walkthrough of common Package Management mistakes and cascading failures in production.
Anti-Primer: Packer · A narrative walkthrough of common Packer mistakes and cascading failures in production.
Anti-Primer: Perf Profiling · A narrative walkthrough of common Perf Profiling mistakes and cascading failures in production.
Anti-Primer: Pipes And Redirection · A narrative walkthrough of common Pipes And Redirection mistakes and cascading failures in production.
Anti-Primer: Platform Engineering · A narrative walkthrough of common Platform Engineering mistakes and cascading failures in production.
Anti-Primer: Policy Engines · A narrative walkthrough of common Policy Engines mistakes and cascading failures in production.
Anti-Primer: Postgresql · A narrative walkthrough of common Postgresql mistakes and cascading failures in production.
Anti-Primer: Postmortem SLO · A narrative walkthrough of common Postmortem Slo mistakes and cascading failures in production.
Anti-Primer: Power · A narrative walkthrough of common Power mistakes and cascading failures in production.
Anti-Primer: Powershell · A narrative walkthrough of common Powershell mistakes and cascading failures in production.
Anti-Primer: Proc Filesystem · A narrative walkthrough of common Proc Filesystem mistakes and cascading failures in production.
Anti-Primer: Process Management · A narrative walkthrough of common Process Management mistakes and cascading failures in production.
Anti-Primer: Progressive Delivery · A narrative walkthrough of common Progressive Delivery mistakes and cascading failures in production.
Anti-Primer: Prometheus Deep Dive · A narrative walkthrough of common Prometheus Deep Dive mistakes and cascading failures in production.
Anti-Primer: Pulumi · A narrative walkthrough of common Pulumi mistakes and cascading failures in production.
Anti-Primer: Python Async Concurrency · A narrative walkthrough of common Python Async Concurrency mistakes and cascading failures in production.
Anti-Primer: Python Debugging · A narrative walkthrough of common Python Debugging mistakes and cascading failures in production.
Anti-Primer: Python Infra · A narrative walkthrough of common Python Infra mistakes and cascading failures in production.
Anti-Primer: Python Packaging · A narrative walkthrough of common Python Packaging mistakes and cascading failures in production.
Anti-Primer: Rabbitmq · A narrative walkthrough of common Rabbitmq mistakes and cascading failures in production.
Anti-Primer: Redfish · A narrative walkthrough of common Redfish mistakes and cascading failures in production.
Anti-Primer: Redis · A narrative walkthrough of common Redis mistakes and cascading failures in production.
Anti-Primer: Regex Text Wrangling · A narrative walkthrough of common Regex Text Wrangling mistakes and cascading failures in production.
Anti-Primer: RHCE · A narrative walkthrough of common Rhce mistakes and cascading failures in production.
Anti-Primer: Ripgrep · A narrative walkthrough of common Ripgrep mistakes and cascading failures in production.
Anti-Primer: Routing · A narrative walkthrough of common Routing mistakes and cascading failures in production.
Anti-Primer: Rsync · A narrative walkthrough of common Rsync mistakes and cascading failures in production.
Anti-Primer: Runbook Craft · A narrative walkthrough of common Runbook Craft mistakes and cascading failures in production.
Anti-Primer: S3 Object Storage · A narrative walkthrough of common S3 Object Storage mistakes and cascading failures in production.
Anti-Primer: Secrets Management · A narrative walkthrough of common Secrets Management mistakes and cascading failures in production.
Anti-Primer: Security Basics · A narrative walkthrough of common Security Basics mistakes and cascading failures in production.
Anti-Primer: Security Scanning · A narrative walkthrough of common Security Scanning mistakes and cascading failures in production.
Anti-Primer: sed · A narrative walkthrough of common Sed mistakes and cascading failures in production.
Anti-Primer: Selinux Apparmor · A narrative walkthrough of common Selinux Apparmor mistakes and cascading failures in production.
Anti-Primer: Server Hardware · A narrative walkthrough of common Server Hardware mistakes and cascading failures in production.
Anti-Primer: Service Mesh · A narrative walkthrough of common Service Mesh mistakes and cascading failures in production.
Anti-Primer: SLO Tooling · A narrative walkthrough of common Slo Tooling mistakes and cascading failures in production.
Anti-Primer: SQL Fundamentals · A narrative walkthrough of common Sql Fundamentals mistakes and cascading failures in production.
Anti-Primer: Sqlite · A narrative walkthrough of common Sqlite mistakes and cascading failures in production.
Anti-Primer: SRE Practices · A narrative walkthrough of common Sre Practices mistakes and cascading failures in production.
Anti-Primer: SSH Deep Dive · A narrative walkthrough of common Ssh Deep Dive mistakes and cascading failures in production.
Anti-Primer: Storage Ops · A narrative walkthrough of common Storage Ops mistakes and cascading failures in production.
Anti-Primer: STP · A narrative walkthrough of common Stp mistakes and cascading failures in production.
Anti-Primer: Strace · A narrative walkthrough of common Strace mistakes and cascading failures in production.
Anti-Primer: Subnetting And IP Addressing · A narrative walkthrough of common Subnetting And Ip Addressing mistakes and cascading failures in production.
Anti-Primer: Supply Chain Security · A narrative walkthrough of common Supply Chain Security mistakes and cascading failures in production.
Anti-Primer: Synthetic Monitoring · A narrative walkthrough of common Synthetic Monitoring mistakes and cascading failures in production.
Anti-Primer: Systemctl Journalctl · A narrative walkthrough of common Systemctl Journalctl mistakes and cascading failures in production.
Anti-Primer: Systems Thinking · A narrative walkthrough of common Systems Thinking mistakes and cascading failures in production.
Anti-Primer: Tailscale · A narrative walkthrough of common Tailscale mistakes and cascading failures in production.
Anti-Primer: Tar And Compression · A narrative walkthrough of common Tar And Compression mistakes and cascading failures in production.
Anti-Primer: TCP/IP Deep Dive · A narrative walkthrough of common Tcp Ip Deep Dive mistakes and cascading failures in production.
Anti-Primer: Terminal Internals · A narrative walkthrough of common Terminal Internals mistakes and cascading failures in production.
Anti-Primer: Terraform · A narrative walkthrough of common Terraform mistakes and cascading failures in production.
Anti-Primer: Terraform Deep Dive · A narrative walkthrough of common Terraform Deep Dive mistakes and cascading failures in production.
Anti-Primer: TLS Certificates Ops · A narrative walkthrough of common Tls Certificates Ops mistakes and cascading failures in production.
Anti-Primer: Tmux And Screen · A narrative walkthrough of common Tmux And Screen mistakes and cascading failures in production.
Anti-Primer: Tracing · A narrative walkthrough of common Tracing mistakes and cascading failures in production.
Anti-Primer: Vendor Management · A narrative walkthrough of common Vendor Management mistakes and cascading failures in production.
Anti-Primer: Virtualization · A narrative walkthrough of common Virtualization mistakes and cascading failures in production.
Anti-Primer: VLANs · A narrative walkthrough of common Vlans mistakes and cascading failures in production.
Anti-Primer: VPN Tunneling · A narrative walkthrough of common Vpn Tunneling mistakes and cascading failures in production.
Anti-Primer: VS Code · A narrative walkthrough of common Vscode mistakes and cascading failures in production.
Anti-Primer: WebAssembly Infrastructure · A narrative walkthrough of common Wasm Infrastructure mistakes and cascading failures in production.
Anti-Primer: Wireshark · A narrative walkthrough of common Wireshark mistakes and cascading failures in production.
Anti-Primer: Xargs · A narrative walkthrough of common Xargs mistakes and cascading failures in production.
Anti-Primer: YAML JSON Config · A narrative walkthrough of common Yaml Json Config mistakes and cascading failures in production.
API Gateways & Ingress · Guide to API gateway and ingress controller patterns for traffic management
API Gateways & Ingress - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Api Gateways.
API Gateways & Ingress Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Api Gateways.
API Gateways β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Api Gateways.
Argo CD & GitOps β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Argocd Gitops.
Argo Workflows · Guide to Argo Workflows for running parallel jobs and CI/CD pipelines on Kubernetes
Argo Workflows Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Argo Workflows.
Argo Workflows β€” Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Argo Workflows.
Argo Workflows β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Argo Workflows.
ArgoCD & GitOps · Guide to ArgoCD for GitOps-based continuous delivery in Kubernetes
ArgoCD & GitOps - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Gitops.
ArgoCD & GitOps Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Argocd Gitops.
ArgoCD & GitOps β€” Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Argocd Gitops.
ARP · Guide to ARP protocol mechanics, cache management, and troubleshooting network connectivity
ARP - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Arp.
ARP Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Arp.
ARP β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Arp.
Audit Logging · Guide to audit logging for compliance, security monitoring, and forensic analysis
Audit Logging - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Audit Logging.
Audit Logging Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Audit Logging.
Audit Logging β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Audit Logging.
Automation · Comprehensive guide to automation with primers, street ops, and common pitfalls
awk β€” Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Awk.
awk β€” Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Awk.
awk β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Awk.
awk: The Record/Field Processor · Guide to AWK for structured text processing, log analysis, and data extraction
AWS CloudWatch · Guide to AWS CloudWatch metrics, alarms, logs, and dashboards for monitoring
AWS CloudWatch - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Aws Cloudwatch.
AWS CloudWatch Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Aws Cloudwatch.
AWS EC2 · Guide to AWS EC2 instance types, lifecycle management, and operational best practices
AWS EC2 - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Aws Ec2.
AWS EC2 Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Aws Ec2.
AWS EC2 β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Aws Ec2.
AWS ECS · Guide to AWS ECS container orchestration with tasks, services, Fargate, and troubleshooting
AWS ECS - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Aws Ecs.
AWS ECS Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Aws Ecs.
AWS IAM · Guide to AWS IAM policies, roles, users, and least-privilege access patterns
AWS IAM - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Aws Iam.
AWS IAM Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Aws Iam.
AWS IAM β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Aws Iam.
AWS Lambda · Guide to AWS Lambda serverless functions, triggers, deployment, and debugging
AWS Lambda - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Aws Lambda.
AWS Lambda Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Aws Lambda.
AWS Lambda β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Aws Lambda.
AWS Networking · Guide to AWS VPCs, subnets, security groups, route tables, and network troubleshooting
AWS Networking - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Aws Networking.
AWS Networking Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Aws Networking.
AWS Networking β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Aws Networking.
AWS Route 53 · Guide to AWS Route 53 DNS service including hosted zones, routing policies, and health checks
AWS Route 53 - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Aws Route53.
AWS Route 53 Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Aws Route53.
AWS S3 Deep Dive · S3 is a key-value object store, not a filesystem. There are no directories -- the / in a key is just a character
AWS S3 Deep Dive - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Aws S3 Deep Dive.
AWS S3 Deep Dive Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Aws S3 Deep Dive.
AWS Troubleshooting · Guide to diagnosing and resolving common AWS service issues across compute and networking
AWS Troubleshooting - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Aws Troubleshooting.
AWS Troubleshooting Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Aws Troubleshooting.
AWS Troubleshooting β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Aws Troubleshooting.
Azure Blob Storage · Guide to Azure Blob Storage resource model, access tiers, and operational best practices
Azure Blob Storage - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Azure Blob Storage.
Azure Blob Storage Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Azure Blob Storage.
Azure Blob Storage β€” Trivia & Quick Facts · Quick-reference facts, limits, and defaults for Azure Blob Storage.
Azure Troubleshooting · Guide to troubleshooting Azure services across compute, networking, and identity
Azure Troubleshooting - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Azure Troubleshooting.
Azure Troubleshooting Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Azure Troubleshooting.
Azure Troubleshooting β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Azure Troubleshooting.
Azure Virtual Machines · Guide to Azure Virtual Machines resource model, lifecycle management, and operational best practices
Azure Virtual Machines - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Azure Virtual Machines.
Azure Virtual Machines Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Azure Virtual Machines.
Azure Virtual Machines β€” Trivia & Quick Facts · Quick-reference facts, limits, and defaults for Azure Virtual Machines.
Azure Virtual Network · Guide to Azure Virtual Network (VNet) subnetting, routing, security, and peering. Partial topic pack.
Azure Virtual Network - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Azure Virtual Network (VNet).
Azure Virtual Network Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Azure Virtual Network (VNet). Partial β€” source research was cut off mid-list.
Backstage & Developer Portals · Guide to Backstage developer portal for service catalogs and infrastructure tooling
Backstage - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Backstage.
Backstage Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Backstage.
Backstage β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Backstage.
Backup & Restore - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Backup Restore.
Backup & Restore Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Backup Restore.
Backup & Restore β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Backup Restore.
Backup Restore · Guide to backup and restore strategies, scheduling, verification, and disaster recovery
Bare Metal Provisioning β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Bare Metal Provisioning.
Bare-Metal Provisioning · Guide to bare-metal server provisioning with PXE, iPXE, and automation tools
Bare-Metal Provisioning - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Bare Metal Provisioning.
Bare-Metal Provisioning Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Bare Metal Provisioning.
BGP EVPN / VXLAN · Guide to BGP routing, EVPN control plane, and VXLAN overlay networking for datacenters
BGP EVPN / VXLAN Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Bgp Evpn Vxlan.
BGP EVPN / VXLAN β€” Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Bgp Evpn Vxlan.
BGP EVPN VXLAN β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Bgp Evpn Vxlan.
Binary & Floats β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Binary And Floats.
Binary and Floating Point Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Binary And Floats.
Binary and Floating Point β€” Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Binary And Floats.
Binary and Floats · Overview and learning resources for binary and floats
Btrfs: subvolume, snapshot, reflink, CoW · Technical explainer covering Btrfs: subvolume, snapshot, reflink, CoW concepts and practical usage.
Capacity Planning · Guide to infrastructure capacity planning, forecasting, and right-sizing resources
Capacity Planning - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Capacity Planning.
Capacity Planning Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Capacity Planning.
Capacity Planning β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Capacity Planning.
Career Engineering Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Career Engineering.
Career Engineering for Ops People · Guide to career growth, skill development, and job strategy for infrastructure engineers
Career Engineering for Ops People - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Career Engineering.
Career Engineering β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Career Engineering.
Ceph Storage · Guide to Ceph distributed storage including RADOS, RBD, CephFS, and cluster operations
Ceph Storage Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Ceph.
Ceph Storage β€” Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Ceph.
Ceph β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Ceph.
cert-manager · Guide to cert-manager for automated TLS certificate management in Kubernetes
cert-manager Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Cert Manager.
cert-manager β€” Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Cert Manager.
cert-manager β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Cert Manager.
Certificates · Comprehensive guide to certificates with primers, street ops, and common pitfalls
cgroups & Linux Namespaces · Guide to Linux cgroups and namespaces, the building blocks of container isolation
cgroups & Linux Namespaces - Street Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Cgroups Namespaces.
cgroups & Namespaces Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Cgroups Namespaces.
Change Management · Guide to change management processes for safe infrastructure and code deployments
Change Management - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Change Management.
Change Management Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Change Management.
Change Management β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Change Management.
Chaos Engineering & Fault Injection · Guide to chaos engineering practices, fault injection tools, and resilience validation
Chaos Engineering & Fault Injection - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Chaos Engineering.
Chaos Engineering Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Chaos Engineering.
Chaos Engineering β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Chaos Engineering.
CI/CD as a System · Technical explainer covering CI/CD as a System concepts and practical usage.
CI/CD Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Cicd.
CI/CD Pipelines & Patterns · Guide to CI/CD pipeline design, build automation, and deployment workflow patterns
CI/CD Pipelines - Street Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Cicd.
CI/CD β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Cicd.
Cilium & eBPF Networking · Guide to Cilium eBPF-based networking, security, and observability for Kubernetes
Cilium & eBPF Networking - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Cilium.
Cilium & eBPF Networking Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Cilium.
Cilium β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Cilium.
Cisco Fundamentals -- Street Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Cisco Fundamentals For Devops.
Cisco Fundamentals Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Cisco Fundamentals For Devops.
Cisco Fundamentals for DevOps · Guide to Cisco IOS basics for DevOps engineers working with network infrastructure
Cisco Fundamentals for DevOps β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Cisco Fundamentals For Devops.
Claude Code · Guide to Claude Code AI assistant for infrastructure automation and DevOps workflows
Claude Code - Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Claude Code.
Claude Code - Street Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Claude Code.
Claude Code β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Claude Code.
Cloud Deep Dive · In-depth guide to cloud architecture, multi-cloud strategies, and cost optimization
Cloud Deep Dive β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Cloud Deep Dive.
Cloud Deep-Dive Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Cloud Deep Dive.
Cloud Operations Basics - Street Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Cloud Ops Basics.
Cloud Ops Basics · Foundational guide to cloud operations covering compute, storage, and networking
Cloud Ops Basics β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Cloud Ops Basics.
Cloud Ops Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Cloud Ops Basics.
Cloud Provider Deep-Dive - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Cloud Deep Dive.
Compliance & Audit Automation · Guide to automating compliance checks, policy enforcement, and audit reporting
Compliance & Audit Automation - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Compliance Automation.
Compliance & Audit Automation Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Compliance Automation.
Compliance Automation β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Compliance Automation.
CompTIA Security+ (SY0-701) Prep · Original SY0-701 practice question bank, high-yield memory tables, and domain-by-domain study material.
Configuration Management · Comprehensive guide to configuration management with primers, street ops, and common pitfalls
Container Base Images β€” Footguns & Pitfalls · Critical mistakes and dangerous pitfalls to avoid when working with Container Images.
Container Base Images β€” Street Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Container Images.
Container Base Images β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Container Images.
Container Images · Guide to container image building, optimization, registries, and security scanning
Container vs VM · Explains the differences and relationships in Container vs VM.
Containers Deep Dive · In-depth guide to container internals, runtimes, image layers, and troubleshooting
Containers Deep Dive - Footguns & Pitfalls · Critical mistakes and dangerous pitfalls to avoid when working with Containers Deep Dive.
Containers Deep Dive - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Containers Deep Dive.
Containers Deep Dive β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Containers Deep Dive.
Continuous Profiling · Guide to continuous profiling for identifying performance bottlenecks in production
Continuous Profiling Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Continuous Profiling.
Continuous Profiling β€” Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Continuous Profiling.
Continuous Profiling β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Continuous Profiling.
Corporate IT Fluency - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Corporate It Fluency.
Corporate IT Fluency Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Corporate It Fluency.
Corporate IT Fluency for Engineers · Guide to corporate IT concepts like AD, GPO, ITIL, and enterprise networking for engineers
Corporate IT Fluency β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Corporate It Fluency.
Cost Optimization & FinOps - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Finops.
CrashLoopBackOff · Guide to diagnosing and fixing CrashLoopBackOff errors in Kubernetes pods
CrashLoopBackOff - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Crashloopbackoff.
CrashLoopBackOff Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Crashloopbackoff.
CrashLoopBackOff β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Crashloopbackoff.
Cron & Job Scheduling · Guide to cron jobs, systemd timers, and task scheduling patterns in Linux
Cron & Job Scheduling - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Cron Scheduling.
Cron & Job Scheduling Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Cron Scheduling.
Cron Scheduling β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Cron Scheduling.
Crossplane · namespace: crossplane-system
Crossplane - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Crossplane.
Crossplane Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Crossplane.
Crossplane β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Crossplane.
CSS Fundamentals · Overview and learning resources for css fundamentals
CSS Fundamentals Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Css Fundamentals.
CSS Fundamentals β€” Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Css Fundamentals.
CSS Fundamentals β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Css Fundamentals.
curl & wget · Guide to curl and wget for HTTP debugging, file transfers, and API testing
curl & wget β€” Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Curl And Wget.
curl & wget β€” Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Curl And Wget.
curl & wget β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Curl And Wget.
Dagger - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Dagger.
Dagger / CI as Code · Guide to Dagger CI/CD pipelines as code using containers and programmable workflows
Dagger Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Dagger.
Dagger β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Dagger.
Data Modeling · Comprehensive guide to data modeling with primers, street ops, and common pitfalls
Database Internals · Replication copies data from a primary (read-write) to one or more replicas (read-only). In PostgreSQL, this is built on WAL (Write-Ahead Log) shipping
Database Internals - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Database Internals.
Database Internals Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Database Internals.
Database Internals β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Database Internals.
Database Operations - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Database Ops.
Database Operations on Kubernetes · Guide to running databases on Kubernetes with StatefulSets, operators, and backup strategies
Database Operations β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Database Ops.
Database Ops Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Database Ops.
Databases · Comprehensive guide to databases with primers, street ops, and common pitfalls
Datacenter & Server Hardware · Guide to datacenter operations including rack layout, cabling, power, and cooling
Datacenter & Server Hardware - Street Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Datacenter.
Datacenter Advanced Operations · Technical explainer covering Datacenter Advanced Operations concepts and practical usage.
Datacenter Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Datacenter.
Datacenter β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Datacenter.
Debian & Ubuntu Ecosystem · Guide to Debian and Ubuntu package management, releases, and server administration
Debian & Ubuntu β€” Footguns & Pitfalls · Critical mistakes and dangerous pitfalls to avoid when working with Debian Ubuntu.
Debian & Ubuntu β€” Street Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Debian Ubuntu.
Debian & Ubuntu β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Debian Ubuntu.
Debugging Methodology · Systematic approach to debugging infrastructure issues from symptoms to root cause
Debugging Methodology - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Debugging Methodology.
Debugging Methodology Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Debugging Methodology.
Debugging Methodology β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Debugging Methodology.
Deep Dive · Advanced Ansible patterns including roles, collections, custom modules, and optimization
Dell PowerEdge Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Dell Poweredge.
Dell PowerEdge Servers · Guide to Dell PowerEdge server management, iDRAC configuration, and hardware troubleshooting
Dell PowerEdge β€” Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Dell Poweredge.
Dell PowerEdge β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Dell Poweredge.
Deployment vs ReplicaSet vs Pod · Explains the differences and relationships in Deployment vs ReplicaSet vs Pod.
Deployments · Comprehensive guide to deployments with primers, street ops, and common pitfalls
DHCP & IP Address Management · Guide to DHCP server configuration, IPAM tools, and IP address lifecycle management
DHCP & IP Address Management - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Dhcp Ipam.
DHCP & IP Address Management Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Dhcp Ipam.
DHCP & IPAM β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Dhcp Ipam.
Disaster Recovery & Backup Engineering · Guide to DR planning, RTO/RPO targets, failover testing, and backup verification
Disaster Recovery & Backup Engineering - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Disaster Recovery.
Disaster Recovery Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Disaster Recovery.
Disaster Recovery β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Disaster Recovery.
Disk & Storage Ops · Guide to disk operations including partitioning, RAID, LVM, and performance tuning
Disk & Storage Ops - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Disk And Storage Ops.
Disk & Storage Ops Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Disk And Storage Ops.
Disk & Storage Ops β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Disk And Storage Ops.
Distributed Systems Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Distributed Systems.
Distributed Systems Fundamentals · Guide to distributed systems concepts including CAP theorem, consensus, and replication
Distributed Systems Fundamentals β€” Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Distributed Systems.
Distributed Systems β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Distributed Systems.
Distributed Tracing - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Tracing.
Distributed Tracing Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Tracing.
DNF Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Dnf.
DNF Package Manager · Overview and learning resources for dnf package manager
DNF β€” Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Dnf.
DNS Deep Dive · In-depth guide to DNS protocol, zone management, resolvers, and troubleshooting
DNS Deep Dive - Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Dns Deep Dive.
DNS Deep Dive - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Dns Deep Dive.
DNS Operations · Guide to DNS zone management, resolver configuration, caching, and troubleshooting
DNS Operations - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Dns Ops.
DNS Operations Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Dns Ops.
DNS Operations β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Dns Ops.
DNS: Stub Resolver vs Recursive Resolver vs Authoritative Server · Explains the differences and relationships in DNS: Stub Resolver vs Recursive Resolver vs Authoritative Server.
DNSSEC & DNS Security · Guide to DNSSEC for cryptographic DNS validation, key rotation, and deployment
DNSSEC & DNS Security Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Dnssec.
DNSSEC & DNS Security β€” Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Dnssec.
DNSSEC β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Dnssec.
Docker · Guide to Docker container operations including builds, networking, volumes, and debugging
Docker / Containers - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Docker.
Docker Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Docker.
Docker β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Docker.
DORA Metrics & DevEx · Guide to DORA metrics for measuring deployment frequency, lead time, and change failure rate
DORA Metrics & DevEx Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Dora Metrics.
DORA Metrics & DevEx β€” Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Dora Metrics.
DORA Metrics β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Dora Metrics.
eBPF & Modern Linux Observability · Guide to eBPF for kernel-level tracing, network monitoring, and security enforcement
eBPF & Modern Linux Observability - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Ebpf Observability.
eBPF & Modern Linux Observability Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Ebpf Observability.
eBPF Observability β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Ebpf Observability.
Edge & IoT Infrastructure · Guide to edge computing architecture, IoT device management, and fleet orchestration
Edge & IoT Infrastructure - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Edge Iot.
Edge & IoT Infrastructure Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Edge Iot.
Edge & IoT β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Edge Iot.
Elasticsearch · Guide to Elasticsearch cluster operations, indexing, querying, and performance tuning
Elasticsearch - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Elasticsearch.
Elasticsearch Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Elasticsearch.
Elasticsearch β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Elasticsearch.
Email Infrastructure · Guide to email infrastructure including MTA configuration, SPF, DKIM, and DMARC
Email Infrastructure Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Email Infrastructure.
Email Infrastructure β€” Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Email Infrastructure.
Email Infrastructure β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Email Infrastructure.
Environment Variables · Guide to environment variables in Linux for configuration, secrets, and shell customization
Environment Variables - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Environment Variables.
Environment Variables Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Environment Variables.
Envoy Proxy · Guide to Envoy proxy configuration, traffic routing, observability, and xDS APIs
Envoy Proxy β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Envoy.
etcd · Guide to etcd distributed key-value store used as the Kubernetes backing store
etcd - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Etcd.
etcd Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Etcd.
etcd β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Etcd.
Falco β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Falco.
fd · Guide to fd, a fast and user-friendly alternative to the find command
fd - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Fd.
fd Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Fd.
fd β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Fd.
Feature Flags · Guide to feature flag systems for progressive rollouts and operational safety
Feature Flags Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Feature Flags.
Feature Flags β€” Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Feature Flags.
Feature Flags β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Feature Flags.
File vs inode vs pathname vs symlink · Explains the differences and relationships in File vs inode vs pathname vs symlink.
find · Guide to the find command for locating files by name, type, size, and modification time
find - Footguns & Pitfalls · Critical mistakes and dangerous pitfalls to avoid when working with Find.
find - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Find.
find β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Find.
FinOps & Cost Optimization · Guide to cloud cost optimization, tagging strategies, and FinOps team practices
FinOps Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Finops.
FinOps β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Finops.
Firewall Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Firewalls.
Firewalls · Guide to firewall configuration, rule management, and network security architecture
Firewalls - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Firewalls.
Firewalls β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Firewalls.
Firmware · Guide to firmware management including BIOS/UEFI updates, BMC, and lifecycle operations
Firmware & BIOS - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Firmware.
Firmware & BIOS Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Firmware.
Firmware β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Firmware.
Fleet Operations at Scale - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Fleet Ops.
Fleet Operations Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Fleet Ops.
Fleet Ops · Guide to managing large server fleets with automation, drift detection, and rollout strategies
Fleet Ops β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Fleet Ops.
Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Consul.
Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Envoy.
Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Graphql.
Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Message Queues.
Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Open Policy Agent.
fzf · Guide to fzf fuzzy finder for interactive filtering of files, history, and command output
fzf - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Fzf.
fzf Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Fzf.
fzf β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Fzf.
GCP Troubleshooting · Guide to Google Cloud Platform troubleshooting across compute, networking, and storage
GCP Troubleshooting - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Gcp Troubleshooting.
GCP Troubleshooting Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Gcp Troubleshooting.
GCP Troubleshooting β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Gcp Troubleshooting.
Git Advanced · Advanced Git techniques including rebasing, bisect, reflog, and complex merge strategies
Git Advanced - Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Git Advanced.
Git Advanced - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Git Advanced.
Git Advanced β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Git Advanced.
Git Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Git.
Git for DevOps · Guide to Git fundamentals including branching, merging, and collaboration workflows
Git for DevOps Engineers - Street Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Git.
Git Workflows & Branching Strategies · Overview and learning resources for git workflows & branching strategies
Git Workflows Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Git Workflows.
Git Workflows β€” Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Git Workflows.
Git β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Git.
Git: commit vs branch vs tag vs HEAD · Explains the differences and relationships in Git: commit vs branch vs tag vs HEAD.
Git: rebase vs merge · Explains the differences and relationships in Git: rebase vs merge.
Git: working tree vs index vs repository · Explains the differences and relationships in Git: working tree vs index vs repository.
GitHub Actions · Guide to GitHub Actions CI/CD workflows, runners, and automation patterns
GitHub Actions - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Github Actions.
GitHub Actions Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Github Actions.
GitHub Actions β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Github Actions.
GitOps · Guide to GitOps principles, workflows, and tools for declarative infrastructure management
GitOps Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Gitops.
GitOps β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Gitops.
GraphQL · Guide to GraphQL APIs for infrastructure tooling, schema design, and query optimization
grep & Regular Expressions · Overview and learning resources for grep & regular expressions
grep & Regular Expressions - Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Grep And Regex.
grep & Regular Expressions - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Grep And Regex.
gRPC & Protocol Buffers · Guide to gRPC protocol buffers, service definitions, streaming, and debugging
gRPC - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Grpc.
gRPC Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Grpc.
gRPC β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Grpc.
HAProxy & Nginx for Ops · Guide to HAProxy and Nginx load balancing, reverse proxy, and health check configuration
HAProxy & Nginx for Ops - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Load Balancing.
HAProxy & Nginx Load Balancing Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Load Balancing.
HashiCorp Consul · Guide to HashiCorp Consul for service discovery, KV configuration, and mesh networking
HashiCorp Vault · Guide to HashiCorp Vault for secrets management, encryption, and access control
HashiCorp Vault - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Hashicorp Vault.
HashiCorp Vault Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Hashicorp Vault.
HashiCorp Vault β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Hashicorp Vault.
Helm · Guide to Helm package manager for Kubernetes charts, releases, and templating
Helm - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Helm.
Helm Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Helm.
Helm β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Helm.
Homelab & Learning Infrastructure · Guide to building a homelab for hands-on practice with servers, networking, and clusters
Homelab & Learning Infrastructure - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Homelab.
Homelab Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Homelab.
Homelab β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Homelab.
HTTP Protocol · Guide to HTTP protocol internals including methods, headers, status codes, and HTTP/2
HTTP Protocol Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Http Protocol.
HTTP Protocol β€” Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Http Protocol.
HTTP Protocol β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Http Protocol.
Image vs Container · Explains the differences and relationships in Image vs Container.
Incident Command & On-Call · Guide to incident command roles, on-call rotations, and escalation procedures
Incident Command & On-Call - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Incident Command.
Incident Command & On-Call Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Incident Command.
Incident Command β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Incident Command.
Incident Postmortem & SLO/SLI - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Postmortem Slo.
Incident Psychology β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Incident Psychology.
Incident Triage · Guide to incident triage methodology for rapid assessment and priority assignment
Incident Triage - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Incident Triage.
Incident Triage Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Incident Triage.
Incident Triage β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Incident Triage.
Infrastructure as Code with Terraform - Street Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Terraform.
Infrastructure Forensics · Guide to post-incident forensic analysis, evidence preservation, and root cause investigation
Infrastructure Forensics - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Infra Forensics.
Infrastructure Forensics Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Infra Forensics.
Infrastructure Forensics β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Infra Forensics.
Infrastructure Testing · Guide to infrastructure testing with Terratest, InSpec, and automated compliance validation
Infrastructure Testing Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Infra Testing.
Infrastructure Testing β€” Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Infra Testing.
Infrastructure Testing β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Infra Testing.
Inode Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Inodes.
Inodes · Overview and learning resources for inodes
Inodes - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Inodes.
Inodes β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Inodes.
IPMI & ipmitool β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Ipmi And Ipmitool.
IPMI and ipmitool · Guide to IPMI protocol and ipmitool for remote server power, sensor, and console management
IPMI and ipmitool -- Street Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Ipmi And Ipmitool.
IPMI and ipmitool Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Ipmi And Ipmitool.
iptables & nftables · Guide to Linux packet filtering with iptables and nftables, including rules, chains, and debugging
iptables & nftables - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Iptables Nftables.
iptables & nftables Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Iptables Nftables.
iptables & nftables β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Iptables Nftables.
Istio Service Mesh · Guide to Istio service mesh for traffic management, mTLS, and observability in Kubernetes
Istio Service Mesh Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Istio.
Istio Service Mesh β€” Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Istio.
Istio β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Istio.
jq · Guide to jq command-line JSON processor for filtering, transforming, and querying data
jq - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Jq.
jq Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Jq.
jq β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Jq.
K8s Concept Chain β€” Footguns · Critical mistakes and dangerous pitfalls to avoid when working with K8S Concept Chain.
K8s Concept Chain β€” Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for K8S Concept Chain.
K8s Ecosystem · Covers the Kubernetes ecosystem tools and includes Kubernetes Operators & CRDs (merged from k8s-operators)
K8s Networking · Guide to Kubernetes networking including CNI plugins, Services, DNS, and network policies
K8s RBAC · Guide to Kubernetes RBAC roles, bindings, and access control best practices
K8s Storage · Guide to Kubernetes storage including PVs, PVCs, StorageClasses, and CSI drivers
Kafka · Guide to Apache Kafka operations including brokers, topics, consumers, and troubleshooting
Kafka - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Kafka.
Kafka Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Kafka.
Kafka β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Kafka.
Kernel Troubleshooting · Guide to Linux kernel troubleshooting including panics, modules, and parameter tuning
Kernel Troubleshooting - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Kernel Troubleshooting.
Kernel Troubleshooting Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Kernel Troubleshooting.
Kernel Troubleshooting β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Kernel Troubleshooting.
Kubernetes Concept Chain · Overview and learning resources for k8s concept chain
Kubernetes Control Plane as Reconciliation Engine · Technical explainer covering Kubernetes Control Plane as Reconciliation Engine concepts and practical usage.
Kubernetes Debugging -- Street Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for K8S Debugging Playbook.
Kubernetes Debugging Footguns · Critical mistakes and dangerous pitfalls to avoid when working with K8S Debugging Playbook.
Kubernetes Debugging Playbook · Systematic playbook for debugging Kubernetes issues from pod failures to cluster problems
Kubernetes Debugging Playbook β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about K8S Debugging Playbook.
Kubernetes Ecosystem - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for K8S Ecosystem.
Kubernetes Ecosystem Footguns · Critical mistakes and dangerous pitfalls to avoid when working with K8S Ecosystem.
Kubernetes Ecosystem β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about K8S Ecosystem.
Kubernetes Networking - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for K8S Networking.
Kubernetes Networking Footguns · Critical mistakes and dangerous pitfalls to avoid when working with K8S Networking.
Kubernetes Networking β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about K8S Networking.
Kubernetes Node Lifecycle · Guide to Kubernetes node provisioning, cordoning, draining, upgrading, and decommissioning
Kubernetes Node Lifecycle & Cluster Upgrades · Covers the stages and operational concerns of Kubernetes Node Lifecycle & Cluster Upgrades.
Kubernetes Node Lifecycle -- Street Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for K8S Node Lifecycle.
Kubernetes Node Lifecycle Footguns · Critical mistakes and dangerous pitfalls to avoid when working with K8S Node Lifecycle.
Kubernetes Node Lifecycle β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about K8S Node Lifecycle.
Kubernetes Ops (Production) · Guide to production Kubernetes operations including upgrades, scaling, and troubleshooting
Kubernetes Ops Footguns · Critical mistakes and dangerous pitfalls to avoid when working with K8S Ops.
Kubernetes Ops β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about K8S Ops.
Kubernetes Pods & Scheduling · Guide to pod lifecycle, resource requests and limits, affinity, and scheduler configuration
Kubernetes Pods & Scheduling - Street Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for K8S Pods And Scheduling.
Kubernetes Pods & Scheduling Footguns · Critical mistakes and dangerous pitfalls to avoid when working with K8S Pods And Scheduling.
Kubernetes RBAC β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about K8S Rbac.
Kubernetes Services & Ingress · Guide to Kubernetes Service types, Ingress controllers, and traffic routing patterns
Kubernetes Services & Ingress - Street Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for K8S Services And Ingress.
Kubernetes Services & Ingress Footguns · Critical mistakes and dangerous pitfalls to avoid when working with K8S Services And Ingress.
Kubernetes Storage - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for K8S Storage.
Kubernetes Storage Footguns · Critical mistakes and dangerous pitfalls to avoid when working with K8S Storage.
Kubernetes Storage β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about K8S Storage.
Kustomize · Guide to Kustomize for Kubernetes manifest customization without templating
Kustomize - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Kustomize.
Kustomize Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Kustomize.
Kustomize β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Kustomize.
LACP · Guide to LACP link aggregation for bandwidth and redundancy across network interfaces
LACP - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Lacp.
LACP Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Lacp.
LACP β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Lacp.
LDAP & Identity Management · Guide to LDAP directory services, identity providers, and centralized authentication
LDAP & Identity Management - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Ldap Identity.
LDAP & Identity Management Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Ldap Identity.
LDAP & Identity β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Ldap Identity.
Legacy Archaeology β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Legacy Archaeology.
Legacy System Archaeology · Guide to safely understanding, documenting, and modernizing undocumented legacy systems
Legacy System Archaeology - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Legacy Archaeology.
Legacy System Archaeology Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Legacy Archaeology.
Linux Boot Process · Guide to the Linux boot process from BIOS/UEFI through bootloader to init
Linux Boot Process β€” Footguns & Pitfalls · Critical mistakes and dangerous pitfalls to avoid when working with Linux Boot Process.
Linux Boot Process β€” Street Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Linux Boot Process.
Linux Boot Process β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Linux Boot Process.
Linux Data Hoarding · Overview and learning resources for linux data hoarding
Linux Data Hoarding - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Linux Data Hoarding.
Linux Data Hoarding Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Linux Data Hoarding.
Linux Deep Triage · Technical explainer covering Linux Deep Triage concepts and practical usage.
Linux Distribution Comparison · Comparison of Linux distributions by use case, package management, and release philosophy
Linux Distribution Comparison β€” Footguns & Pitfalls · Critical mistakes and dangerous pitfalls to avoid when working with Linux Distro Comparison.
Linux Distribution Comparison β€” Street Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Linux Distro Comparison.
Linux Distro Comparison β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Linux Distro Comparison.
Linux Hardening β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Linux Hardening.
Linux Kernel Tuning · Guide to Linux kernel parameter tuning via sysctl for performance and stability
Linux Kernel Tuning - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Linux Kernel Tuning.
Linux Kernel Tuning Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Linux Kernel Tuning.
Linux Logging · Guide to Linux logging with journald, syslog, and structured log management
Linux Logging β€” Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Linux Logging.
Linux Logging β€” Street Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Linux Logging.
Linux Logging β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Linux Logging.
Linux Memory Management · Guide to Linux memory management including virtual memory, swap, OOM killer, and tuning
Linux Memory Management β€” Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Linux Memory Management.
Linux Memory Management β€” Street Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Linux Memory Management.
Linux Memory Management β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Linux Memory Management.
Linux Ops · Foundational Linux operations guide covering common tasks and troubleshooting
Linux Ops Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Linux Ops.
Linux Ops Storage · Overview and learning resources for linux ops storage
Linux Ops Storage β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Linux Ops Storage.
Linux Ops Systemd · Overview and learning resources for linux ops: systemd
Linux Ops β€” systemd β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Linux Ops Systemd.
Linux Ops β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Linux Ops.
Linux Performance Tuning · Guide to Linux performance analysis and tuning for CPU, memory, disk, and network
Linux Performance Tuning - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Linux Performance.
Linux Performance Tuning Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Linux Performance.
Linux Performance β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Linux Performance.
Linux Signals & Process Control · Guide to Linux signals, process groups, job control, and graceful termination
Linux Signals & Process Control - Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Linux Signals And Process Control.
Linux Signals & Process Control - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Linux Signals And Process Control.
Linux Storage Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Linux Ops Storage.
Linux Storage Street Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Linux Ops Storage.
Linux System Administration - Street Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Linux Ops.
Linux Text Processing · Guide to Linux text processing tools including sort, cut, tr, paste, and column
Linux Text Processing - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Linux Text Processing.
Linux Text Processing Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Linux Text Processing.
Linux Text Processing β€” Trivia & History · Surprising history, little-known facts, and interesting details about Linux Text Processing.
Linux Users & Permissions · Guide to Linux user management, file permissions, ACLs, and sudo configuration
Linux Users and Permissions β€” Footguns & Pitfalls · Critical mistakes and dangerous pitfalls to avoid when working with Linux Users And Permissions.
Linux Users and Permissions β€” Street Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Linux Users And Permissions.
Linux Users and Permissions β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Linux Users And Permissions.
Linux: kernel vs userspace vs distro · Explains the differences and relationships in Linux: kernel vs userspace vs distro.
Load Balancing β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Load Balancing.
Load Testing · Guide to load testing tools and methodology for capacity validation and performance tuning
Load Testing Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Load Testing.
Load Testing β€” Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Load Testing.
Load Testing β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Load Testing.
Log Analysis & Alerting Rules - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Alerting Rules.
Log Pipelines · Guide to log collection, parsing, routing, and storage pipeline architecture
Log Pipelines - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Log Pipelines.
Log Pipelines Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Log Pipelines.
Log Pipelines β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Log Pipelines.
Logs vs Metrics vs Traces · Explains the differences and relationships in Logs vs Metrics vs Traces.
LPIC & LFCS β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Lpic Lfcs.
LPIC / LFCS Exam Preparation · Study guide and lab exercises for LPIC and LFCS Linux certification exams
LPIC / LFCS β€” Footguns & Pitfalls · Critical mistakes and dangerous pitfalls to avoid when working with Lpic Lfcs.
LPIC / LFCS β€” Street Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Lpic Lfcs.
Make & Build Systems · Guide to Make and build systems for automation, compilation, and task orchestration
Make & Build Systems β€” Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Make And Build Systems.
Make & Build Systems β€” Street Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Make And Build Systems.
Make & Build Systems β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Make And Build Systems.
Mellanox Switches · Overview and learning resources for mellanox switches
Mellanox Switches β€” Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Mellanox Switches.
Mellanox Switches β€” Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Mellanox Switches.
Mental Models (Core Concepts) · Comprehensive guide to mental models (core concepts) with primers, street ops, and common pitfalls
mergerfs · Guide to mergerfs union filesystem for pooling drives without RAID
mergerfs - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Mergerfs.
mergerfs Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Mergerfs.
Message Queues · Guide to message queue patterns, broker selection, and operational best practices
Modern CLI Tools · Overview of modern CLI replacements for traditional Unix tools
Modern CLI Tools - Street Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Modern Cli.
Modern CLI Tools Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Modern Cli.
Modern Cli Workflows · Overview and learning resources for modern cli workflows
Modern CLI Workflows -- Street Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Modern Cli Workflows.
Modern CLI Workflows Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Modern Cli Workflows.
Modern CLI Workflows β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Modern Cli Workflows.
Modern CLI β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Modern Cli.
MongoDB Operations · Guide to MongoDB operations including replica sets, sharding, and performance optimization
MongoDB Operations Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Mongodb Ops.
MongoDB Operations β€” Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Mongodb Ops.
MongoDB Operations β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Mongodb Ops.
Monitoring Fundamentals · Foundational guide to monitoring concepts, metrics types, and alerting strategies
Monitoring Fundamentals - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Monitoring Fundamentals.
Monitoring Fundamentals Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Monitoring Fundamentals.
Monitoring Fundamentals β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Monitoring Fundamentals.
Monitoring Migration (Legacy to Modern) · Guide to migrating monitoring systems from Nagios-era tools to modern observability stacks
Monitoring Migration (Legacy to Modern) - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Monitoring Migration.
Monitoring Migration Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Monitoring Migration.
Monitoring Migration β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Monitoring Migration.
Mounts & Filesystems - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Mounts Filesystems.
Mounts & Filesystems Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Mounts Filesystems.
Mounts & Filesystems β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Mounts Filesystems.
Mounts Filesystems · Guide to Linux mount operations, filesystem types, fstab configuration, and troubleshooting
MTU · Guide to MTU sizing, path MTU discovery, jumbo frames, and fragmentation troubleshooting
MTU - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Mtu.
MTU Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Mtu.
MTU β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Mtu.
Multi-Cluster & Federation - Exercises & Reference · Quick-reference guide for Multi-Cluster & Federation - Exercises & Reference commands and usage.
Multi-Tenancy Patterns · Guide to multi-tenancy isolation patterns for shared infrastructure and Kubernetes clusters
Multi-Tenancy Patterns - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Multi Tenancy.
Multi-Tenancy Patterns Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Multi Tenancy.
Multi-Tenancy β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Multi Tenancy.
MySQL / MariaDB Operations · Guide to MySQL and MariaDB operations including replication, backups, and performance tuning
MySQL / MariaDB Operations Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Mysql Ops.
MySQL / MariaDB Operations β€” Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Mysql Ops.
MySQL Operations β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Mysql Ops.
NAT · Guide to Network Address Translation including SNAT, DNAT, and masquerading
NAT - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Nat.
NAT Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Nat.
NAT β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Nat.
Network Automation · Guide to network automation with NAPALM, Netmiko, and infrastructure as code
Network Automation Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Network Automation.
Network Automation β€” Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Network Automation.
Network Automation β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Network Automation.
Network Traps & Deep Debugging · Technical explainer covering Network Traps & Deep Debugging concepts and practical usage.
Networking - Street Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Networking.
Networking Deep Dive · In-depth guide to networking concepts from L2 switching through L7 application protocols
Networking Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Networking.
Networking Troubleshooting · Guide to systematic network troubleshooting across layers with diagnostic tools
Networking Troubleshooting Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Networking Troubleshooting.
Networking Troubleshooting Street Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Networking Troubleshooting.
Networking Troubleshooting β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Networking Troubleshooting.
Networking β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Networking.
Nginx & Web Servers · Guide to Nginx configuration, reverse proxy setup, TLS termination, and performance tuning
Nginx & Web Servers - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Nginx Web Servers.
Nginx & Web Servers Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Nginx Web Servers.
nginx Web Servers β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Nginx Web Servers.
Nix / NixOS · Guide to Nix package manager and NixOS for reproducible builds and system configuration
Nix / NixOS - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Nix.
Nix / NixOS Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Nix.
Nix β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Nix.
Node Maintenance · Guide to Kubernetes node maintenance including cordoning, draining, and upgrades
Node Maintenance - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Node Maintenance.
Node Maintenance Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Node Maintenance.
Node Maintenance β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Node Maintenance.
Observability Deep Dive · In-depth guide to observability pillars: metrics, logs, traces, and their correlation
Observability Deep Dive - Street Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Observability Deep Dive.
Observability Deep Dive β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Observability Deep Dive.
Observability Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Observability Deep Dive.
Offensive Security Basics · You can't defend what you don't understand. Every misconfigured firewall rule,
Offensive Security Basics β€” Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Offensive Security Basics.
Offensive Security Basics β€” Street Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Offensive Security Basics.
Offensive Security Basics β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Offensive Security Basics.
On-Call · Comprehensive guide to on-call with primers, street ops, and common pitfalls
OOMKilled · Guide to diagnosing and preventing OOMKilled errors in Kubernetes and Linux
OOMKilled - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Oomkilled.
OOMKilled Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Oomkilled.
OOMKilled β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Oomkilled.
Open Policy Agent · Guide to OPA for policy-as-code enforcement across infrastructure and Kubernetes
Open Policy Agent β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Open Policy Agent.
OpenTelemetry · Guide to OpenTelemetry for unified metrics, traces, and logs instrumentation
OpenTelemetry - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Opentelemetry.
OpenTelemetry Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Opentelemetry.
OpenTelemetry β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Opentelemetry.
OpenTofu & Terraform Ecosystem · Guide to OpenTofu open-source infrastructure as code, a Terraform-compatible alternative
OpenTofu - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Opentofu.
OpenTofu Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Opentofu.
OpenTofu β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Opentofu.
Ops War Stories & Pattern Recognition · Collection of real-world ops failures and the patterns they reveal for prevention
Ops War Stories & Pattern Recognition - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Ops War Stories.
Ops War Stories & Pattern Recognition Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Ops War Stories.
Ops War Stories β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Ops War Stories.
Ops-Focused Security Basics - Street Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Security Basics.
Opsec Mistakes · Guide to common operational security mistakes and how to avoid credential and data leaks
OpSec Mistakes - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Opsec Mistakes.
OpSec Mistakes Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Opsec Mistakes.
OPSEC Mistakes β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Opsec Mistakes.
Package Management · Guide to Linux package management with apt, yum, dnf, and repository configuration
Package Management - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Package Management.
Package Management Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Package Management.
Package Management β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Package Management.
Packer · Guide to HashiCorp Packer for building automated machine images across platforms
Packer β€” Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Packer.
Packer β€” Street Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Packer.
Packer β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Packer.
perf Profiling · Overview and learning resources for perf profiling
perf Profiling Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Perf Profiling.
perf Profiling β€” Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Perf Profiling.
Performance · Comprehensive guide to performance with primers, street ops, and common pitfalls
Performance Profiling β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Perf Profiling.
Permissions: mode bits vs ownership vs ACLs vs capabilities · Explains the differences and relationships in Permissions: mode bits vs ownership vs ACLs vs capabilities.
Persistent Volume vs Persistent Volume Claim · Explains the differences and relationships in Persistent Volume vs Persistent Volume Claim.
Pipes & Redirection · Guide to Unix pipes, file descriptors, redirection, and process substitution
Pipes & Redirection - Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Pipes And Redirection.
Pipes & Redirection - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Pipes And Redirection.
Platform Engineering Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Platform Engineering.
Platform Engineering Patterns · Guide to building internal developer platforms with self-service and golden paths
Platform Engineering Patterns - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Platform Engineering.
Platform Engineering β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Platform Engineering.
Pod vs Container (Kubernetes) · Explains the differences and relationships in Pod vs Container (Kubernetes).
Policy Engine Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Policy Engines.
Policy Engines (OPA / Kyverno) · Guide to policy engines for Kubernetes admission control and infrastructure governance
Policy Engines - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Policy Engines.
Policy Engines β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Policy Engines.
PostgreSQL Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Postgresql.
PostgreSQL Operations · Guide to PostgreSQL database administration, performance tuning, and operations
PostgreSQL Operations - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Postgresql.
PostgreSQL β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Postgresql.
Postmortem & SLO Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Postmortem Slo.
Postmortems & SLOs · Guide to writing blameless postmortems and defining actionable SLOs with error budgets
Postmortems & SLOs β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Postmortem Slo.
Power · Guide to datacenter power systems including PDUs, UPS, redundancy, and capacity planning
Power & UPS - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Power.
Power & UPS Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Power.
Power β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Power.
PowerShell · Guide to PowerShell for cross-platform automation and Windows infrastructure management
PowerShell Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Powershell.
PowerShell Street Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Powershell.
PowerShell β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Powershell.
Practical Kubernetes Ops - Street Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for K8S Ops.
Primer · Core concepts, commands, and practical foundations for Advanced Bash.
Primer · Core concepts, commands, and practical foundations for Ai Devops Tools.
Primer · Core concepts, commands, and practical foundations for Ai Ml Ops.
Primer · Core concepts, commands, and practical foundations for Alerting Rules.
Primer · Core concepts, commands, and practical foundations for Ansible.
Primer · Core concepts, commands, and practical foundations for Ansible Deep Dive.
Primer · Core concepts, commands, and practical foundations for Api Gateways.
Primer · Core concepts, commands, and practical foundations for Argo Workflows.
Primer · Core concepts, commands, and practical foundations for Argocd Gitops.
Primer · Core concepts, commands, and practical foundations for Arp.
Primer · Core concepts, commands, and practical foundations for Audit Logging.
Primer · Core concepts, commands, and practical foundations for Awk.
Primer · Core concepts, commands, and practical foundations for Aws Cloudwatch.
Primer · Core concepts, commands, and practical foundations for Aws Ec2.
Primer · Core concepts, commands, and practical foundations for Aws Ecs.
Primer · Core concepts, commands, and practical foundations for Aws Iam.
Primer · Core concepts, commands, and practical foundations for Aws Lambda.
Primer · Core concepts, commands, and practical foundations for Aws Networking.
Primer · Core concepts, commands, and practical foundations for Aws Route53.
Primer · Core concepts, commands, and practical foundations for Aws S3 Deep Dive.
Primer · Core concepts, commands, and practical foundations for Aws Troubleshooting.
Primer · Core concepts, resource model, and practical foundations for Azure Blob Storage.
Primer · Core concepts, commands, and practical foundations for Azure Troubleshooting.
Primer · Core concepts, resource model, and practical foundations for Azure Virtual Machines.
Primer · Core concepts, resource model, and practical foundations for Azure Virtual Network (VNet).
Primer · Core concepts, commands, and practical foundations for Backstage.
Primer · Core concepts, commands, and practical foundations for Backup Restore.
Primer · Core concepts, commands, and practical foundations for Bare Metal Provisioning.
Primer · Core concepts, commands, and practical foundations for Bgp Evpn Vxlan.
Primer · Core concepts, commands, and practical foundations for Binary And Floats.
Primer · Core concepts, commands, and practical foundations for Capacity Planning.
Primer · Core concepts, commands, and practical foundations for Career Engineering.
Primer · Core concepts, commands, and practical foundations for Ceph.
Primer · Core concepts, commands, and practical foundations for Cert Manager.
Primer · Core concepts, commands, and practical foundations for Cgroups Namespaces.
Primer · Core concepts, commands, and practical foundations for Change Management.
Primer · Core concepts, commands, and practical foundations for Chaos Engineering.
Primer · Core concepts, commands, and practical foundations for Cicd.
Primer · Core concepts, commands, and practical foundations for Cilium.
Primer · Core concepts, commands, and practical foundations for Cisco Fundamentals For Devops.
Primer · Core concepts, commands, and practical foundations for Claude Code.
Primer · Core concepts, commands, and practical foundations for Cloud Deep Dive.
Primer · Core concepts, commands, and practical foundations for Cloud Ops Basics.
Primer · Core concepts, commands, and practical foundations for Compliance Automation.
Primer · Core concepts, commands, and practical foundations for Consul.
Primer · Core concepts, commands, and practical foundations for Container Images.
Primer · Core concepts, commands, and practical foundations for Containers Deep Dive.
Primer · Core concepts, commands, and practical foundations for Continuous Profiling.
Primer · Core concepts, commands, and practical foundations for Corporate It Fluency.
Primer · Core concepts, commands, and practical foundations for Crashloopbackoff.
Primer · Core concepts, commands, and practical foundations for Cron Scheduling.
Primer · Core concepts, commands, and practical foundations for Crossplane.
Primer · Core concepts, commands, and practical foundations for Css Fundamentals.
Primer · Core concepts, commands, and practical foundations for Curl And Wget.
Primer · Core concepts, commands, and practical foundations for Dagger.
Primer · Core concepts, commands, and practical foundations for Database Internals.
Primer · Core concepts, commands, and practical foundations for Database Ops.
Primer · Core concepts, commands, and practical foundations for Datacenter.
Primer · Core concepts, commands, and practical foundations for Debian Ubuntu.
Primer · Core concepts, commands, and practical foundations for Debugging Methodology.
Primer · Core concepts, commands, and practical foundations for Dell Poweredge.
Primer · Core concepts, commands, and practical foundations for Dhcp Ipam.
Primer · Core concepts, commands, and practical foundations for Disaster Recovery.
Primer · Core concepts, commands, and practical foundations for Disk And Storage Ops.
Primer · Core concepts, commands, and practical foundations for Distributed Systems.
Primer · Core concepts, commands, and practical foundations for Dnf.
Primer · Core concepts, commands, and practical foundations for Dns Deep Dive.
Primer · Core concepts, commands, and practical foundations for Dns Ops.
Primer · Core concepts, commands, and practical foundations for Dnssec.
Primer · Core concepts, commands, and practical foundations for Docker.
Primer · Core concepts, commands, and practical foundations for Dora Metrics.
Primer · Core concepts, commands, and practical foundations for Ebpf Observability.
Primer · Core concepts, commands, and practical foundations for Edge Iot.
Primer · Core concepts, commands, and practical foundations for Elasticsearch.
Primer · Core concepts, commands, and practical foundations for Email Infrastructure.
Primer · Core concepts, commands, and practical foundations for Environment Variables.
Primer · Core concepts, commands, and practical foundations for Envoy.
Primer · Core concepts, commands, and practical foundations for Etcd.
Primer · Core concepts, commands, and practical foundations for Falco.
Primer · Core concepts, commands, and practical foundations for Fd.
Primer · Core concepts, commands, and practical foundations for Feature Flags.
Primer · Core concepts, commands, and practical foundations for Find.
Primer · Core concepts, commands, and practical foundations for Finops.
Primer · Core concepts, commands, and practical foundations for Firewalls.
Primer · Core concepts, commands, and practical foundations for Firmware.
Primer · Core concepts, commands, and practical foundations for Fleet Ops.
Primer · Core concepts, commands, and practical foundations for Fzf.
Primer · Core concepts, commands, and practical foundations for Gcp Troubleshooting.
Primer · Core concepts, commands, and practical foundations for Git.
Primer · Core concepts, commands, and practical foundations for Git Advanced.
Primer · Core concepts, commands, and practical foundations for Git Workflows.
Primer · Core concepts, commands, and practical foundations for Github Actions.
Primer · Core concepts, commands, and practical foundations for Gitops.
Primer · Core concepts, commands, and practical foundations for Graphql.
Primer · Core concepts, commands, and practical foundations for Grep And Regex.
Primer · Core concepts, commands, and practical foundations for Grpc.
Primer · Core concepts, commands, and practical foundations for Hashicorp Vault.
Primer · Core concepts, commands, and practical foundations for Helm.
Primer · Core concepts, commands, and practical foundations for Homelab.
Primer · Core concepts, commands, and practical foundations for Http Protocol.
Primer · Core concepts, commands, and practical foundations for Incident Command.
Primer · Core concepts, commands, and practical foundations for Incident Psychology.
Primer · Core concepts, commands, and practical foundations for Incident Triage.
Primer · Core concepts, commands, and practical foundations for Infra Forensics.
Primer · Core concepts, commands, and practical foundations for Infra Testing.
Primer · Core concepts, commands, and practical foundations for Inodes.
Primer · Core concepts, commands, and practical foundations for Ipmi And Ipmitool.
Primer · Core concepts, commands, and practical foundations for Iptables Nftables.
Primer · Core concepts, commands, and practical foundations for Istio.
Primer · Core concepts, commands, and practical foundations for Jq.
Primer · Core concepts, commands, and practical foundations for K8S Concept Chain.
Primer · Core concepts, commands, and practical foundations for K8S Debugging Playbook.
Primer · Core concepts, commands, and practical foundations for K8S Ecosystem.
Primer · Core concepts, commands, and practical foundations for K8S Networking.
Primer · Core concepts, commands, and practical foundations for K8S Node Lifecycle.
Primer · Core concepts, commands, and practical foundations for K8S Ops.
Primer · Core concepts, commands, and practical foundations for K8S Pods And Scheduling.
Primer · Core concepts, commands, and practical foundations for K8S Rbac.
Primer · Core concepts, commands, and practical foundations for K8S Services And Ingress.
Primer · Core concepts, commands, and practical foundations for K8S Storage.
Primer · Core concepts, commands, and practical foundations for Kafka.
Primer · Core concepts, commands, and practical foundations for Kernel Troubleshooting.
Primer · Core concepts, commands, and practical foundations for Kustomize.
Primer · Core concepts, commands, and practical foundations for Lacp.
Primer · Core concepts, commands, and practical foundations for Ldap Identity.
Primer · Core concepts, commands, and practical foundations for Legacy Archaeology.
Primer · Core concepts, commands, and practical foundations for Linux Boot Process.
Primer · Core concepts, commands, and practical foundations for Linux Data Hoarding.
Primer · Core concepts, commands, and practical foundations for Linux Distro Comparison.
Primer · Core concepts, commands, and practical foundations for Linux Hardening.
Primer · Core concepts, commands, and practical foundations for Linux Kernel Tuning.
Primer · Core concepts, commands, and practical foundations for Linux Logging.
Primer · Core concepts, commands, and practical foundations for Linux Memory Management.
Primer · Core concepts, commands, and practical foundations for Linux Ops.
Primer · Core concepts, commands, and practical foundations for Linux Ops Storage.
Primer · Core concepts, commands, and practical foundations for Linux Ops Systemd.
Primer · Core concepts, commands, and practical foundations for Linux Performance.
Primer · Core concepts, commands, and practical foundations for Linux Signals And Process Control.
Primer · Core concepts, commands, and practical foundations for Linux Text Processing.
Primer · Core concepts, commands, and practical foundations for Linux Users And Permissions.
Primer · Core concepts, commands, and practical foundations for Load Balancing.
Primer · Core concepts, commands, and practical foundations for Load Testing.
Primer · Core concepts, commands, and practical foundations for Log Pipelines.
Primer · Core concepts, commands, and practical foundations for Lpic Lfcs.
Primer · Core concepts, commands, and practical foundations for Make And Build Systems.
Primer · Core concepts, commands, and practical foundations for Mellanox Switches.
Primer · Core concepts, commands, and practical foundations for Mergerfs.
Primer · Core concepts, commands, and practical foundations for Message Queues.
Primer · Core concepts, commands, and practical foundations for Modern Cli.
Primer · Core concepts, commands, and practical foundations for Modern Cli Workflows.
Primer · Core concepts, commands, and practical foundations for Mongodb Ops.
Primer · Core concepts, commands, and practical foundations for Monitoring Fundamentals.
Primer · Core concepts, commands, and practical foundations for Monitoring Migration.
Primer · Core concepts, commands, and practical foundations for Mounts Filesystems.
Primer · Core concepts, commands, and practical foundations for Mtu.
Primer · Core concepts, commands, and practical foundations for Multi Tenancy.
Primer · Core concepts, commands, and practical foundations for Mysql Ops.
Primer · Core concepts, commands, and practical foundations for Nat.
Primer · Core concepts, commands, and practical foundations for Network Automation.
Primer · Core concepts, commands, and practical foundations for Networking.
Primer · Core concepts, commands, and practical foundations for Networking Troubleshooting.
Primer · Core concepts, commands, and practical foundations for Nginx Web Servers.
Primer · Core concepts, commands, and practical foundations for Nix.
Primer · Core concepts, commands, and practical foundations for Node Maintenance.
Primer · Core concepts, commands, and practical foundations for Observability Deep Dive.
Primer · Core concepts, commands, and practical foundations for Offensive Security Basics.
Primer · Core concepts, commands, and practical foundations for Oomkilled.
Primer · Core concepts, commands, and practical foundations for Open Policy Agent.
Primer · Core concepts, commands, and practical foundations for Opentelemetry.
Primer · Core concepts, commands, and practical foundations for Opentofu.
Primer · Core concepts, commands, and practical foundations for Ops War Stories.
Primer · Core concepts, commands, and practical foundations for Opsec Mistakes.
Primer · Core concepts, commands, and practical foundations for Package Management.
Primer · Core concepts, commands, and practical foundations for Packer.
Primer · Core concepts, commands, and practical foundations for Perf Profiling.
Primer · Core concepts, commands, and practical foundations for Pipes And Redirection.
Primer · Core concepts, commands, and practical foundations for Platform Engineering.
Primer · Core concepts, commands, and practical foundations for Policy Engines.
Primer · Core concepts, commands, and practical foundations for Postgresql.
Primer · Core concepts, commands, and practical foundations for Postmortem Slo.
Primer · Core concepts, commands, and practical foundations for Power.
Primer · Core concepts, commands, and practical foundations for Powershell.
Primer · Core concepts, commands, and practical foundations for Proc Filesystem.
Primer · Core concepts, commands, and practical foundations for Process Management.
Primer · Core concepts, commands, and practical foundations for Progressive Delivery.
Primer · Core concepts, commands, and practical foundations for Prometheus Deep Dive.
Primer · Core concepts, commands, and practical foundations for Pulumi.
Primer · Core concepts, commands, and practical foundations for Python Async Concurrency.
Primer · Core concepts, commands, and practical foundations for Python Debugging.
Primer · Core concepts, commands, and practical foundations for Python Infra.
Primer · Core concepts, commands, and practical foundations for Python Packaging.
Primer · Core concepts, commands, and practical foundations for Rabbitmq.
Primer · Core concepts, commands, and practical foundations for Redfish.
Primer · Core concepts, commands, and practical foundations for Redis.
Primer · Core concepts, commands, and practical foundations for Regex Text Wrangling.
Primer · Core concepts, commands, and practical foundations for Rhce.
Primer · Core concepts, commands, and practical foundations for Ripgrep.
Primer · Core concepts, commands, and practical foundations for Routing.
Primer · Core concepts, commands, and practical foundations for Rsync.
Primer · Core concepts, commands, and practical foundations for Runbook Craft.
Primer · Core concepts, commands, and practical foundations for S3 Object Storage.
Primer · Core concepts, commands, and practical foundations for Secrets Management.
Primer · Core concepts, commands, and practical foundations for Security Basics.
Primer · Core concepts, commands, and practical foundations for Security Scanning.
Primer · Core concepts, commands, and practical foundations for Sed.
Primer · Core concepts, commands, and practical foundations for Selinux Apparmor.
Primer · Core concepts, commands, and practical foundations for Server Hardware.
Primer · Core concepts, commands, and practical foundations for Service Mesh.
Primer · Core concepts, commands, and practical foundations for Slo Tooling.
Primer · Core concepts, commands, and practical foundations for Sql Fundamentals.
Primer · Core concepts, commands, and practical foundations for Sqlite.
Primer · Core concepts, commands, and practical foundations for Sre Practices.
Primer · Core concepts, commands, and practical foundations for Ssh Deep Dive.
Primer · Core concepts, commands, and practical foundations for Storage Ops.
Primer · Core concepts, commands, and practical foundations for Stp.
Primer · Core concepts, commands, and practical foundations for Strace.
Primer · Core concepts, commands, and practical foundations for Subnetting And Ip Addressing.
Primer · Core concepts, commands, and practical foundations for Supply Chain Security.
Primer · Core concepts, commands, and practical foundations for Synthetic Monitoring.
Primer · Core concepts, commands, and practical foundations for Systemctl Journalctl.
Primer · Core concepts, commands, and practical foundations for Systems Thinking.
Primer · Core concepts, commands, and practical foundations for Tailscale.
Primer · Core concepts, commands, and practical foundations for Tar And Compression.
Primer · Core concepts, commands, and practical foundations for Tcp Ip Deep Dive.
Primer · Core concepts, commands, and practical foundations for Terminal Internals.
Primer · Core concepts, commands, and practical foundations for Terraform.
Primer · Core concepts, commands, and practical foundations for Terraform Deep Dive.
Primer · Core concepts, commands, and practical foundations for Tls Certificates Ops.
Primer · Core concepts, commands, and practical foundations for Tmux And Screen.
Primer · Core concepts, commands, and practical foundations for Tracing.
Primer · Core concepts, commands, and practical foundations for Vendor Management.
Primer · Core concepts, commands, and practical foundations for Virtualization.
Primer · Core concepts, commands, and practical foundations for Vlans.
Primer · Core concepts, commands, and practical foundations for Vpn Tunneling.
Primer · Core concepts, commands, and practical foundations for Vscode.
Primer · Core concepts, commands, and practical foundations for Wasm Infrastructure.
Primer · Core concepts, commands, and practical foundations for Wireshark.
Primer · Core concepts, commands, and practical foundations for Xargs.
Primer · Core concepts, commands, and practical foundations for Yaml Json Config.
Process Management · Guide to Linux process management including ps, top, nice, and process states
Process Management - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Process Management.
Process Management Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Process Management.
Process Management β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Process Management.
Process vs program vs service · Explains the differences and relationships in Process vs program vs service.
Progressive Delivery · Guide to progressive delivery with canary releases, feature flags, and rollback strategies
Progressive Delivery Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Progressive Delivery.
Progressive Delivery β€” Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Progressive Delivery.
Progressive Delivery β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Progressive Delivery.
Prometheus Deep Dive · apiVersion: monitoring.coreos.com/v1
Prometheus Deep Dive - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Prometheus Deep Dive.
Prometheus Deep Dive Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Prometheus Deep Dive.
Pulumi · Guide to Pulumi infrastructure as code using general-purpose programming languages
Pulumi - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Pulumi.
Pulumi Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Pulumi.
Pulumi β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Pulumi.
Python Async & Concurrency · The GIL is a mutex in CPython that allows only one thread to execute Python bytecode at a time. This is the single most important fact about Python conc...
Python Async & Concurrency - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Python Async Concurrency.
Python Async & Concurrency Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Python Async Concurrency.
Python Debugging · Guide to Python debugging tools and techniques including pdb, logging, and profiling
Python Debugging Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Python Debugging.
Python Debugging β€” Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Python Debugging.
Python Debugging β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Python Debugging.
Python for Infrastructure · Guide to Python scripting for infrastructure automation, APIs, and tool development
Python for Infrastructure - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Python Infra.
Python for Infrastructure Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Python Infra.
Python for Infrastructure β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Python Infra.
Python Packaging · Guide to Python packaging with pip, setuptools, wheels, and virtual environments
Python Packaging Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Python Packaging.
Python Packaging β€” Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Python Packaging.
Python Packaging β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Python Packaging.
RabbitMQ & Message Queues · Guide to RabbitMQ message broker operations, clustering, and troubleshooting
RabbitMQ Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Rabbitmq.
RabbitMQ Operations - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Rabbitmq.
RabbitMQ β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Rabbitmq.
RAID vs Backup vs Snapshot · Explains the differences and relationships in RAID vs Backup vs Snapshot.
RBAC - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for K8S Rbac.
RBAC Footguns · Critical mistakes and dangerous pitfalls to avoid when working with K8S Rbac.
Redfish -- Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Redfish.
Redfish -- Street Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Redfish.
Redfish API · Guide to Redfish REST API for modern out-of-band server hardware management
Redfish β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Redfish.
Redis Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Redis.
Redis Operations · Guide to Redis in-memory data store operations, persistence, clustering, and debugging
Redis Operations - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Redis.
Redis β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Redis.
Regex & Text Wrangling · Guide to regular expressions for text processing, log analysis, and data extraction
Regex & Text Wrangling - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Regex Text Wrangling.
Regex & Text Wrangling Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Regex Text Wrangling.
Regex & Text Wrangling β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Regex Text Wrangling.
Reliability Patterns · Comprehensive guide to reliability patterns with primers, street ops, and common pitfalls
Reverse Proxy vs Load Balancer · Explains the differences and relationships in Reverse Proxy vs Load Balancer.
RHCE (EX294) β€” Footguns & Pitfalls · Critical mistakes and dangerous pitfalls to avoid when working with Rhce.
RHCE (EX294) β€” Street Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Rhce.
RHCE Exam Prep · Study guide and practice tasks for the Red Hat Certified Engineer (EX294) exam
RHCE β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Rhce.
ripgrep · Guide to ripgrep (rg) for fast recursive text search across codebases and logs
ripgrep (rg) - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Ripgrep.
ripgrep Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Ripgrep.
ripgrep β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Ripgrep.
Routing · Guide to IP routing fundamentals including static routes, dynamic protocols, and troubleshooting
Routing - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Routing.
Routing Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Routing.
Routing β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Routing.
rsync · Guide to rsync for efficient file synchronization, backups, and remote transfers
rsync - Street Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Rsync.
rsync Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Rsync.
rsync β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Rsync.
Runbook Craft · Guide to writing effective runbooks for operational procedures and incident response
Runbook Craft - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Runbook Craft.
Runbook Craft Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Runbook Craft.
Runbook Craft β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Runbook Craft.
Runtime Security with Falco · Guide to Falco runtime security for detecting anomalous behavior in containers and Kubernetes
Runtime Security with Falco Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Falco.
Runtime Security with Falco β€” Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Falco.
S3 & Object Storage β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about S3 Object Storage.
S3-Compatible Object Storage · Guide to S3-compatible object storage with MinIO, lifecycle policies, and access control
S3-Compatible Object Storage Footguns · Critical mistakes and dangerous pitfalls to avoid when working with S3 Object Storage.
S3-Compatible Object Storage β€” Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for S3 Object Storage.
Secrets Management · Guide to secrets management strategies, tools, and rotation for secure infrastructure
Secrets Management - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Secrets Management.
Secrets Management Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Secrets Management.
Secrets Management β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Secrets Management.
Security Basics (Ops-Focused) · Foundational security guide covering least privilege, patching, and defense in depth
Security Basics β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Security Basics.
Security Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Security Basics.
Security Scanning · Guide to security scanning tools and practices for vulnerability detection in infrastructure
Security Scanning - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Security Scanning.
Security Scanning Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Security Scanning.
Security Scanning β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Security Scanning.
Security+ Domain 1: General Security Concepts · SY0-701 Domain 1 practice questions β€” control types, CIA triad, AAA, zero trust, PKI.
Security+ Domain 2: Threats, Vulnerabilities, and Mitigations · SY0-701 Domain 2 practice questions β€” phishing variants, injection attacks, malware, threat actors, mitigations.
Security+ Domain 3: Security Architecture · SY0-701 Domain 3 practice questions β€” network architecture, cloud/data security, resiliency, hardware roots of trust.
Security+ Domain 4: Security Operations · SY0-701 Domain 4 practice questions β€” SIEM/SOAR, IR lifecycle, forensics, IAM, secure protocols. Largest domain at 28% weight.
Security+ Domain 5: Security Program Management and Oversight · SY0-701 Domain 5 practice questions β€” risk management, compliance, governance. IN PROGRESS: only Q121-123 received so far.
sed β€” Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Sed.
sed β€” Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Sed.
sed β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Sed.
sed: The Stream Editor · Guide to sed for stream text editing, substitution patterns, and in-place file modification
SELinux & AppArmor · Guide to SELinux and AppArmor mandatory access control for Linux security
SELinux & AppArmor - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Selinux Apparmor.
SELinux & AppArmor Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Selinux Apparmor.
SELinux & AppArmor β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Selinux Apparmor.
SELinux & Linux Hardening · Guide to Linux security hardening with SELinux, access controls, and audit policies
SELinux & Linux Hardening - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Linux Hardening.
SELinux & Linux Hardening Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Linux Hardening.
Server Hardware · Guide to server hardware components, specifications, and datacenter rack operations
Server Hardware - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Server Hardware.
Server Hardware Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Server Hardware.
Server Hardware β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Server Hardware.
Serverless Computing · Comprehensive guide to serverless computing with primers, street ops, and common pitfalls
Service Mesh · Guide to service mesh architecture, sidecar proxies, and traffic management patterns
Service Mesh - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Service Mesh.
Service Mesh Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Service Mesh.
Service Mesh β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Service Mesh.
Service vs Ingress (Kubernetes Networking) · Explains the differences and relationships in Service vs Ingress (Kubernetes Networking).
SLO Tooling · Guide to SLO measurement tools, error budgets, and reliability tracking dashboards
SLO Tooling Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Slo Tooling.
SLO Tooling β€” Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Slo Tooling.
SLO Tooling β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Slo Tooling.
SQL Fundamentals · Guide to SQL fundamentals for querying, managing, and optimizing databases
SQL Fundamentals Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Sql Fundamentals.
SQL Fundamentals β€” Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Sql Fundamentals.
SQL Fundamentals β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Sql Fundamentals.
SQLite Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Sqlite.
SQLite Operations & Internals · Guide to SQLite database operations, performance tuning, and use cases in infrastructure
SQLite Operations & Internals - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Sqlite.
SQLite β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Sqlite.
SRE Practices · Guide to Site Reliability Engineering practices including toil, SLOs, and error budgets
SRE Practices - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Sre Practices.
SRE Practices Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Sre Practices.
SRE Practices β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Sre Practices.
SSH Deep Dive · In-depth guide to SSH configuration, tunneling, key management, and troubleshooting
SSH Deep Dive β€” Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Ssh Deep Dive.
SSH Deep Dive β€” Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Ssh Deep Dive.
SSH Deep Dive β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Ssh Deep Dive.
Storage (SAN/NAS/DAS) · Overview of storage architectures including SAN, NAS, and DAS for infrastructure planning
Storage Operations · Guide to storage provisioning, monitoring, capacity management, and troubleshooting
Storage Operations - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Storage Ops.
Storage Operations Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Storage Ops.
Storage Ops β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Storage Ops.
Storage Stack: Disk, Partition, LVM, Filesystem, Mount · Technical explainer covering Storage Stack: Disk, Partition, LVM, Filesystem, Mount concepts and practical usage.
STP (Spanning Tree) · Guide to Spanning Tree Protocol for loop prevention in switched networks
STP - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Stp.
STP Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Stp.
STP β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Stp.
strace · Guide to strace for tracing system calls and debugging application behavior on Linux
strace Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Strace.
strace β€” Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Strace.
strace β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Strace.
Street ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Consul.
Street ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Envoy.
Street ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Graphql.
Street ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Message Queues.
Street ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Open Policy Agent.
Subnetting & IP Addressing · Guide to IP subnetting, CIDR notation, and address planning for network engineers
Subnetting & IP Addressing β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Subnetting And Ip Addressing.
Subnetting and IP Addressing - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Subnetting And Ip Addressing.
Subnetting and IP Addressing Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Subnetting And Ip Addressing.
Supply Chain Security · Guide to software supply chain security including SBOMs, signing, and provenance
Supply Chain Security - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Supply Chain Security.
Supply Chain Security Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Supply Chain Security.
Supply Chain Security β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Supply Chain Security.
Synthetic Monitoring · Guide to synthetic monitoring for proactive availability and performance validation
Synthetic Monitoring Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Synthetic Monitoring.
Synthetic Monitoring β€” Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Synthetic Monitoring.
Synthetic Monitoring β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Synthetic Monitoring.
systemctl & journalctl Deep Dive · In-depth guide to systemd service management with systemctl and log analysis with journalctl
systemctl & journalctl Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Systemctl Journalctl.
systemctl & journalctl Street Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Systemctl Journalctl.
systemd Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Linux Ops Systemd.
systemd Street Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Linux Ops Systemd.
Systemd Units: Unit, Service, Target, Start vs Enable · Explains the differences and relationships in Systemd Units: Unit, Service, Target, Start vs Enable.
Systems Thinking Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Systems Thinking.
Systems Thinking for Engineers · Guide to systems thinking for understanding feedback loops and emergent infrastructure behavior
Systems Thinking for Engineers - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Systems Thinking.
Systems Thinking β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Systems Thinking.
Tailscale & Zero Trust Networking · Guide to Tailscale mesh VPN setup, ACLs, and integration with existing infrastructure
Tailscale - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Tailscale.
Tailscale Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Tailscale.
Tailscale β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Tailscale.
tar & Compression · Guide to tar, gzip, bzip2, and other compression tools for archiving and transfers
tar & Compression - Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Tar And Compression.
tar & Compression - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Tar And Compression.
TCP/IP Deep Dive · In-depth guide to TCP/IP protocol stack, socket states, and network troubleshooting
TCP/IP Deep Dive - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Tcp Ip Deep Dive.
TCP/IP Deep Dive Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Tcp Ip Deep Dive.
Terminal Internals · Overview and learning resources for terminal internals
Terminal Internals - Street Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Terminal Internals.
Terminal Internals Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Terminal Internals.
Terminal Internals β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Terminal Internals.
Terraform / IaC · Guide to Terraform infrastructure as code including providers, state, and module patterns
Terraform Deep Dive · Advanced Terraform patterns including state management, modules, and provider development
Terraform Deep Dive - Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Terraform Deep Dive.
Terraform Deep Dive - Street Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Terraform Deep Dive.
Terraform Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Terraform.
Terraform β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Terraform.
Terraform: Desired State Engine · Technical explainer covering Terraform: Desired State Engine concepts and practical usage.
The Ops of AI/ML Workloads · Guide to operating ML infrastructure including GPU scheduling, model serving, and pipelines
The Ops of AI/ML Workloads - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Ai Ml Ops.
The Psychology of Incidents · Guide to human factors, cognitive biases, and stress management during incidents
The Psychology of Incidents - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Incident Psychology.
The Psychology of Incidents Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Incident Psychology.
Thinking Out Loud: Alerting Rules · Exploratory analysis and deeper reasoning about Alerting Rules concepts and trade-offs.
Thinking Out Loud: Ansible · Exploratory analysis and deeper reasoning about Ansible concepts and trade-offs.
Thinking Out Loud: ArgoCD & GitOps · Exploratory analysis and deeper reasoning about Argocd Gitops concepts and trade-offs.
Thinking Out Loud: Containers Deep Dive · Exploratory analysis and deeper reasoning about Containers Deep Dive concepts and trade-offs.
Thinking Out Loud: CrashLoopBackOff · Exploratory analysis and deeper reasoning about Crashloopbackoff concepts and trade-offs.
Thinking Out Loud: Disk & Storage Ops · Exploratory analysis and deeper reasoning about Disk And Storage Ops concepts and trade-offs.
Thinking Out Loud: DNS Ops · Exploratory analysis and deeper reasoning about Dns Ops concepts and trade-offs.
Thinking Out Loud: Docker · Exploratory analysis and deeper reasoning about Docker concepts and trade-offs.
Thinking Out Loud: Git · Exploratory analysis and deeper reasoning about Git concepts and trade-offs.
Thinking Out Loud: GitHub Actions · Exploratory analysis and deeper reasoning about Github Actions concepts and trade-offs.
Thinking Out Loud: Helm · Exploratory analysis and deeper reasoning about Helm concepts and trade-offs.
Thinking Out Loud: Incident Triage · Exploratory analysis and deeper reasoning about Incident Triage concepts and trade-offs.
Thinking Out Loud: Kubernetes Debugging · Exploratory analysis and deeper reasoning about K8S Debugging Playbook concepts and trade-offs.
Thinking Out Loud: Kubernetes Networking · Exploratory analysis and deeper reasoning about K8S Networking concepts and trade-offs.
Thinking Out Loud: Kubernetes Node Lifecycle · Exploratory analysis and deeper reasoning about K8S Node Lifecycle concepts and trade-offs.
Thinking Out Loud: Kubernetes Ops · Exploratory analysis and deeper reasoning about K8S Ops concepts and trade-offs.
Thinking Out Loud: Kubernetes Pods & Scheduling · Exploratory analysis and deeper reasoning about K8S Pods And Scheduling concepts and trade-offs.
Thinking Out Loud: Kubernetes RBAC · Exploratory analysis and deeper reasoning about K8S Rbac concepts and trade-offs.
Thinking Out Loud: Kubernetes Services & Ingress · Exploratory analysis and deeper reasoning about K8S Services And Ingress concepts and trade-offs.
Thinking Out Loud: Kubernetes Storage · Exploratory analysis and deeper reasoning about K8S Storage concepts and trade-offs.
Thinking Out Loud: Linux Hardening · Exploratory analysis and deeper reasoning about Linux Hardening concepts and trade-offs.
Thinking Out Loud: Linux Logging · Exploratory analysis and deeper reasoning about Linux Logging concepts and trade-offs.
Thinking Out Loud: Linux Memory Management · Exploratory analysis and deeper reasoning about Linux Memory Management concepts and trade-offs.
Thinking Out Loud: Linux Ops · Exploratory analysis and deeper reasoning about Linux Ops concepts and trade-offs.
Thinking Out Loud: Linux Ops β€” systemd · Exploratory analysis and deeper reasoning about Linux Ops Systemd concepts and trade-offs.
Thinking Out Loud: Linux Performance · Exploratory analysis and deeper reasoning about Linux Performance concepts and trade-offs.
Thinking Out Loud: Load Balancing · Exploratory analysis and deeper reasoning about Load Balancing concepts and trade-offs.
Thinking Out Loud: Networking Troubleshooting · Exploratory analysis and deeper reasoning about Networking Troubleshooting concepts and trade-offs.
Thinking Out Loud: OpenTelemetry · Exploratory analysis and deeper reasoning about Opentelemetry concepts and trade-offs.
Thinking Out Loud: Process Management · Exploratory analysis and deeper reasoning about Process Management concepts and trade-offs.
Thinking Out Loud: Prometheus Deep Dive · Exploratory analysis and deeper reasoning about Prometheus Deep Dive concepts and trade-offs.
Thinking Out Loud: Secrets Management · Exploratory analysis and deeper reasoning about Secrets Management concepts and trade-offs.
Thinking Out Loud: Security Basics · Exploratory analysis and deeper reasoning about Security Basics concepts and trade-offs.
Thinking Out Loud: SSH Deep Dive · Exploratory analysis and deeper reasoning about Ssh Deep Dive concepts and trade-offs.
Thinking Out Loud: Terraform · Exploratory analysis and deeper reasoning about Terraform concepts and trade-offs.
Thinking Out Loud: TLS Certificates Ops · Exploratory analysis and deeper reasoning about Tls Certificates Ops concepts and trade-offs.
TLS & Certificates Ops · Guide to TLS certificate lifecycle management, chain debugging, and renewal automation
TLS & Certificates Ops - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Tls Certificates Ops.
TLS & Certificates Ops Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Tls Certificates Ops.
TLS & Certificates β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Tls Certificates Ops.
tmux & screen · Overview and learning resources for tmux & screen
tmux & screen - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Tmux And Screen.
tmux & screen Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Tmux And Screen.
Toil Reduction · Comprehensive guide to toil reduction with primers, street ops, and common pitfalls
Tools reference · Quick-reference guide for Networking Troubleshooting Tools - Primer commands and usage.
Topic Packs · Each topic pack contains a primer, footguns, and street ops for a single subject
Topics · Guide to Ansible automation for configuration management, deployments, and orchestration
Tracing · Guide to distributed tracing with Jaeger, Zipkin, and OpenTelemetry for request flow analysis
Tracing β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Tracing.
Trivia · Surprising history, little-known facts, and interesting details about Consul.
Trivia · Surprising history, little-known facts, and interesting details about Graphql.
Trivia · Surprising history, little-known facts, and interesting details about Message Queues.
Trivia compendium · Comprehensive collection of trivia and historical facts across Ansible topics.
Trivia compendium · Comprehensive collection of trivia and historical facts across Linux Ops topics.
Trivia compendium · Comprehensive collection of trivia and historical facts across Python Infra topics.
Vendor Management & Escalation · Guide to vendor management, support escalation, and RMA processes
Vendor Management & Escalation - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Vendor Management.
Vendor Management & Escalation Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Vendor Management.
Vendor Management β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Vendor Management.
Virtualization · Guide to virtualization technologies including KVM, QEMU, and hypervisor management
Virtualization - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Virtualization.
Virtualization Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Virtualization.
Virtualization β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Virtualization.
VLAN Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Vlans.
VLANs · Guide to VLAN configuration, trunking, and inter-VLAN routing for network segmentation
VLANs - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Vlans.
VLANs β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Vlans.
VMware · Core concepts, commands, and practical foundations for Vmware.
VMware - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Vmware.
VMware Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Vmware.
VPN & Tunneling · Guide to VPN protocols, tunnel configuration, and secure remote access architecture
VPN & Tunneling - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Vpn Tunneling.
VPN & Tunneling Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Vpn Tunneling.
VPN & Tunneling β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Vpn Tunneling.
VS Code Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Vscode.
VS Code for DevOps · Guide to VS Code setup with remote SSH, containers, extensions, and infrastructure workflows
VS Code for DevOps - Street Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Vscode.
VS Code β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Vscode.
WebAssembly for Infrastructure · Guide to WebAssembly runtimes for edge computing, plugins, and serverless workloads
WebAssembly for Infrastructure - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Wasm Infrastructure.
WebAssembly for Infrastructure Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Wasm Infrastructure.
WebAssembly Infrastructure β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Wasm Infrastructure.
Wireshark & Packet Analysis · Guide to Wireshark packet capture and network protocol analysis for troubleshooting
Wireshark / tshark / tcpdump - Street-Level Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Wireshark.
Wireshark / tshark / tcpdump Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Wireshark.
Wireshark β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Wireshark.
xargs · Guide to xargs for building and executing commands from standard input
xargs - Footguns & Pitfalls · Critical mistakes and dangerous pitfalls to avoid when working with Xargs.
xargs - Street Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Xargs.
xargs β€” Trivia & Interesting Facts · Surprising history, little-known facts, and interesting details about Xargs.
YAML, JSON & Config Formats · Guide to YAML, JSON, and configuration formats with parsing, validation, and jq processing
YAML, JSON & Config Formats - Footguns · Critical mistakes and dangerous pitfalls to avoid when working with Yaml Json Config.
YAML, JSON & Config Formats - Street Ops · Real-world operational procedures, quick-reference commands, and field-tested workflows for Yaml Json Config.

Library / Lessons (123)

Ansible - The Complete Guide (Revised, Current, Production-Focused) · Complete Ansible guide from fundamentals through production automation patterns for ops and platform engineers.
Ansible one screen interview quick ref · One-page Ansible interview reference covering core model, variable precedence, safe rollout patterns, and debugging.
Ansible Playbook Debugging · Teaches how to diagnose Ansible playbook failures using variable precedence, check mode, diff mode, verbosity levels, and fact inspection.
Ansible The Complete Guide · Comprehensive guide to Ansible from zero to production, covering architecture, inventory, playbooks, modules, roles, Vault, testing, performance tuning, and Tower/AWX.
Ansible: From Playbook to Production · Walks through building a production Ansible workflow covering inventory, roles, Jinja2 templating, Vault, Molecule testing, and zero-downtime rolling updates.
API Gateways: The Front Door to Your Microservices · Explains API gateway architecture including load balancing, reverse proxies, rate limiting, authentication, and observability for microservices platforms.
AWS EC2: The Virtual Server You Never See · Covers EC2 instance types, EBS storage, security groups, instance metadata, spot instances, auto scaling, and troubleshooting unreachable instances.
AWS IAM: The Permissions Puzzle · Deep dive into IAM users, roles, policies, policy evaluation logic, AssumeRole, cross-account access, permission boundaries, and Access Denied debugging.
AWS Lambda: The Function That Runs Itself · Explains AWS Lambda internals including cold starts, VPC networking, concurrency, event-driven architecture, observability, and cost optimization.
AWS RDS: The Managed Database Tradeoff · Covers AWS RDS and Aurora operations including Multi-AZ failover, read replicas, RDS Proxy, Performance Insights, backup/restore, and connection pooling.
AWS VPC: The Network You Can't See · Builds a production VPC from scratch covering subnets, route tables, security groups, NACLs, NAT gateways, VPC peering, Transit Gateway, and flow logs.
Bash: The Patterns That Matter · Teaches production-grade Bash patterns including process management, signal handling, file descriptors, locking, and POSIX compatibility.
BGP: How the Internet Routes Your Packets · Explains BGP routing, autonomous systems, internet architecture, prefix hijacking, and how BGP applies to datacenter and Kubernetes networking.
Capacity Planning: Math Before Midnight · Covers capacity planning using queueing theory, the USE method, Linux performance metrics, Prometheus queries, and Kubernetes resource management.
cgroups and Namespaces: Containers Are a Lie · Builds a container from scratch using Linux namespaces and cgroups to reveal how Docker and container runtimes actually work under the hood.
Chaos Engineering: Breaking Things on Purpose · Introduces chaos engineering principles including fault injection, game days, Chaos Monkey, Litmus, and steady-state hypothesis testing.
Compliance as Code: Automating the Auditor · Covers automating compliance with policy-as-code tools including OPA/Rego, Kyverno, InSpec, AWS Config, and CIS Benchmarks for CI/CD pipelines.
Connection Refused · Systematic guide to diagnosing 'connection refused' errors across networking, firewalls, DNS, processes, containers, Kubernetes services, and systemd.
Container Registries: Where Your Images Actually Live · Explains container registry internals including OCI distribution spec, content-addressable storage, tags vs digests, authentication, garbage collection, and vulnerability scanning.
Deploy a Web App From Nothing · Deploys a Python web app layer by layer from bare process to systemd, nginx, TLS, Docker, Docker Compose, and Kubernetes, explaining why each layer exists.
DNS Operations: When nslookup Isn't Enough · Covers operational DNS debugging beyond nslookup, including Route 53, CoreDNS in Kubernetes, DNSSEC, TTL propagation, and cross-region resolution issues.
Docker Compose: The Local Cluster · Builds a multi-container local development environment with Docker Compose covering networking, DNS resolution, volumes, health checks, and service orchestration.
eBPF: The Linux Superpower · Introduces eBPF for production tracing, networking, security, and observability including Cilium integration and practical use cases.
Elasticsearch Operations · Operational guide to keeping an inherited Elasticsearch cluster alive, covering shard allocation, disk watermarks, mapping explosions, and cluster health.
Envoy: The Proxy That's Everywhere · Explains Envoy proxy architecture, xDS APIs, L7 proxying, circuit breaking, observability, Istio integration, and WASM extensibility.
etcd: The Database That Runs Kubernetes · Deep dive into etcd covering Raft consensus, the Kubernetes control plane dependency, backup/restore procedures, disk performance tuning, and monitoring.
From Init Scripts to systemd · Traces the evolution of Linux process management from SysV init through Upstart to systemd, explaining the design decisions behind each generation.
Git Internals: The Content-Addressable Filesystem · Builds a Git commit from scratch using plumbing commands to understand the object model, packfiles, reflog, merge internals, and garbage collection.
GitHub Actions: CI/CD That Lives in Your Repo · Builds a GitHub Actions CI/CD pipeline from scratch covering linting, testing, Docker builds, OIDC authentication, caching, and supply chain security.
GitOps: The Repo Is the Truth · Explains GitOps principles and ArgoCD architecture including Kubernetes reconciliation, drift detection, Kustomize/Helm integration, and multi-cluster management.
Grafana: Dashboards That Don't Lie · Covers Grafana dashboard design that surfaces real problems, including PromQL for dashboards, panel types, variable templates, alerting, and observability anti-patterns.
How Incident Response Actually Works · Teaches structured incident response including incident command, triage, communication protocols, cognitive bias awareness, and blameless postmortems.
How to Read a Flame Graph · Teaches how to read, generate, and act on flame graphs for CPU profiling using perf, stack traces, and bottleneck identification.
Interview Cheatsheet Ansible · One-page Ansible interview reference covering architecture, modules, inventory, playbooks, roles, Vault, testing, and common interview questions.
Interview Cheatsheet Linux · One-page Linux interview reference covering kernel internals, process management, networking, storage, systemd, and common interview questions.
Interview Cheatsheet Python · One-page Python interview reference covering data structures, OOP, concurrency, error handling, and common DevOps-focused interview questions.
iptables: Following a Packet · Traces a packet through the iptables chain architecture covering tables, chains, rules, NAT, connection tracking, and Kubernetes networking implications.
Kafka: The Log That Runs Everything · Explains Apache Kafka internals including topics, partitions, consumer groups, replication, exactly-once semantics, and operational troubleshooting.
Kubernetes Debugging: When Pods Won't Behave · Systematic guide to debugging Kubernetes pod failures including CrashLoopBackOff, ImagePullBackOff, pending pods, OOMKilled, and networking issues.
Kubernetes Node Lifecycle: From Provision to Decommission · Covers the full Kubernetes node lifecycle from provisioning through cordoning, draining, and decommissioning, including kubelet registration and node conditions.
Kubernetes Services: How Traffic Finds Your Pod · Explains Kubernetes service networking including ClusterIP, NodePort, LoadBalancer, iptables/IPVS rules, DNS resolution, and Ingress controllers.
Kubernetes: From Scratch to Production Upgrade · Walks through building and upgrading a Kubernetes cluster from scratch, covering control plane components, networking, storage, and production upgrade procedures.
Kustomize: Kubernetes Config Without Templates · Covers Kustomize for Kubernetes configuration management without templates, including overlays, patches, strategic merge patches, and integration with ArgoCD.
Linux - Foundations and Operations Guide · Comprehensive Linux guide from boot process through production operations, covering systemd, storage, networking, and triage.
Linux Hardening: Closing the Doors · Covers Linux security hardening including SSH configuration, firewall rules, user management, filesystem permissions, audit logging, and CIS benchmark compliance.
Linux interview quick ref · One-screen Linux interview reference covering boot chain, systemd, processes, filesystems, networking, and troubleshooting.
Linux Networking: Bridges, Bonds, and VLANs · Explains Linux networking primitives including bridges, bonds, VLANs, network namespaces, and how they map to container and Kubernetes networking.
Linux Storage: LVM, Filesystems, and Beyond · Covers Linux storage stack from block devices through LVM, filesystems, RAID, and how storage concepts apply to cloud and Kubernetes persistent volumes.
Linux The Complete Guide · Comprehensive Linux guide covering kernel internals, process management, networking, storage, systemd, security, and performance tuning for DevOps engineers.
Load Testing: Finding the Breaking Point · Teaches load testing methodology including tool selection, test design, result interpretation, bottleneck identification, and capacity planning validation.
Log Pipelines: From printf to Dashboard · Covers log pipeline architecture from application logging through collection, parsing, shipping, indexing, and dashboard visualization.
Make and Makefiles · Explains Make and Makefiles for build automation including targets, dependencies, variables, pattern rules, and real-world DevOps automation patterns.
MongoDB: The Document Database That Surprised Everyone · Covers MongoDB operations including replica sets, sharding, indexing, aggregation pipeline, backup strategies, and production troubleshooting.
MySQL Operations: The Database You Inherited · Operational guide to managing an inherited MySQL database covering replication, backup/restore, query optimization, schema changes, and monitoring.
Nginx: The Swiss Army Server · Nginx architecture, configuration hierarchy, location matching, reverse proxy, TLS termination, and debugging techniques.
OpenTelemetry: Following a Request Across Services · Distributed tracing with OpenTelemetry including context propagation, collector pipelines, and sampling strategies.
Out of Memory · How the Linux OOM killer works, how cgroups and Kubernetes resource limits interact, and how to prevent OOM kills.
Overview · 109 standalone lessons that follow real problems across multiple domains. Designed for
Packer: Building Machine Images That Don't Lie · Building reproducible machine images with Packer, including CI/CD integration, image testing, and golden image pipelines.
Permission Denied · Diagnosing permission errors across Linux file permissions, sudo, capabilities, SELinux/AppArmor, containers, and K8s RBAC.
Prometheus and the Art of Not Alerting · Designing Prometheus alerting that reduces fatigue using SLOs, error budgets, and meaningful alert rules.
Prometheus: Under the Hood · Prometheus internals including TSDB storage engine, PromQL evaluation, cardinality limits, and high availability.
PXE Boot: From Network to Running Server · Network booting bare-metal servers with PXE, DHCP, TFTP, iPXE, and automated OS installation via kickstart/autoinstall.
Python for Infrastructure Automation · Practical Python guide for infrastructure engineers transitioning from Bash to structured automation.
Python for Ops: The Bash Expert's Bridge · Bridging from Bash to Python for ops engineers covering subprocess, pathlib, requests, argparse, and CLI tools.
Python System Automation Deep Dive · Deep dive into Python system automation β€” replacing bash one-liners with robust Python for process management, file ops, user admin, networking, services, and scheduling.
Python The Complete Guide · Complete Python guide from zero to infrastructure automation covering data structures, APIs, testing, and packaging.
Python.for.infrastructure.interview.one.screen · One-screen Python interview reference for infrastructure roles covering core patterns, safe defaults, and common libraries.
Python: Automating Everything β€” APIs and Infrastructure · Building Python automation for AWS, Kubernetes, Prometheus APIs, Slack webhooks, and production CLI tools.
Python: Data Wrangling for Ops · Replacing Bash log parsing with Python using regex, csv, generators, collections, and data pipelines.
Python: Zero to Script for the Terminal Native · Python basics for terminal-native sysadmins, taught through Bash comparisons and practical scripting exercises.
RAID: Why Your Disks Will Fail · RAID levels, rebuild risks, write hole, URE probability, SMART monitoring, and why modern storage is moving beyond RAID.
Redis in Production · Operating Redis in production covering data structures, persistence modes, replication, and memory management gotchas.
S3: The Object Store That Runs the Internet · S3 deep dive covering object storage internals, pricing traps, lifecycle policies, encryption, IAM, and cost optimization.
Secrets Management Without Tears · Managing secrets across env vars, files, HashiCorp Vault, Kubernetes secrets, external-secrets operator, and rotation.
Server Hardware: When the Blinky Lights Matter · Server hardware fundamentals including IPMI/BMC, Redfish API, SMART monitoring, ECC memory, and remote diagnostics.
SLOs: When Good Enough Is a Number · Defining and operating SLOs, SLIs, SLAs, error budgets, burn rate alerting, and SRE practices for production services.
SSH Is More Than You Think · SSH protocol internals including key exchange, authentication methods, tunneling, agent forwarding, and certificates.
strace: Reading the Matrix · Using strace to debug applications by tracing system calls for file I/O, network connections, and process behavior.
Supply Chain Security: Trusting Your Dependencies · Securing the software supply chain with container scanning, SBOMs, image signing, and admission controllers.
systemd: The Init System You Can't Avoid · Deep dive into systemd covering unit files, journald, timers, cgroups, socket activation, and security hardening.
Terraform Modules: Building Infrastructure LEGOs · Designing reusable Terraform modules with versioning, composition patterns, testing, and governance strategies.
Terraform vs Ansible vs Helm · When to use Terraform, Ansible, or Helm based on tool design, state management, and operational boundaries.
Text Processing: jq, awk, and sed in the Trenches · Practical text processing with jq, awk, and sed for log analysis, JSON wrangling, and pipeline composition.
The /proc Filesystem: Linux's Hidden API · Exploring /proc as Linux's kernel API for process inspection, memory maps, network state, and low-level debugging.
The Art of the Postmortem · Writing blameless postmortems that identify contributing factors and produce action items that prevent repeat incidents.
The Art of the Runbook · Designing runbooks that work at 3am with clear steps, tested procedures, and maintainable operational documentation.
The Backup Nobody Tested · Backup strategies, PITR, restore testing, RPO/RTO planning, Velero, and why untested backups are not backups.
The Cascading Timeout · How a single slow dependency cascades into system-wide failure, and patterns to prevent it: circuit breakers and backpressure.
The Cloud Bill Surprise · FinOps fundamentals covering cost allocation, reserved instances, spot pricing, right-sizing, and preventing bill surprises.
The Container Escape · Container security layers including namespaces, capabilities, seccomp profiles, and Docker isolation boundaries.
The Database That Wouldn't Start · Diagnosing and recovering from database startup failures in PostgreSQL and MySQL including WAL, locks, and permissions.
The Disk That Filled Up · Diagnosing full disks caused by log rotation failures, inode exhaustion, Docker storage, and Kubernetes PVC issues.
The Git Disaster Recovery Guide · Recovering from Git disasters using reflog, reset, rebase recovery, bisect, and fsck to find lost commits.
The Hanging Deploy · Diagnosing stuck deploys by understanding processes, signals, systemd lifecycle, bash job control, and container PID 1.
The Kubernetes Migration That Took a Year · Migration planning lessons from moving to Kubernetes: strangler fig pattern, blast radius control, and realistic timelines.
The Load Balancer Lied · Why load balancer health checks pass while applications fail, and how to design checks that reflect real health.
The Monitoring That Lied · How monitoring deceives through metric lag, counter resets, percentile math errors, and stale scrapes.
The Mysterious Latency Spike · Systematic diagnosis of latency spikes across CPU, disk I/O, garbage collection, network, and noisy neighbor causes.
The Nginx Config That Broke Everything · Common Nginx misconfigurations that pass syntax checks but break routing, and how to debug them systematically.
The Rollback That Wasn't · Why rollbacks fail due to database migrations, config drift, and state changes that outlive the deploy.
The Service Mesh Tax · What service meshes actually do, the operational cost of Istio/Envoy sidecars, and when a mesh is not worth it.
The Split-Brain Nightmare · Distributed consensus failures, network partitions, quorum mechanics, CAP theorem, and fencing strategies.
The Subnet Calculator in Your Head · Mental math for CIDR notation, subnet masks, network/broadcast addresses, and IP planning without a calculator.
The Terraform State Disaster · Terraform state failures including locking, drift, corruption, recovery, and remote backend best practices.
tmux and the Terminal · How terminals, PTYs, and tmux actually work, and why tmux keeps processes alive after SSH disconnects.
Understanding Distributed Systems Without a PhD · Practical guide to CAP theorem, consensus, eventual consistency, replication, and CRDTs for working engineers.
Vault: Secrets That Expire on Purpose · HashiCorp Vault for dynamic credentials, Kubernetes integration, PKI, and secrets that auto-expire by design.
What Happens Inside a Linux Pipe · How Linux pipes work internally through file descriptors, kernel buffers, process scheduling, and backpressure.
What Happens When Kubernetes Evicts Your Pod · Kubernetes pod eviction mechanics including node pressure, eviction signals, QoS classes, and graceful termination.
What Happens When You `docker build` · What docker build does step by step: context, layers, overlay filesystem, build cache, multi-stage builds, and registries.
What Happens When You `git push` to CI · Tracing a git push through hooks, webhooks, CI runners, build pipelines, artifact publishing, and deploy triggers.
What Happens When You `helm install` · What helm install does internally: chart rendering, template evaluation, values merging, hooks, and release management.
What Happens When You `kubectl apply` · Tracing kubectl apply through the API server, etcd, scheduler, kubelet, container runtime, and networking setup.
What Happens When You Click a Link · Full request lifecycle from browser click through DNS, TCP/IP, TLS, HTTP, load balancing, and server response.
What Happens When You Press Power · The Linux boot sequence from power button through BIOS/UEFI, GRUB, kernel, initramfs, and systemd to login prompt.
What Happens When You Type a Regex · How regex engines work internally: NFA/DFA execution, backtracking, catastrophic regex, and performance tradeoffs.
What Happens When Your Certificate Expires · TLS certificate expiry impact, cert-manager automation, ACME/Let's Encrypt, HSTS traps, and expiry monitoring.
When the Queue Backs Up · Diagnosing message queue backpressure in Kafka and RabbitMQ: consumer lag, dead letter queues, and delivery guarantees.
Why DNS Is Always the Problem · DNS from history through modern resolution hierarchy, caching layers, DNSSEC, Kubernetes DNS, and systematic debugging.
Why Everything Uses JSON Now · History and tradeoffs of data serialization formats: SGML, XML, JSON, YAML, Protocol Buffers, and MessagePack.
Why NTP Matters More Than You Think · How clock skew breaks TLS, Kafka, distributed locks, and 2FA, and how to keep time correct with NTP and chrony.
Why YAML Keeps Breaking Your Deploys · YAML's implicit typing footguns, why they exist, and how to defend against them in Helm, Kustomize, and CI configs.

Library / Deep Dives (25)

_Official Sources · This bundle was written against official or primary documentation so the explanations stay close to how the systems actually behave.
Deep Dive: AWS VPC Internals · This document explains AWS VPC the way an infrastructure engineer should understand it. · Topics: Routing, Linux Networking Tools · Aliases: asymmetric-routing, bgp, ethtool, gateway, ip-command, mtr, nmcli, ospf, routes, ss, static-routes, tcpdump
Deep Dive: CI/CD Pipeline Architecture · This document explains CI/CD pipeline architecture from the systems point of view rather than from vendor slogans. · Topics: CI/CD, GitHub Actions · Aliases: actions-runner, ci-github, circleci, continuous-delivery, continuous-integration, gha, github-actions, jenkins, pipelines, workflows, zuul
Deep Dive: Containers How They Really Work · This document explains containers as Linux mechanisms, not as marketing. · Topics: Docker / Containers, Container Runtimes · Aliases: cgroups, container-runtime-debug, containerd, containers, cri-o, docker-compose, dockerfile, images, namespaces, runc
Deep Dive: Dell Linux PowerEdge · Someone with real experience is not just β€œa Linux admin who used a Dell once.” They usually have practical familiarity with all four layers below. · Topics: Server Hardware, Out-of-Band Management · Aliases: bare-metal-provisioning, bmc, cpu, datacenter-provisioning, dimm, dmesg, dmidecode, hba, idrac, ilo, ipmi, lshw
Deep Dive: Docker Image Internals · This document explains Docker image internals and the related OCI image model. · Topics: Docker / Containers, Container Image Optimization (alias β†’ container_images) · Aliases: containers, distroless, docker-compose, docker-slim, dockerfile, image-security, images, layer-caching, multi-stage, slim-images
Deep Dive: eBPF Explained · This document explains eBPF as a Linux capability you can actually reason about, instead of mystical conference smoke. · Topics: eBPF, Linux Networking Tools · Aliases: bcc-tools, bpftrace, ebpf-observability, ethtool, execsnoop, ip-command, kernel-tracing, mtr, nmcli, ss, tcpdump, tcplife
Deep Dive: Kubernetes Networking · This document explains Kubernetes networking from the practical internal view. · Topics: Kubernetes Networking, Linux Networking Tools · Aliases: cni, coredns, ethtool, ingress, ip-command, mtr, network-policies, nmcli, service-mesh, ss, tcpdump, traceroute
Deep Dive: Kubernetes Pod Lifecycle · This document explains what really happens from 'I applied YAML' to 'my Pod is running' and then to termination. · Topics: Kubernetes Core · Aliases: deployments, k8s, kubernetes, openshift, pods, services
Deep Dive: Kubernetes Scheduler · This document explains the Kubernetes scheduler as a real control-plane component, not just 'the thing that picks a node.'. · Topics: Kubernetes Core · Aliases: deployments, k8s, kubernetes, openshift, pods, services
Deep Dive: Linux Boot Sequence · This document walks the Linux boot path from power applied to a usable system. · Topics: Linux Fundamentals, systemd · Aliases: filesystem, journalctl, linux, linux-basics, permissions, services, systemctl, units
Deep Dive: Linux Filesystem Internals · This document explains the filesystem stack as it matters to Linux administrators and DevOps engineers. · Topics: Filesystems & Storage, Linux Fundamentals · Aliases: btrfs, df, disk-debug, disk-troubleshooting, ext4, file-systems, filesystem, fstab, inode-exhaustion, inodes, io-troubleshooting, iostat
Deep Dive: Linux Memory Management · This document explains Linux memory management from the point of view of an admin / DevOps / performance engineer. · Topics: Linux Fundamentals · Aliases: filesystem, linux, linux-basics, permissions
Deep Dive: Linux Network Packet Flow · This document explains what happens to a packet on Linux. · Topics: Linux Networking Tools, Packet Path · Aliases: ethtool, ip-command, mtr, network-path, network-troubleshooting, nmcli, packet-flow, ss, tcpdump, traceroute
Deep Dive: Linux Performance Debugging · This document explains how to debug Linux performance issues systematically. · Topics: Linux Fundamentals, Filesystems & Storage · Aliases: btrfs, df, disk-debug, disk-troubleshooting, ext4, file-systems, filesystem, fstab, inode-exhaustion, inodes, io-troubleshooting, iostat
Deep Dive: Linux Process Scheduler · This document explains Linux scheduling from the viewpoint of operations, performance, and interview readiness. · Topics: Linux Fundamentals · Aliases: filesystem, linux, linux-basics, permissions
Deep Dive: RAID and Storage Internals · This document explains RAID and adjacent Linux storage concepts for operational understanding. · Topics: RAID, Storage (SAN/NAS/DAS) · Aliases: das, disk-arrays, iscsi, megaraid, minio, nas, nfs-storage, portworx, raid-levels, raid-rebuild, san, smart
Deep Dive: Systemd Architecture · This document explains systemd as the init system and service manager that brings a modern Linux system to life after the kernel hands off to PID 1. · Topics: systemd, Linux Fundamentals · Aliases: filesystem, journalctl, linux, linux-basics, permissions, services, systemctl, units
Deep Dive: Systemd Service Design Debugging and Hardening · This document focuses on .service units as they are used in real servers. · Topics: systemd, Linux Hardening · Aliases: cis-benchmark, journalctl, os-hardening, services, system-hardening, systemctl, units
Deep Dive: Systemd Timers Journald Cgroups and Resource Control · [Service] Type=oneshot ExecStart=/usr/local/sbin/do-backup # backup.timer [Timer] OnCalendar=daily Persistent=true journalctl -b journalctl -u. · Topics: systemd, Linux Fundamentals · Aliases: filesystem, journalctl, linux, linux-basics, permissions, services, systemctl, units
Deep Dive: Systemd Units Dependencies and Ordering · This document is the hard part people often fake their way through. · Topics: systemd · Aliases: journalctl, services, systemctl, units
Deep Dive: TCP/IP Deep Dive · Networking is not 'the app talks to the wire.'. · Topics: TCP/IP, Linux Networking Tools · Aliases: ethtool, ip, ip-command, layers, mtr, networking, networking-fundamentals, nmcli, osi-model, ss, tcp, tcpdump
Deep Dive: Terraform State Internals · This document explains Terraform state as a mechanism, not as an annoying file you commit by mistake. · Topics: Terraform · Aliases: hcl, iac, infrastructure-as-code, tfstate
Deep Dive: TLS Handshake · This document focuses on modern TLS, especially TLS 1.3 behavior and the operational concepts you actually need. · Topics: TLS & PKI, TCP/IP · Aliases: ca-chain, cert-manager, cert-rotation, certificates, ip, layers, networking, networking-fundamentals, osi-model, pki, ssl, tcp
DevOps Deep Dives · Focused Linux / DevOps reference documents for interview prep, systems understanding, and production troubleshooting.

Library / Case Studies (547)

Answer Key: The 5% That Can't Resolve · Full solution with system reconstruction and diagnosis for the 10 intermittent dns exercise.
Answer Key: The Alerts That Stopped Firing · Full solution with system reconstruction and diagnosis for the 09 monitoring gap exercise.
Answer Key: The Certificate That Works Sometimes · Full solution with system reconstruction and diagnosis for the 12 tls chain incomplete exercise.
Answer Key: The Cluster That Disagrees With Itself · Full solution with system reconstruction and diagnosis for the 14 split brain etcd exercise.
Answer Key: The Container That Exits Immediately · Full solution with system reconstruction and diagnosis for the 03 docker exec format error exercise.
Answer Key: The Deploy That Didn't Deploy · Full solution with system reconstruction and diagnosis for the 07 stale image tag exercise.
Answer Key: The DR That Looks Ready But Isn't · Full solution with system reconstruction and diagnosis for the 11 dr failover broken exercise.
Answer Key: The Gateway That Returns 502 · Full solution with system reconstruction and diagnosis for the 05 nginx 502 backend exercise.
Answer Key: The Job That Succeeded Wrong · Full solution with system reconstruction and diagnosis for the 08 job wrong database exercise.
Answer Key: The Pods That Won't Schedule · Full solution with system reconstruction and diagnosis for the 13 quota masquerade exercise.
Answer Key: The Replica That Fell Behind · Full solution with system reconstruction and diagnosis for the 04 postgres replica lag exercise.
Answer Key: The Requests That Vanish · Full solution with system reconstruction and diagnosis for the 06 mesh silent drops exercise.
Answer Key: The Service That Won't Start · Full solution with system reconstruction and diagnosis for the 02 systemd permission denied exercise.
Answer Key: The Session Store That Keeps Dying · Full solution with system reconstruction and diagnosis for the 01 redis oom crashloop exercise.
Answer Key: The Slow Death Nobody Noticed · Full solution with system reconstruction and diagnosis for the 15 cascading debug logs exercise.
Case Studies · Incident case studies organized by domain
Case Study: Alert Storm β€” Flapping Health Checks · Overview and file index for the alert storm flapping healthchecks cross-domain incident case study.
Case Study: Ansible Playbook Hangs β€” SSH Agent Forwarding Blocked by Firewall · Overview and file index for the ansible ssh agent firewall cross-domain incident case study.
Case Study: API Latency Spike β€” BGP Route Leak, Fix Is Network ACL · Overview and file index for the api latency bgp route leak acl cross-domain incident case study.
Case Study: ARP Flux Duplicate IP · Overview and file index for the arp flux duplicate ip networking case study.
Case Study: Asymmetric Routing One Direction · Overview and file index for the asymmetric routing one direction networking case study.
Case Study: Backup Job Failing β€” iSCSI Target Unreachable, VLAN Misconfigured · Overview and file index for the backup job iscsi vlan cross-domain incident case study.
Case Study: BGP Peer Flapping · Overview and file index for the bgp peer flapping networking case study.
Case Study: BIOS Settings Reset After CMOS · Overview and file index for the bios settings reset after cmos datacenter operations case study.
Case Study: BMC Clock Skew Cert Failure · Overview and file index for the bmc clock skew cert failure datacenter operations case study.
Case Study: Bonding Failover Not Working · Overview and file index for the bonding failover not working datacenter operations case study.
Case Study: Cable Management Wrong Port · Overview and file index for the cable management wrong port datacenter operations case study.
Case Study: Canary Deploy Routing to Wrong Backend β€” Ingress Misconfigured · Overview and file index for the canary deploy wrong backend ingress cross-domain incident case study.
Case Study: CI Pipeline Fails β€” Docker Layer Cache Corruption · Overview and file index for the ci pipeline docker cache registry cross-domain incident case study.
Case Study: CNI Broken After Restart · Overview and file index for the cni broken after restart Kubernetes operations case study.
Case Study: Container Vuln Scanner False Positive Blocks Deploy · Overview and file index for the container vuln scanner false positive cross-domain incident case study.
Case Study: CoreDNS Timeout Pod DNS · Overview and file index for the coredns timeout pod dns Kubernetes operations case study.
Case Study: CrashLoopBackOff No Logs · Overview and file index for the crashloopbackoff no logs Kubernetes operations case study.
Case Study: DaemonSet Blocks Eviction · Overview and file index for the daemonset blocks eviction Kubernetes operations case study.
Case Study: Database Replication Lag β€” Root Cause Is RAID Degradation · Overview and file index for the database replication lag raid cross-domain incident case study.
Case Study: Deployment Stuck β€” ImagePull Auth Failure, Vault Secret Rotation · Overview and file index for the deployment stuck imagepull vault cross-domain incident case study.
Case Study: DHCP Relay Broken · Overview and file index for the dhcp relay broken networking case study.
Case Study: Disk Full Root Services Down · Overview and file index for the disk full root services down datacenter operations case study.
Case Study: Disk Full β€” Runaway Logs, Fix Is Loki Retention · Overview and file index for the disk full runaway logs loki cross-domain incident case study.
Case Study: DNS Looks Broken β€” TLS Expired, Fix Is Cert-Manager · Overview and file index for the dns tls certmanager cross-domain incident case study.
Case Study: DNS Resolution Slow · Overview and file index for the dns resolution slow networking case study.
Case Study: DNS Split Horizon Confusion · Overview and file index for the dns split horizon confusion networking case study.
Case Study: Drain Blocked by PDB · Overview and file index for the drain blocked by pdb Kubernetes operations case study.
Case Study: Duplex Mismatch Symptoms · Overview and file index for the duplex mismatch symptoms networking case study.
Case Study: Firewall Shadow Rule · Overview and file index for the firewall shadow rule networking case study.
Case Study: Firmware Update Boot Loop · Overview and file index for the firmware update boot loop datacenter operations case study.
Case Study: Grafana Dashboard Empty β€” Prometheus Blocked by NetworkPolicy · Overview and file index for the grafana empty prometheus networkpolicy cross-domain incident case study.
Case Study: HBA Firmware Mismatch · Overview and file index for the hba firmware mismatch datacenter operations case study.
Case Study: HPA Flapping β€” Metrics Server Clock Skew, Fix Is NTP · Overview and file index for the hpa flapping clock skew ntp cross-domain incident case study.
Case Study: iDRAC Unreachable OS Up · Overview and file index for the idrac unreachable os up datacenter operations case study.
Case Study: ImagePullBackOff Registry Auth · Overview and file index for the imagepullbackoff registry auth Kubernetes operations case study.
Case Study: Inode Exhaustion · Overview and file index for the inode exhaustion Linux operations case study.
Case Study: IPTables Blocking Unexpected · Overview and file index for the iptables blocking unexpected Linux operations case study.
Case Study: Job Queue Backlog β€” Worker Pod CPU Throttled by cgroup · Overview and file index for the job queue cpu throttle cgroup cross-domain incident case study.
Case Study: Jumbo Frames Partial · Overview and file index for the jumbo frames partial networking case study.
Case Study: Kernel Soft Lockup · Overview and file index for the kernel soft lockup Linux operations case study.
Case Study: LACP Mismatch One Link Hot · Overview and file index for the lacp mismatch one link hot networking case study.
Case Study: Link Flaps Bad Optic · Overview and file index for the link flaps bad optic datacenter operations case study.
Case Study: Memory ECC Errors Increasing · Overview and file index for the memory ecc errors increasing datacenter operations case study.
Case Study: MTU Blackhole TLS Stalls · Overview and file index for the mtu blackhole tls stalls networking case study.
Case Study: Multicast Not Crossing Router · Overview and file index for the multicast not crossing router networking case study.
Case Study: NAT Exhaustion Intermittent · Overview and file index for the nat exhaustion intermittent networking case study.
Case Study: Network Loop Broadcast Storm · Overview and file index for the network loop broadcast storm networking case study.
Case Study: Node NotReady β€” NIC Firmware Bug, Fix Is Ansible Playbook · Overview and file index for the node notready nic firmware ansible cross-domain incident case study.
Case Study: Node Pressure Evictions · Overview and file index for the node pressure evictions Kubernetes operations case study.
Case Study: NVMe Drive Disappeared · Overview and file index for the nvme drive disappeared datacenter operations case study.
Case Study: OOM Killer Events · Overview and file index for the oom killer events Linux operations case study.
Case Study: OS Install Fails RAID Controller · Overview and file index for the os install fails raid controller datacenter operations case study.
Case Study: OSPF Stuck In Exstart · Overview and file index for the ospf stuck in exstart networking case study.
Case Study: Persistent Volume Stuck Terminating · Overview and file index for the persistent volume stuck terminating Kubernetes operations case study.
Case Study: Pod OOMKilled β€” Memory Leak in Sidecar, Fix Is Helm Values · Overview and file index for the pod oomkilled sidecar helm cross-domain incident case study.
Case Study: Power Supply Redundancy Lost · Overview and file index for the power supply redundancy lost datacenter operations case study.
Case Study: Proxy ARP Causing Issues · Overview and file index for the proxy arp causing issues networking case study.
Case Study: PXE Boot Fails UEFI Mismatch · Overview and file index for the pxe boot fails uefi mismatch datacenter operations case study.
Case Study: Rack PDU Overload Alert · Overview and file index for the rack pdu overload alert datacenter operations case study.
Case Study: RAID Degraded Rebuild Latency · Overview and file index for the raid degraded rebuild latency datacenter operations case study.
Case Study: Resource Quota Blocking Deploy · Overview and file index for the resource quota blocking deploy Kubernetes operations case study.
Case Study: Runaway Logs Fill Disk · Overview and file index for the runaway logs fill disk Linux operations case study.
Case Study: SELinux Denying Service · Overview and file index for the selinux denying service Linux operations case study.
Case Study: Serial Console Garbled · Overview and file index for the serial console garbled datacenter operations case study.
Case Study: Server Intermittent Reboot · Overview and file index for the server intermittent reboot datacenter operations case study.
Case Study: Server Remote Console Lag · Overview and file index for the server remote console lag datacenter operations case study.
Case Study: Service Mesh 503s β€” Envoy Misconfigured, RBAC Policy · Overview and file index for the service mesh 503 envoy rbac cross-domain incident case study.
Case Study: Service No Endpoints · Overview and file index for the service no endpoints Kubernetes operations case study.
Case Study: Source Routing Policy Miss · Overview and file index for the source routing policy miss networking case study.
Case Study: SSH Timeout β€” MTU Mismatch, Fix Is Terraform Variable · Overview and file index for the ssh timeout mtu terraform cross-domain incident case study.
Case Study: SSL Cert Chain Incomplete · Overview and file index for the ssl cert chain incomplete networking case study.
Case Study: Stuck NFS Mount · Overview and file index for the stuck nfs mount Linux operations case study.
Case Study: Systemd Service Flapping · Overview and file index for the systemd service flapping Linux operations case study.
Case Study: TCP RST After Idle · Overview and file index for the tcp rst after idle networking case study.
Case Study: Terraform Apply Fails β€” State Lock Stuck, DynamoDB Throttle · Overview and file index for the terraform state lock dynamodb cross-domain incident case study.
Case Study: Thermal Throttle Fan Failure · Overview and file index for the thermal throttle fan failure datacenter operations case study.
Case Study: Time Sync Skew Breaks App · Overview and file index for the time sync skew breaks app Linux operations case study.
Case Study: User Auth Failing β€” OIDC Cert Expired, Cloud KMS Rotation · Overview and file index for the user auth oidc cert kms cross-domain incident case study.
Case Study: VLAN Trunk Mistag · Overview and file index for the vlan trunk mistag networking case study.
Case Study: Zombie Processes Accumulating · Overview and file index for the zombie processes accumulating Linux operations case study.
Cross-Domain Incident Case Studies · Index of cross-domain incident case studies with incident matrix and navigation.
Datacenter Operations Case Studies · Index of datacenter operations case studies with incident matrix and navigation.
Diagnostic Questions · Diagnostic questions to guide troubleshooting of the alert storm flapping healthchecks incident.
Diagnostic Questions · Diagnostic questions to guide troubleshooting of the ansible ssh agent firewall incident.
Diagnostic Questions · Diagnostic questions to guide troubleshooting of the api latency bgp route leak acl incident.
Diagnostic Questions · Diagnostic questions to guide troubleshooting of the backup job iscsi vlan incident.
Diagnostic Questions · Diagnostic questions to guide troubleshooting of the canary deploy wrong backend ingress incident.
Diagnostic Questions · Diagnostic questions to guide troubleshooting of the ci pipeline docker cache registry incident.
Diagnostic Questions · Diagnostic questions to guide troubleshooting of the container vuln scanner false positive incident.
Diagnostic Questions · Diagnostic questions to guide troubleshooting of the database replication lag raid incident.
Diagnostic Questions · Diagnostic questions to guide troubleshooting of the deployment stuck imagepull vault incident.
Diagnostic Questions · Diagnostic questions to guide troubleshooting of the disk full runaway logs loki incident.
Diagnostic Questions · Diagnostic questions to guide troubleshooting of the dns tls certmanager incident.
Diagnostic Questions · Diagnostic questions to guide troubleshooting of the grafana empty prometheus networkpolicy incident.
Diagnostic Questions · Diagnostic questions to guide troubleshooting of the hpa flapping clock skew ntp incident.
Diagnostic Questions · Diagnostic questions to guide troubleshooting of the job queue cpu throttle cgroup incident.
Diagnostic Questions · Diagnostic questions to guide troubleshooting of the node notready nic firmware ansible incident.
Diagnostic Questions · Diagnostic questions to guide troubleshooting of the pod oomkilled sidecar helm incident.
Diagnostic Questions · Diagnostic questions to guide troubleshooting of the service mesh 503 envoy rbac incident.
Diagnostic Questions · Diagnostic questions to guide troubleshooting of the ssh timeout mtu terraform incident.
Diagnostic Questions · Diagnostic questions to guide troubleshooting of the terraform state lock dynamodb incident.
Diagnostic Questions · Diagnostic questions to guide troubleshooting of the user auth oidc cert kms incident.
Grading Checklist · Self-assessment rubric and scoring criteria for the bios settings reset after cmos case study.
Grading Checklist · Self-assessment rubric and scoring criteria for the bonding failover not working case study.
Grading Checklist · Self-assessment rubric and scoring criteria for the cable management wrong port case study.
Grading Checklist · Self-assessment rubric and scoring criteria for the hba firmware mismatch case study.
Grading Checklist · Self-assessment rubric and scoring criteria for the nvme drive disappeared case study.
Grading Checklist · Self-assessment rubric and scoring criteria for the os install fails raid controller case study.
Grading Checklist · Self-assessment rubric and scoring criteria for the rack pdu overload alert case study.
Grading Checklist · Self-assessment rubric and scoring criteria for the serial console garbled case study.
Grading Checklist · Self-assessment rubric and scoring criteria for the server intermittent reboot case study.
Grading Checklist · Self-assessment rubric and scoring criteria for the server remote console lag case study.
Grading Checklist · Self-assessment rubric and scoring criteria for the cni broken after restart case study.
Grading Checklist · Self-assessment rubric and scoring criteria for the coredns timeout pod dns case study.
Grading Checklist · Self-assessment rubric and scoring criteria for the crashloopbackoff no logs case study.
Grading Checklist · Self-assessment rubric and scoring criteria for the daemonset blocks eviction case study.
Grading Checklist · Self-assessment rubric and scoring criteria for the drain blocked by pdb case study.
Grading Checklist · Self-assessment rubric and scoring criteria for the imagepullbackoff registry auth case study.
Grading Checklist · Self-assessment rubric and scoring criteria for the node pressure evictions case study.
Grading Checklist · Self-assessment rubric and scoring criteria for the persistent volume stuck terminating case study.
Grading Checklist · Self-assessment rubric and scoring criteria for the resource quota blocking deploy case study.
Grading Checklist · Self-assessment rubric and scoring criteria for the service no endpoints case study.
Grading Checklist · Self-assessment rubric and scoring criteria for the inode exhaustion case study.
Grading Checklist · Self-assessment rubric and scoring criteria for the iptables blocking unexpected case study.
Grading Checklist · Self-assessment rubric and scoring criteria for the kernel soft lockup case study.
Grading Checklist · Self-assessment rubric and scoring criteria for the oom killer events case study.
Grading Checklist · Self-assessment rubric and scoring criteria for the runaway logs fill disk case study.
Grading Checklist · Self-assessment rubric and scoring criteria for the selinux denying service case study.
Grading Checklist · Self-assessment rubric and scoring criteria for the stuck nfs mount case study.
Grading Checklist · Self-assessment rubric and scoring criteria for the systemd service flapping case study.
Grading Checklist · Self-assessment rubric and scoring criteria for the time sync skew breaks app case study.
Grading Checklist · Self-assessment rubric and scoring criteria for the zombie processes accumulating case study.
Grading Checklist · Self-assessment rubric and scoring criteria for the arp flux duplicate ip case study.
Grading Checklist · Self-assessment rubric and scoring criteria for the asymmetric routing one direction case study.
Grading Checklist · Self-assessment rubric and scoring criteria for the bgp peer flapping case study.
Grading Checklist · Self-assessment rubric and scoring criteria for the dns split horizon confusion case study.
Grading Checklist · Self-assessment rubric and scoring criteria for the duplex mismatch symptoms case study.
Grading Checklist · Self-assessment rubric and scoring criteria for the firewall shadow rule case study.
Grading Checklist · Self-assessment rubric and scoring criteria for the lacp mismatch one link hot case study.
Grading Checklist · Self-assessment rubric and scoring criteria for the mtu blackhole tls stalls case study.
Grading Checklist · Self-assessment rubric and scoring criteria for the nat exhaustion intermittent case study.
Grading Checklist · Self-assessment rubric and scoring criteria for the vlan trunk mistag case study.
Grading Checklist: BMC Clock Skew - Certificate Failure · Self-assessment rubric and scoring criteria for the bmc clock skew cert failure case study.
Grading Checklist: DHCP Not Working on Remote VLAN · Self-assessment rubric and scoring criteria for the dhcp relay broken case study.
Grading Checklist: Disk Full Root - Services Down · Self-assessment rubric and scoring criteria for the disk full root services down case study.
Grading Checklist: DNS Resolution Taking 5+ Seconds Intermittently · Self-assessment rubric and scoring criteria for the dns resolution slow case study.
Grading Checklist: Firmware Update Boot Loop · Self-assessment rubric and scoring criteria for the firmware update boot loop case study.
Grading Checklist: iDRAC Unreachable, OS Up · Self-assessment rubric and scoring criteria for the idrac unreachable os up case study.
Grading Checklist: Jumbo Frames Enabled But Some Paths Failing · Self-assessment rubric and scoring criteria for the jumbo frames partial case study.
Grading Checklist: Link Flaps - Bad Optic · Self-assessment rubric and scoring criteria for the link flaps bad optic case study.
Grading Checklist: Memory ECC Errors Increasing · Self-assessment rubric and scoring criteria for the memory ecc errors increasing case study.
Grading Checklist: Multicast Traffic Not Crossing Router · Self-assessment rubric and scoring criteria for the multicast not crossing router case study.
Grading Checklist: Network Experiencing Broadcast Storm and High CPU on Switches · Self-assessment rubric and scoring criteria for the network loop broadcast storm case study.
Grading Checklist: OSPF Adjacency Stuck in ExStart/Exchange State · Self-assessment rubric and scoring criteria for the ospf stuck in exstart case study.
Grading Checklist: Power Supply Redundancy Lost · Self-assessment rubric and scoring criteria for the power supply redundancy lost case study.
Grading Checklist: Proxy ARP Causing Unexpected Routing Behavior · Self-assessment rubric and scoring criteria for the proxy arp causing issues case study.
Grading Checklist: PXE Boot Fails - UEFI Mismatch · Self-assessment rubric and scoring criteria for the pxe boot fails uefi mismatch case study.
Grading Checklist: RAID Degraded Rebuild Latency · Self-assessment rubric and scoring criteria for the raid degraded rebuild latency case study.
Grading Checklist: TCP Connections Reset After Idle Period · Self-assessment rubric and scoring criteria for the tcp rst after idle case study.
Grading Checklist: Thermal Throttle - Fan Failure · Self-assessment rubric and scoring criteria for the thermal throttle fan failure case study.
Grading Checklist: TLS Works From Some Clients But Fails From Others · Self-assessment rubric and scoring criteria for the ssl cert chain incomplete case study.
Grading Checklist: Traffic From Specific Source Not Taking Expected Path · Self-assessment rubric and scoring criteria for the source routing policy miss case study.
Grading Rubric · Self-assessment rubric and scoring criteria for the alert storm flapping healthchecks case study.
Grading Rubric · Self-assessment rubric and scoring criteria for the ansible ssh agent firewall case study.
Grading Rubric · Self-assessment rubric and scoring criteria for the api latency bgp route leak acl case study.
Grading Rubric · Self-assessment rubric and scoring criteria for the backup job iscsi vlan case study.
Grading Rubric · Self-assessment rubric and scoring criteria for the canary deploy wrong backend ingress case study.
Grading Rubric · Self-assessment rubric and scoring criteria for the ci pipeline docker cache registry case study.
Grading Rubric · Self-assessment rubric and scoring criteria for the container vuln scanner false positive case study.
Grading Rubric · Self-assessment rubric and scoring criteria for the database replication lag raid case study.
Grading Rubric · Self-assessment rubric and scoring criteria for the deployment stuck imagepull vault case study.
Grading Rubric · Self-assessment rubric and scoring criteria for the disk full runaway logs loki case study.
Grading Rubric · Self-assessment rubric and scoring criteria for the dns tls certmanager case study.
Grading Rubric · Self-assessment rubric and scoring criteria for the grafana empty prometheus networkpolicy case study.
Grading Rubric · Self-assessment rubric and scoring criteria for the hpa flapping clock skew ntp case study.
Grading Rubric · Self-assessment rubric and scoring criteria for the job queue cpu throttle cgroup case study.
Grading Rubric · Self-assessment rubric and scoring criteria for the node notready nic firmware ansible case study.
Grading Rubric · Self-assessment rubric and scoring criteria for the pod oomkilled sidecar helm case study.
Grading Rubric · Self-assessment rubric and scoring criteria for the service mesh 503 envoy rbac case study.
Grading Rubric · Self-assessment rubric and scoring criteria for the ssh timeout mtu terraform case study.
Grading Rubric · Self-assessment rubric and scoring criteria for the terraform state lock dynamodb case study.
Grading Rubric · Self-assessment rubric and scoring criteria for the user auth oidc cert kms case study.
Incident Replay: ARP Flux β€” Duplicate IP Detection · Interactive incident replay with decision points for the arp flux duplicate ip scenario.
Incident Replay: Asymmetric Routing β€” Traffic Works One Direction Only · Interactive incident replay with decision points for the asymmetric routing one direction scenario.
Incident Replay: BGP Peer Flapping · Interactive incident replay with decision points for the bgp peer flapping scenario.
Incident Replay: BIOS Settings Reverted After CMOS Battery Replacement · Interactive incident replay with decision points for the bios settings reset after cmos scenario.
Incident Replay: BMC Clock Skew Causes Certificate Failure · Interactive incident replay with decision points for the bmc clock skew cert failure scenario.
Incident Replay: Cable Plugged Into Wrong Port · Interactive incident replay with decision points for the cable management wrong port scenario.
Incident Replay: CNI Broken After Node Restart · Interactive incident replay with decision points for the cni broken after restart scenario.
Incident Replay: CoreDNS Timeout β€” Pod DNS Resolution Failing · Interactive incident replay with decision points for the coredns timeout pod dns scenario.
Incident Replay: CrashLoopBackOff with No Logs · Interactive incident replay with decision points for the crashloopbackoff no logs scenario.
Incident Replay: DaemonSet Blocks Node Eviction · Interactive incident replay with decision points for the daemonset blocks eviction scenario.
Incident Replay: DHCP Relay Broken · Interactive incident replay with decision points for the dhcp relay broken scenario.
Incident Replay: Disk Full on Root Partition β€” Services Down · Interactive incident replay with decision points for the disk full root services down scenario.
Incident Replay: DNS Resolution Slow · Interactive incident replay with decision points for the dns resolution slow scenario.
Incident Replay: DNS Split-Horizon Confusion · Interactive incident replay with decision points for the dns split horizon confusion scenario.
Incident Replay: Duplex Mismatch Symptoms · Interactive incident replay with decision points for the duplex mismatch symptoms scenario.
Incident Replay: Firewall Shadow Rule · Interactive incident replay with decision points for the firewall shadow rule scenario.
Incident Replay: Firmware Update Causes Boot Loop · Interactive incident replay with decision points for the firmware update boot loop scenario.
Incident Replay: HBA Firmware Mismatch · Interactive incident replay with decision points for the hba firmware mismatch scenario.
Incident Replay: iDRAC Unreachable but OS Running · Interactive incident replay with decision points for the idrac unreachable os up scenario.
Incident Replay: ImagePullBackOff β€” Registry Authentication Failure · Interactive incident replay with decision points for the imagepullbackoff registry auth scenario.
Incident Replay: Inode Exhaustion · Interactive incident replay with decision points for the inode exhaustion scenario.
Incident Replay: iptables Blocking Unexpected Traffic · Interactive incident replay with decision points for the iptables blocking unexpected scenario.
Incident Replay: Jumbo Frames Partial Deployment · Interactive incident replay with decision points for the jumbo frames partial scenario.
Incident Replay: Kernel Soft Lockup · Interactive incident replay with decision points for the kernel soft lockup scenario.
Incident Replay: LACP Mismatch β€” One Link Hot · Interactive incident replay with decision points for the lacp mismatch one link hot scenario.
Incident Replay: Link Flaps from Bad Optic · Interactive incident replay with decision points for the link flaps bad optic scenario.
Incident Replay: Memory ECC Errors Increasing · Interactive incident replay with decision points for the memory ecc errors increasing scenario.
Incident Replay: MTU Blackhole β€” TLS Stalls · Interactive incident replay with decision points for the mtu blackhole tls stalls scenario.
Incident Replay: Multicast Not Crossing Router · Interactive incident replay with decision points for the multicast not crossing router scenario.
Incident Replay: NAT Exhaustion β€” Intermittent Connectivity · Interactive incident replay with decision points for the nat exhaustion intermittent scenario.
Incident Replay: Network Bonding Failover Not Working · Interactive incident replay with decision points for the bonding failover not working scenario.
Incident Replay: Network Loop β€” Broadcast Storm · Interactive incident replay with decision points for the network loop broadcast storm scenario.
Incident Replay: Node Drain Blocked by PDB · Interactive incident replay with decision points for the drain blocked by pdb scenario.
Incident Replay: Node Pressure Evictions · Interactive incident replay with decision points for the node pressure evictions scenario.
Incident Replay: NVMe Drive Disappeared · Interactive incident replay with decision points for the nvme drive disappeared scenario.
Incident Replay: OOM Killer Events · Interactive incident replay with decision points for the oom killer events scenario.
Incident Replay: OS Install Fails β€” RAID Controller Not Detected · Interactive incident replay with decision points for the os install fails raid controller scenario.
Incident Replay: OSPF Stuck in ExStart · Interactive incident replay with decision points for the ospf stuck in exstart scenario.
Incident Replay: Persistent Volume Stuck Terminating · Interactive incident replay with decision points for the persistent volume stuck terminating scenario.
Incident Replay: Power Supply Redundancy Lost · Interactive incident replay with decision points for the power supply redundancy lost scenario.
Incident Replay: Proxy ARP Causing Issues · Interactive incident replay with decision points for the proxy arp causing issues scenario.
Incident Replay: PXE Boot Fails β€” UEFI Mismatch · Interactive incident replay with decision points for the pxe boot fails uefi mismatch scenario.
Incident Replay: Rack PDU Overload Alert · Interactive incident replay with decision points for the rack pdu overload alert scenario.
Incident Replay: RAID Degraded β€” Rebuild Latency · Interactive incident replay with decision points for the raid degraded rebuild latency scenario.
Incident Replay: Resource Quota Blocking Deployment · Interactive incident replay with decision points for the resource quota blocking deploy scenario.
Incident Replay: Runaway Logs Fill Disk · Interactive incident replay with decision points for the runaway logs fill disk scenario.
Incident Replay: SELinux Denying Service · Interactive incident replay with decision points for the selinux denying service scenario.
Incident Replay: Serial Console Output Garbled · Interactive incident replay with decision points for the serial console garbled scenario.
Incident Replay: Server Intermittent Reboots · Interactive incident replay with decision points for the server intermittent reboot scenario.
Incident Replay: Server Remote Console Lag · Interactive incident replay with decision points for the server remote console lag scenario.
Incident Replay: Service Has No Endpoints · Interactive incident replay with decision points for the service no endpoints scenario.
Incident Replay: Source Routing Policy Miss · Interactive incident replay with decision points for the source routing policy miss scenario.
Incident Replay: SSL Certificate Chain Incomplete · Interactive incident replay with decision points for the ssl cert chain incomplete scenario.
Incident Replay: Stuck NFS Mount · Interactive incident replay with decision points for the stuck nfs mount scenario.
Incident Replay: systemd Service Flapping · Interactive incident replay with decision points for the systemd service flapping scenario.
Incident Replay: TCP RST After Idle · Interactive incident replay with decision points for the tcp rst after idle scenario.
Incident Replay: Thermal Throttling from Fan Failure · Interactive incident replay with decision points for the thermal throttle fan failure scenario.
Incident Replay: Time Sync Skew Breaks Application · Interactive incident replay with decision points for the time sync skew breaks app scenario.
Incident Replay: VLAN Trunk Mistag · Interactive incident replay with decision points for the vlan trunk mistag scenario.
Incident Replay: Zombie Processes Accumulating · Interactive incident replay with decision points for the zombie processes accumulating scenario.
Investigation: Alert Storm, Caused by Flapping Health Checks, Fix Is Probe Tuning · Multi-domain investigation walkthrough for the alert storm flapping healthchecks incident.
Investigation: Ansible Playbook Hangs, SSH Agent Forwarding Broken, Root Cause Is Firewall Rule · Multi-domain investigation walkthrough for the ansible ssh agent firewall incident.
Investigation: API Latency Spike, BGP Route Leak, Fix Is Network ACL · Multi-domain investigation walkthrough for the api latency bgp route leak acl incident.
Investigation: Backup Job Failing, iSCSI Target Unreachable, Fix Is VLAN Config · Multi-domain investigation walkthrough for the backup job iscsi vlan incident.
Investigation: Canary Deploy Looks Healthy, Actually Routing to Wrong Backend, Ingress Misconfigured · Multi-domain investigation walkthrough for the canary deploy wrong backend ingress incident.
Investigation: CI Pipeline Fails, Docker Layer Cache Corruption, Fix Is Registry GC · Multi-domain investigation walkthrough for the ci pipeline docker cache registry incident.
Investigation: Container Image Vuln Scanner False Positive, Blocks Deploy Pipeline · Multi-domain investigation walkthrough for the container vuln scanner false positive incident.
Investigation: Database Replication Lag, Root Cause Is RAID Degradation · Multi-domain investigation walkthrough for the database replication lag raid incident.
Investigation: Deployment Stuck, ImagePull Auth Failure, Fix Is Vault Secret Rotation · Multi-domain investigation walkthrough for the deployment stuck imagepull vault incident.
Investigation: Disk Full Alert, Cause Is Runaway Logs, Fix Is Loki Retention · Multi-domain investigation walkthrough for the disk full runaway logs loki incident.
Investigation: DNS Looks Broken, TLS Is Expired, Fix Is in Cert-Manager · Multi-domain investigation walkthrough for the dns tls certmanager incident.
Investigation: Grafana Dashboard Empty, Prometheus Scrape Blocked by NetworkPolicy · Multi-domain investigation walkthrough for the grafana empty prometheus networkpolicy incident.
Investigation: HPA Flapping, Metrics Server Clock Skew, Fix Is NTP Config · Multi-domain investigation walkthrough for the hpa flapping clock skew ntp incident.
Investigation: Job Queue Backlog, Worker Pod CPU Throttled, Fix Is cgroup Config · Multi-domain investigation walkthrough for the job queue cpu throttle cgroup incident.
Investigation: Node NotReady, NIC Firmware Bug, Fix Is Ansible Playbook · Multi-domain investigation walkthrough for the node notready nic firmware ansible incident.
Investigation: Pod OOMKilled, Memory Leak Is in Sidecar, Fix Is Helm Values · Multi-domain investigation walkthrough for the pod oomkilled sidecar helm incident.
Investigation: Service Mesh 503s, Envoy Misconfigured, Root Cause Is RBAC Policy · Multi-domain investigation walkthrough for the service mesh 503 envoy rbac incident.
Investigation: SSH Timeout, MTU Mismatch, Fix Is Terraform Variable · Multi-domain investigation walkthrough for the ssh timeout mtu terraform incident.
Investigation: Terraform Apply Fails, State Lock Stuck, Root Cause Is DynamoDB Throttle · Multi-domain investigation walkthrough for the terraform state lock dynamodb incident.
Investigation: User Auth Failing, OIDC Cert Expired, Fix Is Cloud KMS Rotation · Multi-domain investigation walkthrough for the user auth oidc cert kms incident.
Kubernetes Operations Case Studies · Index of Kubernetes operations case studies with incident matrix and navigation.
Linux Operations Case Studies · Index of Linux operations case studies with incident matrix and navigation.
Networking Case Studies · Index of networking case studies with incident matrix and navigation.
Ops Archaeology: Reverse-Engineering Production Systems · Index of ops archaeology case studies with incident matrix and navigation.
Ops Archaeology: The 5% That Can't Resolve · Initial artifacts and scenario briefing for the 10 intermittent dns archaeology exercise.
Ops Archaeology: The 5% That Can't Resolve · Overview and file index for the 10 intermittent dns ops archaeology case study.
Ops Archaeology: The Alerts That Stopped Firing · Initial artifacts and scenario briefing for the 09 monitoring gap archaeology exercise.
Ops Archaeology: The Alerts That Stopped Firing · Overview and file index for the 09 monitoring gap ops archaeology case study.
Ops Archaeology: The Certificate That Works Sometimes · Initial artifacts and scenario briefing for the 12 tls chain incomplete archaeology exercise.
Ops Archaeology: The Certificate That Works Sometimes · Overview and file index for the 12 tls chain incomplete ops archaeology case study.
Ops Archaeology: The Cluster That Disagrees With Itself · Initial artifacts and scenario briefing for the 14 split brain etcd archaeology exercise.
Ops Archaeology: The Cluster That Disagrees With Itself · Overview and file index for the 14 split brain etcd ops archaeology case study.
Ops Archaeology: The Container That Exits Immediately · Initial artifacts and scenario briefing for the 03 docker exec format error archaeology exercise.
Ops Archaeology: The Container That Exits Immediately · Overview and file index for the 03 docker exec format error ops archaeology case study.
Ops Archaeology: The Deploy That Didn't Deploy · Initial artifacts and scenario briefing for the 07 stale image tag archaeology exercise.
Ops Archaeology: The Deploy That Didn't Deploy · Overview and file index for the 07 stale image tag ops archaeology case study.
Ops Archaeology: The DR That Looks Ready But Isn't · Initial artifacts and scenario briefing for the 11 dr failover broken archaeology exercise.
Ops Archaeology: The DR That Looks Ready But Isn't · Overview and file index for the 11 dr failover broken ops archaeology case study.
Ops Archaeology: The Gateway That Returns 502 · Initial artifacts and scenario briefing for the 05 nginx 502 backend archaeology exercise.
Ops Archaeology: The Gateway That Returns 502 · Overview and file index for the 05 nginx 502 backend ops archaeology case study.
Ops Archaeology: The Job That Succeeded Wrong · Initial artifacts and scenario briefing for the 08 job wrong database archaeology exercise.
Ops Archaeology: The Job That Succeeded Wrong · Overview and file index for the 08 job wrong database ops archaeology case study.
Ops Archaeology: The Pods That Won't Schedule · Initial artifacts and scenario briefing for the 13 quota masquerade archaeology exercise.
Ops Archaeology: The Pods That Won't Schedule · Overview and file index for the 13 quota masquerade ops archaeology case study.
Ops Archaeology: The Replica That Fell Behind · Initial artifacts and scenario briefing for the 04 postgres replica lag archaeology exercise.
Ops Archaeology: The Replica That Fell Behind · Overview and file index for the 04 postgres replica lag ops archaeology case study.
Ops Archaeology: The Requests That Vanish · Initial artifacts and scenario briefing for the 06 mesh silent drops archaeology exercise.
Ops Archaeology: The Requests That Vanish · Overview and file index for the 06 mesh silent drops ops archaeology case study.
Ops Archaeology: The Service That Won't Start · Initial artifacts and scenario briefing for the 02 systemd permission denied archaeology exercise.
Ops Archaeology: The Service That Won't Start · Overview and file index for the 02 systemd permission denied ops archaeology case study.
Ops Archaeology: The Session Store That Keeps Dying · Initial artifacts and scenario briefing for the 01 redis oom crashloop archaeology exercise.
Ops Archaeology: The Session Store That Keeps Dying · Overview and file index for the 01 redis oom crashloop ops archaeology case study.
Ops Archaeology: The Slow Death Nobody Noticed · Initial artifacts and scenario briefing for the 15 cascading debug logs archaeology exercise.
Ops Archaeology: The Slow Death Nobody Noticed · Overview and file index for the 15 cascading debug logs ops archaeology case study.
Progressive Hints · Progressive hints with timed reveals for the 01 redis oom crashloop archaeology exercise.
Progressive Hints · Progressive hints with timed reveals for the 02 systemd permission denied archaeology exercise.
Progressive Hints · Progressive hints with timed reveals for the 03 docker exec format error archaeology exercise.
Progressive Hints · Progressive hints with timed reveals for the 04 postgres replica lag archaeology exercise.
Progressive Hints · Progressive hints with timed reveals for the 05 nginx 502 backend archaeology exercise.
Progressive Hints · Progressive hints with timed reveals for the 06 mesh silent drops archaeology exercise.
Progressive Hints · Progressive hints with timed reveals for the 07 stale image tag archaeology exercise.
Progressive Hints · Progressive hints with timed reveals for the 08 job wrong database archaeology exercise.
Progressive Hints · Progressive hints with timed reveals for the 09 monitoring gap archaeology exercise.
Progressive Hints · Progressive hints with timed reveals for the 10 intermittent dns archaeology exercise.
Progressive Hints · Progressive hints with timed reveals for the 11 dr failover broken archaeology exercise.
Progressive Hints · Progressive hints with timed reveals for the 12 tls chain incomplete archaeology exercise.
Progressive Hints · Progressive hints with timed reveals for the 13 quota masquerade archaeology exercise.
Progressive Hints · Progressive hints with timed reveals for the 14 split brain etcd archaeology exercise.
Progressive Hints · Progressive hints with timed reveals for the 15 cascading debug logs archaeology exercise.
Questions to Determine · Diagnostic questions to guide troubleshooting of the bios settings reset after cmos incident.
Questions to Determine · Diagnostic questions to guide troubleshooting of the bonding failover not working incident.
Questions to Determine · Diagnostic questions to guide troubleshooting of the cable management wrong port incident.
Questions to Determine · Diagnostic questions to guide troubleshooting of the hba firmware mismatch incident.
Questions to Determine · Diagnostic questions to guide troubleshooting of the nvme drive disappeared incident.
Questions to Determine · Diagnostic questions to guide troubleshooting of the os install fails raid controller incident.
Questions to Determine · Diagnostic questions to guide troubleshooting of the rack pdu overload alert incident.
Questions to Determine · Diagnostic questions to guide troubleshooting of the serial console garbled incident.
Questions to Determine · Diagnostic questions to guide troubleshooting of the server intermittent reboot incident.
Questions to Determine · Diagnostic questions to guide troubleshooting of the server remote console lag incident.
Questions to Determine · Diagnostic questions to guide troubleshooting of the cni broken after restart incident.
Questions to Determine · Diagnostic questions to guide troubleshooting of the coredns timeout pod dns incident.
Questions to Determine · Diagnostic questions to guide troubleshooting of the crashloopbackoff no logs incident.
Questions to Determine · Diagnostic questions to guide troubleshooting of the daemonset blocks eviction incident.
Questions to Determine · Diagnostic questions to guide troubleshooting of the drain blocked by pdb incident.
Questions to Determine · Diagnostic questions to guide troubleshooting of the imagepullbackoff registry auth incident.
Questions to Determine · Diagnostic questions to guide troubleshooting of the node pressure evictions incident.
Questions to Determine · Diagnostic questions to guide troubleshooting of the persistent volume stuck terminating incident.
Questions to Determine · Diagnostic questions to guide troubleshooting of the resource quota blocking deploy incident.
Questions to Determine · Diagnostic questions to guide troubleshooting of the service no endpoints incident.
Questions to Determine · Diagnostic questions to guide troubleshooting of the inode exhaustion incident.
Questions to Determine · Diagnostic questions to guide troubleshooting of the iptables blocking unexpected incident.
Questions to Determine · Diagnostic questions to guide troubleshooting of the kernel soft lockup incident.
Questions to Determine · Diagnostic questions to guide troubleshooting of the oom killer events incident.
Questions to Determine · Diagnostic questions to guide troubleshooting of the runaway logs fill disk incident.
Questions to Determine · Diagnostic questions to guide troubleshooting of the selinux denying service incident.
Questions to Determine · Diagnostic questions to guide troubleshooting of the stuck nfs mount incident.
Questions to Determine · Diagnostic questions to guide troubleshooting of the systemd service flapping incident.
Questions to Determine · Diagnostic questions to guide troubleshooting of the time sync skew breaks app incident.
Questions to Determine · Diagnostic questions to guide troubleshooting of the zombie processes accumulating incident.
Questions to Determine · Diagnostic questions to guide troubleshooting of the arp flux duplicate ip incident.
Questions to Determine · Diagnostic questions to guide troubleshooting of the asymmetric routing one direction incident.
Questions to Determine · Diagnostic questions to guide troubleshooting of the bgp peer flapping incident.
Questions to Determine · Diagnostic questions to guide troubleshooting of the dns split horizon confusion incident.
Questions to Determine · Diagnostic questions to guide troubleshooting of the duplex mismatch symptoms incident.
Questions to Determine · Diagnostic questions to guide troubleshooting of the firewall shadow rule incident.
Questions to Determine · Diagnostic questions to guide troubleshooting of the lacp mismatch one link hot incident.
Questions to Determine · Diagnostic questions to guide troubleshooting of the mtu blackhole tls stalls incident.
Questions to Determine · Diagnostic questions to guide troubleshooting of the nat exhaustion intermittent incident.
Questions to Determine · Diagnostic questions to guide troubleshooting of the vlan trunk mistag incident.
Questions: BMC Clock Skew - Certificate Failure · Diagnostic questions to guide troubleshooting of the bmc clock skew cert failure incident.
Questions: DHCP Not Working on Remote VLAN · Diagnostic questions to guide troubleshooting of the dhcp relay broken incident.
Questions: Disk Full Root - Services Down · Diagnostic questions to guide troubleshooting of the disk full root services down incident.
Questions: DNS Resolution Taking 5+ Seconds Intermittently · Diagnostic questions to guide troubleshooting of the dns resolution slow incident.
Questions: Firmware Update Boot Loop · Diagnostic questions to guide troubleshooting of the firmware update boot loop incident.
Questions: iDRAC Unreachable, OS Up · Diagnostic questions to guide troubleshooting of the idrac unreachable os up incident.
Questions: Jumbo Frames Enabled But Some Paths Failing · Diagnostic questions to guide troubleshooting of the jumbo frames partial incident.
Questions: Link Flaps - Bad Optic · Diagnostic questions to guide troubleshooting of the link flaps bad optic incident.
Questions: Memory ECC Errors Increasing · Diagnostic questions to guide troubleshooting of the memory ecc errors increasing incident.
Questions: Multicast Traffic Not Crossing Router · Diagnostic questions to guide troubleshooting of the multicast not crossing router incident.
Questions: Network Experiencing Broadcast Storm and High CPU on Switches · Diagnostic questions to guide troubleshooting of the network loop broadcast storm incident.
Questions: OSPF Adjacency Stuck in ExStart/Exchange State · Diagnostic questions to guide troubleshooting of the ospf stuck in exstart incident.
Questions: Power Supply Redundancy Lost · Diagnostic questions to guide troubleshooting of the power supply redundancy lost incident.
Questions: Proxy ARP Causing Unexpected Routing Behavior · Diagnostic questions to guide troubleshooting of the proxy arp causing issues incident.
Questions: PXE Boot Fails - UEFI Mismatch · Diagnostic questions to guide troubleshooting of the pxe boot fails uefi mismatch incident.
Questions: RAID Degraded Rebuild Latency · Diagnostic questions to guide troubleshooting of the raid degraded rebuild latency incident.
Questions: TCP Connections Reset After Idle Period · Diagnostic questions to guide troubleshooting of the tcp rst after idle incident.
Questions: Thermal Throttle - Fan Failure · Diagnostic questions to guide troubleshooting of the thermal throttle fan failure incident.
Questions: TLS Works From Some Clients But Fails From Others · Diagnostic questions to guide troubleshooting of the ssl cert chain incomplete incident.
Questions: Traffic From Specific Source Not Taking Expected Path · Diagnostic questions to guide troubleshooting of the source routing policy miss incident.
Remediation: Alert Storm, Caused by Flapping Health Checks, Fix Is Probe Tuning · Cross-domain remediation steps and verification for the alert storm flapping healthchecks incident.
Remediation: Ansible Playbook Hangs, SSH Agent Forwarding Broken, Root Cause Is Firewall Rule · Cross-domain remediation steps and verification for the ansible ssh agent firewall incident.
Remediation: API Latency Spike, BGP Route Leak, Fix Is Network ACL · Cross-domain remediation steps and verification for the api latency bgp route leak acl incident.
Remediation: Backup Job Failing, iSCSI Target Unreachable, Fix Is VLAN Config · Cross-domain remediation steps and verification for the backup job iscsi vlan incident.
Remediation: Canary Deploy Looks Healthy, Actually Routing to Wrong Backend, Ingress Misconfigured · Cross-domain remediation steps and verification for the canary deploy wrong backend ingress incident.
Remediation: CI Pipeline Fails, Docker Layer Cache Corruption, Fix Is Registry GC · Cross-domain remediation steps and verification for the ci pipeline docker cache registry incident.
Remediation: Container Image Vuln Scanner False Positive, Blocks Deploy Pipeline · Cross-domain remediation steps and verification for the container vuln scanner false positive incident.
Remediation: Database Replication Lag, Root Cause Is RAID Degradation · Cross-domain remediation steps and verification for the database replication lag raid incident.
Remediation: Deployment Stuck, ImagePull Auth Failure, Fix Is Vault Secret Rotation · Cross-domain remediation steps and verification for the deployment stuck imagepull vault incident.
Remediation: Disk Full Alert, Cause Is Runaway Logs, Fix Is Loki Retention · Cross-domain remediation steps and verification for the disk full runaway logs loki incident.
Remediation: DNS Looks Broken, TLS Is Expired, Fix Is in Cert-Manager · Cross-domain remediation steps and verification for the dns tls certmanager incident.
Remediation: Grafana Dashboard Empty, Prometheus Scrape Blocked by NetworkPolicy · Cross-domain remediation steps and verification for the grafana empty prometheus networkpolicy incident.
Remediation: HPA Flapping, Metrics Server Clock Skew, Fix Is NTP Config · Cross-domain remediation steps and verification for the hpa flapping clock skew ntp incident.
Remediation: Job Queue Backlog, Worker Pod CPU Throttled, Fix Is cgroup Config · Cross-domain remediation steps and verification for the job queue cpu throttle cgroup incident.
Remediation: Node NotReady, NIC Firmware Bug, Fix Is Ansible Playbook · Cross-domain remediation steps and verification for the node notready nic firmware ansible incident.
Remediation: Pod OOMKilled, Memory Leak Is in Sidecar, Fix Is Helm Values · Cross-domain remediation steps and verification for the pod oomkilled sidecar helm incident.
Remediation: Service Mesh 503s, Envoy Misconfigured, Root Cause Is RBAC Policy · Cross-domain remediation steps and verification for the service mesh 503 envoy rbac incident.
Remediation: SSH Timeout, MTU Mismatch, Fix Is Terraform Variable · Cross-domain remediation steps and verification for the ssh timeout mtu terraform incident.
Remediation: Terraform Apply Fails, State Lock Stuck, Root Cause Is DynamoDB Throttle · Cross-domain remediation steps and verification for the terraform state lock dynamodb incident.
Remediation: User Auth Failing, OIDC Cert Expired, Fix Is Cloud KMS Rotation · Cross-domain remediation steps and verification for the user auth oidc cert kms incident.
Solution · Root cause analysis, triage steps, and resolution for the cni broken after restart incident.
Solution · Root cause analysis, triage steps, and resolution for the coredns timeout pod dns incident.
Solution · Root cause analysis, triage steps, and resolution for the crashloopbackoff no logs incident.
Solution · Root cause analysis, triage steps, and resolution for the daemonset blocks eviction incident.
Solution · Root cause analysis, triage steps, and resolution for the drain blocked by pdb incident.
Solution · Root cause analysis, triage steps, and resolution for the imagepullbackoff registry auth incident.
Solution · Root cause analysis, triage steps, and resolution for the node pressure evictions incident.
Solution · Root cause analysis, triage steps, and resolution for the persistent volume stuck terminating incident.
Solution · Root cause analysis, triage steps, and resolution for the resource quota blocking deploy incident.
Solution · Root cause analysis, triage steps, and resolution for the service no endpoints incident.
Solution · Root cause analysis, triage steps, and resolution for the inode exhaustion incident.
Solution · Root cause analysis, triage steps, and resolution for the iptables blocking unexpected incident.
Solution · Root cause analysis, triage steps, and resolution for the kernel soft lockup incident.
Solution · Root cause analysis, triage steps, and resolution for the oom killer events incident.
Solution · Root cause analysis, triage steps, and resolution for the runaway logs fill disk incident.
Solution · Root cause analysis, triage steps, and resolution for the selinux denying service incident.
Solution · Root cause analysis, triage steps, and resolution for the stuck nfs mount incident.
Solution · Root cause analysis, triage steps, and resolution for the systemd service flapping incident.
Solution · Root cause analysis, triage steps, and resolution for the time sync skew breaks app incident.
Solution · Root cause analysis, triage steps, and resolution for the zombie processes accumulating incident.
Solution: ARP Flux / Duplicate IP · Root cause analysis, triage steps, and resolution for the arp flux duplicate ip incident.
Solution: Asymmetric Routing / One-Direction Failure · Root cause analysis, triage steps, and resolution for the asymmetric routing one direction incident.
Solution: BGP Peer Flapping · Root cause analysis, triage steps, and resolution for the bgp peer flapping incident.
Solution: BIOS Settings Reset After CMOS Battery Replacement · Root cause analysis, triage steps, and resolution for the bios settings reset after cmos incident.
Solution: BMC Clock Skew - Certificate Failure · Root cause analysis, triage steps, and resolution for the bmc clock skew cert failure incident.
Solution: Bonding Failover Not Working · Root cause analysis, triage steps, and resolution for the bonding failover not working incident.
Solution: DHCP Not Working on Remote VLAN · Root cause analysis, triage steps, and resolution for the dhcp relay broken incident.
Solution: Disk Full Root - Services Down · Root cause analysis, triage steps, and resolution for the disk full root services down incident.
Solution: DNS Resolution Taking 5+ Seconds Intermittently · Root cause analysis, triage steps, and resolution for the dns resolution slow incident.
Solution: DNS Split-Horizon Confusion · Root cause analysis, triage steps, and resolution for the dns split horizon confusion incident.
Solution: Duplex Mismatch · Root cause analysis, triage steps, and resolution for the duplex mismatch symptoms incident.
Solution: Firewall Shadow Rule · Root cause analysis, triage steps, and resolution for the firewall shadow rule incident.
Solution: Firmware Update Boot Loop · Root cause analysis, triage steps, and resolution for the firmware update boot loop incident.
Solution: HBA Firmware Mismatch Causing I/O Errors · Root cause analysis, triage steps, and resolution for the hba firmware mismatch incident.
Solution: iDRAC Unreachable, OS Up · Root cause analysis, triage steps, and resolution for the idrac unreachable os up incident.
Solution: Jumbo Frames Enabled But Some Paths Failing · Root cause analysis, triage steps, and resolution for the jumbo frames partial incident.
Solution: LACP Mismatch / One Link Hot · Root cause analysis, triage steps, and resolution for the lacp mismatch one link hot incident.
Solution: Link Flaps - Bad Optic · Root cause analysis, triage steps, and resolution for the link flaps bad optic incident.
Solution: Memory ECC Errors Increasing · Root cause analysis, triage steps, and resolution for the memory ecc errors increasing incident.
Solution: MTU Black Hole / TLS Stalls · Root cause analysis, triage steps, and resolution for the mtu blackhole tls stalls incident.
Solution: Multicast Traffic Not Crossing Router · Root cause analysis, triage steps, and resolution for the multicast not crossing router incident.
Solution: NAT Port Exhaustion / Intermittent Failures · Root cause analysis, triage steps, and resolution for the nat exhaustion intermittent incident.
Solution: Network Experiencing Broadcast Storm and High CPU on Switches · Root cause analysis, triage steps, and resolution for the network loop broadcast storm incident.
Solution: NVMe Drive Disappeared After Reboot · Root cause analysis, triage steps, and resolution for the nvme drive disappeared incident.
Solution: OS Install Fails - RAID Controller Driver Missing · Root cause analysis, triage steps, and resolution for the os install fails raid controller incident.
Solution: OSPF Adjacency Stuck in ExStart/Exchange State · Root cause analysis, triage steps, and resolution for the ospf stuck in exstart incident.
Solution: PDU Overload Warning - Phase Imbalance · Root cause analysis, triage steps, and resolution for the rack pdu overload alert incident.
Solution: Power Supply Redundancy Lost · Root cause analysis, triage steps, and resolution for the power supply redundancy lost incident.
Solution: Proxy ARP Causing Unexpected Routing Behavior · Root cause analysis, triage steps, and resolution for the proxy arp causing issues incident.
Solution: PXE Boot Fails - UEFI Mismatch · Root cause analysis, triage steps, and resolution for the pxe boot fails uefi mismatch incident.
Solution: RAID Degraded Rebuild Latency · Root cause analysis, triage steps, and resolution for the raid degraded rebuild latency incident.
Solution: Remote KVM/Console Extremely Laggy · Root cause analysis, triage steps, and resolution for the server remote console lag incident.
Solution: Serial-over-LAN Output Garbled · Root cause analysis, triage steps, and resolution for the serial console garbled incident.
Solution: Server Cabled to Wrong Switch Port · Root cause analysis, triage steps, and resolution for the cable management wrong port incident.
Solution: Server Intermittent Reboots · Root cause analysis, triage steps, and resolution for the server intermittent reboot incident.
Solution: TCP Connections Reset After Idle Period · Root cause analysis, triage steps, and resolution for the tcp rst after idle incident.
Solution: Thermal Throttle - Fan Failure · Root cause analysis, triage steps, and resolution for the thermal throttle fan failure incident.
Solution: TLS Works From Some Clients But Fails From Others · Root cause analysis, triage steps, and resolution for the ssl cert chain incomplete incident.
Solution: Traffic From Specific Source Not Taking Expected Path · Root cause analysis, triage steps, and resolution for the source routing policy miss incident.
Solution: VLAN Trunk Mistag · Root cause analysis, triage steps, and resolution for the vlan trunk mistag incident.
Symptoms · Observable symptoms: After a node reboot (kernel upgrade), pods on `node-4.internal` cannot communicate with pods on other nodes.
Symptoms · Observable symptoms: Application pods intermittently fail to resolve DNS names, both internal services and external domains.
Symptoms · Observable symptoms: Pod `auth-service-8c4f7d6b2-qp5r3` in namespace `prod` is in CrashLoopBackOff.
Symptoms · Observable symptoms: Running `kubectl drain node-5.internal` fails immediately with an error referencing DaemonSet-managed pods.
Symptoms · Observable symptoms: `kubectl drain node-3.internal` has been hanging for over 15 minutes with no progress.
Symptoms · Observable symptoms: New pods for the `checkout-service` deployment are stuck in `ImagePullBackOff`.
Symptoms · Observable symptoms: Multiple pods on `node-7.internal` are being terminated unexpectedly.
Symptoms · Observable symptoms: A PersistentVolume `pv-data-warehouse-01` has been in `Terminating` status for over 3 hours.
Symptoms · Observable symptoms: A new deployment `notification-service` in namespace `staging` is stuck at 0 available replicas.
Symptoms · Observable symptoms: Requests to `http://user-service.prod.svc.cluster.local:8080` from other pods return `connection refused`.
Symptoms · Observable symptoms: Applications on `mail-relay-01` fail with
Symptoms · Observable symptoms: The application on `api-server-02` cannot connect to an external payment gateway at `payments.provider.com:443`.
Symptoms · Observable symptoms: The database server `db-primary-01` intermittently becomes unresponsive for 10-30 seconds.
Symptoms · Observable symptoms: The Java application `order-processor` was killed unexpectedly at 03:22 AM.
Symptoms · Observable symptoms: Monitoring alerts: root filesystem at 98% capacity on `web-prod-03`.
Symptoms · Observable symptoms: The `inventory-api` service fails to start on a RHEL 9 server.
Symptoms · Observable symptoms: Several processes on `app-server-01` are stuck in uninterruptible sleep (D state).
Symptoms · Observable symptoms: The `data-sync` service restarts every 30 seconds and then stops entirely.
Symptoms · Observable symptoms: A 3-node etcd cluster is reporting consensus failures and leader election instability.
Symptoms · Observable symptoms: Monitoring shows the process count on `build-server-01` has been steadily increasing over 3 days.
Symptoms: Alert Storm, Caused by Flapping Health Checks, Fix Is Probe Tuning · Observable symptoms: **Domains:** observability | networking | kubernetes_ops.
Symptoms: Ansible Playbook Hangs, SSH Agent Forwarding Broken, Root Cause Is Firewall Rule · Observable symptoms: **Domains:** devops_tooling | linux_ops | networking.
Symptoms: API Latency Spike, BGP Route Leak, Fix Is Network ACL · Observable symptoms: **Domains:** observability | networking | security.
Symptoms: ARP Flux / Duplicate IP · Observable symptoms: Server app-web-05 (10.30.1.100) experiences intermittent connectivity loss lasting 10-30 seconds at a time.
Symptoms: Asymmetric Routing / One-Direction Failure · Observable symptoms: Host A (10.1.10.25, web-prod-01) can SSH and curl to Host B (10.2.20.40, db-prod-01) without issues.
Symptoms: Backup Job Failing, iSCSI Target Unreachable, Fix Is VLAN Config · Observable symptoms: **Domains:** linux_ops | datacenter_ops | networking.
Symptoms: BGP Peer Flapping · Observable symptoms and initial indicators for the bgp peer flapping incident.
Symptoms: BIOS Settings Reverted After CMOS Battery Replacement · Observable symptoms: A Dell PowerEdge R640 had its CMOS battery replaced due to a
Symptoms: BMC Clock Skew - Certificate Failure · Observable symptoms: The iLO/BMC web interface on `infra-mgmt-11` (HPE ProLiant DL360 Gen10) throws certificate errors.
Symptoms: Canary Deploy Looks Healthy, Actually Routing to Wrong Backend, Ingress Misconfigured · Observable symptoms: **Domains:** devops_tooling | networking | kubernetes_ops.
Symptoms: CI Pipeline Fails, Docker Layer Cache Corruption, Fix Is Registry GC · Observable symptoms: **Domains:** devops_tooling | linux_ops | kubernetes_ops.
Symptoms: Container Image Vuln Scanner False Positive, Blocks Deploy Pipeline · Observable symptoms: **Domains:** security | devops_tooling | kubernetes_ops.
Symptoms: Database Replication Lag, Root Cause Is RAID Degradation · Observable symptoms: **Domains:** kubernetes_ops | linux_ops | datacenter_ops.
Symptoms: Deployment Stuck, ImagePull Auth Failure, Fix Is Vault Secret Rotation · Observable symptoms: **Domains:** kubernetes_ops | security | devops_tooling.
Symptoms: DHCP Not Working on Remote VLAN · Observable symptoms: New hosts on VLAN 50 cannot obtain IP addresses via DHCP.
Symptoms: Disk Full Alert, Cause Is Runaway Logs, Fix Is Loki Retention · Observable symptoms: **Domains:** linux_ops | observability | devops_tooling.
Symptoms: Disk Full Root - Services Down · Observable symptoms: Multiple services on `api-gateway-03` (Ubuntu 22.04) are failing to start or crashing.
Symptoms: DNS Looks Broken, TLS Is Expired, Fix Is in Cert-Manager · Observable symptoms: **Domains:** networking | security | kubernetes_ops.
Symptoms: DNS Resolution Taking 5+ Seconds Intermittently · Observable symptoms: Users and applications report intermittent slowness loading pages and connecting to services.
Symptoms: DNS Split-Horizon Confusion · Observable symptoms and initial indicators for the dns split horizon confusion incident.
Symptoms: Duplex Mismatch · Observable symptoms: Server legacy-app-01 (10.80.1.55) connected to switch sw-access-03 port Gi0/12 shows very poor network throughput.
Symptoms: Firewall Shadow Rule · Observable symptoms and initial indicators for the firewall shadow rule incident.
Symptoms: Firmware Update Boot Loop · Observable symptoms: Server `compute-14` (HPE ProLiant DL380 Gen10) is stuck in a boot loop after a scheduled BIOS update.
Symptoms: Grafana Dashboard Empty, Prometheus Scrape Blocked by NetworkPolicy · Observable symptoms: **Domains:** observability | kubernetes_ops | networking.
Symptoms: HBA Firmware Mismatch Causing I/O Errors · Observable symptoms: A fleet of 24 servers connects to a NetApp SAN via Emulex LPe32002 Fibre Channel HBAs.
Symptoms: HPA Flapping, Metrics Server Clock Skew, Fix Is NTP Config · Observable symptoms: **Domains:** kubernetes_ops | observability | linux_ops.
Symptoms: iDRAC Unreachable, OS Up · Observable symptoms: Server `app-web-22` (Dell PowerEdge R640) has an unreachable iDRAC management interface.
Symptoms: Job Queue Backlog, Worker Pod CPU Throttled, Fix Is cgroup Config · Observable symptoms: **Domains:** observability | kubernetes_ops | linux_ops.
Symptoms: Jumbo Frames Enabled But Some Paths Failing · Observable symptoms: Storage replication between two data centers is failing for large transfers but small ones succeed.
Symptoms: LACP Mismatch / One Link Hot · Observable symptoms: Server db-prod-02 has a 2x10G LACP bond (bond0) to the top-of-rack switch for redundancy and throughput.
Symptoms: Link Flaps - Bad Optic · Observable symptoms: Server `stor-node-05` has intermittent network connectivity on its 10G SFP+ interface (eth2).
Symptoms: Memory ECC Errors Increasing · Observable symptoms: Server `db-replica-04` (Dell PowerEdge R740, 384GB RAM) is logging correctable ECC memory errors.
Symptoms: MTU Black Hole / TLS Stalls · Observable symptoms and initial indicators for the mtu blackhole tls stalls incident.
Symptoms: Multicast Traffic Not Crossing Router · Observable symptoms: A multicast video streaming application works perfectly within the same VLAN.
Symptoms: NAT Port Exhaustion / Intermittent Failures · Observable symptoms: Users on the internal network (10.200.0.0/16) report intermittent failures when accessing external websites and APIs.
Symptoms: Network Bonding Failover Not Working · Observable symptoms: A production web server has a Linux bond (bond0) configured with two interfaces: eno1 and eno2.
Symptoms: Network Experiencing Broadcast Storm and High CPU on Switches · Observable symptoms: Multiple switches in the access layer are showing 90-100% CPU utilization.
Symptoms: Node NotReady, NIC Firmware Bug, Fix Is Ansible Playbook · Observable symptoms: **Domains:** kubernetes_ops | datacenter_ops | devops_tooling.
Symptoms: NVMe Drive Disappeared After Reboot · Observable symptoms: A Dell PowerEdge R750 server was rebooted for scheduled kernel patching.
Symptoms: OS Installation Cannot See Disks · Observable symptoms: Attempting to install RHEL 8.8 on a new Lenovo ThinkSystem SR650 V2 server.
Symptoms: OSPF Adjacency Stuck in ExStart/Exchange State · Observable symptoms: Two routers connected via a point-to-point Ethernet link are unable to form a full OSPF adjacency.
Symptoms: PDU Reporting Overload Warning · Observable symptoms: The monitoring system fires a critical alert: PDU-A in rack C14 reports 87% load (34.8A on a 40A circuit).
Symptoms: Pod OOMKilled, Memory Leak Is in Sidecar, Fix Is Helm Values · Observable symptoms: **Domains:** kubernetes_ops | observability | devops_tooling.
Symptoms: Power Supply Redundancy Lost · Observable symptoms: Server `k8s-worker-17` (HPE ProLiant DL380 Gen10) triggered a BMC alert:
Symptoms: Proxy ARP Causing Unexpected Routing Behavior · Observable symptoms: Hosts on different subnets can communicate even though no explicit route exists between them.
Symptoms: PXE Boot Fails - UEFI Mismatch · Observable symptoms: A batch of 12 new servers (Dell PowerEdge R750) arrived and are being provisioned via PXE boot.
Symptoms: RAID Degraded Rebuild Latency · Observable symptoms: Production database server `db-prod-07` (Dell PowerEdge R740) is experiencing severe I/O latency spikes.
Symptoms: Remote KVM/Console Extremely Laggy · Observable symptoms: A technician needs to troubleshoot a boot issue on an HPE ProLiant DL380 Gen10 via iLO remote console.
Symptoms: Serial-over-LAN Output Garbled · Observable symptoms: A headless Supermicro server in a remote colo is managed exclusively via IPMI Serial-over-LAN (SOL).
Symptoms: Server Cabled to Wrong Switch Port / Wrong VLAN · Observable symptoms: A newly racked server (srv-prod-55) was cabled and powered on, but it cannot reach its expected network gateway.
Symptoms: Server Randomly Rebooting Every Few Hours · Observable symptoms: A Dell PowerEdge R740 production database server has rebooted unexpectedly 6 times in 48 hours.
Symptoms: Service Mesh 503s, Envoy Misconfigured, Root Cause Is RBAC Policy · Observable symptoms: **Domains:** networking | kubernetes_ops | security.
Symptoms: SSH Timeout, MTU Mismatch, Fix Is Terraform Variable · Observable symptoms: **Domains:** linux_ops | networking | cloud.
Symptoms: TCP Connections Reset After Idle Period · Observable symptoms and initial indicators for the tcp rst after idle incident.
Symptoms: Terraform Apply Fails, State Lock Stuck, Root Cause Is DynamoDB Throttle · Observable symptoms: **Domains:** devops_tooling | cloud | observability.
Symptoms: Thermal Throttle - Fan Failure · Observable symptoms: Server `compute-batch-09` (Dell PowerEdge R740) is showing degraded performance for batch processing jobs.
Symptoms: TLS Works From Some Clients But Fails From Others · Observable symptoms and initial indicators for the ssl cert chain incomplete incident.
Symptoms: Traffic From Specific Source Not Taking Expected Path · Observable symptoms and initial indicators for the source routing policy miss incident.
Symptoms: User Auth Failing, OIDC Cert Expired, Fix Is Cloud KMS Rotation · Observable symptoms: **Domains:** security | kubernetes_ops | cloud.
Symptoms: VLAN Trunk Mistag · Observable symptoms: A new VMware ESXi host (esxi-node-07) was added to the cluster and connected to switch sw-dist-02 via a trunk port.

Library / Runbooks (57)

Operational Runbooks · Step-by-step procedures for common operational incidents.
Runbook · 1. Read the error output β€” Ansible prints the failing task, host, and error message 2. Check connectivity: ansible all -m ping -i inventory 3. Check. · Topics: Ansible · Aliases: ansible-galaxy, chef, config-management, inventory, playbooks, puppet, roles, saltstack
Runbook: Alert Storm (Flapping / Too Many Alerts) · Why: An alert storm is almost never 20 independent problems happening simultaneously. · Topics: Alerting Rules, Prometheus · Aliases: alert-rules, alerting-rules, alertmanager, datadog, metrics, monitoring, on-call, pagerduty, prometheusstack, promql, service-monitor
Runbook: ArgoCD Out of Sync · Symptoms: spec.replicas in diff. · Topics: GitOps · Aliases: argo, argocd, desired-state, drift-detection, flux
Runbook: Build Failure Triage · Why: CI pipelines often fail at a late step due to an error introduced in an earlier step β€” always diagnose from the first failure, not the last. · Topics: CI/CD, CI/CD Pipelines Realities · Aliases: cicd-real-world, circleci, continuous-delivery, continuous-integration, jenkins, pipeline-gotchas, pipeline-realities, pipelines, zuul
Runbook: Certificate Renewal Failed · Symptoms: CertificateRequest pending, Challenge in state pending or invalid. · Topics: TLS & PKI · Aliases: ca-chain, cert-manager, cert-rotation, certificates, pki, ssl, tls, x509
Runbook: Cloud Capacity Limit Hit · Why: Cloud providers have hundreds of different quotas. · Topics: Cloud Deep Dive, Terraform · Aliases: aws, azure, cloud, gcp, hcl, iac, iam, infrastructure-as-code, multi-cloud, tfstate, vpc
Runbook: Container Registry Pull Failure · Why: ImagePullBackOff can mean several different things β€” missing image tag, wrong registry URL, expired credentials, or network issue. · Topics: Docker / Containers, CI/CD · Aliases: circleci, containers, continuous-delivery, continuous-integration, docker-compose, dockerfile, images, jenkins, pipelines, zuul
Runbook: Credential Rotation (Exposed Secret) · Why: Every second the credential is active is a second an attacker can use it. · Topics: Secrets Management · Aliases: external-secrets, sealed-secrets, sops
Runbook: CVE Response (Critical Vulnerability) · Why: Not all critical CVEs are equally dangerous. · Topics: Security Scanning, Incident Triage · Aliases: cyber-security, incident-assessment, security, severity-classification, triage, trivy, vulnerability-scan
Runbook: Deploy Rollback · Why: Rolling back when the deployment is not the cause wastes time and may make things worse β€” confirm the error timeline matches the deployment time. · Topics: CI/CD, GitOps · Aliases: argo, argocd, circleci, continuous-delivery, continuous-integration, desired-state, drift-detection, flux, jenkins, pipelines, zuul
Runbook: Deployment Stuck / Rollout Stalled · Why: The rollout status command gives a human-readable summary of what the deployment controller is waiting for, which narrows the investigation. · Topics: Kubernetes Core · Aliases: deployments, k8s, kubernetes, openshift, pods, services
Runbook: Disaster Recovery · This runbook covers recovery from major failures. · Topics: Kubernetes Core · Aliases: deployments, k8s, kubernetes, openshift, pods, services
Runbook: Disk Full · Why: A system has multiple filesystems. · Topics: Filesystems & Storage · Aliases: btrfs, df, disk-debug, disk-troubleshooting, ext4, file-systems, filesystem, fstab, inode-exhaustion, inodes, io-troubleshooting, iostat
Runbook: DNS Resolution Failure · Why: DNS failures outside the cluster mean nothing β€” Kubernetes DNS only applies inside the cluster network. · Topics: DNS, Kubernetes Networking, Networking Troubleshooting · Aliases: cni, connectivity-issues, coredns, dig, dns-ops, domain-name-system, ingress, latency, network-debug, network-policies, nslookup, packet-loss
Runbook: etcd Backup & Restore · WARNING: Restore replaces ALL cluster data with the snapshot. · Topics: etcd · Aliases: etcd-backup, etcd-cluster, etcd-restore
Runbook: etcd High Latency / Slow Operations · Why: etcd exposes Prometheus metrics that show exactly which operation type is slow (WAL fsync, backend commit, apply). · Topics: etcd · Aliases: etcd-backup, etcd-cluster, etcd-restore
Runbook: Grafana Dashboard Blank / No Data · Why: If Prometheus has no data, the problem is upstream of Grafana. · Topics: Grafana · Aliases: dashboards, data-sources, panels
Runbook: Helm Upgrade Failed · 1. Bad values β€” invalid YAML, wrong image tag, missing required field 2. Template rendering error β€” Helm template syntax issue 3. Resource conflict β€”. · Topics: Helm · Aliases: helm-charts, helm-rollback, helm-upgrade, values-files
Runbook: High CPU (Runaway Process) · Why: You need to know exactly which process or processes are consuming CPU and whether this is expected load or a runaway before taking any action. · Topics: Linux Performance Tuning, Process Management · Aliases: flamegraph, job-control, kernel-tuning, kill, nohup, orphan, perf, ps, signals, strace, sysctl, tuning
Runbook: HPA Not Scaling · 1. metrics-server not installed β€” HPA needs metrics API 2. No resource requests defined β€” HPA percentage-based scaling requires CPU requests 3. · Topics: HPA / Autoscaling · Aliases: autoscaling, horizontal-pod-autoscaler, metrics-server
Runbook: HPA Thrashing (Rapid Scale Up/Down) · Why: The HPA status and events show the exact scaling decisions being made and why. · Topics: HPA / Autoscaling, Kubernetes Core · Aliases: autoscaling, deployments, horizontal-pod-autoscaler, k8s, kubernetes, metrics-server, openshift, pods, services
Runbook: ImagePullBackOff · 1. Wrong image tag β€” typo or nonexistent version 2. Image not imported into k3s β€” local images need docker save | sudo k3s ctr images import - 3. · Topics: Docker / Containers, Kubernetes Core · Aliases: containers, deployments, docker-compose, dockerfile, images, k8s, kubernetes, openshift, pods, services
Runbook: Ingress 404 · 1. Wrong path or host in ingress spec β€” path doesn't match app routes 2. Service name mismatch β€” ingress backend points to wrong service 3. No. · Topics: Kubernetes Networking · Aliases: cni, coredns, ingress, network-policies, service-mesh
Runbook: Ingress 502 Bad Gateway · Why: A 502 on one path means one backend service is broken. · Topics: Kubernetes Networking, Kubernetes Services & Ingress · Aliases: clusterip, cni, coredns, gateway-api, ingress, loadbalancer, network-policies, network-policy, nodeport, service-mesh, services
Runbook: Istio 503 Errors · [!WARNING] This is the most common cause of 503s after mesh enablement and the easiest to miss. · Topics: Service Mesh · Aliases: envoy, istio, linkerd, mtls, sidecar
Runbook: Kyverno Blocking Workloads · [!WARNING] Switching a policy to Audit mode unblocks deployments but also masks real violations. · Topics: Policy Engines · Aliases: admission-control, gatekeeper, kyverno, opa
Runbook: Load Balancer Health Check Failure · Why: The LoadBalancer service must have an assigned external IP or hostname before any health checks can run. · Topics: Load Balancing, Networking Troubleshooting · Aliases: connectivity-issues, haproxy, latency, load-balancer, network-debug, nginx-lb, packet-loss, reverse-proxy
Runbook: Log Pipeline Backpressure / Logs Not Appearing · Why: If only one service's logs are missing, the problem is likely that service's log format or labels. · Topics: Log Pipelines, Loki · Aliases: fluentbit, fluentd, log-aggregation, log-routing, logql, logs, promtail, structured-logging, vector
Runbook: Loki No Logs · 1. Promtail not running β€” DaemonSet crashed or misconfigured 2. Promtail can't reach Loki β€” wrong endpoint URL or network issue 3. Label mismatch β€”. · Topics: Loki · Aliases: log-aggregation, logql, logs, promtail
Runbook: Long-Running Query / Lock Contention · Why: You need to distinguish between a long-running query (bad plan, missing index) and a long-running idle transaction (application bug, connection. · Topics: PostgreSQL Operations, Database Locking · Aliases: deadlock, optimistic-locking, pg, pg-stat, pgbouncer, postgres, postgresql, replication-slots, row-lock, table-lock, vacuuming, wal
Runbook: MTU Mismatch · Why: MTU mismatches produce a very specific pattern β€” small requests succeed, large transfers silently stall or fail. · Topics: MTU, Networking Troubleshooting · Aliases: connectivity-issues, fragmentation, jumbo-frames, latency, mtu-blackhole, network-debug, packet-loss, path-mtu
Runbook: Network Partition (Split Brain / Partial Connectivity) · Why: A partition is not all-or-nothing. · Topics: Networking Troubleshooting · Aliases: connectivity-issues, latency, network-debug, packet-loss
Runbook: NetworkPolicy Block · 1. Deny-all policy applied β€” blocks all ingress/egress 2. Missing egress rule for DNS β€” pods can't resolve names (need port 53 to kube-system) 3. Label. · Topics: Kubernetes Networking · Aliases: cni, coredns, ingress, network-policies, service-mesh
Runbook: Node NotReady · Why: Knowing the count, zone, and node pool helps determine if this is a single hardware failure or a broader infrastructure problem. · Topics: Node Lifecycle & Maintenance · Aliases: cordon, daemonsets, drain, kubelet, node-drain, node-maintenance, node-upgrade, pdb, uncordon
Runbook: OOM Killer Activated · Why: The kernel logs exactly which process was killed, its PID, how much memory it used, and what the system-wide memory state was at the time of the kill. · Topics: Linux Memory Management · Aliases: cgroup-memory, hugepages, memory, numa, oom-killer, overcommit, swap, vmstat
Runbook: OOMKilled Container · Why: Exit code 137 is the definitive signal of an OOMKill (SIGKILL sent by the Linux OOM killer). · Topics: Kubernetes Core, OOMKilled · Aliases: deployments, k8s, kubernetes, memory, oom, openshift, out-of-memory, pods, resource-limits, services
Runbook: Pipeline Stuck / Hung Job · Why: A pipeline that is 'stuck' could be genuinely hung (infinite loop, deadlock, waiting for user input) or just slow (large test suite, slow network). · Topics: CI/CD · Aliases: circleci, continuous-delivery, continuous-integration, jenkins, pipelines, zuul
Runbook: Pod CrashLoopBackOff · Why: Confirms the CrashLoopBackOff and shows how long it has been failing, which sets urgency. · Topics: CrashLoopBackOff, Kubernetes Core · Aliases: crash-loop, crashloop, deployments, k8s, kubernetes, openshift, pod-restart, pods, services
Runbook: Pod Eviction · 1. Ephemeral storage exhaustion β€” container logs, emptyDir, or tmp files filling node disk 2. Memory pressure β€” node running out of allocatable memory. · Topics: Kubernetes Core · Aliases: deployments, k8s, kubernetes, openshift, pods, services
Runbook: PostgreSQL Connection Exhaustion · Why: You need to know exactly how full the connection pool is before deciding how aggressively to act. · Topics: PostgreSQL Operations · Aliases: pg, pg-stat, pgbouncer, postgres, postgresql, replication-slots, vacuuming, wal
Runbook: PostgreSQL Disk Space Critical · Why: You need to know which databases are consuming the most space so you can target investigation at the right place rather than guessing. · Topics: PostgreSQL Operations, Filesystems & Storage · Aliases: btrfs, df, disk-debug, disk-troubleshooting, ext4, file-systems, filesystem, fstab, inode-exhaustion, inodes, io-troubleshooting, iostat
Runbook: PostgreSQL Replication Lag · Why: The primary tracks exactly how far behind each replica is. · Topics: PostgreSQL Operations, Database Replication · Aliases: db-replication, pg, pg-stat, pgbouncer, postgres, postgresql, primary-replica, replication-lag, replication-slots, vacuuming, wal
Runbook: Prometheus Target Down · Why: Prometheus shows the exact scrape error for each down target. · Topics: Prometheus · Aliases: alertmanager, datadog, metrics, monitoring, prometheusstack, promql, service-monitor
Runbook: PVC Stuck in Pending · Why: The PVC events section is almost always the fastest path to the root cause β€” it shows provisioner messages, StorageClass errors, and topology. · Topics: Kubernetes Storage · Aliases: csi, pv, pvc, storage-class
Runbook: RBAC Forbidden · 1. Missing RoleBinding β€” Role exists but not bound to the service account 2. Wrong namespace β€” RoleBinding in different namespace than the resource 3. · Topics: RBAC · Aliases: clusterroles, rbac, roles, service-accounts
Runbook: Readiness Probe Failed · 1. Wrong probe path β€” endpoint doesn't exist or returns non-200 2. App not listening on expected port β€” check containerPort vs probe port 3. Slow. · Topics: Probes (Liveness/Readiness), Kubernetes Core · Aliases: deployments, healthcheck, k8s, kubernetes, liveness, openshift, pods, readiness, services, startup-probe
Runbook: Secret Rotation · Secret Rotation (Zero Downtime). · Topics: Secrets Management · Aliases: external-secrets, sealed-secrets, sops
Runbook: Systemd Service Crash Loop · Why: systemctl status gives you the current state, the last few log lines, the restart count, and the PID β€” all in one command. · Topics: systemd · Aliases: journalctl, services, systemctl, units
Runbook: Tempo No Traces · [!NOTE] Tempo is receive-only β€” unlike Promtail (which actively pulls logs from nodes), Tempo passively waits for traces to be pushed via OTLP. · Topics: Tempo · Aliases: distributed-tracing, spans, traces
Runbook: Terraform Drift Detection Response · Why: Before making any decisions, you need a full picture of every drifted resource. · Topics: Terraform, Terraform Deep Dive · Aliases: hcl, hcl-advanced, iac, infrastructure-as-code, remote-state, terraform-import, terraform-modules, terraform-state, tfstate, workspaces
Runbook: Terraform State Lock Stuck · Why: The lock error message contains the Lock ID, the holder, and the operation in progress. · Topics: Terraform · Aliases: hcl, iac, infrastructure-as-code, tfstate
Runbook: TLS Certificate Expiry · Why: Confirms whether the certificate is already expired (P1) or just expiring soon (P2) and which exact resources are affected. · Topics: TLS & PKI · Aliases: ca-chain, cert-manager, cert-rotation, certificates, pki, ssl, tls, x509
Runbook: Unauthorized Access Investigation · Why: You cannot contain what you haven't scoped. · Topics: Incident Triage, Audit Logging · Aliases: audit-trail, auditd, incident-assessment, security-logging, severity-classification, triage
Runbook: Velero Backup & Restore · [!WARNING] Cross-cluster restores are destructive. · Topics: Kubernetes Core · Aliases: deployments, k8s, kubernetes, openshift, pods, services
Runbook: VPC IP Exhaustion · Immediate fix: Enable prefix delegation (16 IPs per slot instead of 1). · Topics: Cloud Deep Dive · Aliases: aws, azure, cloud, gcp, iam, multi-cloud, vpc
Runbook: Zombie Processes Accumulating · Why: The number of zombies tells you severity. · Topics: Process Management, Linux Signals & Process Control · Aliases: disown, job-control, kill, nohup, orphan, pkill, process-states, ps, sigkill, signals, sigterm, zombie

Library / Cheatsheets (35)

Alerting Rules Cheatsheet · The for duration in an alert rule is the most important field for reducing false positives.
Bash Cheatsheet · Every production bash script starts with set -euo pipefail.
Cheatsheet · Ansible was created by Michael DeHaan in 2012.
Cicd Cheatsheet · CI/CD stands for Continuous Integration / Continuous Delivery (or Deployment).
Cloud Deep Dive Cheatsheet · IRSA (AWS) and Workload Identity (GCP) both work by projecting a signed service account token into the pod.
Cloud Ops Cheatsheet · One-liner: aws sts get-caller-identity is the cloud equivalent of whoami β€” run it first in every troubleshooting session to confirm which account.
Container Runtime Debug Cheatsheet · The debug hierarchy matches the abstraction layers: kubectl (K8s level) > crictl (container runtime level) > nsenter (Linux namespace level) > /proc.
Database Ops Cheatsheet · StatefulSet pod names are predictable (postgres-0, postgres-1, etc.) and PVCs are bound 1:1 to pods.
Datacenter Cheatsheet · POST stands for 'Power-On Self-Test' β€” the BIOS/UEFI firmware's built-in diagnostic that runs before any operating system loads.
Docker Cheatsheet · Each Dockerfile instruction creates a new image layer.
Etcd Operations Cheatsheet · etcd = 'et cetera daemon' β€” it stores configuration data, like the /etc directory in Unix but as a distributed, consistent key-value store.
Finops Cheatsheet · CPU and Memory limits behave differently: CPU is throttled (slowed down) when exceeding its limit, but Memory is OOMKilled (terminated).
Git Cheatsheet · Git was created by Linus Torvalds in 2005 in just two weeks, after the Linux kernel project lost access to the proprietary BitKeeper VCS.
Gitops Argocd Cheatsheet · GitOps was coined by Alexis Richardson (Weaveworks CEO) in 2017.
Helm Cheatsheet · Helm means 'ship's steering wheel' β€” fitting the nautical Kubernetes naming theme (kubernetes = Greek for 'helmsman').
K8S Operators Cheatsheet · The 'Operator pattern' was coined by CoreOS in 2016 to describe software that encodes human operational knowledge into Kubernetes controllers.
K8S Yaml Patterns Cheatsheet · Copy-paste-ready patterns for common Kubernetes resources.
Kubernetes Core Cheatsheet · Kubernetes (K8s) is Greek for 'helmsman' or 'pilot.' The 8 replaces the eight letters between K and s.
Linux Ops Cheatsheet · Signal numbers to memorize: SIGHUP (1) = reload config, SIGTERM (15) = graceful shutdown, SIGKILL (9) = force kill.
Modern Cli Cheatsheet · Most modern CLI tools (fd, ripgrep, bat, eza, dust) are written in Rust, which is why they are significantly faster than their C/shell predecessors.
Networking Cheatsheet · ClusterIP creates virtual IP + iptables/IPVS rules.
Observability Cheatsheet · Two complementary monitoring frameworks: RED for services (Request rate, Error rate, Duration) and USE for infrastructure (Utilization, Saturation.
Overview · Quick-reference cards for each topic area.
Phone Interview Devops Linux Cheatsheet · One-page-ish quick reference for phone screens.
Policy Engines Cheatsheet · Kyverno vs OPA Gatekeeper decision shortcut: if your team already knows Rego, use Gatekeeper.
Postmortem Slo Cheatsheet · SLO comes from Google's Site Reliability Engineering (SRE) book (2016).
Python Devops Cheatsheet · Python's json.tool module is a built-in JSON pretty-printer with no dependencies: python3 -m json.tool < file.json.
Secrets Management Cheatsheet · Kubernetes Secrets are base64-encoded, NOT encrypted.
Security Cheatsheet · The 'minimum viable security' for a Kubernetes pod: runAsNonRoot: true, readOnlyRootFilesystem: true, allowPrivilegeEscalation: false.
Service Mesh Cheatsheet · Istio is Greek for 'sail' β€” continuing the Kubernetes nautical naming theme.
Ssh Cheatsheet · SSH (Secure Shell) was created by Tatu Ylonen in 1995 at Helsinki University of Technology after a password-sniffing attack on the university network.
Systemd Cheatsheet · systemd = 'system daemon.' Created by Lennart Poettering and Kay Sievers (Red Hat) in 2010 to replace SysVinit.
Terraform Cheatsheet · Terraform = 'to transform the earth' (Latin: terra + forma).
Tls Pki Cheatsheet · TLS = Transport Layer Security, the successor to SSL (Secure Sockets Layer).
Troubleshooting Flows Cheatsheet · Quick decision trees for common DevOps problems.

Library / Skillchecks (33)

DevOps Skill Check Pack (internals-first, visual, jargon-explained) · 1. linux.fundamentals.md 2. bash.skillcheck.md 3. git.skillcheck.md 4. networking.fundamentals.md 5. cloud.basics.md 6. python.automation.md 7.
Skill Checks · Self-assessment quizzes per topic. See the full skill check pack for details
Skillcheck · Ansible is: controller connects over SSH/WinRM, transfers a module, runs it, and returns structured results. · Topics: Ansible · Aliases: ansible-galaxy, chef, config-management, inventory, playbooks, puppet, roles, saltstack
Skillcheck: Alerting Rules · Good alerts tell you about user-impacting problems, not infrastructure noise. · Topics: Alerting Rules, Prometheus · Aliases: alert-rules, alerting-rules, alertmanager, datadog, metrics, monitoring, on-call, pagerduty, prometheusstack, promql, service-monitor
Skillcheck: Bash · Bash takes text and does parsing + expansions + word splitting + globbing, then execs programs. · Topics: Bash / Shell Scripting · Aliases: advanced-bash, bash, scripting, shell, shell-scripting
Skillcheck: CI/CD · CI/CD is an event-driven DAG of jobs running on runners, producing artifacts and optionally deploying them. · Topics: CI/CD · Aliases: circleci, continuous-delivery, continuous-integration, jenkins, pipelines, zuul
Skillcheck: Cloud Basics · Cloud is: API control plane creating/controlling data plane resources (compute/storage/network).
Skillcheck: Cloud Providers · Managed Kubernetes adds cloud-specific glue. · Topics: Cloud Deep Dive · Aliases: aws, azure, cloud, gcp, iam, multi-cloud, vpc
Skillcheck: Container Runtime Debug · When kubectl isn't enough, go deeper. · Topics: Container Runtimes · Aliases: cgroups, container-runtime-debug, containerd, cri-o, namespaces, runc
Skillcheck: Database Ops · Databases are stateful workloads. · Topics: Database Operations · Aliases: backup-restore, database, mysql, statefulsets
Skillcheck: Datacenter · A data center is a physical building that provides power, cooling, and network to racks of servers. · Topics: Rack & Stack, Out-of-Band Management, RAID, Server Hardware · Aliases: airflow, bare-metal-provisioning, bmc, cabling, cpu, datacenter-provisioning, dimm, disk-arrays, dmesg, dmidecode, hba, idrac
Skillcheck: DevOps Roadmap (Expanded) · This is the roadmap (in the order mentioned in the transcript), with 10 bullet questions per topic and short indented answers, roughly easy -> harder.
Skillcheck: Docker · Containers are just processes with namespaces (isolation) and cgroups (limits). · Topics: Docker / Containers · Aliases: containers, docker-compose, dockerfile, images
Skillcheck: etcd · Rate yourself 0-2 on each item: 0 = never done, 1 = done with help, 2 = confident. · Topics: etcd · Aliases: etcd-backup, etcd-cluster, etcd-restore
Skillcheck: FinOps · You pay for requests, not usage. · Topics: FinOps · Aliases: cloud-billing, cost-optimization, reserved-instances, spot
Skillcheck: Git · Git is a content-addressed object database plus refs (names) that point at commits. · Topics: Git · Aliases: branching, merging, rebase, version-control
Skillcheck: GitOps · Git is the single source of truth. · Topics: GitOps · Aliases: argo, argocd, desired-state, drift-detection, flux
Skillcheck: Helm & Release Ops · Rate yourself 0-2 on each item: 0 = never done, 1 = done with help, 2 = confident. · Topics: Helm · Aliases: helm-charts, helm-rollback, helm-upgrade, values-files
Skillcheck: Kubernetes · Kubernetes is a set of controllers that keep actual cluster state matching desired state stored in the API server. · Topics: Kubernetes Core, Probes (Liveness/Readiness), HPA / Autoscaling · Aliases: autoscaling, deployments, healthcheck, horizontal-pod-autoscaler, k8s, kubernetes, liveness, metrics-server, openshift, pods, readiness, services
Skillcheck: Kubernetes Operators · A CRD extends the K8s API with custom types. · Topics: K8s Ecosystem · Aliases: cncf, k8s-landscape, kubernetes-ecosystem
Skillcheck: Kubernetes Under the Covers · A practical, internals-first learning guide (with diagrams, β€œwhat actually happens”, and drills). · Topics: Kubernetes Core, Kubernetes Networking, Node Lifecycle & Maintenance · Aliases: cni, cordon, coredns, daemonsets, deployments, drain, ingress, k8s, kubelet, kubernetes, network-policies, node-drain
Skillcheck: Linux Fundamentals · Linux is a kernel that offers syscalls. · Topics: Linux Fundamentals, systemd, Package Management · Aliases: apt, dpkg, filesystem, journalctl, linux, linux-basics, permissions, rpm, services, systemctl, units, yum
Skillcheck: Modern CLI Tools · Classic Unix tools (find, grep, cat, ls, du, cd) are universal but show their age. · Topics: jq / JSON Processing, ripgrep (rg), fzf, fd, Modern CLI Tools · Aliases: bat, cli-tools, code-search, dust, eza, fast-find, fast-grep, fd-find, file-search, fuzzy-finder, interactive-selection, jq-filters
Skillcheck: Networking Fundamentals · Networking is layers: bits -> frames -> packets -> streams -> HTTP. · Topics: TCP/IP, DNS, VLANs, Routing · Aliases: 802.1q, asymmetric-routing, bgp, coredns, dig, dns-ops, domain-name-system, gateway, ip, layers, networking, networking-fundamentals
Skillcheck: Observability · Observability is signals: metrics, logs, traces. · Topics: Prometheus, Grafana, Loki, Tempo · Aliases: alertmanager, dashboards, data-sources, datadog, distributed-tracing, log-aggregation, logql, logs, metrics, monitoring, panels, prometheusstack
Skillcheck: Policy Engines · RBAC controls who can act. · Topics: Policy Engines · Aliases: admission-control, gatekeeper, kyverno, opa
Skillcheck: Postmortems & SLOs · SLIs measure user experience. · Topics: Postmortems & SLOs · Aliases: blameless, devops, error-budget, postmortem, sla, sli, slo
Skillcheck: Python Automation · Python is an interpreter + standard library that's great at orchestration and data shaping. · Topics: Python Automation · Aliases: automation, perl, python, scripting, software-development
Skillcheck: Secrets Management · K8s Secrets are base64, not encrypted. · Topics: Secrets Management · Aliases: external-secrets, sealed-secrets, sops
Skillcheck: Security (Expanded) · Rate yourself 0-2 on each item: 0 = never done, 1 = done with help, 2 = confident. · Topics: Security Scanning · Aliases: cyber-security, security, trivy, vulnerability-scan
Skillcheck: Service Mesh · A service mesh is a sidecar proxy per pod that handles mTLS, retries, traffic splitting, and observability without app changes. · Topics: Service Mesh · Aliases: envoy, istio, linkerd, mtls, sidecar
Skillcheck: Terraform / IaC · Terraform builds a dependency graph, then compares desired config vs state vs real APIs to create a plan. · Topics: Terraform · Aliases: hcl, iac, infrastructure-as-code, tfstate
Skillcheck: TLS & PKI · TLS encrypts traffic using certificates signed by a certificate authority (CA). · Topics: TLS & PKI · Aliases: ca-chain, cert-manager, cert-rotation, certificates, pki, ssl, tls, x509

Library / Guides (16)

AI-Assisted DevOps Cookbook · Concrete recipes for common DevOps tasks using AI tools (ChatGPT, Codex, Claude Code). · Topics: AI Tools for DevOps · Aliases: ai, ai-tools, chatgpt, claude, codex, copilot, generativeai, llm, prompt-engineering
Bare-Metal Provisioning · Reference guide for out-of-band management, automated OS deployment, and server lifecycle operations on Dell PowerEdge hardware.
CI Pipeline Documentation · The GitHub Actions CI pipeline runs on every push to main/develop and on pull requests to main. · Topics: CI/CD · Aliases: circleci, continuous-delivery, continuous-integration, jenkins, pipelines, zuul
Cluster Management Guide · The default inventory (hosts.local.yml) uses connection: local for a single-node k3s cluster on the current machine.
Cluster Upgrade Exercise · Practice a rolling k3s upgrade using Ansible, following production-safe patterns.
Dell Server Management · Reference guide for managing Dell PowerEdge servers in a DevOps context.
DevOps Learning Roadmap · Developer pushes code to GitHub.
Gitops Example · ArgoCD watches a Git repository for changes and automatically syncs the desired state (Helm chart + values) to the Kubernetes cluster.
Mental-Model-First Learning Guide · A learning method that builds a reduced, accurate internal model of a topic before introducing details. · Topics: Mental Models (Core) · Aliases: conceptual-models, mental-model-first, mental-models
Modern Cli Tools · Reference guide for the new generation of command-line tools that replace classic Linux utilities.
Observability Architecture · The GrokDevOps observability stack provides metrics, logging, and tracing using open-source tools deployed on Kubernetes. · Topics: Prometheus, Grafana, Loki, Tempo · Aliases: alertmanager, dashboards, data-sources, datadog, distributed-tracing, log-aggregation, logql, logs, metrics, monitoring, panels, prometheusstack
Rack & Data Center Operations · Reference guide for physical rack layout, cabling, power distribution, thermal management, and asset inventory.
Security Scanning · Container images can include vulnerable packages inherited from base images or installed dependencies.
Tools Reference · Reference guide for all CLI tools, validation scripts, and infrastructure scripts in grokdevops.
Topic Tag Cloud & Content Index · Click any area to see everything in it β€” topic packs, exercises, flashcard decks, scenarios, runbooks, and more. Bigger tag = more content.
Troubleshooting · Common causes: - Image not imported into k3s (docker save | sudo k3s ctr images import -) - Wrong image tag in values file - Port conflict.

Library / War Stories (51)

DNS: The Eternal Enemy · I was a junior SRE at a healthcare SaaS company, about 500 employees.
From Monolith to Misery · I was an SRE at a B2B SaaS company with a Rails monolith that had been growing for seven years.
It Was Always DNS · It was a Tuesday morning in November, and our e-commerce platform was bleeding money.
One Character from Disaster · We had 1,400 servers across three data centers.
The 3 AM Cert Expiry · I was the sole SRE at a 40-person fintech startup.
The Auth System Swap · I was the security-focused SRE at a healthcare scheduling platform β€” 45,000 daily active users, about 800 API requests per second.
The Autoscaler That Almost Bankrupted Us · We ran a 60-node GKE cluster on Google Cloud.
The Backup We Never Tested · Our PostgreSQL 14 database backed a B2B SaaS platform with 1,400 paying customers.
The Cascading Timeout · I was an SRE at an online food delivery platform β€” about 400 engineers, 12 million orders per month, running roughly 45 microservices on Kubernetes.
The Case of the Missing Packets · We ran a microservices platform on bare-metal Linux servers behind a pair of HA firewalls.
The CI/CD Pipeline Rewrite · I was the build engineer (fancy title for 'the person Jenkins calls at 2 AM') at a company with 35 development teams.
The Clock Skew Catastrophe · I was a mid-level SRE at a financial data aggregation company.
The Cloud Bill Surprise · I was the senior SRE at a healthcare analytics company.
The Config Management Lie · We had 140 servers managed by Ansible.
The Container That Worked on My Machine · We had a Go service that processed PDF reports -- nothing exotic, just a wrapper around a C library via cgo.
The Cost of No Staging · We were a 30-person startup.
The Database Migration Weekend · I was the DBA-slash-SRE at a logistics company.
The Database That Wasn't Backing Up · Enterprise logistics company, about 2,000 employees.
The Datacenter Exit · I was the infrastructure lead at a media company when the landlord dropped the bomb: our datacenter lease wasn't being renewed.
The Deploy That Ate Prod · Mid-size e-commerce company, 200 engineers, running about 80 microservices on EKS.
The DNS Provider Switch · I was one of two SREs at a 200-person company running a SaaS product for real estate agents.
The Documentation That Didn't Exist · We ran a bespoke event-streaming platform built on Kafka, Flink, and a custom ingestion layer that Marcus had written.
The Firewall Rule That Blocked Itself · I was a network engineer at a regional hospital system, about 1,200 employees.
The Git Deploy That Deployed Nothing · We had a fairly standard GitOps-style deploy pipeline: push to main, GitHub Actions builds the image, tags it with the output of git describe --tags.
The Intern and the DROP TABLE · It was Jake's second week.
The Kubernetes Migration That Took a Year · I was the lead SRE at a mid-size e-commerce company running 60 services on a fleet of EC2 instances managed by Ansible.
The Leap Second Incident · It was June 30th, and I was on-call for a fleet of about 1,200 Linux servers running a real-time ad-bidding platform.
The Load Balancer Lie · Series B startup, real-time analytics platform.
The Log That Filled the Disk · Small B2B SaaS company, about 25 engineers.
The Memory Leak Marathon · I was the SRE lead at a ride-sharing company in a mid-size city.
The Metrics That Lied · We had beautiful dashboards.
The Monitoring Save · Friday, 4:47 PM. Most of the team had already mentally checked out. I was closing browser tabs and thinking about dinner when my phone buzzed: a.
The Monitoring We Ignored · We had 847 alerts configured in PagerDuty across 12 services.
The Network Change Window · We were migrating from a flat /16 network to a segmented hub-and-spoke topology across three AWS VPCs.
The Observability Migration · I was the observability lead at an insurance tech company.
The Permissions Avalanche · Large media company, about 3,000 employees.
The Phantom Latency Spike · Our API gateway served around 12,000 requests per second for a fintech platform.
The Postmortem Nobody Read · We wrote excellent postmortems.
The Rollback That Wasn't · I was the tech lead at a mid-size travel booking platform, about 150 engineers.
The Secret Rotation We Postponed · Our main API key for the Stripe payment integration was 26 months old.
The Secrets in the Repo · I'd just gotten a new MacBook Pro.
The Single Point of Failure · We had a server called legacy-core-01.
The Split-Brain Nightmare · I was a database engineer at a financial services company β€” about 600 employees, heavily regulated, zero tolerance for data inconsistency.
The SSL Handshake Timeout · We ran a B2B SaaS platform serving about 300 enterprise customers.
The Technical Debt Interest Payment · Two years ago, we needed to ship a feature fast.
The Terraform Plan That Would Have Destroyed Prod · We were running Terraform 1.5 managing about 280 AWS resources: VPCs, RDS instances, ECS clusters, ALBs, S3 buckets β€” the full stack for a B2B SaaS.
The Terraform State Disaster · I was the platform engineer at a fintech startup that had grown from 5 to 80 engineers in two years.
The Test We Never Wrote · It was a Rails monolith serving about 40,000 requests per minute.
The Zombie Cron Job · I was an SRE at a mid-size e-learning platform β€” about 300 employees, 2 million active students.
War Stories Collection · First-person accounts of production incidents, migrations, mysteries, close calls, and hard lessons.
When the Queue Backed Up · I was a senior SRE at an insurance claims processing company.

Library / Comparisons (25)

Comparison: Alerting & Paging · Last meaningful update consideration: 2026-03 Verdict (opinionated): PagerDuty for mature orgs that need reliable escalation and analytics.
Comparison: Caching · Last meaningful update consideration: 2026-03 Verdict (opinionated): Redis unless you only need simple key-value caching with no data structures, no.
Comparison: CI Platforms · Last meaningful update consideration: 2026-03 Verdict (opinionated): GitHub Actions for most teams β€” the ecosystem integration is unbeatable.
Comparison: Cloud Compute (EC2 vs Azure VMs) · AWS EC2 vs Azure Virtual Machines β€” service crosswalk for engineers who know one cloud and need the other.
Comparison: Cloud Storage (S3 vs Blob Storage) · AWS S3 vs Azure Blob Storage β€” service crosswalk for engineers who know one cloud and need the other.
Comparison: CNI Plugins · Last meaningful update consideration: 2026-03 Verdict (opinionated): Cilium for new clusters β€” eBPF-based networking is the future and it comes with.
Comparison: Configuration Management · Last meaningful update consideration: 2026-03 Verdict (opinionated): Ansible unless you need continuous convergence on long-lived servers (then Puppet).
Comparison: Container Orchestrators · Last meaningful update consideration: 2026-03 Verdict (opinionated): Kubernetes unless your team is under 5 engineers and AWS-only β€” then ECS is.
Comparison: GitOps CD · Last meaningful update consideration: 2026-03 Verdict (opinionated): ArgoCD for most Kubernetes teams.
Comparison: Image Scanners · Last meaningful update consideration: 2026-03 Verdict (opinionated): Trivy for the broadest OSS scanning (images, IaC, SBOM, secrets).
Comparison: Infrastructure as Code Tools · Last meaningful update consideration: 2026-03 Verdict (opinionated): Terraform (or OpenTofu) for multi-cloud.
Comparison: Ingress Controllers · Last meaningful update consideration: 2026-03 Verdict (opinionated): Ingress-NGINX for most K8s deployments.
Comparison: Kubernetes Templating · Last meaningful update consideration: 2026-03 Verdict (opinionated): Helm for packaging and distributing applications.
Comparison: Local Dev for Kubernetes · Last meaningful update consideration: 2026-03 Verdict (opinionated): Tilt for most teams β€” the live-reload workflow and Tiltfile flexibility are unmatched.
Comparison: Local Kubernetes Clusters · Last meaningful update consideration: 2026-03 Verdict (opinionated): k3d for speed and lightweight development clusters.
Comparison: Logging Platforms · Last meaningful update consideration: 2026-03 Verdict (opinionated): Loki for K8s-native teams on a budget β€” it is the Prometheus of logs.
Comparison: Managed Kubernetes · Last meaningful update consideration: 2026-03 Verdict (opinionated): GKE for best Kubernetes experience and features.
Comparison: Messaging · Last meaningful update consideration: 2026-03 Verdict (opinionated): SQS for simple AWS queues β€” zero ops, just works.
Comparison: Metrics Platforms · Last meaningful update consideration: 2026-03 Verdict (opinionated): Prometheus + Grafana Cloud for cost control and ecosystem fit.
Comparison: Policy Engines · Last meaningful update consideration: 2026-03 Verdict (opinionated): Kyverno for simplicity and YAML-native policies β€” most teams start here.
Comparison: Relational Databases · Last meaningful update consideration: 2026-03 Verdict (opinionated): PostgreSQL unless Aurora's auto-scaling storage model or MySQL read replica.
Comparison: Secrets Management · Last meaningful update consideration: 2026-03 Verdict (opinionated): HashiCorp Vault for multi-cloud and complex secret lifecycle needs.
Comparison: Service Meshes · Last meaningful update consideration: 2026-03 Verdict (opinionated): No mesh until you actually need mTLS at scale or fine-grained traffic management.
Comparison: Tracing Platforms · Last meaningful update consideration: 2026-03 Verdict (opinionated): Tempo if you are in the Grafana ecosystem β€” it is the cheapest to operate and.
Tool Comparison Matrices · Honest, opinionated tool comparisons from an operator's perspective.

Library / Curriculum (23)

Coverage Gaps Analysis · Based on concept coverage across docs, runbooks, labs, incidents, skillchecks, and cards.
K8s learning plan · A hands-on project to learn Kubernetes, CI/CD, Terraform, and cloud-native infrastructure.
Level 1: Foundations · Linux, Bash, Git, Python, Networking basics.
Level 2: Container Platform · Docker, Dockerfile, images, registries, scanning.
Level 3: Production Kubernetes · Core K8s objects, probes, rollouts, services, DNS, scaling.
Level 4: Operations & Observability · Helm, Prometheus, Loki, Tempo, Grafana, CI/CD, GitOps, Ansible, Terraform.
Level 5: SRE & Incident Response · Incidents, chaos engineering, forensics, postmortems, interview preparation.
Level 6: Advanced Platform Engineering · Service mesh, GitOps, operators, policy engines, secrets management.
Level 7: SRE & Cloud Operations · SLOs, alerting, cloud providers, FinOps, database operations, etcd, disaster recovery.
Master Curriculum: 40 Weeks · This is the comprehensive reference plan β€” the full day-by-day schedule covering every topic, case study, and scenario.
Track: Advanced Platform Engineering · This track covers the tools and patterns that turn a basic Kubernetes cluster into a production-grade platform.
Track: Cloud & FinOps · This track covers cloud provider specifics, cost optimization, database operations, and multi-cluster management.
Track: Containers · Docker fundamentals, image building, runtime security. · Topics: Docker / Containers · Aliases: containers, docker-compose, dockerfile, images
Track: Foundations · Bash, Linux, Git, Python basics. · Topics: Linux Fundamentals, Bash / Shell Scripting, Git, Docker / Containers · Aliases: advanced-bash, bash, branching, containers, docker-compose, dockerfile, filesystem, images, linux, linux-basics, merging, permissions
Track: Helm & Release Ops · Values, templating, upgrades, rollbacks, release lifecycle. · Topics: Helm, GitOps · Aliases: argo, argocd, desired-state, drift-detection, flux, helm-charts, helm-rollback, helm-upgrade, values-files
Track: Incident Response · Incidents, forensics, runbooks, postmortems, interview scenarios. · Topics: Incident Response · Aliases: chaos-engineering, forensics, incidents, postmortem, rca, triage
Track: Infrastructure · Physical infrastructure, Dell server management, bare-metal provisioning, and data center operations. · Topics: Ansible, Terraform · Aliases: ansible-galaxy, chef, config-management, hcl, iac, infrastructure-as-code, inventory, playbooks, puppet, roles, saltstack, tfstate
Track: Kubernetes Core · Pods, deployments, services, config, DNS, probes, scaling, RBAC. · Topics: Kubernetes Core, Probes (Liveness/Readiness), Kubernetes Networking, Kubernetes Storage, RBAC · Aliases: clusterroles, cni, coredns, csi, deployments, healthcheck, ingress, k8s, kubernetes, liveness, network-policies, openshift
Track: Modern CLI Tools · Fast, composable command-line tools for searching, filtering, and transforming data.
Track: Observability · Prometheus, Loki, Tempo, Grafana. · Topics: Prometheus, Grafana, Loki, Tempo · Aliases: alertmanager, dashboards, data-sources, datadog, distributed-tracing, log-aggregation, logql, logs, metrics, monitoring, panels, prometheusstack
Track: Professional Skills · Career development and organizational dynamics: engineering your career, vendor management, and change management.
Track: SRE & Reliability Engineering · This track covers the practices, tools, and mindset for running reliable production systems.
Training Curriculum · Structured learning paths through existing repo content.

Library / Domains (9)

Overview · > Domain guide: browse all CLI tools content organized as a learning sequence
Overview · > Domain guide: browse all cloud content organized as a learning sequence
Overview · > Domain guide: browse all datacenter and hardware content organized as a learning sequence
Overview · > Domain guide: browse all DevOps tooling content organized as a learning sequence
Overview · > Domain guide: browse all Kubernetes content organized as a learning sequence
Overview · > Domain guide: browse all Linux content organized as a learning sequence
Overview · > Domain guide: browse all networking content organized as a learning sequence
Overview · > Domain guide: browse all observability content organized as a learning sequence
Overview · > Domain guide: browse all security content organized as a learning sequence

Library / Other (374)

Adversarial Interview Gauntlet (30 sequences) · 30 escalating interview sequences where each question builds on the previous answer, going 5 rounds deep.
Alerting Rules Drills · Good alerts have three properties: Actionable (someone can do something about it), Relevant (it signals real user impact), Contextualized (includes. · Topics: Alerting Rules, Prometheus · Aliases: alert-rules, alerting-rules, alertmanager, datadog, metrics, monitoring, on-call, pagerduty, prometheusstack, promql, service-monitor
Architecture & Design Models · Structural patterns for resilience and maintainability
Certification Exam Prep · Study guides mapping wiki content to industry certification objectives
Certification Prep: AWS SAA β€” Solutions Architect Associate · Study guide and preparation resources for the AWS Solutions Architect Associate exam
Certification Prep: CKA β€” Certified Kubernetes Administrator · Study guide and preparation resources for the Certified Kubernetes Administrator exam
Certification Prep: CKAD β€” Certified Kubernetes Application Developer · Study guide and preparation resources for the Certified Kubernetes Application Developer exam
Certification Prep: CKS β€” Certified Kubernetes Security Specialist · Study guide and preparation resources for the Certified Kubernetes Security Specialist exam
Certification Prep: HashiCorp Terraform Associate · Study guide and preparation resources for the HashiCorp Terraform Associate exam
Certification Prep: PCA β€” Prometheus Certified Associate · Study guide and preparation resources for the Prometheus Certified Associate exam
CI/CD Drill Answers · Look for the jobs: section to see defined jobs and their steps.
CI/CD Drills · 10 drills for CI pipeline and security scanning operations. · Topics: CI/CD · Aliases: circleci, continuous-delivery, continuous-integration, jenkins, pipelines, zuul
Cloud Deep Dive Drills · Cloud workload identity follows the same pattern across providers: a Kubernetes ServiceAccount is mapped to a cloud IAM identity via OIDC federation. · Topics: Cloud Deep Dive · Aliases: aws, azure, cloud, gcp, iam, multi-cloud, vpc
Cloud Ops Drills · AWS IAM troubleshooting order: Role attached? · Topics: Cloud Deep Dive · Aliases: aws, azure, cloud, gcp, iam, multi-cloud, vpc
Container Runtime Drills · On Kubernetes nodes, crictl is the standard CLI for CRI-compatible runtimes (containerd, CRI-O). · Topics: Container Runtimes · Aliases: cgroups, container-runtime-debug, containerd, cri-o, namespaces, runc
Database Ops Drills · The database backup hierarchy: Logical (pg_dump β€” portable, slow for large DBs), Physical (pg_basebackup β€” fast, same PG version only), Continuous (WAL. · Topics: Database Operations · Aliases: backup-restore, database, mysql, statefulsets
Datacenter Drills · The physical troubleshooting order: Power (is it plugged in and on?) -> POST (does the BIOS initialize?) -> Network (can you reach the BMC/OS?) ->. · Topics: Rack & Stack, Out-of-Band Management · Aliases: airflow, bare-metal-provisioning, bmc, cabling, datacenter-provisioning, idrac, ilo, ipmi, labeling, oob, oob-provisioning, pdu
Debugging & Diagnosis Models · Mental models for finding the cause of problems
Decision Tree: A Secret Was Exposed · Starting Question: 'A secret (credential, key, token) was exposed β€” what do I do?' Estimated traversal: 2-5 minutes.
Decision Tree: Alert Fired β€” Is This Real? · Starting Question: 'An alert fired β€” is this a real incident or noise?' Estimated traversal: 2-4 minutes.
Decision Tree: Certificate Is Expiring β€” What Do I Do? · Starting Question: 'A TLS certificate is expiring β€” what's the renewal path?' Estimated traversal: 3-5 minutes.
Decision Tree: Container Running as Root · Starting Question: 'I found a container running as root β€” what's the risk and what do I do?' Estimated traversal: 2-5 minutes.
Decision Tree: Dependency Has a CVE · Starting Question: 'A dependency we use has a published CVE β€” what do I do?' Estimated traversal: 2-5 minutes.
Decision Tree: Deployment Is Stuck · Starting Question: 'A Kubernetes deployment is stuck β€” what's blocking it?' Estimated traversal: 2-5 minutes.
Decision Tree: Disk Is Filling Up · Starting Question: 'Disk usage is high or growing β€” what's consuming space?' Estimated traversal: 2-5 minutes.
Decision Tree: Do I Need a Service Mesh? · Starting Question: 'Should we adopt a service mesh for this system?' Estimated traversal: 3-5 minutes.
Decision Tree: How to Handle This Config Change? · Starting Question: 'I need to make a config change to a running system β€” what process?' Estimated traversal: 3-5 minutes.
Decision Tree: I Found a Vulnerability · Starting Question: 'I found a security vulnerability β€” what do I do?' Estimated traversal: 2-5 minutes.
Decision Tree: Latency Has Increased · Starting Question: 'Response latency has spiked β€” where is it?' Estimated traversal: 3-5 minutes.
Decision Tree: Managed vs Self-Hosted Service · Starting Question: 'Should we use a managed service or self-host?' Estimated traversal: 3-5 minutes.
Decision Tree: Memory Usage Is High · Starting Question: 'Memory usage is high on a host or in a pod β€” why?' Estimated traversal: 2-4 minutes.
Decision Tree: Monolith vs Microservices · Starting Question: 'Should we decompose this into microservices?' Estimated traversal: 4-5 minutes.
Decision Tree: Node Is NotReady · Starting Question: 'A Kubernetes node is in NotReady state β€” what's wrong?' Estimated traversal: 3-5 minutes.
Decision Tree: Pod Won't Start · Starting Question: 'A pod is stuck and won't start β€” what state is it in?' Estimated traversal: 2-4 minutes.
Decision Tree: Roll Back or Fix Forward? · Starting Question: 'Something broke after a deployment β€” should I roll back or fix forward?' Estimated traversal: 2-5 minutes.
Decision Tree: Scale Up or Optimize First? · Starting Question: 'The system is struggling under load β€” should I scale up or optimize?' Estimated traversal: 3-5 minutes.
Decision Tree: Service Returning 5xx Errors · Starting Question: 'My service is returning 5xx errors β€” where do I start?' Estimated traversal: 2-5 minutes.
Decision Tree: Should I Automate This? · Starting Question: 'I keep doing this manual task β€” should I automate it?' Estimated traversal: 3-5 minutes.
Decision Tree: Should I Page Someone? · Starting Question: 'Something looks wrong β€” should I page the on-call?' Estimated traversal: 2-5 minutes.
Decision Tree: Suspicious Activity Detected · Starting Question: 'Something looks like a security incident β€” is it?' Estimated traversal: 2-5 minutes.
Decision Tree: Sync vs Async Communication · Starting Question: 'Should service A communicate with service B synchronously or asynchronously?' Estimated traversal: 3-4 minutes.
Decision Tree: Where Should This Run? · Starting Question: 'Where should we deploy this workload?' Estimated traversal: 3-5 minutes.
Decision Tree: Which Database for This Workload? · Starting Question: 'What type of database should we use for this workload?' Estimated traversal: 4-5 minutes.
Decision Trees · Flowcharts for common operational decisions.
Docker Drills · Docker exit codes tell you what happened: 0 = clean exit, 1 = application error, 137 = OOM-killed or SIGKILL (128+9), 139 = segfault (128+11), 143 =. · Topics: Docker / Containers · Aliases: containers, docker-compose, dockerfile, images
Drill: Advanced Stash Usage · Use git stash to save work-in-progress changes with messages, stash specific files, and understand pop vs apply.
Drill: Analyze Network Path Quality with mtr · Use mtr to combine traceroute and ping functionality to identify problem hops and measure path quality.
Drill: Basic Port Scanning with nmap · Use nmap to perform basic port scans and service detection to verify what is actually listening and reachable.
Drill: Build an Event Timeline for Debugging · Use kubectl events and describe to build a chronological timeline of cluster events for incident investigation.
Drill: Capture and Filter Traffic with tcpdump · Use tcpdump to capture network traffic and apply filters by host, port, and protocol for targeted debugging.
Drill: Check Resource Quotas and Limit Ranges · Inspect resource quotas and limit ranges in a namespace to understand resource constraints and capacity.
Drill: Cherry-Pick Specific Commits Between Branches · Use git cherry-pick to selectively apply specific commits from one branch to another without merging all changes.
Drill: Code Search with ripgrep (rg) · Use ripgrep for fast code search with file type filters, context lines, multiline patterns, and glob-based exclusions.
Drill: Create a Systemd Drop-in Override · Create a systemd drop-in override to modify a service's configuration without editing the original unit file.
Drill: Debug a Pod Stuck in Pending State · Systematically diagnose why a pod is stuck in Pending state by checking scheduling constraints, resources, and events.
Drill: Debug DNS with dig · Use dig to perform DNS lookups, trace delegation chains, query specific servers, and perform reverse lookups.
Drill: Debug Network Interfaces with ethtool · Use ethtool to check link status, negotiated speed, error counters, and driver information for network interfaces.
Drill: Effective kubectl get Output Formats · Use kubectl get with various output formats including wide, yaml, jsonpath, and custom-columns for efficient inspection.
Drill: Examine TCP Connection States with ss · Use ss to inspect TCP connection states, identify CLOSE_WAIT and TIME_WAIT accumulation, and diagnose connection issues.
Drill: Explore Systemd Dependency Tree · Use systemctl list-dependencies to understand how systemd units relate to and depend on each other.
Drill: Explore the /proc Filesystem for a Process · Read /proc/PID entries to inspect a running process's command line, environment, file descriptors, memory maps, and status.
Drill: fd as a Modern find Replacement · Use fd for fast, user-friendly file finding with type filters, regex, exec, and exclusion patterns.
Drill: Filter Journal Entries by Time, Unit, and Priority · Use journalctl to filter systemd journal entries by time range, service unit, and log priority level.
Drill: Find a Regression with git bisect · Use git bisect to perform a binary search through commit history and find the exact commit that introduced a bug.
Drill: Find Open Files with lsof · Use lsof to find open files, detect deleted files still holding disk space, and identify network connections by process.
Drill: Find What Is Consuming Disk Space · Use du, sort, and ncdu to identify directories and files consuming the most disk space.
Drill: fzf Integration Patterns · Use fzf for interactive fuzzy selection in common workflows: file selection, history search, git branches, and process killing.
Drill: fzf Interactive Selection · Use fzf's interactive selection features β€” single select, multi-select, preview, and custom key bindings β€” to build efficient selection workflows from.
Drill: Get Logs from Multi-Container Pods · Retrieve logs from specific containers in multi-container pods, including init containers and previous crashed instances.
Drill: HTTP Debugging with curl · Use curl with verbose flags, custom resolution, and connection overrides to debug HTTP/HTTPS issues.
Drill: Inspect and Decode ConfigMaps and Secrets · Inspect configmaps and decode secrets to verify application configuration in a Kubernetes cluster.
Drill: Inspect and Manage the ARP Table · Inspect the ARP table to verify MAC-to-IP mappings, detect anomalies, and understand local network neighbor resolution.
Drill: Inspect Cgroup Resource Limits and Usage · Inspect cgroup resource limits and current usage for a systemd service to understand resource constraints.
Drill: Interactive Rebase to Clean Up Commits · Use interactive rebase to squash, reorder, and edit commits before pushing to create a clean commit history.
Drill: jq Recipes for JSON Processing · Use jq to select, filter, map, and transform JSON output from APIs and Kubernetes commands.
Drill: Manage Network Connections with nmcli · Use nmcli to view, create, modify, and troubleshoot NetworkManager-managed connections.
Drill: Read and Analyze Saved pcap Files · Use tcpdump to read saved pcap files and apply display filters to analyze previously captured traffic.
Drill: Read and Manipulate the Routing Table · Use ip route to inspect, query, and manipulate the Linux routing table for network debugging.
Drill: Recover Lost Commits with Reflog · Use git reflog to find and recover commits that were lost after a reset, rebase, or other destructive operation.
Drill: Safely Drain a Kubernetes Node · Perform a safe node drain workflow: cordon the node, check PodDisruptionBudgets, drain, and verify workloads moved.
Drill: tmux Session Management · Use tmux to manage persistent terminal sessions with windows, panes, and detach/attach workflows.
Drill: Trace Process Relationships · Use pstree and ps to visualize parent-child process relationships and understand process hierarchies.
Drill: Trace Syscalls with strace · Use strace to observe system calls made by a process to diagnose failures, permission issues, and missing files.
Drill: Understanding git diff Variants · Understand the difference between git diff, git diff --staged, and git diff HEAD to inspect changes at each stage.
Drill: Use Port-Forward to Test Services and Pods · Use kubectl port-forward to directly access pods and services for debugging without exposing them externally.
Drill: Work on Multiple Branches with git worktree · Use git worktree to check out multiple branches simultaneously in separate directories without stashing or switching.
Drills · Ansible module categories for ad-hoc commands: command (raw shell, no pipes), shell (supports pipes and redirects), copy (files to remote), service. · Topics: Ansible · Aliases: ansible-galaxy, chef, config-management, inventory, playbooks, puppet, roles, saltstack
Drills · Quick muscle-memory exercises.
etcd Drills · Every etcdctl command in a TLS-secured cluster needs three certificate flags: --cacert (CA certificate), --cert (client certificate), --key (client key). · Topics: etcd · Aliases: etcd-backup, etcd-cluster, etcd-restore
Evolution Guides β€” "How We Got Here" · Historical context for modern DevOps tooling.
Failure Pattern Catalog · A taxonomy of recurring failure shapes across production systems.
FinOps Drills · The FinOps cycle: Inform (visibility β€” who spends what) -> Optimize (right-size, reserved instances, spot) -> Operate (governance, budgets, alerts). · Topics: FinOps · Aliases: cloud-billing, cost-optimization, reserved-instances, spot
Git Drills · Git undo safety levels: --soft (safest, keeps everything staged), --mixed (default, unstages but keeps files), --hard (destructive, discards everything). · Topics: Git · Aliases: branching, merging, rebase, version-control
GitOps & ArgoCD Drills · GitOps has two sync strategies: automated (ArgoCD applies changes immediately when Git changes) and manual (human clicks sync). · Topics: GitOps · Aliases: argo, argocd, desired-state, drift-detection, flux
Hands-On Labs · Self-contained exercises that run on your local k3s cluster.
Helm Debugging Decision Flow · Systematic approach to diagnosing Helm release issues.
Helm Drill Answers · Shows STATUS (deployed/failed/pending-upgrade), REVISION, chart version.
Helm Drills · 15 drills for Helm chart management muscle memory. · Topics: Helm · Aliases: helm-charts, helm-rollback, helm-upgrade, values-files
How We Got Here: Application Architecture · Arc: Platform Eras covered: 5 Timeline: ~2000-2025.
How We Got Here: Artifact Management · Arc: Deployment Eras covered: 6 Timeline: ~2000-2025.
How We Got Here: CI/CD Evolution · Arc: Deployment Eras covered: 5 Timeline: ~2000-2025.
How We Got Here: Configuration Management · Arc: Infrastructure Eras covered: 6 Timeline: ~2005-2025.
How We Got Here: Container Evolution · Arc: Infrastructure Eras covered: 6 Timeline: ~1979-2025.
How We Got Here: Deployment Strategies · Arc: Deployment Eras covered: 6 Timeline: ~2005-2025.
How We Got Here: Developer Experience · Arc: Platform Eras covered: 6 Timeline: ~2005-2025.
How We Got Here: From Bare Metal to Serverless · Arc: Infrastructure Eras covered: 6 Timeline: ~2000-2025.
How We Got Here: Incident Management · Arc: Observability Eras covered: 5 Timeline: ~2005-2025.
How We Got Here: Infrastructure as Code · Arc: Infrastructure Eras covered: 6 Timeline: ~2010-2025.
How We Got Here: Kubernetes Itself · Arc: Platform Eras covered: 5 Timeline: ~2010-2025.
How We Got Here: Logging Evolution · Arc: Observability Eras covered: 5 Timeline: ~2005-2025.
How We Got Here: Monitoring Evolution · Arc: Observability Eras covered: 6 Timeline: ~2005-2025.
How We Got Here: Secrets Management · Arc: Security Eras covered: 5 Timeline: ~2010-2025.
How We Got Here: Service Communication · Arc: Networking Eras covered: 5 Timeline: ~2005-2025.
How We Got Here: Supply Chain Security · Arc: Security Eras covered: 5 Timeline: ~2015-2025.
Human Factors Models · How people interact with complex systems
Interview Gauntlet: Alerts Firing but System Seems Fine · Interviewer: 'Alerts are firing for high error rate and elevated latency across multiple services.
Interview Gauntlet: Ansible Playbook 9x Slower · Interviewer: 'An Ansible playbook that used to run in 5 minutes now takes 45 minutes.
Interview Gauntlet: API Returning 503s · Interviewer: 'Your API is returning 503 errors.
Interview Gauntlet: CI/CD for a Monorepo · Interviewer: 'Design a CI/CD pipeline for a monorepo that contains 8 microservices, shared libraries, and infrastructure-as-code.
Interview Gauntlet: Container Image Build and Distribution Pipeline · Interviewer: 'Design a container image build and distribution pipeline for an organization with 20 services.
Interview Gauntlet: Container Using 2x Expected Memory · Interviewer: 'A Java service in Kubernetes is using 2 GB of memory but the application only needs about 1 GB based on its heap configuration.
Interview Gauntlet: Customer Reports Data Inconsistency · Interviewer: 'A customer opens a support ticket: 'I updated my profile 2 hours ago but the old information is still showing.' Support confirms the.
Interview Gauntlet: Deploy Succeeded but Old Version Visible · Interviewer: 'A deployment completes successfully β€” Argo CD shows synced, all pods are running the new image, health checks pass.
Interview Gauntlet: Disagreeing with a Technical Decision · Interviewer: 'Describe a technical decision you disagreed with.
Interview Gauntlet: Disk Usage on Prod Database · Interviewer: 'You get a disk usage alert at 3 AM β€” the production PostgreSQL database server is at 92% disk utilization.
Interview Gauntlet: eBPF for Observability · Interviewer: 'Your team is evaluating eBPF-based observability tools β€” things like Pixie, Cilium Hubble, or Tetragon.
Interview Gauntlet: Flaky CI Build · Interviewer: 'Your build passes locally every time but fails in CI about 30% of the time.
Interview Gauntlet: GitOps or Traditional CI/CD? · Interviewer: 'Your team is debating between GitOps (using Argo CD or Flux) and traditional CI/CD push-based deployments (GitHub Actions triggering.
Interview Gauntlet: Handling a Production Incident · Interviewer: 'Tell me about a time you handled a production incident.
Interview Gauntlet: Improving Team Development Workflow · Interviewer: 'Describe how you've improved a team's development workflow.
Interview Gauntlet: Intermittent gRPC Failures · Interviewer: 'gRPC calls between Service A and Service B fail intermittently β€” about 5% of calls return UNAVAILABLE.
Interview Gauntlet: Kubernetes or Simpler Orchestrator? · Interviewer: 'A startup with 3 services and 4 engineers is considering Kubernetes for their production infrastructure.
Interview Gauntlet: Learning Something Quickly · Interviewer: 'Tell me about a time you had to learn a technology quickly for a project.
Interview Gauntlet: Log Aggregation Pipeline · Interviewer: 'Design a log aggregation pipeline for a Kubernetes cluster running 10,000 pods.
Interview Gauntlet: Managed Database or Self-Hosted? · Interviewer: 'Your team needs a PostgreSQL database for a new production service.
Interview Gauntlet: Monitoring Stack from Scratch · Interviewer: 'You're joining a startup with 15 microservices, no monitoring, and a history of finding outages from customer complaints.
Interview Gauntlet: Monolith or Microservices? · Interviewer: 'You're starting a new project β€” an internal platform for managing employee onboarding workflows.
Interview Gauntlet: Multi-Region Kubernetes Deployment · Interviewer: 'Design a multi-region Kubernetes deployment for an e-commerce platform.
Interview Gauntlet: Network Latency Spikes Every 30 Seconds · Interviewer: 'Your service is experiencing network latency spikes every 30 seconds.
Interview Gauntlet: Pods Crash-Looping · Interviewer: 'Your monitoring shows pods in CrashLoopBackOff.
Interview Gauntlet: Secrets Management System · Interviewer: 'Design a secrets management system for an organization running 30 microservices across Kubernetes and some legacy VMs.
Interview Gauntlet: Should We Use a Service Mesh? · Interviewer: 'Your team is considering adopting a service mesh.
Interview Gauntlet: Terraform Plan Shows 47 Resources to Destroy/Recreate · Interviewer: 'You run terraform plan and it shows 47 resources will be destroyed and recreated.
Interview Gauntlet: When Automation Went Wrong · Interviewer: 'Tell me about a time automation went wrong.
Interview Gauntlet: Your Approach to On-Call · Interviewer: 'Describe your approach to on-call.
Interview Scenarios · DevOps/SRE interview scenarios tied to the runtime labs and runbooks in this repository.
Interview: Certificate Expired · 'Our production API started returning TLS errors 30 minutes ago. · Topics: TLS & PKI · Aliases: ca-chain, cert-manager, cert-rotation, certificates, pki, ssl, tls, x509
Interview: CI Vuln Scan Failed · 'Our CI pipeline started failing this morning. · Topics: Security Scanning, CI/CD · Aliases: circleci, continuous-delivery, continuous-integration, cyber-security, jenkins, pipelines, security, trivy, vulnerability-scan, zuul
Interview: Config Drift Detected · 'Our GitOps tool (ArgoCD) shows the production deployment as 'OutOfSync'. · Topics: GitOps · Aliases: argo, argocd, desired-state, drift-detection, flux
Interview: Cost Spike Investigation · 'Our AWS bill jumped 40% this month compared to last month. · Topics: FinOps · Aliases: cloud-billing, cost-optimization, reserved-instances, spot
Interview: Database Failover During Deploy · 'During a routine deployment, your PostgreSQL primary pod was evicted for a node drain. · Topics: Database Operations · Aliases: backup-restore, database, mysql, statefulsets
Interview: Deployment Stuck Progressing · 'We deployed a new version of our application 20 minutes ago. · Topics: Kubernetes Core, Probes (Liveness/Readiness) · Aliases: deployments, healthcheck, k8s, kubernetes, liveness, openshift, pods, readiness, services, startup-probe
Interview: Docker Container Debugging · 'We push a new Docker image to production and the container keeps crashing. · Topics: Docker / Containers, Container Runtimes · Aliases: cgroups, container-runtime-debug, containerd, containers, cri-o, docker-compose, dockerfile, images, namespaces, runc
Interview: etcd Space Exceeded · 'The Kubernetes API server is returning errors. · Topics: etcd · Aliases: etcd-backup, etcd-cluster, etcd-restore
Interview: GitOps Drift Detected · 'ArgoCD shows our production app as 'OutOfSync' with a 'Degraded' health status. · Topics: GitOps · Aliases: argo, argocd, desired-state, drift-detection, flux
Interview: Helm Upgrade Broke Prod · 'A colleague ran helm upgrade in production with a new values file. · Topics: Helm · Aliases: helm-charts, helm-rollback, helm-upgrade, values-files
Interview: HPA Not Scaling · 'We have an HPA configured for our deployment with target CPU at 50%. · Topics: HPA / Autoscaling · Aliases: autoscaling, horizontal-pod-autoscaler, metrics-server
Interview: Ingress 404 · 'Users report getting 404 errors intermittently when accessing our app through the ingress. · Topics: Kubernetes Networking · Aliases: cni, coredns, ingress, network-policies, service-mesh
Interview: Kyverno Blocking Deploys · 'Nobody can deploy anything to the production namespace. · Topics: Policy Engines · Aliases: admission-control, gatekeeper, kyverno, opa
Interview: Linux Server Slow · 'Users are reporting that the application is slow. · Topics: Linux Fundamentals · Aliases: filesystem, linux, linux-basics, permissions
Interview: Loki Logs Disappeared · 'Our developers report that application logs stopped appearing in Grafana about 30 minutes ago. · Topics: Loki · Aliases: log-aggregation, logql, logs, promtail
Interview: Pods OOMKilled · 'Our application pods keep getting OOMKilled during peak traffic hours. · Topics: OOMKilled · Aliases: memory, oom, out-of-memory, resource-limits
Interview: Prometheus Target Down · 'Our Grafana dashboards suddenly show 'No data' for application metrics. · Topics: Prometheus · Aliases: alertmanager, datadog, metrics, monitoring, prometheusstack, promql, service-monitor
Interview: RBAC Forbidden · 'A new engineer joined the team and can't deploy to our Kubernetes cluster. · Topics: RBAC · Aliases: clusterroles, rbac, roles, service-accounts
Interview: Secret Leaked to Git · 'A developer just pinged you - they accidentally committed a database password to a public GitHub repository. · Topics: Secrets Management, Git · Aliases: branching, external-secrets, merging, rebase, sealed-secrets, sops, version-control
Interview: Server Won't POST · 'You arrive at the data center for a scheduled maintenance window. · Topics: Firmware / BIOS / UEFI, Server Hardware · Aliases: bios, cpu, dimm, dmesg, dmidecode, firmware-update, hba, lshw, nic, post, uefi
Interview: Service Mesh 503s · 'We enabled Istio on the grokdevops namespace. · Topics: Service Mesh · Aliases: envoy, istio, linkerd, mtls, sidecar
Interview: Vault Token Expired · 'All our microservices suddenly started failing with 'permission denied' errors when trying to read secrets from Vault. · Topics: Secrets Management · Aliases: external-secrets, sealed-secrets, sops
Knowledge Architecture · Concept graphs, failure taxonomies, and command debugging flows
kubectl Debugging Decision Flow · Systematic approach to diagnosing Kubernetes issues.
kubectl Drill Answers · Look for STATUS != Running or READY showing 0/1.
kubectl Drills · 25 drills for Kubernetes CLI muscle memory. · Topics: Kubernetes Core · Aliases: deployments, k8s, kubernetes, openshift, pods, services
Kubernetes Operators Drills · An operator is a CRD + a controller. · Topics: K8s Ecosystem · Aliases: cncf, k8s-landscape, kubernetes-ecosystem
Lab 10: RBAC & Security · Your organization has three teams sharing a Kubernetes cluster: the development team, the operations team, and an external auditor.
Lab 11: Monitoring Stack · Your company just experienced a 4-hour outage that nobody detected until customers complained on social media.
Lab 12: CI/CD Pipeline · Your team is deploying a Python web application manually.
Lab 13: Terraform IaC · Your infrastructure team has been provisioning resources manually through cloud consoles.
Lab 14: Log Analysis · At 3:47 AM, your monitoring system detected elevated error rates on the payment service.
Lab 15: Incident Response · It is 2:00 PM on a Tuesday.
Lab 16: Chaos Engineering · Your team claims the application stack is 'highly available,' but nobody has tested what happens when things actually fail.
Lab 17: Performance Tuning · The product team reports that page load times have increased from 200ms to over 3 seconds in the past month.
Lab 18: Zero-Downtime Migration · The database team needs to migrate the primary PostgreSQL database from version 15 to version 16.
Lab 19: Multi-Cluster · Your company is expanding to a second region for disaster recovery.
Lab 1: Linux Triage · You have just joined an on-call rotation at a mid-size SaaS company.
Lab 20: Platform Engineering · Your company has 30 development teams, each deploying their own microservices.
Lab 21: Production Readiness Review · A development team is requesting production deployment approval for their new 'Order Processing Service.' They claim it is production-ready, but your.
Lab 22: Incident Simulation · This is a full-scale incident simulation.
Lab 23: Architecture Review · Your company acquired a startup and inherited their application.
Lab 24: On-Call Shift · Welcome to your first simulated on-call shift.
Lab 25: Tech Lead Challenge · You have been promoted to Tech Lead for the platform team.
Lab 2: Container Basics · Your team has inherited a legacy Node.js application that was hastily containerized by a contractor who left no documentation.
Lab 3: Networking Fundamentals · Your company runs a microservices stack using Docker Compose on a staging server.
Lab 4: Git Operations · It is Friday afternoon and four different developers have managed to create four different git disasters in the same repository.
Lab 5: Shell Scripting · Your team runs five critical services on a bare-metal server: a web server, an API gateway, a cache (Redis), a database (PostgreSQL), and a message.
Lab 6: Deploy & Scale · Your company is migrating a three-tier web application from VMs to Kubernetes.
Lab 7: Pod Debugging · You inherit a Kubernetes namespace from a departing colleague.
Lab 8: Service Networking · Your platform team runs a multi-namespace architecture: frontend-ns for public-facing services, backend-ns for APIs, and data-ns for databases.
Lab 9: Storage & State · The data engineering team needs a PostgreSQL instance running on Kubernetes with persistent storage that survives pod restarts.
Learning Paths · Curated sequences for common career goals. Each path lists topics in optimal order
Linux Ops Drills · The five essential Linux diagnostic commands: top (CPU/memory live), df -h (disk space), free -h (memory), ss -tlnp (listening ports), journalctl -p. · Topics: Linux Fundamentals, Bash / Shell Scripting · Aliases: advanced-bash, bash, filesystem, linux, linux-basics, permissions, scripting, shell, shell-scripting
LogQL Drills · 20 drills for Loki query language muscle memory. · Topics: Loki · Aliases: log-aggregation, logql, logs, promtail
Mental Model Library · Thinking frameworks for production systems β€” the models that experienced engineers use to reason about behavior, diagnose problems, and make decisions.
Mental Model: 12-Factor App · Origin: Adam Wiggins and the Heroku engineering team (2011) One-liner: Twelve concrete constraints that make an application portable, scalable, and.
Mental Model: Alert Fatigue · Origin: The term has roots in clinical medicine (nurse call-bell fatigue, cardiac monitor alarms in ICUs) and was imported into software operations and.
Mental Model: Amdahl's Law · Origin: Gene Amdahl, 1967 (presented at the AFIPS Spring Joint Computer Conference) One-liner: The maximum speedup from parallelism is strictly bounded.
Mental Model: Automation Complacency · Origin: The concept emerged from aviation human factors research in the 1980s-90s, particularly from studies of glass-cockpit aircraft (Wiener, 1988.
Mental Model: Bisect · Origin: Binary search algorithm (classical computer science); applied to version control debugging, most famously formalized as git bisect (Linus.
Mental Model: Blameless Postmortem · Origin: Google Site Reliability Engineering team, popularized in the SRE Book (2016); roots in Sidney Dekker's 'Just Culture' and human factors.
Mental Model: Blast Radius · Origin: Borrowed from military/explosive engineering into software reliability; popularized in SRE practice through Netflix Chaos Engineering (circa.
Mental Model: Bulkhead · Origin: Naval architecture (watertight compartments); applied to software by Michael Nygard, Release It!
Mental Model: CAP Theorem · Origin: Eric Brewer, 2000 (conjecture at PODC symposium); formally proven by Gilbert and Lynch, 2002 One-liner: A distributed system can guarantee at.
Mental Model: Cattle vs Pets · Origin: Bill Baker, Distinguished Engineer at Microsoft, coined the metaphor around 2011–2012; popularized in the cloud-native and DevOps communities.
Mental Model: Circuit Breaker · Origin: Michael Nygard, Release It!
Mental Model: Correlation vs Causation · Origin: Classical statistics and scientific method (Francis Bacon, 17th century); the specific formulation 'correlation does not imply causation' is.
Mental Model: Differential Diagnosis · Origin: Clinical medicine (formalized in the 19th century); adopted into engineering debugging culture through DevOps and SRE practices, popularized by.
Mental Model: Error Budget · Origin: Google Site Reliability Engineering, formalized in the SRE Book (2016); attributed to Ben Treynor Sloss and the original Google SRE team.
Mental Model: Event Sourcing · Origin: Domain-Driven Design community; Greg Young is credited with formalizing the pattern (~2010); informed by Martin Fowler's earlier writings on.
Mental Model: Failure Domains · Origin: Systems engineering and high-availability design practice; institutionalized in cloud architecture through AWS Availability Zones (2006) and.
Mental Model: Five Whys · Origin: Sakichi Toyoda, Toyota Production System (1930s–1950s); popularized in DevOps/SRE through the Site Reliability Engineering book and incident.
Mental Model: Graceful Degradation · Origin: Fault-tolerant computing and aerospace systems engineering; widely formalized in web architecture by Michael Nygard ('Release It!', 2007) and.
Mental Model: Hindsight Bias · Origin: First experimentally documented by Baruch Fischhoff in 1975 ('Hindsight β‰  Foresight: The Effect of Outcome Knowledge on Judgment Under.
Mental Model: Idempotency · Origin: Mathematics (Peirce, 1867); applied to distributed systems design throughout the 1970s–2000s; formalized in REST API design by Roy Fielding.
Mental Model: Immutable Infrastructure · Origin: Chad Fowler's 2013 blog post 'Trash Your Servers and Burn Your Code'; the principle has roots in functional programming's immutability concept.
Mental Model: Little's Law · Origin: John D.C. Little, 1961 (proven formally in 'A Proof for the Queuing Formula: L = Ξ»W') One-liner: The average number of items in a system equals.
Mental Model: Normalization of Deviance · Origin: Diane Vaughan, sociologist β€” coined after her exhaustive sociological study of the 1986 Space Shuttle Challenger disaster, published in The.
Mental Model: OODA Loop · Origin: Colonel John Boyd, USAF β€” developed in the 1970s to describe fighter pilot decision-making; later generalized to all competitive and crisis.
Mental Model: PACELC · Origin: Daniel Abadi, 2012 ('Consistency Tradeoffs in Modern Distributed Database System Design') One-liner: During a partition, choose between.
Mental Model: Queueing Theory · Origin: Agner Krarup Erlang, 1909 (telephone network traffic analysis); formalized by Leonard Kleinrock in the 1960s One-liner: As utilization.
Mental Model: RED Method · Origin: Tom Wilkie (Grafana Labs, formerly Weaveworks), ~2015 One-liner: For every service, monitor Rate (requests/sec), Errors (failed requests), and.
Mental Model: Runbook-Driven Recovery · Origin: Operations and systems administration practice; formalized in SRE and DevOps culture; the term 'runbook' originates from mainframe operations.
Mental Model: Shift Left · Origin: Larry Smith's 2001 article 'Shift-Left Testing' in STQE Magazine; the model has since expanded beyond testing to encompass security.
Mental Model: Sidecar Pattern · Origin: Broadly attributed to the service mesh and container communities; formalized in the context of Kubernetes and Envoy proxy (~2016–2018).
Mental Model: Strangler Fig · Origin: Martin Fowler (2004), inspired by the strangler fig tree (Ficus aurea) One-liner: Incrementally replace a legacy system by building new.
Mental Model: Swiss Cheese Model · Origin: James Reason, 1990 ('Human Error,' Cambridge University Press); widely adopted in aviation and healthcare safety One-liner: Accidents occur not.
Mental Model: Toil vs Automation ROI · Origin: Google Site Reliability Engineering, formalized in the SRE Book (2016); Chapter 5 'Eliminating Toil' by Sara Smollett and Sherry Moore.
Mental Model: USE Method · Origin: Brendan Gregg (Netflix performance engineer), formalized ~2012 One-liner: For every resource in the system, check Utilization, Saturation, and.
Modern CLI Drills · Modern CLI tool replacements: find -> fd, grep -> rg (ripgrep), cat -> bat, ls -> eza, du -> dust/ncdu, cd -> zoxide, curl -> httpie. · Topics: Modern CLI Tools, jq / JSON Processing, ripgrep (rg), fzf · Aliases: bat, cli-tools, code-search, dust, eza, fast-grep, fuzzy-finder, interactive-selection, jq-filters, json, modern-unix, rg
Networking Drills · The network debugging order: DNS -> Routing -> Firewall -> Application. · Topics: TCP/IP, DNS, Linux Networking Tools · Aliases: coredns, dig, dns-ops, domain-name-system, ethtool, ip, ip-command, layers, mtr, networking, networking-fundamentals, nmcli
Observability Debugging Decision Flow · Systematic approach to diagnosing metrics, logs, and traces issues.
Observability Drill Answers · If the API call returns data, metrics-server is working.
Observability Drills · The observability debugging flow: Alert fires -> check Grafana dashboard (what metric is off?) -> check Prometheus (what changed?) -> check Loki logs. · Topics: Prometheus, Loki · Aliases: alertmanager, datadog, log-aggregation, logql, logs, metrics, monitoring, prometheusstack, promql, promtail, service-monitor
On-Call Survival Guides · Pocket-card reference for your first on-call rotation.
On-Call Survival: CI/CD · Print this. Pin it. Read it at 3 AM.
On-Call Survival: Cloud/Infrastructure · Print this. Pin it. Read it at 3 AM.
On-Call Survival: Databases (PostgreSQL) · Print this. Pin it. Read it at 3 AM.
On-Call Survival: Kubernetes · Print this. Pin it. Read it at 3 AM.
On-Call Survival: Linux/OS · Print this. Pin it. Read it at 3 AM.
On-Call Survival: Networking · Print this. Pin it. Read it at 3 AM.
On-Call Survival: Observability · Print this. Pin it. Read it at 3 AM.
On-Call Survival: Security · Print this. Pin it. Read it at 3 AM.
Operational Reasoning Models · Mental models for how to run and improve systems over time
Pattern: Alerting on Restart (Not Root Cause) · A Kubernetes alert fires every time a pod restarts.
Pattern: Apply-Without-Reading Manifest · An engineer applies a Kubernetes manifest from an external source (email, paste, Stack Overflow, a colleague's gist) without fully reading it.
Pattern: Cache Stampede · A cached item expires or is evicted.
Pattern: Cgroup Soft/Hard Limit Confusion · Cgroups have two memory limit mechanisms: memory.high (soft β€” throttles allocation, triggers reclaim but doesn't kill) and memory.limit_in_bytes /.
Pattern: Clock Skew Ordering · Distributed systems that use wall-clock timestamps (from different servers) to order events assume all clocks agree.
Pattern: Connection Pool Exhaustion · A pool of persistent connections (to a database, cache, or downstream service) is finite.
Pattern: Deep Health Check Cascade · A readiness probe checks not just whether the service itself is alive, but also whether its downstream dependencies are reachable (database ping, cache.
Pattern: Deleted-But-Open File · A file is deleted (rm) while a process still holds an open file descriptor to it.
Pattern: Dependency Chain Collapse · Service A depends on Service B, which depends on Service C.
Pattern: Device Name Confusion · Linux device names (like /dev/sdb, /dev/sdc) are not persistent β€” they're assigned by the kernel at boot based on detection order.
Pattern: Disk Full (Reserved Blocks Gone) · ext4 and other Linux filesystems reserve 5% of blocks for the root user by default.
Pattern: Dual-Write Divergence · An application writes the same logical update to two separate stores (e.g., a database and a cache, or two databases in different regions) in a.
Pattern: Hardcoded Namespace Override · A Kubernetes manifest explicitly sets namespace: in the metadata.
Pattern: Health Check Lying · A service's health check endpoint returns 200 OK while the service is silently failing to perform its core function.
Pattern: Inode Exhaustion · A filesystem tracks two independent resources: blocks (raw storage) and inodes (file metadata slots).
Pattern: latest Tag in Production · Using a mutable image tag (latest, stable, main) in a production deployment means the deployed image can change without any deployment action.
Pattern: Memory Limit Equals Request · Setting resources.limits.memory equal to resources.requests.memory in Kubernetes provides no headroom for memory spikes.
Pattern: Metric Cardinality Explosion · Prometheus (and similar time-series databases) store one time series per unique combination of metric name + label values.
Pattern: Missing absent() Alert · Most Prometheus alerts are threshold-based: 'alert if metric > N.' These alerts require the metric to exist.
Pattern: Missing Backpressure · A producer generates work faster than a consumer can process it.
Pattern: Missing Escalation Criteria · An on-call engineer is paged for an incident.
Pattern: Missing Point-in-Time Recovery · An application has daily full backups but no continuous WAL/binlog archiving that would allow recovery to an arbitrary point in time.
Pattern: ndots:5 Query Amplification · Kubernetes defaults ndots:5 in pod DNS configuration, meaning any name with fewer than 5 dots triggers search-domain expansion before the literal name.
Pattern: No Circuit Breaker · A service calls a downstream dependency synchronously.
Pattern: OOM Without Swap Buffer · When physical memory is exhausted and swap is disabled (or absent), the OOM killer fires immediately and terminates processes without warning.
Pattern: Percentile Blindness · Dashboards and alerts that use mean (average) latency hide the true user experience.
Pattern: PID Exhaustion via Zombies · A parent process forks children but never calls waitpid() to reap them.
Pattern: Port-Forward as Permanent Fix · During an incident, an engineer uses kubectl port-forward to bypass a broken Ingress, Service, or load balancer.
Pattern: PVC Reclaim Policy Delete · A Kubernetes PersistentVolume's reclaim policy (Retain or Delete) controls what happens to the underlying storage when the PVC is deleted.
Pattern: RAID Rebuild I/O Saturation · After a disk failure, RAID array reconstruction reads every block from surviving disks and writes them to the replacement.
Pattern: rate() Over Too-Short Window · Prometheus rate() function requires at least 2 data points in the range window to calculate a rate.
Pattern: Replication Lag at Failover · A database replica is typically seconds or minutes behind the primary.
Pattern: Restart Avalanche · Multiple pods (or services) restart simultaneously β€” either due to a rolling deploy with high maxSurge, a node restart bringing all pods back at once.
Pattern: Retry Amplification · A multi-tier system where each tier retries independently.
Pattern: Retry Storm · When a downstream service becomes slow or unavailable, all upstream callers begin retrying simultaneously.
Pattern: Rollout Hang (Zero Surge + Zero Unavailable) · A Kubernetes Deployment with maxSurge: 0 and maxUnavailable: 0 can never progress.
Pattern: Runbook with No Contacts · A runbook for a critical operation says 'escalate to the database team' or 'contact the security team' without naming specific people, providing.
Pattern: Simultaneous Timer Expiry · When many instances of a service are deployed at the same time (rolling deploy, cluster restart), time-based events (lease renewals, token refreshes.
Pattern: Stale Image Tag · A container image tag (especially latest or a mutable semantic version) is used in production.
Pattern: Stale Leader · A leader/primary node is partitioned from the cluster but not immediately aware of it.
Pattern: StatefulSet OrderedReady Deadlock · StatefulSets with podManagementPolicy: OrderedReady (the default) bring pods up and down sequentially and wait for each pod to be Ready before.
Pattern: STP Disabled + Loop Created · Spanning Tree Protocol (STP) prevents Layer 2 loops by selectively blocking redundant network paths.
Pattern: Thread Pool Exhaustion · A service uses a fixed-size thread pool (or goroutine pool, or connection pool) to handle concurrent requests.
Pattern: Timeout Assumed = Not Executed · A client sends a request and receives a timeout (or network error).
Pattern: tmpfs Consuming Hidden RAM · tmpfs filesystems (including /tmp, /dev/shm, /run, and container overlay layers) are backed by RAM and swap, not disk.
Pattern: Transaction ID Wraparound · PostgreSQL uses 32-bit transaction IDs (XIDs) to track visibility of rows.
Pattern: Two-Node Quorum Trap · Distributed consensus protocols (Raft, Paxos, etcd) require a quorum β€” a majority of nodes β€” to make decisions.
Pattern: Unstructured Logging · Logs written as free-form strings ('Processing order 12345 for user 67890') cannot be reliably parsed, aggregated, or alerted on.
Pattern: Untested Backup · Automated backups complete successfully for months or years.
Pattern: Untested Rollback Procedure · A deployment causes a production issue.
Pattern: Wrong Terminal Tab · An engineer has two terminal tabs (or tmux panes): one connected to production, one to staging.
Pattern: Zombie Process Accumulation · Zombie processes (defunct) accumulate over time when parent processes don't call waitpid() to reap children.
Policy Engine Drills · Policy engines enforce guardrails at admission time β€” before a resource is persisted to etcd. · Topics: Policy Engines · Aliases: admission-control, gatekeeper, kyverno, opa
Postmortem & SLO Drills · SLI -> SLO -> SLA, from measurement to promise. · Topics: Postmortems & SLOs · Aliases: blameless, devops, error-budget, postmortem, sla, sli, slo
Postmortem Anthology · Realistic postmortem documents as they would appear in a company's internal wiki.
Postmortem: Alert Routing Sends All Pages to Decommissioned Channel · On Friday 2025-10-17 at 17:02 UTC, an Alertmanager configuration migration from Slack webhooks to PagerDuty receivers was deployed by SRE engineer.
Postmortem: Ansible Playbook Targets Production Instead of Staging · On 2025-04-08 at 14:22 UTC, an engineer on the Platform Engineering team ran an NTP configuration playbook against production hosts instead of staging.
Postmortem: AWS AZ Network Degradation Triggers Cascading Health Check Failures · On 2025-06-03, AWS us-east-1a experienced elevated packet loss (measured at 1.8–2.3%) due to an underlying network issue AWS later attributed to a.
Postmortem: AWS Credentials Committed to Public Repo β€” Caught by Pre-Commit Hook · On 2025-03-12 at 14:07 UTC, a Platform Engineering engineer staged a .env file containing production AWS credentials with AdministratorAccess.
Postmortem: BGP Route Leak Sends Customer Traffic Through Monitoring VLAN · On July 22, 2025, a network engineer added a new BGP neighbor for the monitoring infrastructure but misconfigured the outbound route-map, causing the.
Postmortem: Core Switch Firmware Bug Causes Cascading Network Partition · On 2025-09-03, a Juniper QFX5120-32C core switch running firmware version 21.4R3-S4 triggered an STP (Spanning Tree Protocol) recalculation storm.
Postmortem: Custom Controller Missing Backoff Overwhelms API Server · On November 3, 2025, the db-provisioner Kubernetes operator β€” an in-house controller that manages database provisioning resources β€” entered a tight.
Postmortem: Debug Build Deployed to Production via Copy-Paste Error · On 2025-03-11, an engineer on the Platform Engineering team copy-pasted a Helm deployment command from an earlier debug session into a production.
Postmortem: DNS CNAME Chain Breaks After Load Balancer Rename · On 2025-05-07 at 10:14 UTC, a DevOps engineer renamed an internal AWS Application Load Balancer from api-internal-legacy to api-internal as part of a.
Postmortem: Expired Wildcard TLS Certificate Causes Full API Gateway Outage · At 03:12 UTC on 2025-07-09, the wildcard TLS certificate for *.api.meridiancloud.io expired.
Postmortem: Go Dependency Update Silently Changes Default Timeout β€” Caught in Canary · On 2025-07-15 at 19:42 UTC, a canary deployment of the order-service at 5% traffic began showing an elevated API error rate driven by HTTP client.
Postmortem: Helm Values Mismatch Routes Staging Traffic to Production Database · On 2025-04-22, a Helm values refactoring effort accidentally introduced a production PostgreSQL connection string into the storefront-staging values file.
Postmortem: Kernel TCP Regression After Security Patch · On 2025-09-03 at 02:00 UTC, a scheduled monthly security kernel patching job rolled out Linux kernel 5.15.149 across the production fleet, replacing.
Postmortem: Memory Leak in Log Shipping Agent Causes Fleet-Wide OOM Kills · On 2025-05-14, a newly deployed Python microservice (catalog-indexer) began emitting approximately 5% malformed JSON log lines due to a bug in its.
Postmortem: Missing Circuit Breaker Lets Redis Failure Cascade to All Services · On March 4, 2025, the Redis primary instance for session caching exhausted its memory allocation and entered a hard-blocked state due to an noeviction.
Postmortem: Missing Runbook Extends CrashLoopBackOff Recovery by 45 Minutes · On 2025-07-08, a ConfigMap update to the auth-service introduced invalid YAML (an unquoted string containing a colon), causing all auth-service pods to.
Postmortem: No Review Gate on Terraform Destroy Leads to Wrong Account Teardown · On May 13, 2025, an infrastructure engineer ran terraform destroy to tear down a personal dev environment but had AWS_PROFILE set to staging-us-east.
Postmortem: On-Call Handoff Gap Leaves Alerts Unacknowledged for 3 Hours · On Friday 2025-07-11 at 18:00 UTC, the on-call rotation handed off from Alice Marchetti (outgoing) to Bob Nguyen (incoming).
Postmortem: Production Database Deleted by Terraform Apply on Wrong Workspace · On 2025-03-14 at 14:22 UTC, a Platform Engineering engineer ran terraform apply in the production Terraform workspace while intending to apply changes.
Postmortem: Prometheus Cardinality Explosion from Debug Labels · On 2025-07-22 at 11:03 UTC, a new version of the checkout-service was deployed to production containing a Prometheus histogram label, request_id, that.
Postmortem: Race Condition in Distributed Lock Manager Corrupts Shared State · On 2025-05-21, a race condition in the company's hand-rolled distributed lock manager caused two concurrent writers to simultaneously acquire the same.
Postmortem: Resource Quota Misconfiguration Blocks All Deployments · On 2025-06-03 starting at 09:41 UTC, the Platform Engineering team applied a batch of 15 namespace resource quota changes as part of a cost.
Postmortem: S3 Bucket Policy Change Nearly Deletes All Backup Archives · On 2025-09-03 at 11:18 UTC, Infrastructure Engineer Dariusz Kowalczyk was running a terraform plan for a routine S3 bucket encryption upgrade when he.
Postmortem: Single etcd Member Disk Full Degrades Control Plane · On 2025-03-11 at 14:22 UTC, one member of a production 3-node etcd cluster (etcd-2) ran out of disk space after etcd compaction was never configured.
Postmortem: SSD Firmware Bug Causes Silent Bit Corruption · On 2025-09-14, the weekly scheduled ZFS scrub on the arc-pool-01 through arc-pool-24 storage nodes detected checksum mismatches on 3 of 24 nodes β€” all.
Postmortem: Stale Docker Base Image Ships Known CVE to Production · On 2025-10-15, catalog-service was built and deployed to production using a Docker base image pinned to python:3.11.4-slim.
Postmortem: Unbounded Kafka Topic Exhausts Broker Disk · On September 9, 2025, a Kafka consumer group for the user-activity-raw topic fell 6 hours behind after a deployment introduced a deserialization bug.
Postmortem: Unbounded Retry Storm Takes Down Payment Processing · On 2025-11-12, the company's payment gateway operator (Clearpath Payments) performed a planned 20-minute maintenance window.
Postmortem: UPS Battery Degradation Causes Rack Power Loss During Utility Blip · On 2025-11-19 at 03:47 UTC, a 5-second utility power dip (a routine event that occurs 2-3 times per year at the Meridian Systems primary datacenter).
Postmortem: Wildcard Ingress Rule Nearly Exposes Internal Admin Panel · On 2025-05-08 at 10:23 UTC, a Kubernetes Ingress resource with a wildcard host: '*' field was deployed to the staging cluster by a Platform Engineering.
Production Readiness Assessment · You are about to go on-call for the Meridian platform for the first time.
Production Readiness Assessment · Evaluate whether a system is ready for production
Production Readiness Review: Answer Key · This document provides model answers at the Strong (3) level for all 50 questions in the assessment.
Production Readiness Review: Scoring Guide · For each of the 50 questions in the assessment, rate yourself honestly.
Production Readiness Review: Study Plans · These study plans are generated based on your assessment results and scoring.
Production Readiness Review: System Architecture · You are the newest member of the platform engineering team at Meridian, a mid-scale B2B SaaS company that provides real-time inventory management and.
PromQL Drills · 25 drills for Prometheus query language muscle memory. · Topics: Prometheus · Aliases: alertmanager, datadog, metrics, monitoring, prometheusstack, promql, service-monitor
Python Drills · subprocess.run('kubectl get pods | grep Error', shell=True) works but is a security risk β€” shell injection via untrusted input. · Topics: Python Automation · Aliases: automation, perl, python, scripting, software-development
Reference · Knowledge architecture -- concept graphs, failure taxonomies, and command flows
Scenario: Asymmetric Routing · Hands-on troubleshooting scenario: Asymmetric Routing Through a Stateful Firewall · Topics: Routing, Linux Networking Tools · Aliases: asymmetric-routing, bgp, ethtool, gateway, ip-command, mtr, nmcli, ospf, routes, ss, static-routes, tcpdump
Scenario: DNS Looks Fine but App Fails · Hands-on troubleshooting scenario: DNS Resolves Correctly but Application Fails to Connect · Topics: DNS, TLS & PKI · Aliases: ca-chain, cert-manager, cert-rotation, certificates, coredns, dig, dns-ops, domain-name-system, nslookup, pki, resolve, ssl
Scenario: Duplex Mismatch · Hands-on troubleshooting scenario: Duplex Mismatch Causing Slow Transfers and Late Collisions · Topics: Linux Networking Tools · Aliases: ethtool, ip-command, mtr, nmcli, ss, tcpdump, traceroute
Scenario: etcd Troubleshooting · Hands-on troubleshooting scenario: etcd Troubleshooting Scenarios · Topics: etcd · Aliases: etcd-backup, etcd-cluster, etcd-restore
Scenario: MTU Blackhole · Hands-on troubleshooting scenario: MTU Black Hole β€” Large Packets Silently Dropped · Topics: MTU, Linux Networking Tools · Aliases: ethtool, fragmentation, ip-command, jumbo-frames, mtr, mtu-blackhole, nmcli, path-mtu, ss, tcpdump, traceroute
Scenario: NIC Flapping / LACP Mismatch · Hands-on troubleshooting scenario: NIC Flapping in LACP Bond · Topics: LACP / Link Aggregation, Server Hardware, Linux Networking Tools · Aliases: bonding, cpu, dimm, dmesg, dmidecode, etherchannel, ethtool, hba, ip-command, link-aggregation, lshw, mtr
Scenario: OOB Unreachable but Host Responds · Hands-on troubleshooting scenario: Out-of-Band Management (iDRAC) Unreachable · Topics: Out-of-Band Management, Linux Networking Tools · Aliases: bare-metal-provisioning, bmc, datacenter-provisioning, ethtool, idrac, ilo, ip-command, ipmi, mtr, nmcli, oob, oob-provisioning
Scenario: RAID Array Degraded · Hands-on troubleshooting scenario: RAID 5 Array Degraded with Predictive Failure · Topics: RAID, Server Hardware · Aliases: cpu, dimm, disk-arrays, dmesg, dmidecode, hba, lshw, megaraid, nic, raid-levels, raid-rebuild, smart
Scenario: Server Won't Boot After Update · Hands-on troubleshooting scenario: Server Won't Boot After Firmware Update · Topics: Firmware / BIOS / UEFI, Out-of-Band Management · Aliases: bare-metal-provisioning, bios, bmc, datacenter-provisioning, firmware-update, idrac, ilo, ipmi, oob, oob-provisioning, post, remote-provisioning
Scenario: Thermal Throttling · Hands-on troubleshooting scenario: CPU Thermal Throttling During Peak Hours · Topics: Server Hardware, Rack & Stack · Aliases: airflow, cabling, cpu, dimm, dmesg, dmidecode, hba, labeling, lshw, nic, pdu, power
Scenario: VLAN Trunk Mismatch · Hands-on troubleshooting scenario: VLAN Trunk Mismatch β€” Server Cannot Reach Its Gateway · Topics: VLANs, Cisco CLI · Aliases: 802.1q, cisco-basics, cisco-fundamentals-for-devops, ios, show-commands, tagged, trunk, untagged, vlan
Scenarios · Structured incident scenarios organized by domain
Secrets Management Drills · Secrets management layers: Storage (where secrets live β€” Vault, AWS Secrets Manager, SOPS), Delivery (how they reach the app β€” env vars, volume mounts. · Topics: Secrets Management · Aliases: external-secrets, sealed-secrets, sops
Security Drills · The container security checklist: Non-root user, Read-only filesystem, Drop all capabilities, No privilege escalation, Specific image tag (not :latest). · Topics: Security Scanning · Aliases: cyber-security, security, trivy, vulnerability-scan
Service Mesh Drills · A service mesh adds three capabilities to your cluster: mTLS (encryption between services), Observability (automatic metrics, traces, and access logs. · Topics: Service Mesh · Aliases: envoy, istio, linkerd, mtls, sidecar
Skill Tree · A visual representation of the training curriculum as an interconnected skill tree
Solution: Lab Runtime 01 -- Readiness Probe Failure · SPOILER WARNING: Try to solve it yourself first.
Solution: Lab Runtime 02 -- HPA Live Scaling · SPOILER WARNING: Try to solve it yourself first.
Solution: Lab Runtime 03 -- Observability Target Down · SPOILER WARNING: Try to solve it yourself first.
Solution: Lab Runtime 04 -- Loki No Logs · SPOILER WARNING: Try to solve it yourself first.
Solution: Lab Runtime 05 -- Helm Upgrade Rollback · SPOILER WARNING: Try to solve it yourself first.
Solution: Lab Runtime 06 -- Trivy Fail to Green · SPOILER WARNING: Try to solve it yourself first.
Solution: Lab Runtime 07 -- GitOps Sync and Drift · SPOILER WARNING: Try to solve it yourself first.
Solution: Lab Runtime 08 -- Resource Limits OOM · SPOILER WARNING: Try to solve it yourself first.
Solutions · Hint ladders and answer keys for labs and runbooks.
System Behavior Models · Mental models for understanding how systems act under load and failure
Terraform Drills · Terraform's core loop: init (download providers + backend) -> plan (dry-run diff) -> apply (execute). · Topics: Terraform · Aliases: hcl, iac, infrastructure-as-code, tfstate
TLS & PKI Drills · The TLS handshake in 4 steps: ClientHello (supported ciphers + SNI hostname) -> ServerHello (chosen cipher + certificate) -> Key Exchange (agree on. · Topics: TLS & PKI · Aliases: ca-chain, cert-manager, cert-rotation, certificates, pki, ssl, tls, x509
Training Dependency Graph · A prerequisite map for all training content. Use this to sequence your learning

Interactive (99)

Anki Starter Decks · Lightweight Anki-importable flashcard exports for each major topic area
Assessments · Hands-on skill assessments with rubrics and scoring
Chaos Scripts · Fault injection scripts for chaos engineering practice
Common Mistakes - Level 11: Deployment Rollback · Common Mistakes - Level 11: Deployment Rollback.
Common Mistakes - Level 1: CrashLoopBackOff · Common Mistakes - Level 1: CrashLoopBackOff.
Generic Hints · Fallback hints when no incident-specific hint file exists
Hints: crashloopbackoff · Progressive investigation hints for diagnosing crashloopbackoff issues
Hints: dns-resolution · Progressive investigation hints for diagnosing dns resolution issues
Hints: helm-upgrade-failed · Progressive investigation hints for diagnosing helm upgrade failed issues
Hints: hpa-not-scaling · Progressive investigation hints for diagnosing hpa not scaling issues
Hints: imagepullbackoff · Progressive investigation hints for diagnosing imagepullbackoff issues
Hints: latency-spike · Progressive investigation hints for diagnosing latency spike issues
Hints: loki-no-logs · Progressive investigation hints for diagnosing loki no logs issues
Hints: networkpolicy-block · Progressive investigation hints for diagnosing networkpolicy block issues
Hints: oomkilled · Progressive investigation hints for diagnosing oomkilled issues
Hints: probe-failure · Progressive investigation hints for diagnosing probe failure issues
Hints: prometheus-target-down · Progressive investigation hints for diagnosing prometheus target down issues
Hints: service-down · Progressive investigation hints for diagnosing service down issues
Index · Chaos engineering exercises for practicing fault injection and incident response
Index · Learn by fixing broken things.
index
index
Index · Hands-on incident simulations for practicing triage, debugging, and resolution
Index · Guided investigation engine for structured troubleshooting practice
Index · Flashcard decks and interview preparation tools
Index · Runtime lab exercise: Lab Runtime 01 β€” Rollout Probe Failure
Index · Runtime lab exercise: Lab Runtime 02 β€” HPA Live Scaling
Index · Runtime lab exercise: Lab Runtime 03 β€” Observability Target Down
Index · Runtime lab exercise: Lab Runtime 04 β€” Loki No Logs
Index · Runtime lab exercise: Lab Runtime 05 β€” Helm Upgrade & Rollback
Index · Runtime lab exercise: Lab Runtime 06 β€” Trivy Vulnerability Scanning β€” Fail to Green
Index · Runtime lab exercise: Lab Runtime 07 β€” GitOps Sync and Drift
Index · Runtime lab exercise: Lab Runtime 08 β€” Resource Limits and OOMKilled
Interactive Training · Hands-on, runnable content. These require a working environment (k3s cluster, Docker, or a Linux VM)
Level 24: Ingress Path Mismatch - Mission Debrief · Objective: Fix an Ingress configuration where the path routing was incorrectly configured, causing 404 errors for all requests to the application.
Level 25: NetworkPolicy Too Restrictive - Mission Debrief · Objective: Fix an overly restrictive NetworkPolicy that was blocking legitimate traffic between frontend and backend pods.
Level 26: Session Affinity Missing - Mission Debrief · Objective: Fix a stateful application that was losing user sessions by configuring session affinity on the Kubernetes Service.
Level 27: Cross-namespace Service Communication - Mission Debrief · Objective: Fix a frontend application that couldn't communicate with a backend service in a different namespace by using the proper DNS FQDN format.
Old Notes · Written on 2020-01-15, these notes cover basics
Runtime Labs · Labs that deploy real workloads and verify fixes on a live cluster
Sample Flashcard Source · Q: What is the default data home directory?
Scenario Drills · Structured incident-response scenarios for practicing triage, containment, and resolution
Scoring Rubric - General Criteria · This rubric applies to all topics. Each topic may also have specific criteria
Scoring Rubric - Kubernetes · Scoring rubric for kubernetes skill assessments
Scoring Rubric - Linux & DevOps CLI · Scoring rubric for linux skill assessments
Scoring Rubric - Networking · Scoring rubric for networking skill assessments
Scripting Rosetta β€” Lesson 1: Text Processing · Bundle: Bash + Python + CLI Tools + Regex
Scripting Rosetta β€” Lesson 2: File Operations · Bundle: Bash + Python + CLI Tools
Skills Scorecard · Rate yourself 0-9 for each topic. Be honest - the point is to find gaps, not impress yourself
Story Arc Flashcard Index · Guide for story arc flashcard index
Training Metadata Conventions · Metadata and tooling conventions for the learner asset system
Training Registry · Single source of truth for all learning assets in this repo
Typing Trainer · Guide for typing trainer
πŸŽ“ Level 21 Debrief: Service Selector Mismatch · You just fixed a service selector mismatch - one of the most common networking issues in Kubernetes!
πŸŽ“ Level 22 Debrief: NodePort Configuration · You fixed a NodePort service configuration issue!
πŸŽ“ Level 23 Debrief: DNS Resolution in Kubernetes · You fixed a DNS resolution failure where a pod couldn't connect to a service because it was using the wrong hostname!
πŸŽ“ LEVEL 32 DEBRIEF: Volume Mount Path Configuration · Congratulations! You've successfully fixed a volume mount path misconfiguration. This is one of the most common storage errors in Kubernetes!
πŸŽ“ LEVEL 33 DEBRIEF: PV/PVC Access Modes · Congratulations! You've successfully fixed the access mode configuration! Understanding access modes is critical for shared storage scenarios.
πŸŽ“ LEVEL 34 DEBRIEF: StatefulSet Volume Templates · Congratulations! You've mastered StatefulSet volumeClaimTemplates - the key to providing persistent, per-pod storage for stateful applications!
πŸŽ“ LEVEL 35 DEBRIEF: StorageClass Configuration · Congratulations! You've mastered StorageClass troubleshooting - the foundation of dynamic storage provisioning in Kubernetes!
πŸŽ“ LEVEL 36 DEBRIEF: ConfigMap Key Management · Congratulations! You've mastered ConfigMap key references - essential for application configuration in Kubernetes!
πŸŽ“ LEVEL 37 DEBRIEF: Base64 Encoding β‰  Encryption · Congratulations! You've discovered a critical security lesson about Kubernetes Secrets!
πŸŽ“ LEVEL 38 DEBRIEF: Volume Permissions & fsGroup · Congratulations! You've mastered volume permissions - critical for secure container file access!
πŸŽ“ LEVEL 39 DEBRIEF: PV Reclaim Policies · Congratulations! You've mastered reclaim policies - essential for preventing accidental data loss!
πŸŽ“ LEVEL 40 DEBRIEF: emptyDir vs PersistentVolumeClaim · Congratulations! You've completed World 4 and mastered the difference between ephemeral and persistent storage!
πŸŽ“ LEVEL 41 DEBRIEF: Kubernetes RBAC (Role-Based Access Control) · Congratulations! You've mastered Kubernetes RBAC - the foundation of security and access control in production clusters!
πŸŽ“ LEVEL 42 DEBRIEF: Container SecurityContext & Privilege Escalation · Congratulations! You've secured containers using Kubernetes SecurityContext - a critical skill for production deployments!
πŸŽ“ LEVEL 43 DEBRIEF: Kubernetes ResourceQuota & Resource Management · Congratulations! You've mastered ResourceQuota - essential for multi-tenant clusters and cost control!
πŸŽ“ LEVEL 44 DEBRIEF: Kubernetes NetworkPolicy · Congratulations! You've mastered NetworkPolicy - the firewall for your Kubernetes pods!
πŸŽ“ LEVEL 45 DEBRIEF: Node Affinity & Advanced Scheduling · Congratulations! You've mastered NodeAffinity - the key to intelligent pod placement in Kubernetes!
πŸŽ“ LEVEL 46 DEBRIEF: Taints & Tolerations · Congratulations! You've mastered taints and tolerations - the gatekeepers of node scheduling!
πŸŽ“ LEVEL 47 DEBRIEF: PodDisruptionBudget · Congratulations! You've mastered PodDisruptionBudgets - maintaining availability during disruptions!
πŸŽ“ LEVEL 48 DEBRIEF: Pod Security Standards · Congratulations! You've mastered Pod Security Standards!
πŸŽ“ LEVEL 49 DEBRIEF: PriorityClass · Congratulations! PriorityClass mastered!
πŸŽ“ LEVEL 50 DEBRIEF: CHAOS FINALE - The Perfect Storm · You've conquered the CHAOS FINALE - the ultimate test combining ALL World 5 concepts!
πŸŽ“ Mission Debrief: Blue-Green Deployment Gone Wrong · You deployed a new version of your application (GREEN) using a blue-green deployment strategy, but users were still seeing the old version (BLUE).
πŸŽ“ Mission Debrief: Canary Weight Imbalance · Your canary deployment had a 50/50 traffic split (5 stable pods, 5 canary pods) instead of the intended 90/10 split.
πŸŽ“ Mission Debrief: Deployment Update Stuck · Your deployment tried to roll out a new version with image nginx:nonexistent-v2.0-xyz that doesn't exist in Docker Hub.
πŸŽ“ Mission Debrief: Fix the Crashing Pod · Your pod was crashing because it tried to run a command called nginxzz - but that command doesn't exist in the nginx container image.
πŸŽ“ Mission Debrief: Fix the Deployment · Your deployment was configured with replicas: 0, which tells Kubernetes: 'I want ZERO instances of this application running.'.
πŸŽ“ Mission Debrief: HPA Can't Scale · Your HorizontalPodAutoscaler (HPA) was configured correctly, but it couldn't scale because metrics-server was not installed.
πŸŽ“ Mission Debrief: ImagePullBackOff Mystery · Your pod was stuck in ImagePullBackOff status because Kubernetes couldn't pull the container image nginx:nonexistent-tag-xyz-123 from Docker Hub.
πŸŽ“ Mission Debrief: Init Container Gridlock · Your init container was waiting for a service that doesn't exist.
πŸŽ“ Mission Debrief: Lost Connection - Labels & Selectors · Your Service had a selector for app: frontend, but your Pod had the label app: backend.
πŸŽ“ Mission Debrief: Namespace Confusion · Your resources were deployed to the 'default' namespace instead of 'k8squest'.
πŸŽ“ Mission Debrief: PDB Blocks All Evictions · Your PodDisruptionBudget (PDB) was configured with minAvailable: 3 for a deployment with 3 replicas.
πŸŽ“ Mission Debrief: Pending Pod Problem · Your pod was stuck in Pending status because it requested 999 CPUs and 999Gi of memoryβ€”far more than any node in your cluster can provide.
πŸŽ“ Mission Debrief: Pod Logs Mystery · The PostgreSQL container needed the POSTGRES_PASSWORD environment variable to initialize, but it wasn't provided.
πŸŽ“ Mission Debrief: Port Mismatch Mayhem · Your Service was forwarding traffic to port 8080 on the container, but the NGINX container actually listens on port 80.
πŸŽ“ Mission Debrief: ReplicaSet Without Deployment · You had a standalone ReplicaSet, which is a low-level Kubernetes resource that doesn't provide update management capabilities.
πŸŽ“ Mission Debrief: Sidecar Sabotage · Your pod had two containers: a main app and a log-sidecar.
πŸŽ“ Mission Debrief: Stateful App Data Loss · You were using a Deployment for a database, which is designed for stateless applications.
πŸŽ“ Mission Debrief: The Restart Loop · Your pods were stuck in a restart loop because the liveness probe was checking endpoint /nonexistent-healthz which returned 404 (Not Found).
πŸŽ“ Mission Debrief: Traffic to Unready Pods · Your pods were receiving traffic before they were ready to handle it, causing 502 Bad Gateway errors for users.
πŸŽ“ Mission Debrief: Zero-Downtime Deployment Failure · Your deployment was configured with maxUnavailable: 100% and maxSurge: 0, which allowed Kubernetes to terminate all pods simultaneously during a.
🎯 Level 28 Debrief: Service Endpoints & Readiness Probes · Congratulations! You've mastered one of the most critical concepts in production Kubernetes: readiness probes and service endpoint management. This.
🎯 Level 29 Debrief: LoadBalancer vs NodePort Service Types · Congratulations! You've learned the crucial differences between Kubernetes service types and why LoadBalancer services don't work in local development.
🎯 Level 30 Debrief: Headless Services & StatefulSet DNS · Congratulations! You've completed World 3: Networking & Services by mastering headless services and StatefulSet DNS! This is one of the most important.
🎯 Level 31 Debrief: PersistentVolumes & PersistentVolumeClaims · Congratulations! You've mastered the fundamental concept of persistent storage in Kubernetes: the relationship between PersistentVolumes (PV) and.

Atoms (7)

Atoms · Canonical atomic knowledge decks β€” 7,056 deduplicated concepts extracted from the full grokdevops corpus.
grokdevops atoms
GrokDevOps Wiki
GrokDevOps Wiki
GrokDevOps Wiki
GrokDevOps Wiki
GrokDevOps Wiki

Other Pages (9)

Asset Registry Inventory · Every registered learning asset, grouped by domain and type.
Coverage Priorities · Generated from make report-topic-coverage-grid on 2026-03-15
GrokDevOps Flashcards
How to Contribute · Thanks for your interest in improving GrokDevOps Wiki! Here are ways you can help
kubectl Debugging Cheatsheet · Quick-reference cheatsheet for kubectl debugging commands and common workflows · Topics: Kubernetes Core · Aliases: deployments, k8s, kubernetes, openshift, pods, services
Page Not Found · The page you're looking for doesn't exist or has been moved
Start Here · Onboarding guide for the GrokDevOps training system β€” exercises, labs, flashcards, and incident scenarios for DevOps skill building.
Topic Buildout Process · Step-by-step process for building out a new training topic from zero to full coverage
Training Library · Static, browsable reference and study content. Read these to build knowledge before (or alongside) hands-on labs