---
tags:
- devops
- l1
- flashcard-deck
- fleet-ops
---
<!-- wiki:breadcrumb:start -->
[Portal](../../../../library/portal/index.md) | **Level:** [L1: Foundations](../../../../library/portal/levels.md) | **Topics:** [Fleet Operations](../../../../library/portal/topics.md) | **Domain:** DevOps & Tooling
<!-- wiki:breadcrumb:end -->

id	category	difficulty	tags	question	answer	source_path
fleet-ops/9e722f2247fa	fleet-ops	easy	fleet-ops, mindset, cattle	What is the "cattle, not pets" principle in fleet operations?	Servers should be identical, automated, and interchangeable (cattle). They are replaced when sick, not repaired. This contrasts with pet servers that are unique, hand-configured, and treated as irreplaceable — which does not scale.	training/library/topics/fleet-ops/primer.md
fleet-ops/6243f11f21eb	fleet-ops	easy	fleet-ops, parallel, execution	Why is serial execution impractical for fleet operations?	Running a command serially on 1,500 hosts at 6 seconds each takes 2.5 hours. Parallel execution tools like Ansible forks, GNU parallel, or xargs -P run commands on many hosts simultaneously, reducing the time dramatically.	training/library/topics/fleet-ops/primer.md
fleet-ops/34b5139ee899	fleet-ops	easy	fleet-ops, inventory, management	What is the purpose of a fleet inventory and why do static files not scale?	An inventory is the source of truth for what exists, where it is, and what role it plays. Static files do not scale because they become stale; instead, generate inventory dynamically from a CMDB, cloud APIs, or Kubernetes.	training/library/topics/fleet-ops/primer.md
fleet-ops/6a0b68bf0bb0	fleet-ops	medium	fleet-ops, rolling, deployment	Describe the recommended rolling operation batch progression for a fleet of 1,500 servers.	Start with a canary batch of 1 server, wait 30 minutes and validate. Then 15 servers (1%), wait 15 minutes. Then 150 servers (10%), wait 10 minutes. Then remaining servers in batches of 150 with 5-minute gaps. This limits blast radius while still completing in reasonable time.	training/library/topics/fleet-ops/primer.md
fleet-ops/c76b371fc723	fleet-ops	medium	fleet-ops, ansible, rolling	How does Ansible implement rolling updates with automatic abort on failures?	Use the serial directive with escalating batch sizes (e.g., 1, then 5%, then 25%) and max_fail_percentage (e.g., \n2) to abort if more than 2% of hosts fail. Combine with pre_tasks to drain from load balancer and post_tasks to validate health and re-add.	training/library/topics/fleet-ops/primer.md
fleet-ops/3cffff3ea458	fleet-ops	medium	fleet-ops, drift, detection	How do you detect configuration drift across a fleet?	Compare package versions across the fleet (e.g., ansible webservers -m command -a 'rpm -q nginx' | sort | uniq -c) or compare config file checksums (e.g., ansible -m stat -a 'path=/etc/nginx/nginx.conf' | grep checksum | sort | uniq -c). Variations indicate drift.	training/library/topics/fleet-ops/primer.md
fleet-ops/91f184a3a21a	fleet-ops	medium	fleet-ops, communication, architecture	What is a phone-home architecture and what are its advantages?	Servers push status reports to a central collector via HTTP POST rather than being polled. Advantages: scales better (servers push, collector receives), works through NAT/firewalls, and missing reports serve as a dead-man's switch indicating the server is down.	training/library/topics/fleet-ops/primer.md
fleet-ops/6cb295fadc46	fleet-ops	hard	fleet-ops, observability, health	Describe a fleet-wide aggregate health check script pattern.	Use GNU parallel to SSH into all hosts concurrently (e.g., 50 at a time). For each host, collect load average, memory percent, and disk percent. Classify hosts as ok, warn (disk > 80% or mem > 85%), crit (disk > 90% or mem > 95%), or unreachable. Aggregate results into a summary dashboard and run via cron every 5 minutes.	training/library/topics/fleet-ops/primer.md
fleet-ops/2c8c4f07d475	fleet-ops	hard	fleet-ops, command-bus, architecture	How does a fleet command bus (pull-based) pattern work?	An operator posts a command to a message queue tagged with a role or group. Agents on servers poll the queue for commands matching their role. Agents execute the command and post results back to the queue. The operator views aggregated results. This is the pattern used by Salt, MCollective, and Bolt.	training/library/topics/fleet-ops/primer.md
fleet-ops/23b4d7fd2367	fleet-ops	hard	fleet-ops, rollback, change-management	What should a fleet change rollback strategy include?	Before the change: snapshot the current state (e.g., capture all package versions with ansible -m command -a 'rpm -qa' --tree /tmp/fleet-snapshot/). Define rollback commands in advance. Classify the change (standard, normal, emergency). After problems: execute rollback commands (e.g., downgrade packages), validate health, and document the failure in a post-mortem.	training/library/topics/fleet-ops/primer.md

<!-- wiki:related:start -->
---

## Wiki Navigation

### Related Content

- [Fleet Operations at Scale](../../../../library/topics/fleet-ops/index.md) (Topic Pack, L2) — Fleet Operations

<!-- wiki:related:end -->
