---
tags:
- datacenter
- l1
- flashcard-deck
- datacenter-oob-and-provisioning
---
<!-- wiki:breadcrumb:start -->
[Portal](../../../../library/portal/index.md) | **Level:** [L1: Foundations](../../../../library/portal/levels.md) | **Topics:** [Out-of-Band Management](../../../../library/portal/topics.md) | **Domain:** Datacenter & Hardware
<!-- wiki:breadcrumb:end -->

id	category	difficulty	tags	question	answer	source_path
datacenter/racadm-001	datacenter	medium	datacenter, idrac, racadm	What is RACADM and when do you use it?	RACADM (Remote Access Controller Admin) is Dell's CLI for managing iDRAC. Use it for headless server configuration: setting IP addresses, resetting iDRAC, updating firmware, and querying hardware inventory. Example: racadm getconfig -g cfgLanNetworking\n\nRemember: out-of-band management (iDRAC/iLO/IPMI) works even when the OS is down. It has its own network interface, IP address, and web UI — like a KVM switch built into the server.\n\nGotcha: always verify network changes with a second connection (iDRAC/console) before applying. A bad IP change on the only interface locks you out remotely.	training/library/cheatsheets/datacenter.cheatsheet.md
datacenter/racadm-002	datacenter	medium	datacenter, racadm	How do you reset an iDRAC using RACADM?	racadm racreset soft (graceful reset) or racadm racreset hard (forced reset). Use when iDRAC web UI is unresponsive. Can also run locally: racadm -r <idrac-ip> -u admin -p pass racreset\n\nRemember: out-of-band management (iDRAC/iLO/IPMI) works even when the OS is down. It has its own network interface, IP address, and web UI — like a KVM switch built into the server.\n\nGotcha: always verify network changes with a second connection (iDRAC/console) before applying. A bad IP change on the only interface locks you out remotely.	training/library/cheatsheets/datacenter.cheatsheet.md
datacenter/racadm-003	datacenter	medium	datacenter, racadm	How do you set the iDRAC IP address via RACADM from the host OS?	racadm setniccfg -s 10.0.0.50 255.255.255.0 10.0.0.1. Can also use: racadm set iDRAC.IPv4.Address 10.0.0.50. Requires the local RACADM package installed on the host.\n\nRemember: out-of-band management (iDRAC/iLO/IPMI) works even when the OS is down. It has its own network interface, IP address, and web UI — like a KVM switch built into the server.\n\nGotcha: always verify network changes with a second connection (iDRAC/console) before applying. A bad IP change on the only interface locks you out remotely.	training/library/cheatsheets/datacenter.cheatsheet.md
datacenter/redfish-001	datacenter	medium	datacenter, redfish, api	What is Redfish and why is it replacing IPMI?	Redfish is a RESTful API standard (DMTF) for server management using JSON over HTTPS. Replaces IPMI's binary protocol. Benefits: TLS encryption, role-based access, scriptable with curl/Python, richer data model.\n\nRemember: out-of-band management (iDRAC/iLO/IPMI) works even when the OS is down. It has its own network interface, IP address, and web UI — like a KVM switch built into the server.\n\nRemember: Redfish = modern replacement for IPMI. RESTful JSON over HTTPS vs. IPMI's binary protocol. Scriptable with curl: curl -k -u admin:pass https://bmc-ip/redfish/v1/Systems/1.\n\nGotcha: always verify network changes with a second connection (iDRAC/console) before applying. A bad IP change on the only interface locks you out remotely.	training/library/cheatsheets/datacenter.cheatsheet.md
datacenter/redfish-002	datacenter	hard	datacenter, redfish	How do you power cycle a server using the Redfish API?	"curl -k -u admin:pass -X POST https://<idrac-ip>/redfish/v1/Systems/System.Embedded.1/Actions/ComputerSystem.Reset -d '{\ResetType\"":\""ForceRestart\""}' -H 'Content-Type: application/json'""\n\nRemember: Redfish = modern replacement for IPMI. RESTful JSON over HTTPS vs. IPMI's binary protocol. Scriptable with curl: curl -k -u admin:pass https://bmc-ip/redfish/v1/Systems/1.\n\nGotcha: always calculate power at full load, not idle. A server may idle at 200W but spike to 700W under CPU stress. Plan PDU capacity for peak, not average.\n\nGotcha: always verify network changes with a second connection (iDRAC/console) before applying. A bad IP change on the only interface locks you out remotely."	training/library/cheatsheets/datacenter.cheatsheet.md
datacenter/redfish-003	datacenter	medium	datacenter, redfish	How do you query server health via Redfish?	curl -k -u admin:pass https://<idrac-ip>/redfish/v1/Systems/System.Embedded.1 | jq '{Health: .Status.Health, Power: .PowerState, Model: .Model}'. Returns JSON with health status, power state, and hardware model.\n\nRemember: Redfish = modern replacement for IPMI. RESTful JSON over HTTPS vs. IPMI's binary protocol. Scriptable with curl: curl -k -u admin:pass https://bmc-ip/redfish/v1/Systems/1.\n\nGotcha: always verify network changes with a second connection (iDRAC/console) before applying. A bad IP change on the only interface locks you out remotely.	training/library/cheatsheets/datacenter.cheatsheet.md
datacenter/idrac-001	datacenter	easy	datacenter, idrac	What is iDRAC and what can you do with it?	iDRAC (Integrated Dell Remote Access Controller) is Dell's BMC. Provides: remote KVM console, power control, hardware monitoring, virtual media (mount ISO remotely), firmware updates, and serial-over-LAN. Accessible even when OS is down.\n\nRemember: out-of-band management (iDRAC/iLO/IPMI) works even when the OS is down. It has its own network interface, IP address, and web UI — like a KVM switch built into the server.	training/library/cheatsheets/datacenter.cheatsheet.md
datacenter/idrac-002	datacenter	medium	datacenter, idrac	What is the iDRAC Lifecycle Controller?	Lifecycle Controller is embedded firmware for hardware deployment and updates. Provides: OS deployment via virtual media, firmware update repository, hardware configuration export/import (SCP profiles), and hardware diagnostics — all without a running OS.\n\nRemember: out-of-band management (iDRAC/iLO/IPMI) works even when the OS is down. It has its own network interface, IP address, and web UI — like a KVM switch built into the server.	training/library/drills/datacenter_drills.md
datacenter/idrac-003	datacenter	hard	datacenter, idrac	How do you export and import server configuration profiles (SCP) with iDRAC?	Export: racadm get -t xml -f server_config.xml. Import: racadm set -t xml -f server_config.xml. SCP captures BIOS, RAID, NIC, and iDRAC settings. Use for fleet-wide consistent configuration.\n\nRemember: out-of-band management (iDRAC/iLO/IPMI) works even when the OS is down. It has its own network interface, IP address, and web UI — like a KVM switch built into the server.	training/library/drills/datacenter_drills.md
datacenter/perc-001	datacenter	medium	datacenter, raid, perc	What is a PERC controller?	PERC (PowerEdge RAID Controller) is Dell's hardware RAID controller. Manages disk arrays with hardware acceleration. Configure via: BIOS (Ctrl+R at boot), RACADM, or storcli/perccli from the OS.\n\nGotcha: in datacenter operations, always verify changes with out-of-band access (iDRAC, iLO, serial console) before and after applying them.	training/library/cheatsheets/datacenter.cheatsheet.md
datacenter/perc-002	datacenter	medium	datacenter, raid, perc	How do you check RAID status on a Dell server with PERC?	perccli /c0 show (or storcli /c0 show). Shows virtual drives, physical drives, and their state (Online, Degraded, Rebuild). Also: perccli /c0/v0 show for virtual drive details.\n\nRemember: RAID levels — RAID 0 (stripe, no redundancy), RAID 1 (mirror), RAID 5 (stripe + 1 parity), RAID 6 (stripe + 2 parity), RAID 10 (mirror + stripe). Mnemonic: '0=Zero safety, 1=One copy, 5=Five minus one survives, 10=Ten for speed+safety.'	training/library/cheatsheets/datacenter.cheatsheet.md
datacenter/raid-001	datacenter	easy	datacenter, raid	What are the key RAID levels and their trade-offs?	RAID 0: striping, no redundancy, max speed. RAID 1: mirror, 50% capacity, good for boot. RAID 5: parity, survives 1 disk loss, write penalty. RAID 6: dual parity, survives 2 disks. RAID 10: mirror+stripe, best for databases.\n\nRemember: RAID levels — RAID 0 (stripe, no redundancy), RAID 1 (mirror), RAID 5 (stripe + 1 parity), RAID 6 (stripe + 2 parity), RAID 10 (mirror + stripe). Mnemonic: '0=Zero safety, 1=One copy, 5=Five minus one survives, 10=Ten for speed+safety.'\n\nGotcha: always verify network changes with a second connection (iDRAC/console) before applying. A bad IP change on the only interface locks you out remotely.	training/library/cheatsheets/datacenter.cheatsheet.md
datacenter/raid-002	datacenter	medium	datacenter, raid	What is a hot spare and when does it activate?	A hot spare is an unused disk assigned to a RAID controller that automatically replaces a failed drive. When a disk fails, the controller starts rebuilding onto the hot spare immediately — no human intervention needed. Reduces exposure window.\n\nGotcha: in datacenter operations, always verify changes with out-of-band access (iDRAC, iLO, serial console) before and after applying them.	training/library/drills/datacenter_drills.md
datacenter/raid-003	datacenter	hard	datacenter, raid	What is the RAID write penalty for RAID 5 vs RAID 10?	RAID 5 write penalty = 4 (2 reads + 2 writes per logical write for parity). RAID 10 write penalty = 2 (1 write to each mirror). RAID 5 is poor for write-heavy workloads like databases.\n\nRemember: RAID levels — RAID 0 (stripe, no redundancy), RAID 1 (mirror), RAID 5 (stripe + 1 parity), RAID 6 (stripe + 2 parity), RAID 10 (mirror + stripe). Mnemonic: '0=Zero safety, 1=One copy, 5=Five minus one survives, 10=Ten for speed+safety.'	training/library/drills/datacenter_drills.md
datacenter/firmware-001	datacenter	hard	datacenter, firmware	What is the recommended order for firmware updates on Dell servers?	1) iDRAC/BMC first (management plane). 2) BIOS. 3) RAID controller (PERC). 4) NIC firmware. 5) Drive firmware last. Always read release notes — some updates require specific ordering or reboot between steps.\n\nGotcha: firmware updates often require a reboot and can occasionally brick a system. Always have an out-of-band management path and a rollback plan before flashing firmware.	training/library/drills/datacenter_drills.md
datacenter/firmware-002	datacenter	hard	datacenter, firmware	How do you update firmware on Dell servers at scale?	Dell Repository Manager (DRM) creates custom update repositories. Dell System Update (DSU) applies updates from repos. For automation: Redfish API firmware update endpoints, or Ansible with dellemc.openmanage collection.\n\nGotcha: firmware updates often require a reboot and can occasionally brick a system. Always have an out-of-band management path and a rollback plan before flashing firmware.	training/library/drills/datacenter_drills.md
datacenter/lifecycle-001	datacenter	medium	datacenter, lifecycle	What does the Dell Lifecycle Controller provide?	Embedded management firmware for: OS deployment (virtual media / PXE), firmware updates without OS, hardware diagnostics, RAID configuration, system configuration backup/restore (SCP profiles), and driver packs.\n\nGotcha: in datacenter operations, always verify changes with out-of-band access (iDRAC, iLO, serial console) before and after applying them.	training/library/drills/datacenter_drills.md
datacenter/bios-001	datacenter	easy	datacenter, bios, boot	What is the difference between UEFI and Legacy BIOS boot?	Legacy BIOS: MBR-based, 2TB disk limit, 16-bit. UEFI: GPT-based, no practical disk limit, Secure Boot support, faster boot, network stack. Modern servers default to UEFI. PXE works with both but uses different bootloaders.\n\nGotcha: firmware updates often require a reboot and can occasionally brick a system. Always have an out-of-band management path and a rollback plan before flashing firmware.	training/library/cheatsheets/datacenter.cheatsheet.md
datacenter/bios-002	datacenter	hard	datacenter, bios	What BIOS/UEFI settings matter most for server performance?	1) Virtualization extensions (VT-x/AMD-V) enabled. 2) Power profile set to Performance (not Balanced). 3) C-states and P-states configured per workload. 4) NUMA enabled. 5) Boot order set correctly (PXE for provisioning, disk for production).\n\nGotcha: firmware updates often require a reboot and can occasionally brick a system. Always have an out-of-band management path and a rollback plan before flashing firmware.	training/library/drills/datacenter_drills.md
datacenter/pxe-001	datacenter	medium	datacenter, pxe, provisioning	What are the components of a PXE boot infrastructure?	1) DHCP server (options 66/67 for TFTP server and bootfile). 2) TFTP server (serves bootloader like pxelinux.0 or grubx64.efi). 3) HTTP server (kickstart/preseed/cloud-init configs and OS images). 4) Provisioning VLAN.\n\nGotcha: in datacenter operations, always verify changes with out-of-band access (iDRAC, iLO, serial console) before and after applying them.	training/library/cheatsheets/datacenter.cheatsheet.md
datacenter/pxe-002	datacenter	hard	datacenter, pxe	A server won't PXE boot. What do you check?	1) NIC is first boot device in BIOS. 2) DHCP is reachable and responding with next-server. 3) TFTP service is running. 4) Bootloader matches boot mode (UEFI vs Legacy). 5) Server is on the provisioning VLAN. 6) Check DHCP lease logs.\n\nGotcha: in datacenter operations, always verify changes with out-of-band access (iDRAC, iLO, serial console) before and after applying them.	training/library/drills/datacenter_drills.md
datacenter/troubleshoot-001	datacenter	hard	datacenter, troubleshooting	Server shows correctable ECC memory errors. What do you do?	1) Check edac-util -s or mcelog for error counts and DIMM location. 2) Check iDRAC/BMC event log. 3) dmidecode -t memory to identify the slot. Correctable errors are handled by ECC but a rising count means the DIMM is degrading — schedule replacement.\n\nGotcha: in datacenter operations, always verify changes with out-of-band access (iDRAC, iLO, serial console) before and after applying them.	training/library/drills/datacenter_drills.md
datacenter/troubleshoot-002	datacenter	medium	datacenter, troubleshooting	A server keeps thermal-throttling. What do you check?	1) Inlet temperature (should be 18-27C per ASHRAE). 2) Fan RPM (ipmitool sensor list | grep -i fan — 0 RPM = dead fan). 3) Airflow (blanking panels installed? cables blocking?). 4) Dust on heatsinks. 5) Hot/cold aisle containment intact.\n\nRemember: ASHRAE recommended inlet temperature: 18-27C (64-80F). Every 10C above optimal roughly halves component lifespan.\n\nGotcha: always verify network changes with a second connection (iDRAC/console) before applying. A bad IP change on the only interface locks you out remotely.	training/library/drills/datacenter_drills.md
datacenter/troubleshoot-003	datacenter	easy	datacenter, troubleshooting	How do you remotely access a server when the OS is completely hung?	Use out-of-band management (iDRAC/iLO/IPMI). Connect via web console for virtual KVM, or ipmitool -I lanplus for CLI. You can view the screen, send Ctrl-Alt-Del, power cycle, or activate serial-over-LAN — all independent of the OS.\n\nGotcha: always verify network changes with a second connection (iDRAC/console) before applying. A bad IP change on the only interface locks you out remotely.	training/library/cheatsheets/datacenter.cheatsheet.md
datacenter/troubleshoot-004	datacenter	medium	datacenter, troubleshooting, network	A server NIC shows CRC errors and frame errors. What is likely wrong?	Physical layer issue: bad cable, damaged SFP/transceiver, dirty fiber connector, or cable exceeding max distance. Check with ethtool -S eth0. Swap cable first. CRC/frame errors almost always point to physical layer, not software.\n\nGotcha: in datacenter operations, always verify changes with out-of-band access (iDRAC, iLO, serial console) before and after applying them.	training/library/drills/datacenter_drills.md
datacenter/15d01e301dcb	datacenter	medium	datacenter-ops,capacity,planning	Describe a scenario where you had to execute a full-scale disaster recovery plan, including failover and failback procedures.	Capacity planning: monitor current utilization (CPU, memory, storage, network, power), project growth trends, plan for N+1 redundancy, consider lead times for procurement, model seasonal peaks, and maintain buffer capacity (typically 20-30% headroom). Tools: DCIM software, custom dashboards.\n\nRemember: out-of-band management (iDRAC/iLO/IPMI) works even when the OS is down. It has its own network interface, IP address, and web UI — like a KVM switch built into the server.	projects/knowledge/interview/miscellaneous/096-describe-a-scenario-where-you-had-to-execute-a-ful.txt
datacenter/1d40789ec14b	datacenter	easy	datacenter-ops,monitoring,basics	How do you handle the decommissioning of servers and equipment in a data center?	Essential DC monitoring: server health (CPU, memory, disk via SNMP/agents), network (bandwidth, errors, latency), environmental (temperature, humidity sensors), power (PDU metrics, UPS status), and application-level metrics. Tools: Prometheus, Nagios, DCIM, BMS for environmental.\n\nRemember: decommissioning checklist: data wipe (NIST 800-88), certificate of destruction, asset tag removal, inventory update. Failing to wipe drives = data breach risk.	projects/knowledge/interview/miscellaneous/087-how-do-you-handle-the-decommissioning-of-servers-a.txt
datacenter/1d8d013997c2	datacenter	hard	datacenter-ops,security,physical	How do you troubleshoot server OS boot issues?	Physical DC security: multi-factor access control (badge + biometric), mantrap/vestibule entry, CCTV surveillance with retention, visitor escort policy, rack-level locks, cabinet-level access logging, background checks for staff, security zones (public, private, restricted), and regular security audits.\n\nGotcha: in datacenter operations, always verify changes with out-of-band access (iDRAC, iLO, serial console) before and after applying them.	projects/knowledge/interview/miscellaneous/029-how-do-you-troubleshoot-server-os-boot-issues.txt
datacenter/276dbfefb646	datacenter	easy	datacenter-ops,hardware,lifecycle	Explain the importance of redundancy in a data center environment.	Hardware lifecycle: procurement -> receiving/asset tagging -> burn-in testing -> rack/stack/cable -> production deployment -> maintenance/patching -> capacity monitoring -> decommission -> secure data wipe -> disposal/recycling. Track in CMDB/asset management system.\n\nGotcha: always verify network changes with a second connection (iDRAC/console) before applying. A bad IP change on the only interface locks you out remotely.	projects/knowledge/interview/miscellaneous/002-explain-the-importance-of-redundancy-in-a-data-cen.txt
datacenter/2bd3a7407143	datacenter	easy	datacenter-ops,power,ups	What are IPMI best practices?	UPS (Uninterruptible Power Supply) provides battery backup during power outages. Types: online (double-conversion, best protection), line-interactive, standby. Sizing: calculate total rack load in kVA, add 20-30% headroom. Runtime: typically 10-15 minutes to allow graceful shutdown or generator start.\n\nRemember: out-of-band management (iDRAC/iLO/IPMI) works even when the OS is down. It has its own network interface, IP address, and web UI — like a KVM switch built into the server.	projects/knowledge/interview/miscellaneous/103-ipmi-best-practices.txt
datacenter/39f12c841235	datacenter	medium	datacenter-ops,networking,vlans	How do you validate new hardware before production?	VLANs segment broadcast domains logically. Common DC VLANs: management (IPMI/iLO), production, storage, backup, DMZ. Trunk ports carry multiple VLANs between switches. Access ports connect servers to a single VLAN. Use 802.1Q tagging. Benefits: security isolation, reduced broadcast traffic, flexible topology.\n\nGotcha: always verify network changes with a second connection (iDRAC/console) before applying. A bad IP change on the only interface locks you out remotely.	projects/knowledge/interview/miscellaneous/102-how-do-you-validate-new-hardware.txt
datacenter/5bed3e35219a	datacenter	hard	datacenter-ops,networking,load-balancing	What is the difference between disaster recovery and business continuity?	DC load balancing: L4 (TCP/UDP, fastest, DSR for asymmetric traffic) vs L7 (HTTP-aware, SSL termination, content routing). HA: active-passive or active-active pairs. Health checks: TCP, HTTP, custom scripts. Session persistence: source IP, cookie-based. Tools: F5, HAProxy, Nginx, cloud LBs.\n\nGotcha: always verify network changes with a second connection (iDRAC/console) before applying. A bad IP change on the only interface locks you out remotely.	projects/knowledge/interview/miscellaneous/041-what-is-the-difference-between-disaster-recovery-a.txt
datacenter/5f6b86b3bc6f	datacenter	easy	datacenter-ops,networking,dns	How do you validate the effectiveness of a disaster recovery plan through testing and simulations?	DNS in datacenter: internal DNS for service discovery and name resolution, split-horizon DNS (different answers for internal/external), forward and reverse zones, low TTLs for services that move. Redundancy: primary + secondary DNS servers. Tools: BIND, PowerDNS, Infoblox for IPAM+DNS.\n\nGotcha: always verify network changes with a second connection (iDRAC/console) before applying. A bad IP change on the only interface locks you out remotely.	projects/knowledge/interview/miscellaneous/100-how-do-you-validate-the-effectiveness-of-a-disaste.txt
datacenter/6910740b7a5e	datacenter	medium	datacenter-ops,networking,firewall	What is the role of a Data Center Engineer, and what are the key responsibilities?	DC firewall deployment: perimeter firewalls (north-south traffic), internal firewalls between security zones, host-based firewalls for defense in depth. Rule management: least-privilege, documented, regularly audited, change-controlled. Modern: micro-segmentation with distributed firewalls for east-west traffic control.\n\nGotcha: in datacenter operations, always verify changes with out-of-band access (iDRAC, iLO, serial console) before and after applying them.	projects/knowledge/interview/miscellaneous/001-what-is-the-role-of-a-data-center-engineer-and-wha.txt
datacenter/77f796472027	datacenter	easy	datacenter-ops,power,pdu	What fails most in datacenters?	PDU (Power Distribution Unit) distributes power within a rack. Types: basic (power distribution only), metered (per-outlet monitoring), switched (remote on/off per outlet), intelligent (monitoring + switching + environmental sensors). Mount vertically in rack, use A/B PDUs from separate circuits for redundancy.\n\nGotcha: in datacenter operations, always verify changes with out-of-band access (iDRAC, iLO, serial console) before and after applying them.	projects/knowledge/interview/miscellaneous/101-what-fails-most-in-datacenters.txt
datacenter/96cd46afd13c	datacenter	medium	datacenter-ops,power,monitoring	Explain the process of installing and configuring a new server.	Power monitoring: track per-rack power consumption via smart PDUs, monitor UPS load and battery health, track PUE trending, alert on circuit approaching capacity (>80%), log historical data for capacity planning. Tools: DCIM software, PDU SNMP polling, BMS integration for facility power.\n\nGotcha: in datacenter operations, always verify changes with out-of-band access (iDRAC, iLO, serial console) before and after applying them.	projects/knowledge/interview/miscellaneous/007-explain-the-process-of-installing-and-configuring-.txt
datacenter/afa9ae005ff7	datacenter	hard	datacenter-ops,automation,infrastructure-as-code	How do you perform server hardware troubleshooting?	Infrastructure as Code in DC: Terraform for provisioning (VMs, networks, storage), Ansible for configuration, Packer for golden images, Git for version control. Benefits: reproducible environments, change tracking, peer review via PRs, automated testing. Treat infrastructure definitions like application code.\n\nGotcha: in datacenter operations, always verify changes with out-of-band access (iDRAC, iLO, serial console) before and after applying them.	projects/knowledge/interview/miscellaneous/006-how-do-you-perform-server-hardware-troubleshooting.txt
datacenter/c4729d8f5be9	datacenter	medium	datacenter-ops,hardware,procurement	What are the main differences between a Tier 1 and Tier 4 data center?	Hardware procurement: define specs (CPU, RAM, storage, NIC requirements), get vendor quotes (Dell, HPE, Lenovo, Supermicro), compare TCO (not just purchase price — include power, cooling, support), negotiate volume discounts and SLAs, verify lead times (4-12 weeks typical), plan for standardization.\n\nGotcha: in datacenter operations, always verify changes with out-of-band access (iDRAC, iLO, serial console) before and after applying them.	projects/knowledge/interview/miscellaneous/005-what-are-the-main-differences-between-a-tier-1-and.txt
datacenter/cb4de3448458	datacenter	hard	datacenter-ops,security,zero-trust	Discuss the role of backup rotation strategies in disaster recovery.	Zero-trust in datacenter: assume no implicit trust, verify every request. Implementation: micro-segmentation (per-workload firewall rules), mTLS between services, identity-based access (not network-based), continuous authentication, least-privilege access, encrypted east-west traffic, and comprehensive logging for all access.\n\nGotcha: in datacenter operations, always verify changes with out-of-band access (iDRAC, iLO, serial console) before and after applying them.	projects/knowledge/interview/miscellaneous/043-discuss-the-role-of-backup-rotation-strategies-in-.txt
datacenter/d7872b63f372	datacenter	medium	datacenter-ops,networking,vpn	Discuss the steps involved in applying patches and updates to a server OS.	DC VPN types: site-to-site (IPSec tunnels between DCs or DC-to-cloud), remote access (engineer VPN for management), DMVPN (dynamic mesh between multiple sites). Key considerations: bandwidth sizing, encryption overhead, split vs full tunnel, redundant VPN endpoints, and monitoring tunnel health.\n\nGotcha: firmware updates often require a reboot and can occasionally brick a system. Always have an out-of-band management path and a rollback plan before flashing firmware.\n\nGotcha: always verify network changes with a second connection (iDRAC/console) before applying. A bad IP change on the only interface locks you out remotely.	projects/knowledge/interview/miscellaneous/027-discuss-the-steps-involved-in-applying-patches-and.txt
datacenter/e8dc5878b193	datacenter	medium	datacenter-ops,storage,virtualization	Discuss the considerations for migrating servers or workloads to the cloud.	Storage virtualization abstracts physical storage into logical pools. Benefits: simplified management, thin provisioning, non-disruptive migration, automated tiering. Technologies: SAN volume controllers, VM datastores (vSAN, Ceph), software-defined storage. Enables storage mobility and capacity optimization.\n\nGotcha: in datacenter operations, always verify changes with out-of-band access (iDRAC, iLO, serial console) before and after applying them.	projects/knowledge/interview/miscellaneous/062-discuss-the-considerations-for-migrating-servers-o.txt
datacenter/f1f56edf0a8c	datacenter	easy	datacenter-ops,hardware,asset-management	What is server virtualization, and how does it benefit data center operations?	Asset management: tag all hardware with unique asset IDs, record in CMDB (serial number, location, owner, status), track lifecycle stage, conduct periodic physical inventory audits, update records on moves/adds/changes. Tools: NetBox, Device42, ServiceNow CMDB. Accurate asset data enables capacity planning and compliance.\n\nGotcha: in datacenter operations, always verify changes with out-of-band access (iDRAC, iLO, serial console) before and after applying them.	projects/knowledge/interview/miscellaneous/008-what-is-server-virtualization-and-how-does-it-bene.txt
datacenter/fd563bb24a70	datacenter	hard	datacenter-ops,security,encryption	Discuss the importance of backup and disaster recovery planning in a data center.	DC encryption strategy: at rest (disk encryption, SED drives, database TDE), in transit (TLS 1.2+ for all services, IPSec for inter-DC, mTLS for service mesh), key management (HSM for master keys, automated rotation, separation of duties). Certificate management: PKI infrastructure, automated renewal, short-lived certs.\n\nGotcha: always verify network changes with a second connection (iDRAC/console) before applying. A bad IP change on the only interface locks you out remotely.	projects/knowledge/interview/miscellaneous/023-discuss-the-importance-of-backup-and-disaster-reco.txt
datacenter/00e3c08956cb	datacenter	medium	miscellaneous, backup, control-flow, functions	How do you handle stressful situations, such as a critical system failure or a major security incident in the data center?	• Maintain Calmness: Stay calm and composed to make well-informed decisions under pressure. • Incident Response Plan: Activate the predefined incident response plan to ensure a structured and coordinated approach to addressing the issue. • Communication: Communicate transparently with relevant stakeholders, providing updates on the situation, progress, and expected resolution timelines. • Priority Setting: Prioritize tasks based on the severity and impact of the incident to address critical issues first.\n\nRemember: defense in depth — layer multiple security controls. No single mechanism is sufficient. Assume breach and design for containment.	
datacenter/12a317f6277e	datacenter	medium	miscellaneous, backup, control-flow, logging	What scripting languages are you proficient in, and how have you used them in a data center environment?	I am proficient in scripting languages such as PowerShell, Python, and Bash. In a data center environment, these languages have been instrumental in automating various tasks: **PowerShell:* • Used for automating Windows-based tasks, such as server provisioning, configuration management, and Active Directory operations. **Python:* • Applied for cross-platform automation, scripting, and developing custom tools for data center monitoring, log analysis, and reporting.\n\nGotcha: always verify network changes with a second connection (iDRAC/console) before applying. A bad IP change on the only interface locks you out remotely.	
datacenter/2323b2cd4af9	datacenter	medium	miscellaneous, backup, control-flow, functions	How do you balance the need for innovation with maintaining a stable and reliable data center environment?	• Risk Assessment: Conduct a thorough risk assessment to evaluate the potential impact of innovations on data center stability. • Pilot Programs: Implement innovations through pilot programs in controlled environments to assess their impact before widespread adoption. • Gradual Integration: Integrate innovations gradually, allowing for careful monitoring of performance and stability. • Compatibility Testing: Test innovations for compatibility with existing infrastructure, applications, and workflows to avoid disruptions.	
datacenter/ed0935f4193e	datacenter	medium	miscellaneous, backup, control-flow, functions	Provide an example of a time when you had to quickly adapt to a changing situation or unexpected challenge in the data center.	During a scheduled maintenance window, unexpected issues arose, causing a critical application to go offline. The situation required swift adaptation and resolution. **Steps Taken:* •  • Immediate Triage: Conducted an immediate triage to identify the cause of the application outage. • Communication: Communicated transparently with stakeholders, notifying them of the issue and setting realistic expectations for resolution timelines. • Emergency Response Plan: Activated the emergency response plan to prioritize critical applications and services.	
datacenter/a1b2c3d4	datacenter	easy	bmc,ipmi,basics	What is a BMC and why is it always powered on?	A BMC (Baseboard Management Controller) is an independent embedded computer within a server with its own CPU (ARM), RAM, NIC, and flash storage. It runs on standby power, so it is always active as long as the server has AC power -- even when the main system is completely off. It provides out-of-band management including remote console, power control, and hardware monitoring.\n\nRemember: out-of-band management (iDRAC/iLO/IPMI) works even when the OS is down. It has its own network interface, IP address, and web UI — like a KVM switch built into the server.\n\nGotcha: always calculate power at full load, not idle. A server may idle at 200W but spike to 700W under CPU stress. Plan PDU capacity for peak, not average.	training/library/topics/ipmi-and-ipmitool/primer.md
datacenter/b2c3d4e5	datacenter	easy	bmc,ipmi,vendors	What are the vendor-specific names for BMC products, and what protocol do they all share?	Dell calls theirs iDRAC, HP/HPE uses iLO, Supermicro uses a generic IPMI BMC, and Lenovo uses XClarity/IMM. All of them speak the IPMI protocol, so ipmitool works with all vendors. Vendor UIs add proprietary features on top of the IPMI common denominator.\n\nRemember: out-of-band management (iDRAC/iLO/IPMI) works even when the OS is down. It has its own network interface, IP address, and web UI — like a KVM switch built into the server.\n\nGotcha: always verify network changes with a second connection (iDRAC/console) before applying. A bad IP change on the only interface locks you out remotely.	training/library/topics/ipmi-and-ipmitool/primer.md
datacenter/c3d4e5f6	datacenter	medium	ipmi,ipmitool,lanplus	What is the difference between IPMI 1.5 (lan) and IPMI 2.0 (lanplus) transport?	"IPMI 1.5 uses RMCP with weak MD5 authentication and no encryption. IPMI 2.0 (lanplus) uses RMCP+ with RAKP key exchange, AES-128-CBC encryption, and HMAC-SHA integrity. Always use ""ipmitool -I lanplus"" for remote operations. Both use UDP port 623.\n\nRemember: out-of-band management (iDRAC/iLO/IPMI) works even when the OS is down. It has its own network interface, IP address, and web UI — like a KVM switch built into the server.\n\nGotcha: always verify network changes with a second connection (iDRAC/console) before applying. A bad IP change on the only interface locks you out remotely."	training/library/topics/ipmi-and-ipmitool/primer.md
datacenter/d4e5f6a7	datacenter	medium	ipmi,sol,console	What is Serial-over-LAN (SOL) and when would you use it?	"SOL redirects the server's serial console through the BMC to your terminal over the network. Use it to see boot output (POST, GRUB, kernel messages), diagnose kernel panics, and interact with the system when the OS network stack is down. Connect with ""ipmitool -I lanplus -H <bmc> -U admin -P pass sol activate"" and disconnect with ~. (tilde-dot).\n\nGotcha: always verify network changes with a second connection (iDRAC/console) before applying. A bad IP change on the only interface locks you out remotely."	training/library/topics/ipmi-and-ipmitool/primer.md
datacenter/e5f6a7b8	datacenter	medium	ipmi,sel,events	What is the System Event Log (SEL) and why must you export it regularly?	"The SEL is a circular buffer in the BMC's non-volatile storage that records hardware events: temperature threshold crossings, fan failures, PSU faults, ECC errors, and boot events. It is small (typically 512-2048 entries). When full, events are either dropped or overwritten. Export with ""ipmitool sel elist"" and clear with ""ipmitool sel clear"" to free space.\n\nGotcha: always verify network changes with a second connection (iDRAC/console) before applying. A bad IP change on the only interface locks you out remotely."	training/library/topics/ipmi-and-ipmitool/primer.md
datacenter/f6a7b8c9	datacenter	medium	ipmi,sensors,thresholds	What are the six IPMI sensor threshold levels, and what happens when a reading crosses one?	"From lowest to highest: Lower Non-Recoverable (LNR), Lower Critical (LC), Lower Non-Critical (LNC), Upper Non-Critical (UNC), Upper Critical (UC), Upper Non-Recoverable (UNR). When a reading crosses a threshold, the BMC logs an event to the SEL and can send an SNMP trap or PET alert. View thresholds with ""ipmitool sensor get <name>"".\n\nRemember: out-of-band management (iDRAC/iLO/IPMI) works even when the OS is down. It has its own network interface, IP address, and web UI — like a KVM switch built into the server.\n\nGotcha: always verify network changes with a second connection (iDRAC/console) before applying. A bad IP change on the only interface locks you out remotely."	training/library/topics/ipmi-and-ipmitool/primer.md
datacenter/a7b8c9da	datacenter	easy	ipmi,power,commands	What ipmitool commands control server power, and what is the correct order for an unresponsive server?	"Key commands: power status, power soft (ACPI graceful), power off (hard), power on, power cycle, power reset. For an unresponsive server: try ""power soft"" first, wait 60 seconds, check SOL for shutdown progress, then escalate to ""power cycle"" only after confirming the OS is truly hung.\n\nRemember: out-of-band management (iDRAC/iLO/IPMI) works even when the OS is down. It has its own network interface, IP address, and web UI — like a KVM switch built into the server.\n\nGotcha: always calculate power at full load, not idle. A server may idle at 200W but spike to 700W under CPU stress. Plan PDU capacity for peak, not average."	training/library/topics/ipmi-and-ipmitool/street_ops.md
datacenter/b8c9daeb	datacenter	hard	ipmi,security,rakp	What is the IPMI RAKP authentication vulnerability (CVE-2013-4786) and why can't it be patched?	During IPMI 2.0's RAKP handshake, the BMC returns a salted HMAC of the password to any unauthenticated client who sends a session request. An attacker can capture this hash and crack it offline (Hashcat mode 7300). This is a protocol design flaw, not an implementation bug, so it cannot be patched. The only mitigation is isolating BMC traffic on a dedicated management VLAN.\n\nRemember: out-of-band management (iDRAC/iLO/IPMI) works even when the OS is down. It has its own network interface, IP address, and web UI — like a KVM switch built into the server.\n\nGotcha: always verify network changes with a second connection (iDRAC/console) before applying. A bad IP change on the only interface locks you out remotely.	training/library/topics/ipmi-and-ipmitool/primer.md
datacenter/c9daebfc	datacenter	hard	ipmi,security,cipher0	What is IPMI cipher suite 0 and why is it dangerous?	"Cipher suite 0 means no authentication at all. Some BMCs accept it by default, allowing anyone on the network to execute IPMI commands without credentials. Test with ""ipmitool -I lanplus -H <bmc> -C 0 -U '' -P '' chassis status"" -- if it succeeds, the BMC is wide open. Disable cipher 0 via vendor-specific commands (e.g., racadm on Dell).\n\nRemember: out-of-band management (iDRAC/iLO/IPMI) works even when the OS is down. It has its own network interface, IP address, and web UI — like a KVM switch built into the server.\n\nGotcha: always verify network changes with a second connection (iDRAC/console) before applying. A bad IP change on the only interface locks you out remotely."	training/library/topics/ipmi-and-ipmitool/primer.md
datacenter/daebfcad	datacenter	medium	bmc,network,config	How do you configure the BMC network settings using ipmitool?	"Use ""ipmitool lan print 1"" to view current settings. Set static IP with: ""ipmitool lan set 1 ipsrc static"", then set ipaddr, netmask, and defgw ipaddr. For DHCP: ""ipmitool lan set 1 ipsrc dhcp"". Set management VLAN with ""ipmitool lan set 1 vlan id 100"". These commands work both in-band (local) and over-LAN (remote).\n\nRemember: out-of-band management (iDRAC/iLO/IPMI) works even when the OS is down. It has its own network interface, IP address, and web UI — like a KVM switch built into the server.\n\nGotcha: always verify network changes with a second connection (iDRAC/console) before applying. A bad IP change on the only interface locks you out remotely."	training/library/topics/ipmi-and-ipmitool/primer.md
datacenter/ebfcadbe	datacenter	medium	bmc,reset,firmware	When should you perform a BMC cold reset and what does it do?	ipmitool mc reset cold restarts the BMC firmware without affecting the host OS. Use it when the BMC web UI is unresponsive, sensor readings are stale, or SOL sessions won't connect. If the BMC is completely unresponsive to IPMI, a cold reset won't help -- you need an AC power cycle (pull power cables, wait 30 seconds).\n\nRemember: out-of-band management (iDRAC/iLO/IPMI) works even when the OS is down. It has its own network interface, IP address, and web UI — like a KVM switch built into the server.\n\nGotcha: always verify network changes with a second connection (iDRAC/console) before applying. A bad IP change on the only interface locks you out remotely.	training/library/topics/ipmi-and-ipmitool/primer.md
datacenter/fcadbeaf	datacenter	easy	redfish,api,basics	How does Redfish differ from IPMI for server management?	Redfish uses HTTPS with JSON payloads over TCP 443, replacing IPMI's binary protocol over UDP 623. It provides TLS encryption, token-based session auth (no RAKP hash leak), a richer data model covering storage/BIOS/NICs/firmware, and is scriptable with any HTTP client (curl, Python, Ansible).\n\nRemember: out-of-band management (iDRAC/iLO/IPMI) works even when the OS is down. It has its own network interface, IP address, and web UI — like a KVM switch built into the server.\n\nRemember: Redfish = modern replacement for IPMI. RESTful JSON over HTTPS vs. IPMI's binary protocol. Scriptable with curl: curl -k -u admin:pass https://bmc-ip/redfish/v1/Systems/1.	training/library/topics/redfish/primer.md
datacenter/adbecfb0	datacenter	hard	redfish,virtual-media,provisioning	How do you mount a remote ISO via the Redfish API for OS provisioning?	"POST to the VirtualMedia InsertMedia action endpoint with the image URL: curl -sk -u admin:pass -X POST https://<bmc>/redfish/v1/Managers/iDRAC.Embedded.1/VirtualMedia/CD/Actions/VirtualMedia.InsertMedia -H ""Content-Type: application/json"" -d '{""Image"": ""https://iso-repo/rhel9.iso""}'. Then set one-time boot to CD via PATCH to the Systems endpoint and power cycle.\n\nRemember: Redfish = modern replacement for IPMI. RESTful JSON over HTTPS vs. IPMI's binary protocol. Scriptable with curl: curl -k -u admin:pass https://bmc-ip/redfish/v1/Systems/1."	training/library/topics/redfish/primer.md
datacenter/becfb0c1	datacenter	medium	redfish,events,webhook	How does Redfish eventing replace IPMI's SNMP traps?	Redfish supports push-based eventing via Server-Sent Events (SSE) streams or webhook subscriptions. Create a subscription by POSTing to /redfish/v1/EventService/Subscriptions with a destination URL, protocol, and event types. This integrates directly with modern alerting stacks like Alertmanager, replacing IPMI's SNMP trap / PET alert mechanism.\n\nRemember: out-of-band management (iDRAC/iLO/IPMI) works even when the OS is down. It has its own network interface, IP address, and web UI — like a KVM switch built into the server.\n\nRemember: Redfish = modern replacement for IPMI. RESTful JSON over HTTPS vs. IPMI's binary protocol. Scriptable with curl: curl -k -u admin:pass https://bmc-ip/redfish/v1/Systems/1.	training/library/topics/redfish/primer.md
datacenter/cfb0c1d2	datacenter	hard	ipmi,dcmi,power	How do you read and limit server power consumption using IPMI DCMI commands?	"DCMI (Data Center Management Interface) extends IPMI for datacenter power management. Read power with ""ipmitool dcmi power reading"" to get instantaneous, minimum, maximum, and average watts. Set a power cap with ""ipmitool dcmi power set_limit action 1 limit 350"" then ""ipmitool dcmi power activate"". This lets you enforce rack-level power budgets.\n\nRemember: out-of-band management (iDRAC/iLO/IPMI) works even when the OS is down. It has its own network interface, IP address, and web UI — like a KVM switch built into the server.\n\nGotcha: always calculate power at full load, not idle. A server may idle at 200W but spike to 700W under CPU stress. Plan PDU capacity for peak, not average."	training/library/topics/ipmi-and-ipmitool/primer.md
datacenter/a1c3e5f7	datacenter	easy	raid,storage,basics	What does RAID stand for, and what is it NOT?	RAID stands for Redundant Array of Independent Disks. It combines multiple block devices to improve redundancy, capacity, or performance. RAID is NOT backup -- it does not protect against deletion, corruption, ransomware, operator error, or site loss.\n\nRemember: RAID levels — RAID 0 (stripe, no redundancy), RAID 1 (mirror), RAID 5 (stripe + 1 parity), RAID 6 (stripe + 2 parity), RAID 10 (mirror + stripe). Mnemonic: '0=Zero safety, 1=One copy, 5=Five minus one survives, 10=Ten for speed+safety.'\n\nGotcha: always verify network changes with a second connection (iDRAC/console) before applying. A bad IP change on the only interface locks you out remotely.	training/library/deep-dives/raid.and.storage.internals.md
datacenter/b2d4f6a8	datacenter	easy	raid,raid0,striping	What is RAID 0 and what happens when one disk fails?	RAID 0 uses striping with no redundancy. Data is distributed across disks for maximum performance and capacity. If any single disk dies, the entire array is lost because there is no parity or mirror copy.\n\nRemember: RAID levels — RAID 0 (stripe, no redundancy), RAID 1 (mirror), RAID 5 (stripe + 1 parity), RAID 6 (stripe + 2 parity), RAID 10 (mirror + stripe). Mnemonic: '0=Zero safety, 1=One copy, 5=Five minus one survives, 10=Ten for speed+safety.'\n\nGotcha: always verify network changes with a second connection (iDRAC/console) before applying. A bad IP change on the only interface locks you out remotely.	training/library/deep-dives/raid.and.storage.internals.md
datacenter/c3e5a7b9	datacenter	easy	raid,raid1,mirroring	What is RAID 1 and what is its primary tradeoff?	RAID 1 mirrors data across two or more disks, providing full redundancy. Read performance can benefit from multiple copies. The primary tradeoff is storage efficiency -- you lose 50% of total capacity to mirroring.\n\nRemember: RAID levels — RAID 0 (stripe, no redundancy), RAID 1 (mirror), RAID 5 (stripe + 1 parity), RAID 6 (stripe + 2 parity), RAID 10 (mirror + stripe). Mnemonic: '0=Zero safety, 1=One copy, 5=Five minus one survives, 10=Ten for speed+safety.'\n\nGotcha: always verify network changes with a second connection (iDRAC/console) before applying. A bad IP change on the only interface locks you out remotely.	training/library/deep-dives/raid.and.storage.internals.md
datacenter/d4f6b8ca	datacenter	medium	raid,raid5,parity	How does RAID 5 distribute parity and how many disk failures can it survive?	RAID 5 uses striping with distributed parity spread across all member disks (minimum 3). It can survive exactly one disk failure. Parity writes incur overhead because the controller must read old data/parity, compute new parity, then write both.\n\nRemember: RAID levels — RAID 0 (stripe, no redundancy), RAID 1 (mirror), RAID 5 (stripe + 1 parity), RAID 6 (stripe + 2 parity), RAID 10 (mirror + stripe). Mnemonic: '0=Zero safety, 1=One copy, 5=Five minus one survives, 10=Ten for speed+safety.'\n\nGotcha: always verify network changes with a second connection (iDRAC/console) before applying. A bad IP change on the only interface locks you out remotely.	training/library/deep-dives/raid.and.storage.internals.md
datacenter/e5a7c9db	datacenter	medium	raid,raid6,parity	How does RAID 6 differ from RAID 5 and when should you prefer it?	RAID 6 uses dual distributed parity, requiring a minimum of 4 disks and tolerating up to 2 simultaneous disk failures. Prefer RAID 6 over RAID 5 for large arrays where rebuild times are long (8TB+ disks can take 12-24 hours), reducing the risk of a second failure during rebuild.\n\nRemember: RAID levels — RAID 0 (stripe, no redundancy), RAID 1 (mirror), RAID 5 (stripe + 1 parity), RAID 6 (stripe + 2 parity), RAID 10 (mirror + stripe). Mnemonic: '0=Zero safety, 1=One copy, 5=Five minus one survives, 10=Ten for speed+safety.'	training/library/deep-dives/raid.and.storage.internals.md
datacenter/f6b8daec	datacenter	medium	raid,raid10,performance	What is RAID 10 and why is it preferred for write-heavy workloads like databases?	RAID 10 combines striping and mirroring (striped mirrors). It requires a minimum of 4 disks and can tolerate one failure per mirror pair. It is preferred for databases because mirror writes are simpler than parity computation, providing better write performance with strong redundancy.\n\nRemember: RAID levels — RAID 0 (stripe, no redundancy), RAID 1 (mirror), RAID 5 (stripe + 1 parity), RAID 6 (stripe + 2 parity), RAID 10 (mirror + stripe). Mnemonic: '0=Zero safety, 1=One copy, 5=Five minus one survives, 10=Ten for speed+safety.'\n\nGotcha: always verify network changes with a second connection (iDRAC/console) before applying. A bad IP change on the only interface locks you out remotely.	training/library/topics/datacenter/primer.md
datacenter/a7c9ebfd	datacenter	hard	raid,rebuild,degraded	Why is a RAID rebuild a high-risk period, especially on large-capacity disks?	During rebuild, redundancy margin is reduced, remaining disks are stressed harder with additional I/O, and performance degrades. On large disks (8TB+), rebuilds can take 12-24 hours. Another disk failure during this window can cause complete data loss on RAID 5 or exceed RAID 6 tolerance.\n\nRemember: RAID levels — RAID 0 (stripe, no redundancy), RAID 1 (mirror), RAID 5 (stripe + 1 parity), RAID 6 (stripe + 2 parity), RAID 10 (mirror + stripe). Mnemonic: '0=Zero safety, 1=One copy, 5=Five minus one survives, 10=Ten for speed+safety.'	training/library/deep-dives/raid.and.storage.internals.md
datacenter/b8dafcae	datacenter	hard	raid,write-hole,consistency	What is the RAID 5 write hole and why does it matter?	The RAID 5 write hole occurs when a crash or power loss interrupts a parity update: some data/parity blocks are written but others are not, leaving parity inconsistent. This can cause silent data corruption during degraded operation or rebuild. Mitigations include battery-backed cache, write-intent bitmaps, and partial parity logs (PPL).\n\nRemember: RAID levels — RAID 0 (stripe, no redundancy), RAID 1 (mirror), RAID 5 (stripe + 1 parity), RAID 6 (stripe + 2 parity), RAID 10 (mirror + stripe). Mnemonic: '0=Zero safety, 1=One copy, 5=Five minus one survives, 10=Ten for speed+safety.'	training/library/deep-dives/raid.and.storage.internals.md
datacenter/c9ebadbe	datacenter	medium	raid,cache,bbu	What is the difference between write-back and write-through cache on a RAID controller?	Write-back cache acknowledges writes as soon as data hits the controller's RAM cache, giving much better performance. Write-through cache waits until data is written to disk before acknowledging, which is slower but safer. Write-back requires a Battery Backup Unit (BBU) to protect cached data during power loss.\n\nRemember: RAID levels — RAID 0 (stripe, no redundancy), RAID 1 (mirror), RAID 5 (stripe + 1 parity), RAID 6 (stripe + 2 parity), RAID 10 (mirror + stripe). Mnemonic: '0=Zero safety, 1=One copy, 5=Five minus one survives, 10=Ten for speed+safety.'	training/library/topics/datacenter/primer.md
datacenter/dafc0ecf	datacenter	medium	raid,spare,hot-spare	What is the difference between a hot spare and a cold spare in a RAID array?	A hot spare is a disk installed in the system and pre-assigned to the RAID controller; it automatically begins rebuilding when a member disk fails, minimizing the degraded window. A cold spare is a replacement disk kept on the shelf that requires manual intervention to install and initiate the rebuild.\n\nRemember: RAID levels — RAID 0 (stripe, no redundancy), RAID 1 (mirror), RAID 5 (stripe + 1 parity), RAID 6 (stripe + 2 parity), RAID 10 (mirror + stripe). Mnemonic: '0=Zero safety, 1=One copy, 5=Five minus one survives, 10=Ten for speed+safety.'	training/library/topics/datacenter/primer.md
datacenter/ebad1fd0	datacenter	easy	raid,monitoring,mdstat	How do you check the status of a Linux software RAID array using mdadm?	"Run ""cat /proc/mdstat"" for a quick overview of all arrays including sync/rebuild progress. Use ""mdadm --detail /dev/md0"" for detailed info on a specific array including state, member disks, and rebuild status. Also check dmesg and journalctl -k for RAID-related kernel messages.\n\nRemember: RAID levels — RAID 0 (stripe, no redundancy), RAID 1 (mirror), RAID 5 (stripe + 1 parity), RAID 6 (stripe + 2 parity), RAID 10 (mirror + stripe). Mnemonic: '0=Zero safety, 1=One copy, 5=Five minus one survives, 10=Ten for speed+safety.'"	training/library/deep-dives/raid.and.storage.internals.md
datacenter/fcbe20e1	datacenter	hard	raid,storcli,perccli	What are perccli and storcli, and when would you use them instead of mdadm?	perccli (Dell PERC) and storcli (LSI MegaRAID/Broadcom) are CLI tools for managing hardware RAID controllers. Use them instead of mdadm when the server has a hardware RAID controller rather than Linux software RAID. They manage virtual drives, check physical disk status, configure hot spares, and monitor rebuild progress.\n\nGotcha: in datacenter operations, always verify changes with out-of-band access (iDRAC, iLO, serial console) before and after applying them.	training/library/topics/datacenter/primer.md
datacenter/adcf31f2	datacenter	medium	raid,degraded,failure	What are common RAID failure patterns that operators should watch for?	Key failure patterns include: replacing the wrong disk (classic mistake), array degraded unnoticed for too long before a second failure, running rebuilds under heavy I/O load, assuming parity protects against all corruption, ignoring SMART errors on underlying drives, and having no backup despite RAID redundancy.\n\nRemember: RAID levels — RAID 0 (stripe, no redundancy), RAID 1 (mirror), RAID 5 (stripe + 1 parity), RAID 6 (stripe + 2 parity), RAID 10 (mirror + stripe). Mnemonic: '0=Zero safety, 1=One copy, 5=Five minus one survives, 10=Ten for speed+safety.'	training/library/deep-dives/raid.and.storage.internals.md
datacenter/beda42f3	datacenter	hard	raid,smart,disk-health	What SMART attributes indicate an imminent disk failure in a RAID array?	"Critical SMART indicators include: Reallocated Sector Count (bad sectors remapped -- rising count signals failure), Current Pending Sector (sectors awaiting remap), and elevated temperature over sustained periods. Monitor with ""smartctl -a /dev/sdX"". A rising reallocated sector count is the strongest predictor of impending drive failure.\n\nRemember: RAID levels — RAID 0 (stripe, no redundancy), RAID 1 (mirror), RAID 5 (stripe + 1 parity), RAID 6 (stripe + 2 parity), RAID 10 (mirror + stripe). Mnemonic: '0=Zero safety, 1=One copy, 5=Five minus one survives, 10=Ten for speed+safety.'"	training/library/topics/datacenter/primer.md
datacenter/cfeb53f4	datacenter	hard	raid,parity,write-penalty	Why are small random writes particularly expensive on parity RAID (5/6)?	Each small random write on RAID 5/6 triggers a read-modify-write cycle: the controller must read the old data block and old parity, compute new parity via XOR, then write both the new data and new parity. This means each logical write generates 4 I/O operations, making parity RAID a poor choice for random write-heavy workloads like OLTP databases.\n\nRemember: RAID levels — RAID 0 (stripe, no redundancy), RAID 1 (mirror), RAID 5 (stripe + 1 parity), RAID 6 (stripe + 2 parity), RAID 10 (mirror + stripe). Mnemonic: '0=Zero safety, 1=One copy, 5=Five minus one survives, 10=Ten for speed+safety.'	training/library/deep-dives/raid.and.storage.internals.md
datacenter/a1f2e3d4	datacenter	easy	pxe,boot,sequence	Describe the PXE boot sequence from power-on to OS installer start.	Server powers on, NIC firmware sends a DHCP DISCOVER, DHCP server responds with an IP address plus next-server (TFTP IP) and filename (bootloader path). Server downloads the bootloader via TFTP, bootloader loads the kernel and initrd, kernel starts the installer (Kickstart/Preseed/Autoinstall) which partitions disks and installs the OS.\n\nGotcha: always calculate power at full load, not idle. A server may idle at 200W but spike to 700W under CPU stress. Plan PDU capacity for peak, not average.\n\nGotcha: always verify network changes with a second connection (iDRAC/console) before applying. A bad IP change on the only interface locks you out remotely.	training/library/topics/bare-metal-provisioning/primer.md
datacenter/b2a3f4e5	datacenter	medium	pxe,dhcp,options	What are DHCP options 66 and 67 (next-server and filename) and why are they critical for PXE?	Option 66 (next-server) tells the PXE client the IP address of the TFTP server to fetch the bootloader from. Option 67 (filename) specifies the bootloader file path, such as pxelinux.0 for BIOS or ipxe/snponly.efi for UEFI. Without these options, a PXE client receives an IP but has no bootloader to download and simply skips network boot.\n\nGotcha: always verify network changes with a second connection (iDRAC/console) before applying. A bad IP change on the only interface locks you out remotely.	training/library/topics/bare-metal-provisioning/primer.md
datacenter/c3b4a5f6	datacenter	easy	pxe,tftp,protocol	Why is TFTP used in PXE boot and what is its major limitation?	TFTP (Trivial File Transfer Protocol) is used because it is simple enough to implement in NIC firmware ROM with minimal code. Its major limitation is performance: it transfers data in 512-byte blocks with no windowing, making large file transfers (like 50-100MB initrd images) extremely slow and prone to timeouts on congested networks.	training/library/topics/bare-metal-provisioning/primer.md
datacenter/d4c5b6a7	datacenter	medium	pxe,ipxe,chainloading	What is iPXE chainloading and why is it preferred over plain PXE?	iPXE chainloading means PXE first downloads a small iPXE binary via TFTP, then iPXE takes over and downloads the kernel and initrd via HTTP instead of TFTP. HTTP is 10-50x faster, supports retries gracefully, and allows dynamic boot scripts. The iPXE script can customize boot behavior per MAC address using variables.\n\nGotcha: always verify network changes with a second connection (iDRAC/console) before applying. A bad IP change on the only interface locks you out remotely.	training/library/topics/bare-metal-provisioning/primer.md
datacenter/e5d6c7b8	datacenter	hard	pxe,uefi,bios	How do you configure DHCP to serve different bootloaders for UEFI vs Legacy BIOS clients?	"Use DHCP option architecture matching: if option arch equals 00:07 (EFI x64), serve ""ipxe/snponly.efi"" or ""grubx64.efi""; if 00:00 (BIOS), serve ""pxelinux.0"". Modern servers default to UEFI, so a PXE setup that only serves pxelinux.0 will silently fail for UEFI clients -- they will skip PXE and boot from disk instead.\n\nGotcha: firmware updates often require a reboot and can occasionally brick a system. Always have an out-of-band management path and a rollback plan before flashing firmware.\n\nGotcha: always verify network changes with a second connection (iDRAC/console) before applying. A bad IP change on the only interface locks you out remotely."	training/library/topics/bare-metal-provisioning/primer.md
datacenter/f6e7d8c9	datacenter	medium	kickstart,automation,install	What is Kickstart and what are its key sections for automated OS installation?	Kickstart is Red Hat/CentOS's automated installer. Key sections include: install source (url), language/keyboard/timezone, rootpw, network config, disk layout (zerombr, clearpart, autopart), %packages (package selection), and %post (post-install scripts for bootstrap, SSH keys, config management registration). A fully configured Kickstart enables unattended installation.\n\nGotcha: always verify network changes with a second connection (iDRAC/console) before applying. A bad IP change on the only interface locks you out remotely.	training/library/topics/bare-metal-provisioning/primer.md
datacenter/a7f8e9da	datacenter	medium	provisioning,workflow,pipeline	What are the key infrastructure components in a bare-metal provisioning architecture?	A provisioning architecture requires: an isolated provisioning VLAN for PXE traffic, a DHCP+TFTP server (often dnsmasq) for IP assignment and bootloader delivery, an HTTP server (nginx) hosting OS images and kickstart files, a configuration management system (Ansible) for post-install config, and a CMDB/inventory system to track provisioned assets.\n\nGotcha: always verify network changes with a second connection (iDRAC/console) before applying. A bad IP change on the only interface locks you out remotely.	training/library/topics/bare-metal-provisioning/primer.md
datacenter/b8a9fade	datacenter	easy	provisioning,lifecycle,stages	What are the stages of the bare-metal provisioning lifecycle?	The lifecycle stages are: Rack and Cable, OOB Setup (BMC/IPMI config), BIOS Configuration, PXE Boot, OS Install, Post-Install validation, Production deployment, and eventually Decommission/Reprovision. The goal is zero-touch provisioning where a newly racked server bootstraps itself to production-ready state automatically.\n\nGotcha: always verify network changes with a second connection (iDRAC/console) before applying. A bad IP change on the only interface locks you out remotely.	training/library/topics/bare-metal-provisioning/primer.md
datacenter/c9bafaef	datacenter	hard	provisioning,validation,post-install	What should a post-install validation script check before a server enters production?	Validate hardware (CPU count, RAM size, disk presence), OS services (sshd, chronyd running, NTP synchronized), network connectivity (gateway ping, DNS resolution, provisioning server reachable), and security (SELinux enforcing, no leftover private keys). Exit non-zero on any failure so automation can catch provisioning errors before workload placement.\n\nGotcha: always verify network changes with a second connection (iDRAC/console) before applying. A bad IP change on the only interface locks you out remotely.	training/library/topics/bare-metal-provisioning/primer.md
datacenter/dabafba0	datacenter	hard	pxe,console,serial	Why must you configure serial console parameters when PXE-booting headless servers?	Headless servers have no video output. Without serial console parameters (console=tty0 console=ttyS1,115200n8) in the kernel command line, the installer runs on an invisible display and may hang waiting for interactive input. You must also ensure the Kickstart is fully unattended and configure GRUB and systemd getty for serial output.\n\nGotcha: in datacenter operations, always verify changes with out-of-band access (iDRAC, iLO, serial console) before and after applying them.	training/library/topics/bare-metal-provisioning/street_ops.md
datacenter/ebcbacb1	datacenter	medium	pxe,ipmi,bootdev	How do you set a server to PXE boot on next reboot using IPMI and Redfish?	"Via IPMI: ""ipmitool -I lanplus -H <bmc> -U admin -P pass chassis bootdev pxe"" sets one-time PXE boot. Via Redfish: PATCH /redfish/v1/Systems/<id> with {""Boot"": {""BootSourceOverrideTarget"": ""Pxe"", ""BootSourceOverrideEnabled"": ""Once""}}. Both set a one-time override that reverts to normal boot order after the next reboot.\n\nRemember: out-of-band management (iDRAC/iLO/IPMI) works even when the OS is down. It has its own network interface, IP address, and web UI — like a KVM switch built into the server.\n\nRemember: Redfish = modern replacement for IPMI. RESTful JSON over HTTPS vs. IPMI's binary protocol. Scriptable with curl: curl -k -u admin:pass https://bmc-ip/redfish/v1/Systems/1."	training/library/topics/bare-metal-provisioning/primer.md
datacenter/fcdbade2	datacenter	medium	pxe,per-host,template	How do you generate per-host Kickstart files for a fleet of servers?	Create a Kickstart template with placeholders (%%HOSTNAME%%, %%IP%%, %%GATEWAY%%) and a CSV inventory of hostnames, MACs, IPs, and gateways. A script iterates over the inventory, substituting values with sed, and writes each Kickstart file named by MAC address to the HTTP server. The iPXE script then requests the Kickstart URL using the client's MAC.\n\nGotcha: always verify network changes with a second connection (iDRAC/console) before applying. A bad IP change on the only interface locks you out remotely.	training/library/topics/bare-metal-provisioning/primer.md
datacenter/a1b2c3d4e5f6	datacenter	easy	pxe, dhcp, boot	What two DHCP options are critical for PXE boot, and what does each specify?	Option 66 (next-server) specifies the TFTP server IP, and option 67 (bootfile name) specifies the path to the bootloader file (e.g., pxelinux.0 or ipxe.efi).\n\nGotcha: always verify network changes with a second connection (iDRAC/console) before applying. A bad IP change on the only interface locks you out remotely.	training/library/topics/bare-metal-provisioning/primer.md
datacenter/b2c3d4e5f6a7	datacenter	easy	pxe, provisioning, lifecycle	What is the goal of zero-touch provisioning?	Rack a server, plug in power and network, and it bootstraps itself to a production-ready state without any manual intervention.\n\nGotcha: in datacenter operations, always verify changes with out-of-band access (iDRAC, iLO, serial console) before and after applying them.	training/library/topics/bare-metal-provisioning/primer.md
datacenter/c3d4e5f6a7b8	datacenter	easy	pxe, boot, uefi	How does the DHCP server differentiate between UEFI and Legacy BIOS PXE clients?	It checks the architecture option (option arch). If the value is 00:07, the client is UEFI and receives an EFI bootloader (e.g., ipxe/snponly.efi). Otherwise it receives a legacy bootloader (e.g., pxelinux.0).\n\nGotcha: firmware updates often require a reboot and can occasionally brick a system. Always have an out-of-band management path and a rollback plan before flashing firmware.\n\nGotcha: always verify network changes with a second connection (iDRAC/console) before applying. A bad IP change on the only interface locks you out remotely.	training/library/topics/bare-metal-provisioning/primer.md
datacenter/d4e5f6a7b8c9	datacenter	medium	pxe, ipxe, chainloading	Why do modern PXE setups chainload from PXE to iPXE, and what protocol does iPXE use instead of TFTP?	iPXE supports HTTP-based boot, which is significantly faster than TFTP. Chainloading lets the initial PXE ROM hand off to iPXE so that the kernel and initrd can be downloaded over HTTP.\n\nRemember: UPS = bridge power during utility failure until generators start (10-30 seconds). Battery runtime: typically 5-15 minutes. Not for extended outages — that's the generator's job.	training/library/topics/bare-metal-provisioning/primer.md
datacenter/e5f6a7b8c9d0	datacenter	medium	pxe, kickstart, automation	What does the %post section of a Kickstart file do, and why is it important for provisioning?	The %post section runs shell commands after the OS is installed — typically registering with config management, deploying SSH keys, and phoning home to the provisioning server to signal completion. It bridges the gap between bare OS install and production readiness.\n\nGotcha: in datacenter operations, always verify changes with out-of-band access (iDRAC, iLO, serial console) before and after applying them.	training/library/topics/bare-metal-provisioning/primer.md
datacenter/f6a7b8c9d0e1	datacenter	medium	pxe, ipmi, bmc	What ipmitool command forces a server to PXE boot on its next power cycle?	ipmitool -I lanplus -H <BMC_IP> -U admin -P secret chassis bootdev pxe\n\nRemember: out-of-band management (iDRAC/iLO/IPMI) works even when the OS is down. It has its own network interface, IP address, and web UI — like a KVM switch built into the server.\n\nGotcha: always calculate power at full load, not idle. A server may idle at 200W but spike to 700W under CPU stress. Plan PDU capacity for peak, not average.\n\nGotcha: always verify network changes with a second connection (iDRAC/console) before applying. A bad IP change on the only interface locks you out remotely.	training/library/topics/bare-metal-provisioning/primer.md
datacenter/a7b8c9d0e1f2	datacenter	medium	pxe, redfish, api	How does the Redfish API set a one-time PXE boot on a server?	"By sending a PATCH request to /redfish/v1/Systems/1 with the body {""Boot"": {""BootSourceOverrideTarget"": ""Pxe"", ""BootSourceOverrideEnabled"": ""Once""}}. Redfish is the modern REST-based successor to IPMI.\n\nRemember: Redfish = modern replacement for IPMI. RESTful JSON over HTTPS vs. IPMI's binary protocol. Scriptable with curl: curl -k -u admin:pass https://bmc-ip/redfish/v1/Systems/1.\n\nGotcha: always verify network changes with a second connection (iDRAC/console) before applying. A bad IP change on the only interface locks you out remotely."	training/library/topics/bare-metal-provisioning/primer.md
datacenter/b8c9d0e1f2a3	datacenter	hard	pxe, onie, switches	What is ONIE and how does it differ from standard PXE?	ONIE (Open Network Install Environment) is the PXE equivalent for network switches (Cumulus Linux, SONiC). Unlike server PXE which boots an OS installer via TFTP/HTTP, ONIE discovers a network OS installer via DHCP options, HTTP discovery, USB, or TFTP fallback, then installs the NOS and reboots.\n\nGotcha: in datacenter operations, always verify changes with out-of-band access (iDRAC, iLO, serial console) before and after applying them.	training/library/topics/bare-metal-provisioning/primer.md
datacenter/c9d0e1f2a3b4	datacenter	hard	pxe, provisioning, architecture	What are the key components of a provisioning network architecture, and why is the provisioning network isolated?	Key components: DHCP+TFTP server (e.g., dnsmasq), HTTP server (nginx) for OS images and kickstart files, config management (Ansible) for post-install, and a CMDB for inventory tracking. The provisioning network is isolated on a separate VLAN to prevent PXE traffic from interfering with production and to secure the out-of-band management plane.\n\nGotcha: always verify network changes with a second connection (iDRAC/console) before applying. A bad IP change on the only interface locks you out remotely.	training/library/topics/bare-metal-provisioning/primer.md
datacenter/d0e1f2a3b4c5	datacenter	hard	pxe, validation, post-install	What types of checks should a post-install validation script perform before marking a server as production-ready?	Hardware checks (CPU count, RAM, disk presence), OS checks (SSH and NTP services running, time synchronization), network checks (gateway reachable, provisioning server accessible, DNS resolution), and security checks (SELinux enforcing, no stray private keys). The script should count failures and exit with the error count.\n\nGotcha: always verify network changes with a second connection (iDRAC/console) before applying. A bad IP change on the only interface locks you out remotely.	training/library/topics/bare-metal-provisioning/primer.md
datacenter/a1d2e3f4	datacenter	easy	power,pdu,types	What are the four main types of PDUs and what does each add over a basic PDU?	Basic PDU: simple power distribution, no monitoring. Metered PDU: shows per-outlet or per-phase power draw. Switched PDU: adds remote on/off control per outlet. ATS (Automatic Transfer Switch): automatically fails over between two input power feeds if one goes down.\n\nRemember: PDU = Power Distribution Unit. Rack PDU distributes power to servers. Monitored PDUs show per-outlet power draw. Redundant PDUs (A+B feeds) prevent single-feed failure.	training/library/topics/datacenter/primer.md
datacenter/b2e3f4a5	datacenter	medium	power,redundancy,feeds	What is A+B feed redundancy and why is it critical in datacenter racks?	A+B feed means each rack receives power from two independent circuits connected to separate UPS systems and ideally separate utility feeds. Servers with dual PSUs connect one to each feed. If one feed fails (UPS failure, breaker trip, maintenance), the other feed keeps all equipment running with zero downtime.\n\nGotcha: always verify network changes with a second connection (iDRAC/console) before applying. A bad IP change on the only interface locks you out remotely.	training/library/topics/datacenter/primer.md
datacenter/c3f4a5b6	datacenter	medium	power,ups,sizing	What is UPS and what happens when battery replacement is missed?	A UPS (Uninterruptible Power Supply) provides battery backup during utility power outages, bridging the gap until generators start (typically 10-30 seconds). UPS batteries degrade over time and must be replaced on schedule (usually every 3-5 years). Missing replacement means the UPS may not hold load during an outage, causing an unplanned shutdown.\n\nRemember: UPS = bridge power during utility failure until generators start (10-30 seconds). Battery runtime: typically 5-15 minutes. Not for extended outages — that's the generator's job.	training/library/topics/datacenter/primer.md
datacenter/d4a5b6c7	datacenter	hard	power,budget,rack	How do you calculate a rack power budget and what is a typical limit?	Sum the rated wattage of all equipment in the rack (servers, switches, storage) and compare against the PDU capacity (commonly 5-10 kW per feed). Account for peak vs average draw -- servers under full CPU load draw significantly more than idle. Maintain headroom (typically 80% of circuit capacity) to avoid breaker trips. Monitor actual draw with metered PDUs.\n\nGotcha: always calculate power at full load, not idle. A server may idle at 200W but spike to 700W under CPU stress. Plan PDU capacity for peak, not average.\n\nGotcha: always verify network changes with a second connection (iDRAC/console) before applying. A bad IP change on the only interface locks you out remotely.	training/library/topics/datacenter/primer.md
datacenter/e5b6c7d8	datacenter	easy	cooling,hot-cold-aisle,containment	What is hot aisle / cold aisle containment and why does it matter?	Servers intake cool air from the front (cold aisle) and exhaust hot air from the rear (hot aisle). Containment uses physical barriers (doors, curtains, panels) to prevent hot exhaust from recirculating to server intakes. Without containment, mixing reduces cooling efficiency, raises inlet temperatures, and can cause thermal throttling or shutdowns.\n\nRemember: cold aisle faces rack fronts (air intake). Hot aisle faces rack backs (exhaust). Containment prevents mixing. This is the #1 cooling efficiency optimization.	training/library/topics/datacenter/primer.md
datacenter/f6c7d8e9	datacenter	medium	cooling,crac,crah	What is the difference between a CRAC and a CRAH unit?	A CRAC (Computer Room Air Conditioner) uses a compressor-based refrigeration cycle to cool air directly. A CRAH (Computer Room Air Handler) uses chilled water from a central plant and fan coils to cool air. CRAHs are more energy-efficient at scale and allow centralized chiller management, making them preferred in large datacenters.\n\nRemember: comparison questions are best answered with a structured format: name the key dimensions (use case, performance, complexity, cost) and compare each.	training/library/topics/datacenter/primer.md
datacenter/a7d8e9fa	datacenter	easy	cooling,blanking-panels,airflow	What are blanking panels and why should empty rack U-spaces never be left open?	Blanking panels are covers installed in empty rack unit spaces. Without them, hot exhaust air from the rear recirculates through the open gaps to the cold aisle, bypassing servers and creating hot spots. This reduces cooling efficiency and can cause adjacent servers to overheat even when overall room temperature is within spec.	training/library/topics/datacenter/primer.md
datacenter/b8e9faab	datacenter	medium	monitoring,environmental,ashrae	What is the ASHRAE A1 recommended temperature range for datacenter inlet air?	ASHRAE A1 recommends 18-27 degrees Celsius (64-80 degrees Fahrenheit) for server inlet temperature. Operating outside this range increases hardware failure rates and can void warranties. Monitor inlet temps with BMC sensors (ipmitool sdr type Temperature) and environmental monitoring systems. Humidity should be 40-60% relative humidity.\n\nRemember: ASHRAE recommended inlet temperature: 18-27C (64-80F). Every 10C above optimal roughly halves component lifespan.\n\nGotcha: always verify network changes with a second connection (iDRAC/console) before applying. A bad IP change on the only interface locks you out remotely.	training/library/topics/datacenter/primer.md
datacenter/c9faabbc	datacenter	hard	power,n-plus-1,redundancy	What does N+1 redundancy mean for power and cooling, and how does it differ from 2N?	N+1 means you have one additional unit beyond the minimum needed (e.g., 3 cooling units when 2 would suffice). 2N means fully duplicated capacity (e.g., two independent UPS systems each able to handle the full load). 2N is more resilient but costs twice as much. N+1 is the minimum for production datacenters; 2N is standard for Tier III/IV facilities.\n\nGotcha: always calculate power at full load, not idle. A server may idle at 200W but spike to 700W under CPU stress. Plan PDU capacity for peak, not average.\n\nRemember: ASHRAE recommended inlet temperature: 18-27C (64-80F). Every 10C above optimal roughly halves component lifespan.	training/library/topics/datacenter/primer.md
datacenter/daabbbcd	datacenter	hard	power,whip,outlet	What is the difference between a power whip and a standard outlet in datacenter power distribution?	A power whip is a hard-wired, high-amperage cable connection from the overhead or underfloor busway directly to the PDU, typically used for high-density racks (30A-60A circuits). Standard outlets (C13/C14, C19/C20) are plug-based connections for individual devices. Whips provide higher capacity and more reliable connections for heavy loads.\n\nRemember: PDU = Power Distribution Unit. Rack PDU distributes power to servers. Monitored PDUs show per-outlet power draw. Redundant PDUs (A+B feeds) prevent single-feed failure.	training/library/topics/datacenter/primer.md
datacenter/1c5aed88bbc5	datacenter	easy	power, ups, types	What are the three types of UPS and which is standard for datacenters?	Offline/Standby (5-12ms transfer time), Line-Interactive (2-4ms), and Online/Double-Conversion (0ms, always on inverter). Double-conversion is the datacenter standard because it provides zero transfer time and complete isolation from utility power quality issues.\n\nRemember: UPS = bridge power during utility failure until generators start (10-30 seconds). Battery runtime: typically 5-15 minutes. Not for extended outages — that's the generator's job.	training/library/topics/power/primer.md
datacenter/f9b8cb337b60	datacenter	easy	power, psu, redundancy	What is 1+1 PSU redundancy?	1+1 means two power supplies in a server — one active, one standby. If the active PSU fails, the standby takes over immediately. Each PSU should be connected to a different PDU/power feed for full redundancy.\n\nGotcha: always calculate power at full load, not idle. A server may idle at 200W but spike to 700W under CPU stress. Plan PDU capacity for peak, not average.	training/library/topics/power/primer.md
datacenter/d839d519b968	datacenter	easy	power, pdu, distribution	What is a PDU and what types exist?	A PDU (Power Distribution Unit) distributes power from UPS to server racks. Types: Basic (power strip), Metered (shows power draw), Monitored (network-connected, per-outlet monitoring), Switched (remote power cycling per outlet), Intelligent (monitoring + switching + environmental sensors).\n\nRemember: PDU = Power Distribution Unit. Rack PDU distributes power to servers. Monitored PDUs show per-outlet power draw. Redundant PDUs (A+B feeds) prevent single-feed failure.	training/library/topics/power/primer.md
datacenter/b2e91f5042f7	datacenter	medium	power, dual-feed, rack	How is dual-feed power redundancy implemented at the rack level?	Each rack has two PDUs: PDU A from Power Feed A (UPS A) and PDU B from Power Feed B (UPS B). Each server connects one PSU to PDU A and one to PDU B. Either PDU can handle the full rack load if the other fails. This eliminates single points of failure from UPS to server.\n\nGotcha: always calculate power at full load, not idle. A server may idle at 200W but spike to 700W under CPU stress. Plan PDU capacity for peak, not average.	training/library/topics/power/primer.md
datacenter/867cf8c6dfa1	datacenter	medium	power, monitoring, nut	How do you monitor UPS status on Linux?	Using NUT (Network UPS Tools): upsc myups (full status), upsc myups ups.status (OL=online, OB=on battery, LB=low battery), upsc myups battery.charge (percentage), upsc myups battery.runtime (seconds). For APC: apcaccess status. Server power draw: ipmitool dcmi power reading.\n\nRemember: UPS = bridge power during utility failure until generators start (10-30 seconds). Battery runtime: typically 5-15 minutes. Not for extended outages — that's the generator's job.	training/library/topics/power/primer.md
datacenter/67dbdd73ee47	datacenter	medium	power, pue, efficiency	What is PUE and what values are considered good?	PUE (Power Usage Effectiveness) = Total Facility Power / IT Equipment Power. PUE 1.0 is perfect (impossible), 1.2 is excellent, 1.5 is average, 2.0 is poor (half the power goes to cooling/overhead). It measures datacenter power efficiency.\n\nRemember: PUE = Total Facility Power / IT Equipment Power. PUE 1.0 = perfect (impossible). PUE 1.2 = excellent. PUE 2.0 = half your power goes to cooling/overhead.	training/library/topics/power/primer.md
datacenter/eef71d863af2	datacenter	medium	power, budgeting, rack	How do you calculate power requirements for a rack?	Sum the average power draw of all equipment. Example: 20 servers x 500W = 10kW, plus 2 switches x 150W = 300W. Total ~10.3kW. With 1+1 PSU redundancy, each PDU must handle the full 10.3kW. Typical rack capacity is 5-10kW (standard) up to 30+kW (high density).\n\nGotcha: always calculate power at full load, not idle. A server may idle at 200W but spike to 700W under CPU stress. Plan PDU capacity for peak, not average.\n\nGotcha: always verify network changes with a second connection (iDRAC/console) before applying. A bad IP change on the only interface locks you out remotely.	training/library/topics/power/primer.md
datacenter/f6b64d694263	datacenter	hard	power, shutdown, graceful	What is the correct shutdown order during a power emergency?	1. Applications (drain connections, flush buffers). 2. Virtual machines. 3. Hypervisors/bare-metal OS. 4. Storage arrays (after all servers are down). 5. Network switches (last — needed for management). Configure NUT/apcupsd for automatic shutdown when battery is low.\n\nGotcha: always calculate power at full load, not idle. A server may idle at 200W but spike to 700W under CPU stress. Plan PDU capacity for peak, not average.	training/library/topics/power/primer.md
datacenter/6984a041ae4a	datacenter	hard	power, alerting, thresholds	What power-related alert thresholds should be configured for datacenter monitoring?	UPS on battery: immediate page. Battery < 50%: warning. Battery < 20%: critical, initiate shutdown. PDU circuit > 80% capacity: warning (prevent overload trips). Inlet temperature > 35C: warning. Also monitor UPS input/output voltage, frequency, and server power consumption (watts).\n\nRemember: good alerting follows the RED method for services (Rate, Errors, Duration) and USE method for resources (Utilization, Saturation, Errors).	training/library/topics/power/primer.md
datacenter/4b0b464a4cec	datacenter	hard	power, generator, ats	How does the power failover chain work from utility loss to generator?	Utility power fails. ATS (Automatic Transfer Switch) detects loss. Diesel generator starts (10-30 second startup). UPS batteries bridge the gap during transfer. ATS switches to generator feed once stable. Generator runs until utility is restored. UPS capacity must exceed generator startup time.\n\nRemember: diesel generators provide long-term backup power. UPS bridges the 10-30 second startup gap. Tier 3+ datacenters have N+1 generator redundancy.	training/library/topics/power/primer.md

<!-- wiki:related:start -->
---

## Wiki Navigation

### Related Content

- [Bare-Metal Provisioning](../../../../library/topics/bare-metal-provisioning/index.md) (Topic Pack, L2) — Out-of-Band Management
- [Case Study: BMC Clock Skew Cert Failure](../../../../library/case-studies/datacenter_ops/bmc-clock-skew-cert-failure/README.md) (Case Study, L2) — Out-of-Band Management
- [Case Study: Serial Console Garbled](../../../../library/case-studies/datacenter_ops/serial-console-garbled/README.md) (Case Study, L1) — Out-of-Band Management
- [Case Study: Server Remote Console Lag](../../../../library/case-studies/datacenter_ops/server-remote-console-lag/README.md) (Case Study, L1) — Out-of-Band Management
- [Case Study: iDRAC Unreachable OS Up](../../../../library/case-studies/datacenter_ops/idrac-unreachable-os-up/README.md) (Case Study, L1) — Out-of-Band Management
- [Datacenter & Server Hardware](../../../../library/topics/datacenter/index.md) (Topic Pack, L1) — Out-of-Band Management
- [Datacenter Drills](../../../../library/drills/datacenter_drills.md) (Drill, L1) — Out-of-Band Management
- [Deep Dive: Dell Linux PowerEdge](../../../../library/deep-dives/dell.linux.poweredge.deep.dive.md) (deep_dive, L2) — Out-of-Band Management
- [Dell PowerEdge Servers](../../../../library/topics/dell-poweredge/index.md) (Topic Pack, L1) — Out-of-Band Management
- [IPMI and ipmitool](../../../../library/topics/ipmi-and-ipmitool/index.md) (Topic Pack, L1) — Out-of-Band Management

<!-- wiki:related:end -->
