---
tags:
- datacenter
- l1
- flashcard-deck
- server-hardware
---
<!-- wiki:breadcrumb:start -->
[Portal](../../../../library/portal/index.md) | **Level:** [L1: Foundations](../../../../library/portal/levels.md) | **Topics:** [Server Hardware](../../../../library/portal/topics.md) | **Domain:** Datacenter & Hardware
<!-- wiki:breadcrumb:end -->

id	category	difficulty	tags	question	answer	source_path
server-hardware/07f8f4b43700	server-hardware	easy	hardware, diagnostics, lshw	What command gives a quick hardware inventory on Linux?	"lshw -short. It lists all hardware classes (CPU, memory, network, disk, etc.) in a concise table. Filter by class: lshw -class memory, lshw -class network. Output as JSON: lshw -json -short.\n\nRemember: ""RAID protects against disk failure, not data corruption."" RAID is not a backup.\n\nExample: RAID 1 = mirror (2 disks), RAID 5 = striping with parity (min 3 disks), RAID 10 = mirror + stripe (min 4 disks).\n\nGotcha: lshw requires root for full output. Without root, it hides some details. Use `sudo lshw -short` for the complete picture."	training/library/topics/server-hardware/primer.md
server-hardware/9b6eb2fee8e3	server-hardware	easy	hardware, dmidecode, system-info	How do you read hardware information from BIOS/UEFI tables on Linux?	"dmidecode reads SMBIOS/DMI data. Key types: dmidecode -t system (manufacturer, model, serial), dmidecode -t memory (DIMM details), dmidecode -t processor (CPU), dmidecode -t bios (BIOS version). Quick serial: dmidecode -s system-serial-number.\n\nRemember: ""ECC RAM detects and corrects single-bit errors."" Servers use ECC; desktops usually don't.\n\nFun fact: Google research found 8% of DIMMs have at least one error per year.\n\nName origin: DMI = Desktop Management Interface. SMBIOS = System Management BIOS. Both are BIOS/UEFI tables describing hardware to the OS."	training/library/topics/server-hardware/primer.md
server-hardware/d46d31ed48d3	server-hardware	easy	hardware, dmesg, kernel	How do you check for hardware errors in the kernel message buffer?	"dmesg -T -l err,crit,alert,emerg shows only error-level and above messages with human-readable timestamps. Filter for hardware: dmesg | grep -i ""error\|fault\|fail\|hardware\|mce"". Watch live: dmesg -Tw.\n\nRemember: ""IPMI/BMC = out-of-band management."" Access the server even when the OS is down. iDRAC (Dell), iLO (HP), IMM (Lenovo).\n\nGotcha: BMC has its own IP and web interface — secure it! Default passwords are a common attack vector.\n\nDebug clue: MCE (Machine Check Exception) in dmesg = CPU/memory hardware error. I/O errors = disk/controller. Link down = NIC/cable."	training/library/topics/server-hardware/primer.md
server-hardware/f7eed6e2f424	server-hardware	medium	hardware, ecc, memory	What is ECC memory and how do you check for memory errors on Linux?	"ECC (Error-Correcting Code) memory detects and corrects single-bit errors (correctable, CE) and detects double-bit errors (uncorrectable, UE — usually causes kernel panic). Check with: edac-util -s (summary), edac-util -l (per-DIMM errors). Increasing CE counts on a DIMM indicate it is failing.\n\nRemember: ""Hot-swap = replace without downtime."" Drives, power supplies, and fans are typically hot-swappable in servers. CPUs and RAM are not.\n\nNumber anchor: Google research found 8% of DIMMs experience at least one correctable error per year. ECC catches these silently."	training/library/topics/server-hardware/primer.md
server-hardware/2bf0d992aa6e	server-hardware	medium	hardware, smart, disks	What SMART attributes indicate a failing disk?	"Reallocated Sector Count (bad sectors remapped — increasing means drive is dying), Current Pending Sector Count (sectors waiting to be remapped), and Uncorrectable Error Count (read errors that could not be recovered). Check with: smartctl -a /dev/sda. Run self-test: smartctl -t short /dev/sda.\n\nRemember: ""Reallocated + Pending + Uncorrectable = the three horsemen."" Any non-zero value warrants investigation."	training/library/topics/server-hardware/primer.md
server-hardware/346858cf2380	server-hardware	medium	hardware, nic, ethtool	How do you diagnose NIC hardware problems?	"ethtool eth0 (link status, speed, duplex), ethtool -i eth0 (driver, firmware), ethtool -S eth0 | grep -E ""error|drop|miss|crc"" (error counters). Increasing CRC errors indicate cable or hardware issues. Link flapping and rx/tx drops suggest a failing NIC.\n\nDebug clue: Increasing CRC errors = bad cable or connector. rx_missed_errors = NIC ring buffer overflow (increase with ethtool -G)."	training/library/topics/server-hardware/primer.md
server-hardware/6525cde9717b	server-hardware	medium	hardware, numa, cpu	What is NUMA and why does it matter for server performance?	NUMA (Non-Uniform Memory Access) means each CPU socket has local memory that is faster to access than remote memory (other socket's memory). Applications should be pinned to use local memory. Check topology: numactl --hardware or lscpu | grep NUMA. Misaligned NUMA access causes latency.\n\nName origin: NUMA = Non-Uniform Memory Access. The alternative is UMA (Uniform Memory Access), which doesn\'t scale beyond ~4 sockets.	training/library/topics/server-hardware/primer.md
server-hardware/a18aca11b0a6	server-hardware	hard	hardware, mce, kernel-panic	What is an MCE and how do you investigate one?	"MCE (Machine Check Exception) is a CPU-reported hardware error. Uncorrectable MCEs cause kernel panics. Investigate: journalctl | grep -i ""mce\|machine check"", mcelog --client (if mcelog daemon runs). Common causes: failing DIMMs, CPU cache errors, overheating. Check edac-util for memory-related MCEs.\n\nName origin: MCE = Machine Check Exception. CPUs report internal hardware errors via this mechanism. Uncorrectable MCEs crash the system."	training/library/topics/server-hardware/primer.md
server-hardware/808e99c8c3b4	server-hardware	hard	hardware, thermal, throttling	What causes CPU thermal throttling and how do you detect it?	Throttling occurs when CPU exceeds thermal limits: failed fans, blocked airflow, ambient temperature too high (CRAC failure), dust buildup. Detect: sensors (lm-sensors), ipmitool sensor list | grep -i temp, cpupower frequency-info (check if frequency is reduced). Check dmesg for thermal throttling messages.\n\nNumber anchor: Most server CPUs throttle at 95-100°C. Each 10°C increase above design temp roughly doubles component failure rate.	training/library/topics/server-hardware/primer.md
server-hardware/69934ebc8b65	server-hardware	hard	hardware, diagnostics, workflow	What is the recommended workflow for diagnosing a suspected hardware issue?	"1. Check dmesg -T -l err,crit (MCE, I/O errors, link down). 2. Check ipmitool sel elist (BMC events). 3. Check smartctl -a /dev/sdX (disk health). 4. Check edac-util -s (memory errors). 5. Check ethtool -S eth0 (NIC errors). 6. Check ipmitool sensor list (thermal, voltage, fan). 7. Check vendor BMC UI (iDRAC/iLO) for detailed diagnostics.\n\nRemember: ""Outside-in: BMC logs → kernel logs → device-specific tools."" Start with the broadest view (BMC) and narrow down."	training/library/topics/server-hardware/primer.md

<!-- wiki:related:start -->
---

## Wiki Navigation

### Related Content

- [Bare-Metal Provisioning](../../../../library/topics/bare-metal-provisioning/index.md) (Topic Pack, L2) — Server Hardware
- [Case Study: BIOS Settings Reset After CMOS](../../../../library/case-studies/datacenter_ops/bios-settings-reset-after-cmos/README.md) (Case Study, L1) — Server Hardware
- [Case Study: Cable Management Wrong Port](../../../../library/case-studies/datacenter_ops/cable-management-wrong-port/README.md) (Case Study, L1) — Server Hardware
- [Case Study: Database Replication Lag — Root Cause Is RAID Degradation](../../../../library/case-studies/cross-domain/database-replication-lag-raid/README.md) (Case Study, L2) — Server Hardware
- [Case Study: Firmware Update Boot Loop](../../../../library/case-studies/datacenter_ops/firmware-update-boot-loop/README.md) (Case Study, L2) — Server Hardware
- [Case Study: Link Flaps Bad Optic](../../../../library/case-studies/datacenter_ops/link-flaps-bad-optic/README.md) (Case Study, L1) — Server Hardware
- [Case Study: Memory ECC Errors Increasing](../../../../library/case-studies/datacenter_ops/memory-ecc-errors-increasing/README.md) (Case Study, L1) — Server Hardware
- [Case Study: Power Supply Redundancy Lost](../../../../library/case-studies/datacenter_ops/power-supply-redundancy-lost/README.md) (Case Study, L1) — Server Hardware
- [Case Study: Serial Console Garbled](../../../../library/case-studies/datacenter_ops/serial-console-garbled/README.md) (Case Study, L1) — Server Hardware
- [Case Study: Server Intermittent Reboot](../../../../library/case-studies/datacenter_ops/server-intermittent-reboot/README.md) (Case Study, L2) — Server Hardware

<!-- wiki:related:end -->
