---
tags:
- devops
- l1
- flashcard-deck
- debugging-methodology
---
<!-- wiki:breadcrumb:start -->
[Portal](../../../../library/portal/index.md) | **Level:** [L1: Foundations](../../../../library/portal/levels.md) | **Topics:** [Debugging Methodology](../../../../library/portal/topics.md) | **Domain:** DevOps & Tooling
<!-- wiki:breadcrumb:end -->

id	category	difficulty	tags	question	answer	source_path
debugging-methodology/a1b2c3d4e5f4	debugging-methodology	easy	debugging-methodology, scientific-method	What are the five steps of the scientific method applied to debugging?	1) Observe: what is actually happening (symptoms, not assumptions). 2) Hypothesize: what could cause this. 3) Predict: if hypothesis X is true, what else should be true. 4) Test: check the prediction, change one variable. 5) Conclude: confirmed or eliminated, then repeat.\n\nRemember: mnemonic OH-PTC — Observe, Hypothesize, Predict, Test, Conclude. The same method finds bugs in DNA and bugs in code.	training/library/topics/debugging-methodology/primer.md
debugging-methodology/b2c3d4e5f4a5	debugging-methodology	easy	debugging-methodology, divide-and-conquer	What is divide-and-conquer debugging and why is it more efficient than linear search?	Divide-and-conquer bisects the system at the midpoint and tests there, cutting the problem space in half with each test. For a 10-component pipeline, you need at most 4 tests (log2 of 10) instead of 10. It is binary search applied to infrastructure.\n\nExample: 10-stage pipeline broken? Test stage 5. Works? Bug is in 6-10. Half the problem eliminated in one test.\n\nRemember: binary search applied to systems — log2(N) tests instead of N sequential checks.	training/library/topics/debugging-methodology/primer.md
debugging-methodology/c3d4e5f4a5b6	debugging-methodology	easy	debugging-methodology, five-whys	What is the Five Whys technique and why does the first answer rarely give the root cause?	The Five Whys is a root cause analysis technique where you keep asking "why" until you reach the systemic cause. The first "why" typically gives the symptom fix (e.g., kill the slow query). The fifth "why" gives the systemic fix (e.g., add automated index validation in migration pipeline) that prevents recurrence.\n\nExample: Why down? Bad deploy. Why? Wrong config. Why? No validation. Why? No CI gate. Why? Never built one. Fix: add config validation to CI.	training/library/topics/debugging-methodology/primer.md
debugging-methodology/d4e5f4a5b6c7	debugging-methodology	medium	debugging-methodology, hypothesis	What five questions should you ask to generate debugging hypotheses?	1) What changed recently? (deploys, config, infra, traffic). 2) What is different about the failing cases? (users, regions, endpoints, time). 3) What resources could be exhausted? (CPU, memory, disk, FDs, connections). 4) What dependencies could be failing? (DBs, caches, APIs, DNS, certs). 5) What has failed like this before? (incident history, postmortems).\n\nRemember: WIRED — What changed, Is it intermittent, Resources exhausted, External deps failing, Done this before? Five hypothesis generators.	training/library/topics/debugging-methodology/primer.md
debugging-methodology/e5f4a5b6c7d8	debugging-methodology	medium	debugging-methodology, traps	What is "shotgun debugging" and why is changing one variable at a time critical?	Shotgun debugging is changing multiple things at once hoping one helps. The problem is that if the issue resolves, you do not know which change actually fixed it, so you cannot prevent recurrence. Always change one variable at a time so you know which change had the effect.\n\nGotcha: changing 3 things and seeing a fix means 3 possible causes and zero understanding. Always change one variable at a time.	training/library/topics/debugging-methodology/primer.md
debugging-methodology/f4a5b6c7d8e9	debugging-methodology	medium	debugging-methodology, correlation-causation	How do you distinguish correlation from causation when a deployment and a failure happen near the same time?	Three tests: 1) Revert the suspected change — if the problem goes away, strong evidence of causation. 2) Reproduce in isolation — can you trigger the failure by making only that change? 3) Explain the mechanism — can you trace from the change to the symptom step by step? Correlation is temporal proximity; causation requires a verifiable mechanism.\n\nExample: CPU spikes at 3 PM when the cron runs, but the real cause is the database backup that also starts at 3 PM. Temporal proximity is not causation.	training/library/topics/debugging-methodology/primer.md
debugging-methodology/05a6b7c8d9e0	debugging-methodology	medium	debugging-methodology, layer-model	How does the network layer model help narrow down connectivity problems?	Test from L7 down: curl for application (L7), telnet/nc for transport (L4), ping for network (L3). If L3 works (ping succeeds) but L4 fails (cannot connect to port), the problem is narrowed to: firewall, security group, service not listening, or wrong port. Each layer test eliminates a class of causes.\n\nRemember: test top-down — L7 curl, L4 nc/telnet, L3 ping. Each success eliminates a class of problems.	training/library/topics/debugging-methodology/primer.md
debugging-methodology/16b7c8d9e0f1	debugging-methodology	hard	debugging-methodology, git-bisect	How does git bisect use binary search to find a breaking commit, and what is its worst-case efficiency?	git bisect start, mark HEAD as bad and a known-good commit as good. Git checks out the midpoint; you test and mark it good or bad. Repeat until the exact breaking commit is found. For N commits, worst case is log2(N) tests — 25 commits need at most 5 tests instead of 25.\n\nExample: 1000 commits? git bisect finds the breaking one in ~10 tests. Automate: git bisect run ./test.sh for fully hands-off binary search.	training/library/topics/debugging-methodology/primer.md
debugging-methodology/27c8d9e0f1a2	debugging-methodology	hard	debugging-methodology, traps, cognitive-bias	What are tunnel vision and confirmation bias in debugging, and how do you counter them?	Tunnel vision is fixating on one hypothesis and ignoring contradictory evidence. Confirmation bias is only looking for evidence that supports your theory. Counter both by: writing down at least 3 hypotheses before testing any, and actively seeking disconfirming evidence for your leading theory.\n\nRemember: write THREE hypotheses before investigating ANY. Forces broader thinking and prevents latching onto the first idea.	training/library/topics/debugging-methodology/primer.md
debugging-methodology/38d9e0f1a2b3	debugging-methodology	hard	debugging-methodology, checklist	What eight questions should you walk through before starting to debug, according to the debugging checklist?	1) What is the actual symptom? 2) When did it start (exact timestamp)? 3) What changed around that time? 4) Who/what is affected (scope)? 5) Is it consistent or intermittent? 6) What have you already tried? 7) What are your hypotheses (list at least 3)? 8) What is the fastest test to eliminate a hypothesis?\n\nRemember: SWITCH-HF — Symptom, When, whIch changed, sCope, Half-intermittent, tried, Hypotheses (3+), Fastest test.	training/library/topics/debugging-methodology/primer.md
debugging-methodology/z1a2b3c4d5e6	debugging-methodology	easy	debugging-methodology,crime-scene,evidence	Why should you preserve evidence before attempting a fix?	Restarting services, clearing logs, or redeploying destroys the information needed to understand root cause. Save logs, record versions, snapshot config, and keep failing input before touching anything.\n\nRemember: CSI rule — Capture State Immediately. Restarting destroys the crime scene. Save logs, snapshots, and config before touching anything.\n\nExample: kubectl logs pod > /tmp/crash.log && kubectl describe pod > /tmp/describe.txt — THEN restart.	zines/the.pocket.guide.to.debugging.wizard.zines.cleaned.notes
debugging-methodology/z2b3c4d5e6f7	debugging-methodology	easy	debugging-methodology,error-message,reading	Why should you read an error message twice?	The first reading is emotional (panic, frustration). The second reading is analytical: extract the exact text, file/line/function, component name, timing, and whether the error is primary or secondary fallout.\n\nRemember: first reading is emotional (panic). Second is analytical: extract file, line, component, timing, and whether the error is primary or secondary.	zines/the.pocket.guide.to.debugging.wizard.zines.cleaned.notes
debugging-methodology/z3c4d5e6f7a8	debugging-methodology	medium	debugging-methodology,reproduce,leverage	Why is reproduction the most important step in debugging?	Reproduction creates leverage: you can test hypotheses, validate fixes, and write regression tests. Without reproduction, debugging is guesswork and you cannot confirm the fix actually works.\n\nGotcha: without reproduction, debugging is guesswork. Reproduction turns debugging from art into science — you can verify fixes and write regression tests.	zines/the.pocket.guide.to.debugging.wizard.zines.cleaned.notes
debugging-methodology/z4d5e6f7a8b9	debugging-methodology	medium	debugging-methodology,suspects,brainstorm	What categories of causes should you brainstorm when debugging?	Config changes, dependency behavior changes, recent code/deploy changes, race conditions, stale data or caches, time zone issues, permission changes, and bad assumptions in your mental model.\n\nRemember: CCRD-CTPB — Config, Code/deploy, Race conditions, Dependencies, Caches/stale data, Time zones, Permission changes, Bad assumptions.	zines/the.pocket.guide.to.debugging.wizard.zines.cleaned.notes
debugging-methodology/z5e6f7a8b9c0	debugging-methodology	medium	debugging-methodology,simplify,reproducer	Why is writing a tiny reproducer valuable during debugging?	It strips away irrelevant complexity, making the bug's mechanism visible. A minimal reproducer also serves as the basis for a regression test and is easier to share when asking for help.\n\nExample: reduce 500 lines to 10 that still fail. Now you can share it, understand it, and write a regression test from it.	zines/the.pocket.guide.to.debugging.wizard.zines.cleaned.notes
debugging-methodology/z6f7a8b9c0d1	debugging-methodology	medium	debugging-methodology,bisect,version	How does finding a version that works help debugging?	A known-good baseline lets you narrow the search to what changed between working and broken. Use git bisect, deploy history, or version comparison to find the exact change that introduced the failure.\n\nExample: "Worked yesterday" → git bisect between yesterday's and today's deploy. "Works on A not B" → diff their configs.	zines/the.pocket.guide.to.debugging.wizard.zines.cleaned.notes
debugging-methodology/z7a8b9c0d1e2	debugging-methodology	hard	debugging-methodology,after-fix,prevention	What five things should you do after fixing a bug?	1. Write a regression test.\n2. Document the root cause.\n3. Improve observability (logging, metrics, alerts).\n4. Remove misleading logs or dead code.\n5. Ask: what would have made this easier to diagnose?\n\nRemember: RIDOC — Regression test, Improve observability, Document root cause, Obsolete misleading code, Consider what would have helped diagnose faster.	zines/the.pocket.guide.to.debugging.wizard.zines.cleaned.notes
debugging-methodology/z8b9c0d1e2f3	debugging-methodology	easy	debugging-methodology,print-statements,instrumentation	Why are print statements still an effective debugging technique?	They are cheap, local, require no setup, and brutally effective at revealing values, branches taken, timing, and request flow. They can be added anywhere instantly and removed just as easily.\n\nGotcha: use structured logging in production. But locally, print is zero-setup instant insight. Remove before committing.	zines/the.pocket.guide.to.debugging.wizard.zines.cleaned.notes
debugging-methodology/z9c0d1e2f3a4	debugging-methodology	hard	debugging-methodology,unstuck,strategies	Name five strategies for getting unstuck during debugging.	Take a break (diffuse thinking), pair with someone, timebox the rabbit hole, explain the bug out loud (rubber duck), and verify the code running is actually the code you changed (stale deploys, wrong branch).\n\nRemember: BPTED — Break (rest), Pair (fresh eyes), Timebox (15 min), Explain (rubber duck), Deploy check (right code running?).	zines/the.pocket.guide.to.debugging.wizard.zines.cleaned.notes
debugging-methodology/zad1e2f3a4b5	debugging-methodology	medium	debugging-methodology,one-variable,change	Why should you change only one variable at a time when debugging?	If you change multiple things and the problem resolves, you do not know which change fixed it. You cannot prevent recurrence, write a targeted test, or explain the root cause to others.\n\nRemember: the scientific method demands controlled experiments. Multi-variable changes create mystery fixes you cannot explain or reproduce.	zines/the.pocket.guide.to.debugging.wizard.zines.cleaned.notes
debugging-methodology/a1b2c3d4e5f6	debugging-methodology	easy	linux-debugging,dstat,overview	Why is dstat a good first tool when a machine is misbehaving?	dstat shows CPU, disk, network, and memory stats updating in real time in a single view. It quickly answers "is the machine CPU-bound, disk-bound, memory-bound, or network-bound?" without switching between tools.\n\nRemember: DCNM — dstat shows CPU, Disk, Network, Memory in one real-time view. Instantly answers "what kind of bottleneck?"\n\nSee also: modern alternatives include glances (Python), btop (interactive), dool (dstat fork on newer distros).	zines/linux.debugging.tools.you.ll.love.zine.cleaned.notes
debugging-methodology/b2c3d4e5f6a7	debugging-methodology	easy	linux-debugging,lsof,files	What three questions can lsof answer?	1. What files is this process holding open? (lsof -p <pid>)\n2. Who is listening on this port? (lsof -i :<port>)\n3. Why won't this filesystem unmount? (lsof /mount/point)\n\nRemember: LSOF = List Some Open Files. On Linux everything is a file — sockets, pipes, devices. lsof sees them all.	zines/linux.debugging.tools.you.ll.love.zine.cleaned.notes
debugging-methodology/c3d4e5f6a7b8	debugging-methodology	medium	linux-debugging,strace,syscalls	What kinds of problems is strace best at revealing?	File access errors (ENOENT, EACCES), network connection failures, missing libraries, permission problems, signal delivery, and slow syscalls. It shows every syscall a process makes with arguments and return values.\n\nExample: strace -e trace=file reveals ENOENT. strace -e trace=network reveals connection failures. strace -c shows syscall time stats.	zines/linux.debugging.tools.you.ll.love.zine.cleaned.notes
debugging-methodology/d4e5f6a7b8c9	debugging-methodology	medium	linux-debugging,perf,cpu	What does perf help you understand that top does not?	perf shows WHERE CPU time is going at the function level (hot functions, call stacks) and can trace kernel and user-space events. top only shows per-process CPU percentage.\n\nExample: perf top = live function-level hotspots. perf record -g = call graph capture. perf stat = hardware counters (cache misses, mispredicts).	zines/linux.debugging.tools.you.ll.love.zine.cleaned.notes
debugging-methodology/e5f6a7b8c9d0	debugging-methodology	medium	linux-debugging,proc,inspection	A process is leaking file descriptors. How would you confirm this and find what it is opening?	Check /proc/<pid>/fd — each entry is a symlink to an open file or socket. Count them over time (ls /proc/<pid>/fd | wc -l) to confirm growth. Read the symlink targets to see what is being leaked (ls -la /proc/<pid>/fd). lsof -p <pid> gives the same data with more detail.\n\nGotcha: default fd limit is often 1024 (ulimit -n). FD leaks hit this limit causing "Too many open files" even with free disk and memory.	zines/linux.debugging.tools.you.ll.love.zine.cleaned.notes
debugging-methodology/f6a7b8c9d0e1	debugging-methodology	easy	linux-debugging,ss,sockets	How do you list all TCP connections with process info using ss?	ss -tp shows established TCP connections with the owning process. Add -l for listening sockets (ss -tlnp), -u for UDP (ss -unp).\n\nRemember: ss = Socket Statistics. ss -tlnp = TCP Listening Numeric Process. Replaced netstat on modern Linux.	zines/linux.debugging.tools.you.ll.love.zine.cleaned.notes
debugging-methodology/a7b8c9d0e1f2	debugging-methodology	medium	linux-debugging,workflow,triage	A service is slow but top shows low CPU usage. What does that tell you and what do you check next?	Low CPU with high latency means the process is waiting, not computing. It is likely I/O-bound or blocked on a lock/network call. Check: iostat (disk I/O saturation), ss -tp (connection state — many CLOSE_WAIT or SYN_SENT?), strace -p <pid> -e trace=network (what syscall is it stuck on?). If strace shows futex or poll, the process is idle waiting for something external.\n\nRemember: low CPU + high latency = WAITING, not computing. Check I/O (iostat), connections (ss), blocked syscalls (strace). Process is stuck externally.	zines/linux.debugging.tools.you.ll.love.zine.cleaned.notes
debugging-methodology/b8c9d0e1f2a3	debugging-methodology	hard	linux-debugging,mistakes,anti-patterns	What are common debugging mistakes that Linux tools can prevent?	Reaching for application logs only (instead of OS-level data), assuming "slow" means CPU (could be disk or network), assuming "network issue" without packet capture, ignoring /proc, and restarting services before gathering evidence.\n\nRemember: ALARM — App logs only, Latency assumed CPU, Assuming network, Restart before evidence, Missing /proc.	zines/linux.debugging.tools.you.ll.love.zine.cleaned.notes
debugging-methodology/c9d0e1f2a3b4	debugging-methodology	medium	linux-debugging,vmstat,iostat	What do vmstat and iostat show that top does not?	vmstat shows memory, swap, I/O, and CPU stats per interval (reveals swapping and I/O wait). iostat shows per-device disk I/O statistics (throughput, queue depth, utilization per disk).\n\nExample: vmstat 1 per-second: si/so columns reveal swapping. iostat -x 1 shows per-device %util and await (I/O latency ms).	zines/linux.debugging.tools.you.ll.love.zine.cleaned.notes
debugging-methodology/d0e1f2a3b4c5	debugging-methodology	easy	linux-debugging,dmesg,kernel	When should you check dmesg during debugging and what would you look for?	Check dmesg when processes are killed unexpectedly (OOM killer messages), hardware errors are suspected (disk I/O errors, NIC failures), or containers crash without application logs. Look for: "Out of memory: Killed process", "I/O error", "segfault at", or "hardware error".\n\nExample: dmesg -T for human timestamps. Grep: "Out of memory", "I/O error", "segfault" — kernel evidence invisible to application logs.	zines/linux.debugging.tools.you.ll.love.zine.cleaned.notes
debugging-methodology/e1f2a3b4c5d6	debugging-methodology	hard	linux-debugging,strace,production	You need to strace a production service handling 10K requests/sec. What precautions do you take?	strace uses ptrace which stops the process on every traced syscall — at 10K req/s the overhead is severe. Precautions: (1) Filter aggressively with -e trace=network or -e trace=file to minimize intercepted calls. (2) Write to file with -o, never to terminal. (3) Limit duration to 10-30 seconds max. (4) Consider perf trace instead (kernel tracepoints, ~5x less overhead). (5) If possible, trace a single worker thread with -p <tid> rather than the whole process with -f.\n\nRemember: strace overhead = ptrace per syscall x request rate. At 10K req/s severe slowdown. Always filter (-e), limit duration, prefer perf trace.	zines/linux.debugging.tools.you.ll.love.zine.cleaned.notes
debugging-methodology/f2a3b4c5d6e7	debugging-methodology	medium	linux-debugging,lsof,port	How do you find which process is preventing a port from being reused?	lsof -i :<port> shows all processes with that port open. This reveals zombie listeners, processes in TIME_WAIT, or unexpected services that grabbed the port first.	zines/linux.debugging.tools.you.ll.love.zine.cleaned.notes

<!-- wiki:related:start -->
---

## Wiki Navigation

### Related Content

- [Debugging Methodology](../../../../library/topics/debugging-methodology/index.md) (Topic Pack, L1) — Debugging Methodology

<!-- wiki:related:end -->
