---
tags:
- devops
- l1
- flashcard-deck
- observability
---
<!-- wiki:breadcrumb:start -->
[Portal](../../../../library/portal/index.md) | **Level:** [L1: Foundations](../../../../library/portal/levels.md) | **Topics:** [Observability Deep Dive](../../../../library/portal/topics.md) | **Domain:** DevOps & Tooling
<!-- wiki:breadcrumb:end -->

id	category	difficulty	tags	question	answer	source_path
obs/001	observability	easy	observability, pillars	What are the three pillars of observability?	Metrics (numeric time-series), Logs (discrete events with context), Traces (request flow across services). Each answers different questions: metrics show what is broken, logs show why, traces show where in the call chain.\n\nRemember: Three pillars: Metrics, Logs, Traces. "MLT."	training/library/deep-dives/observability
obs/002	observability	easy	observability, metrics, logs	When should you use a metric vs a log?	Use metrics for aggregatable numeric data you alert on (request rate, error rate, latency). Use logs for discrete events with rich context (stack traces, request payloads). Metrics are cheap to query at scale; logs are expensive.\n\nRemember: Counter(cumulative), Gauge(current), Histogram(distribution).	training/library/deep-dives/observability
obs/003	observability	medium	observability, traces	What problem do distributed traces solve that metrics and logs alone cannot?	Traces show causality across service boundaries. When a request touches 5 services, traces reveal which hop is slow or failing, even when each service's metrics look fine individually.\n\nRemember: Distributed tracing = follow request across services. Span = segment.\n\nExample: Jaeger, Zipkin, OpenTelemetry. Trace ID in HTTP headers.	training/library/deep-dives/observability
obs/004	observability	easy	prometheus, scrape	How does Prometheus collect metrics?	Prometheus uses a pull (scrape) model. It periodically HTTP GETs the /metrics endpoint of each target. Targets expose metrics in the Prometheus exposition format. This is the opposite of push-based systems like StatsD.\n\nRemember: Counter(cumulative), Gauge(current), Histogram(distribution).	training/library/deep-dives/observability
obs/005	observability	medium	prometheus, exporter	What is a Prometheus exporter and when do you need one?	An exporter is a sidecar or standalone process that translates metrics from a non-Prometheus system into the Prometheus exposition format. You need one when the application cannot natively expose /metrics (e.g., node_exporter for OS metrics, blackbox_exporter for probing).\n\nRemember: Observability ≠ monitoring. Monitoring=known questions. Observability=ask new questions.	training/library/deep-dives/observability
obs/006	observability	medium	prometheus, cardinality	What is label cardinality and why is it dangerous in Prometheus?	Cardinality is the number of unique time-series created by label combinations. High-cardinality labels (user IDs, request paths) create millions of series, consuming memory and slowing queries. Fix with relabeling, recording rules, or dropping unbounded labels at the source.\n\nGotcha: High cardinality = #1 perf killer. Never user IDs as metric labels.	training/library/deep-dives/observability
obs/007	observability	medium	prometheus, metric-types	What is the difference between a counter and a gauge?	A counter only goes up (resets on restart): total requests, total errors. A gauge goes up and down: current memory, queue depth. Always use rate() on counters; use gauge values directly.\n\nRemember: Observability ≠ monitoring. Monitoring=known questions. Observability=ask new questions.	training/library/deep-dives/observability
obs/008	observability	medium	prometheus, histogram	When do you use a histogram vs a summary?	Use histograms for latency percentiles when you need server-side aggregation across instances (histogram_quantile). Summaries compute quantiles client-side so they cannot be aggregated. Prefer histograms in almost all cases.\n\nRemember: Observability ≠ monitoring. Monitoring=known questions. Observability=ask new questions.	training/library/deep-dives/observability
obs/009	observability	easy	sli, slo, sla	What is the difference between SLI, SLO, and SLA?	SLI (indicator): a measured metric like request latency p99. SLO (objective): target value for the SLI, e.g., p99 < 300ms. SLA (agreement): contractual commitment with consequences if SLO is breached. SLI measures, SLO sets the goal, SLA is the contract.\n\nRemember: Observability ≠ monitoring. Monitoring=known questions. Observability=ask new questions.	training/library/deep-dives/observability
obs/010	observability	medium	slo, error-budget	What is an error budget and how does it guide decision-making?	Error budget = 1 minus SLO target (e.g., 99.9% SLO gives 0.1% budget). When budget is healthy, ship features fast. When budget is nearly spent, freeze risky changes and focus on reliability. It aligns dev velocity with reliability.\n\nRemember: Observability ≠ monitoring. Monitoring=known questions. Observability=ask new questions.	training/library/deep-dives/observability
obs/011	observability	medium	alerting, fatigue	What is alert fatigue and how do you combat it?	Alert fatigue occurs when too many non-actionable alerts cause responders to ignore or mute them. Fix: alert on symptoms not causes, require every alert to be actionable and have a runbook, tune thresholds, deduplicate via grouping, and regularly prune alerts that nobody acts on.\n\nRemember: Observability ≠ monitoring. Monitoring=known questions. Observability=ask new questions.	training/library/deep-dives/observability
obs/012	observability	medium	alerting, routing	How does alert routing work in Alertmanager?	Alertmanager receives alerts from Prometheus and routes them based on label matchers. Routes form a tree: alerts match the most specific route. Each route can set a receiver (Slack, PagerDuty), group_by labels, and timing (group_wait, group_interval, repeat_interval).\n\nRemember: Observability ≠ monitoring. Monitoring=known questions. Observability=ask new questions.	training/library/deep-dives/observability
obs/013	observability	hard	alerting, best-practices	What makes a good production alert?	1) Actionable: someone must do something now. 2) Urgent: cannot wait until business hours. 3) Real: low false-positive rate. 4) Symptom-based: alert on error rate, not CPU. 5) Includes runbook link. 6) Has clear ownership.\n\nRemember: Observability ≠ monitoring. Monitoring=known questions. Observability=ask new questions.	training/library/deep-dives/observability
obs/014	observability	easy	dashboards, alerts	When should you use a dashboard vs an alert?	Dashboards are for investigation and trend analysis (pull). Alerts are for immediate notification of problems (push). Good rule: if it needs human attention right now, alert. If it provides context during investigation, dashboard.\n\nRemember: Observability ≠ monitoring. Monitoring=known questions. Observability=ask new questions.	training/library/deep-dives/observability
obs/015	observability	medium	blackbox, whitebox	What is the difference between blackbox and whitebox monitoring?	Blackbox monitoring tests externally visible behavior (HTTP probe returns 200, TCP connect succeeds). Whitebox monitoring uses internal instrumentation (request latency histogram, error counters). Use blackbox for user-facing SLIs; whitebox for debugging and capacity planning.\n\nRemember: Observability ≠ monitoring. Monitoring=known questions. Observability=ask new questions.	training/library/deep-dives/observability
obs/016	observability	medium	prometheus, blackbox	How does blackbox_exporter work?	Prometheus scrapes blackbox_exporter with a target parameter. The exporter probes the target (HTTP, TCP, ICMP, DNS) and returns success/failure, latency, TLS expiry, etc. as metrics. Useful for monitoring endpoints you don't control.\n\nRemember: Observability ≠ monitoring. Monitoring=known questions. Observability=ask new questions.	training/library/deep-dives/observability
obs/017	observability	medium	grafana, troubleshooting	Grafana dashboard shows No Data. What do you check?	1) Data source connection: Settings > Data Sources > Test. 2) Time range: too narrow or shifted. 3) Query syntax: run in Explore tab. 4) Metric name changed or labels don't match. 5) Prometheus retention: data may have expired.\n\nRemember: Observability ≠ monitoring. Monitoring=known questions. Observability=ask new questions.	training/library/deep-dives/observability
obs/018	observability	hard	prometheus, troubleshooting	Prometheus memory is growing fast. How do you diagnose?	Check TSDB status page (/tsdb-status) for high-cardinality metrics. Look for labels with unbounded values. Use promtool tsdb analyze on data directory. Fix: drop labels via metric_relabel_configs, use recording rules to pre-aggregate, or increase scrape interval for noisy targets.\n\nRemember: Observability ≠ monitoring. Monitoring=known questions. Observability=ask new questions.	training/library/deep-dives/observability
obs/019	observability	medium	prometheus, staleness	A Prometheus target shows as DOWN but the app is healthy. What do you check?	1) ServiceMonitor selector vs Service labels. 2) Port name in ServiceMonitor matches Service port name. 3) Metrics endpoint path is correct (/metrics). 4) Network policy blocking Prometheus from scraping. 5) Prometheus namespace selector config.\n\nRemember: Observability ≠ monitoring. Monitoring=known questions. Observability=ask new questions.	training/library/deep-dives/observability
obs/020	observability	medium	prometheus, recording-rules	What are recording rules and why use them?	Recording rules pre-compute frequently used or expensive PromQL expressions and store the result as a new time-series. Use them to speed up dashboard queries, reduce query-time load, and create stable metric names for alerting rules.\n\nRemember: Observability ≠ monitoring. Monitoring=known questions. Observability=ask new questions.	training/library/deep-dives/observability
obs/021	observability	hard	tracing, sampling	What is head-based vs tail-based sampling in distributed tracing?	Head-based: decide at trace start whether to sample (simple, but may miss interesting traces). Tail-based: decide after all spans arrive (can keep error/slow traces, discard boring ones). Tail-based is better for debugging but needs a collector buffer.\n\nRemember: Distributed tracing = follow request across services. Span = segment.\n\nExample: Jaeger, Zipkin, OpenTelemetry. Trace ID in HTTP headers.	training/library/deep-dives/observability
obs/022	observability	medium	logging, structured	Why use structured logging over plain text?	Structured logs (JSON) have consistent fields that can be parsed, indexed, and queried without regex. Enables filtering by severity, request ID, user, service. Plain text requires fragile pattern matching and breaks when log format changes.\n\nRemember: Observability ≠ monitoring. Monitoring=known questions. Observability=ask new questions.	training/library/deep-dives/observability
obs/023	observability	hard	prometheus, federation	When do you need Prometheus federation or remote write?	When you have multiple Prometheus servers (per-cluster) and need a global view. Federation pulls selected series from leaf Prometheus into a global one. Remote write pushes to long-term storage (Thanos, Cortex, Mimir). Use remote write for durability; federation for lightweight aggregation.\n\nRemember: Observability ≠ monitoring. Monitoring=known questions. Observability=ask new questions.	training/library/deep-dives/observability
obs/024	observability	medium	observability, red-use	Explain the RED and USE methods for monitoring.	RED (for services): Rate, Errors, Duration per service endpoint. USE (for resources): Utilization, Saturation, Errors per resource (CPU, disk, network). RED tells you about user experience; USE tells you about infrastructure health.\n\nRemember: Observability ≠ monitoring. Monitoring=known questions. Observability=ask new questions.	training/library/deep-dives/observability
obs/025	observability	hard	alerting, on-call	How do you structure on-call alerting to reduce burnout?	1) Symptom-based alerts only (no cause-based noise). 2) Every alert has a runbook. 3) Track alert-to-action ratio; prune noisy alerts quarterly. 4) Use escalation policies with timeout. 5) Follow-the-sun rotation across time zones if possible. 6) Blameless postmortems improve signal over time.\n\nRemember: Observability ≠ monitoring. Monitoring=known questions. Observability=ask new questions.	training/library/deep-dives/observability

<!-- wiki:related:start -->
---

## Wiki Navigation

### Related Content

- [Mental Models (Core Concepts)](../../../../library/topics/mental-models-core/index.md) (Topic Pack, L0) — Observability Deep Dive

<!-- wiki:related:end -->
