---
tags:
- observability
- l1
- flashcard-deck
- prometheus
---
<!-- wiki:breadcrumb:start -->
[Portal](../../../../library/portal/index.md) | **Level:** [L1: Foundations](../../../../library/portal/levels.md) | **Topics:** [Prometheus](../../../../library/portal/topics.md) | **Domain:** Observability
<!-- wiki:breadcrumb:end -->

id	category	difficulty	tags	question	answer	source_path
datadog/1d9dfdefb429	datadog	medium	datadog, use-cases, overview	Describe at least three use cases for using something like Datadog. Can be as specific as you would like	* Monitor instances/servers downtime\n* Detect anomalies and send an alert when it happens\n* Service request or response latency\n\nRemember: Datadog's three pillars: metrics (numeric time-series), traces (request flows across services), and logs (text events). Correlate all three for fast incident resolution.\n\nGotcha: Datadog pricing is per host, per million log events, and per indexed span. Monitor your Datadog usage to avoid surprise bills — observability tools can be expensive at scale.	projects/knowledge/interview/datadog/001-describe-at-least-three-use-cases-for-using-someth.txt
datadog/289f4d3d63ee	datadog	medium	datadog, integrations, setup	What can you tell about Datadog integrations?	- Datadog has many supported integrations with different services, platforms, etc.\n- Each integration includes information on how to apply it, how to use it and what configuration options it supports\n\nRemember: consistent tagging (env, service, version) across metrics, traces, and logs is what makes Datadog powerful. Without consistent tags, correlation is impossible.	projects/knowledge/interview/datadog/007-what-can-you-tell-about-datadog-integrations.txt
datadog/34879ad187d2	datadog	medium	datadog, agent, fundamentals	What are the components of a Datadog agent?	* Collector: its role is to collect data from the host on which it's installed. The default period of time as of today is every 15 seconds.\n* Forwarder: responsible for sending the data to Datadog over HTTPS\n\nRemember: Datadog Agent collects metrics, traces, and logs from the host. Runs as a service (systemd) or DaemonSet in Kubernetes. Configurable via datadog.yaml.	projects/knowledge/interview/datadog/006-what-are-the-component-of-a-datadog-agent.txt
datadog/548505bd2305	datadog	medium	datadog, integrations, setup	"When opening some of the integrations windows/pages, there is a section called ""Monitors"". What can be found there?"	Usually you can find there some anomaly types that Datadog suggests to monitor and track.\n\nRemember: Datadog's three pillars: metrics (numeric time-series), traces (request flows across services), and logs (text events). Correlate all three for fast incident resolution.\n\nGotcha: Datadog pricing is per host, per million log events, and per indexed span. Monitor your Datadog usage to avoid surprise bills — observability tools can be expensive at scale.	projects/knowledge/interview/datadog/008-what-opening-some-of-the-integrations-windowspages.txt
datadog/9cc53f5b7ab6	datadog	easy	datadog, agent, fundamentals	What is a Datadog agent?	A software runs on a Datadog host. Its purpose is to collect data from the host and sent it to Datadog (data like metrics, logs, etc.)\n\nRemember: Datadog Agent collects metrics, traces, and logs from the host. Runs as a service (systemd) or DaemonSet in Kubernetes. Configurable via datadog.yaml.	projects/knowledge/interview/datadog/004-what-is-a-datadog-agent.txt
datadog/c2b59aa89898	datadog	medium	datadog, data-collection, methods	What ways are there to collect or send data to Datadog?	* Datadog agent installed on the device or location which you would like to monitor\n* Using Datadog API\n* Built-in integrations\n\nRemember: consistent tagging (env, service, version) across metrics, traces, and logs is what makes Datadog powerful. Without consistent tags, correlation is impossible.	projects/knowledge/interview/datadog/002-what-ways-are-there-to-collect-or-send-data-to-dat.txt
datadog/d9391e3d7c89	datadog	medium	datadog, tags, configuration	What are Datadog tags?	"Datadog tags are used to mark different information with unique properties. For example, you might want to tag some data with ""environment: production"" while tagging information from staging or dev environment with ""environment: staging"".\n\nRemember: tags are key:value pairs (env:production, service:web-api). They enable filtering, grouping, and aggregation across all Datadog data. Tag everything consistently."	projects/knowledge/interview/datadog/005-what-are-datadog-tags.txt
datadog/ec2cd38ab24e	datadog	easy	datadog, infrastructure, hosts	What is a host in regards to Datadog?	Any physical or virtual instance that is monitored with Datadog. Few examples:\n\n- Cloud Instance, Virtual Machine\n- Bare metal node\n- Platform or service specific nodes like Kubernetes node\n\nBasically any device or location that has Datadog agent installed and running on.\n\nRemember: consistent tagging (env, service, version) across metrics, traces, and logs is what makes Datadog powerful. Without consistent tags, correlation is impossible.	projects/knowledge/interview/datadog/003-what-is-a-host-in-regards-to-datadog.txt
datadog/dd-agent-arch	datadog	medium	datadog, agent, architecture	Describe the Datadog Agent architecture and its main processes.	The Datadog Agent runs as a service on hosts and consists of: 1) Core Agent: collects system metrics, handles check scheduling, and manages the event pipeline; 2) Trace Agent (APM): receives traces from instrumented applications and forwards them to Datadog; 3) Process Agent: collects live process and container data; 4) Log Agent: tails log files, listens on TCP/UDP, and ships logs. All communicate with Datadog via HTTPS on port 443.\n\nRemember: Datadog Agent collects metrics, traces, and logs from the host. Runs as a service (systemd) or DaemonSet in Kubernetes. Configurable via datadog.yaml.	training/interactive/knowledge/data/cards/datadog.tsv
datadog/dd-custom-metrics	datadog	medium	datadog, metrics, custom	How do you submit custom metrics to Datadog?	Methods: 1) DogStatsD: a StatsD-compatible UDP server bundled with the agent (send from app code via client libraries); 2) Agent checks: Python scripts in the agent that collect and submit metrics; 3) Datadog API: POST metrics directly via REST API. Metric types: count, gauge, rate, histogram, distribution. Custom metrics are billed per unique metric name + tag combination, so tag cardinality matters.\n\nRemember: consistent tagging (env, service, version) across metrics, traces, and logs is what makes Datadog powerful. Without consistent tags, correlation is impossible.	training/interactive/knowledge/data/cards/datadog.tsv
datadog/dd-apm-traces	datadog	hard	datadog, apm, tracing	How does Datadog APM work and what is a trace vs a span?	A trace represents a single request flowing through a distributed system. Each trace is composed of spans, where each span represents a unit of work (e.g., an HTTP handler, a database query, an external call). Spans have start time, duration, tags, and parent-child relationships. The tracing library instruments your code, sends spans to the Trace Agent (localhost:8126), which forwards them to Datadog. Traces enable latency analysis, error tracking, and service dependency mapping.\n\nRemember: Datadog APM instruments your code to trace requests across services. Shows latency, error rates, and dependency maps. Supports auto-instrumentation for Python, Java, Go, etc.	training/interactive/knowledge/data/cards/datadog.tsv
datadog/dd-log-pipelines	datadog	hard	datadog, logs, pipelines	What are Datadog log pipelines and how do they process logs?	Log pipelines are ordered sets of processors that parse and enrich raw log data. Key processors: 1) Grok Parser: extract structured fields from unstructured log lines; 2) Date Remapper: set the official log timestamp; 3) Status Remapper: set log severity; 4) Attribute Remapper: rename or transform fields; 5) Category Processor: classify logs by rules. Pipelines are matched to logs by filter queries. They transform raw text into structured, searchable, queryable data.\n\nRemember: Datadog Log Management: collect -> parse (pipelines) -> index (include/exclude) -> analyze. Use facets and saved views for fast investigation. Archive to S3 for long-term retention.	training/interactive/knowledge/data/cards/datadog.tsv
datadog/dd-monitors	datadog	medium	datadog, monitors, alerting	What types of monitors does Datadog support and how do alerts work?	Monitor types: 1) Metric: threshold or anomaly on numeric metrics; 2) Service Check: agent or integration health status; 3) Log: alert on log patterns or counts; 4) APM: latency, error rate, or throughput thresholds; 5) Composite: combine multiple monitors with boolean logic; 6) Forecast: predict future metric values. Alerts flow through notification channels (Slack, PagerDuty, email). Monitors support warn/alert thresholds, recovery conditions, and no-data handling.\n\nRemember: good alerting follows the RED method for services (Rate, Errors, Duration) and USE method for resources (Utilization, Saturation, Errors).\n\nRemember: Datadog monitors are alert rules. Types: metric, anomaly, forecast, outlier, log, APM, composite. Each can notify Slack, PagerDuty, email, webhooks.	training/interactive/knowledge/data/cards/datadog.tsv
datadog/dd-dashboards	datadog	easy	datadog, dashboards, visualization	What dashboard types does Datadog offer and when would you use each?	Timeboards: all widgets share the same time scope, useful for troubleshooting and correlation (synchronized zoom/pan). Screenboards: free-form layout with widgets at independent time scopes, useful for status boards and executive views. Widgets include: timeseries, query value, top list, heatmap, distribution, log stream, service map, SLO summary. Template variables allow filtering dashboards by tag values.\n\nExample: a good service dashboard: request rate, error rate, latency (p50/p95/p99), saturation (CPU/memory), and deployment markers. The RED method in one view.	training/interactive/knowledge/data/cards/datadog.tsv
datadog/dd-slos	datadog	medium	datadog, slo, reliability	How do you define and track SLOs in Datadog?	Datadog SLOs track service reliability against targets. Types: 1) Metric-based: percentage of time a metric meets a threshold (e.g., latency p99 < 500ms); 2) Monitor-based: percentage of time a monitor is in OK state. Configure a target (e.g., 99.9%), time window (7d, 30d, 90d), and error budget. SLO widgets on dashboards show remaining error budget. Alerts can fire when error budget is being consumed too fast.\n\nRemember: consistent tagging (env, service, version) across metrics, traces, and logs is what makes Datadog powerful. Without consistent tags, correlation is impossible.	training/interactive/knowledge/data/cards/datadog.tsv
datadog/dd-synthetics	datadog	medium	datadog, synthetics, testing	What is Datadog Synthetic Monitoring?	Synthetic monitoring runs automated tests against your services from global locations. Types: 1) API tests: HTTP, SSL, DNS, TCP, gRPC checks to verify availability and response correctness; 2) Browser tests: record and replay user journeys in a headless browser to catch UI regressions; 3) Multistep API tests: chain multiple API calls. Tests run on a schedule and alert on failures. They provide uptime tracking and help catch issues before users do.\n\nRemember: Datadog = SaaS monitoring platform. Metrics, traces, logs, and security in one pane. Agent-based collection, cloud-native integrations.	training/interactive/knowledge/data/cards/datadog.tsv
datadog/dd-integrations	datadog	easy	datadog, integrations, setup	How do Datadog integrations work and name five common ones.	Integrations connect Datadog to external services and technologies. Types: 1) Agent-based: the agent runs a check (e.g., postgres, nginx, redis checks); 2) API/webhook-based: services push data to Datadog (e.g., AWS, GCP, PagerDuty); 3) Library-based: instrumentation SDKs (e.g., ddtrace for APM). Common integrations: AWS (CloudWatch metrics), Kubernetes (pod/node metrics), PostgreSQL, Nginx, Redis. Each integration provides out-of-the-box dashboards and monitors.\n\nRemember: consistent tagging (env, service, version) across metrics, traces, and logs is what makes Datadog powerful. Without consistent tags, correlation is impossible.	training/interactive/knowledge/data/cards/datadog.tsv
datadog/dd-tagging	datadog	easy	datadog, tags, best-practices	What are Datadog tagging best practices?	Tags are key:value pairs that enable filtering, grouping, and aggregation. Best practices: 1) Use consistent naming (env:production, service:api, team:platform); 2) Tag at the source (agent config, cloud provider tags auto-imported); 3) Keep cardinality reasonable (avoid per-request or per-user tags on metrics); 4) Use reserved tags: env, service, version (Unified Service Tagging) for correlation across metrics, traces, and logs; 5) Document your tagging convention.\n\nRemember: tags are key:value pairs (env:production, service:web-api). They enable filtering, grouping, and aggregation across all Datadog data. Tag everything consistently.	training/interactive/knowledge/data/cards/datadog.tsv

<!-- wiki:related:start -->
---

## Wiki Navigation

### Related Content

- [Adversarial Interview Gauntlet (30 sequences)](../../../../library/interview-scenarios/gauntlet/README.md) (Scenario, L2) — Prometheus
- [Alerting Rules](../../../../library/topics/alerting-rules/index.md) (Topic Pack, L2) — Prometheus
- [Alerting Rules Drills](../../../../library/drills/alerting_rules_drills.md) (Drill, L2) — Prometheus
- [Capacity Planning](../../../../library/topics/capacity-planning/index.md) (Topic Pack, L2) — Prometheus
- [Case Study: Disk Full — Runaway Logs, Fix Is Loki Retention](../../../../library/case-studies/cross-domain/disk-full-runaway-logs-loki/README.md) (Case Study, L2) — Prometheus
- [Case Study: Grafana Dashboard Empty — Prometheus Blocked by NetworkPolicy](../../../../library/case-studies/cross-domain/grafana-empty-prometheus-networkpolicy/README.md) (Case Study, L2) — Prometheus
- Incident Simulator (18 scenarios) *(CLI)* (Exercise Set, L2) — Prometheus
- [Interview: Prometheus Target Down](../../../../library/interview-scenarios/03-prometheus-target-down.md) (Scenario, L2) — Prometheus
- Lab: Prometheus Target Down *(CLI)* (Lab, L2) — Prometheus
- Monitoring Flashcards *(CLI)* (flashcard_deck, L1) — Prometheus

<!-- wiki:related:end -->
