---
tags:
- devops
- l1
- flashcard-deck
- ml-ops
---
<!-- wiki:breadcrumb:start -->
[Portal](../../../../library/portal/index.md) | **Level:** [L1: Foundations](../../../../library/portal/levels.md) | **Topics:** [AI/ML Infrastructure Ops](../../../../library/portal/topics.md) | **Domain:** DevOps & Tooling
<!-- wiki:breadcrumb:end -->

id	category	difficulty	tags	question	answer	source_path
ml-ops/0998be8a2c47	ml-ops	easy	mlops, gpu, monitoring	What are the six key GPU metrics that ops engineers need to monitor?	"GPU utilization (%), GPU memory used (GB), GPU temperature (C), GPU power draw (W), PCIe throughput (GB/s), and NVLink throughput (GB/s). Use nvidia-smi for snapshots and DCGM exporter with Prometheus for continuous monitoring.\n\nRemember: MLOps = DevOps for ML. Training, versioning, deployment, monitoring. ""CI/CD for models.""\n\nDebug clue: `nvidia-smi dmon -s pucvmet` gives a continuous monitoring stream of all six metrics in one command."	training/library/topics/ai-ml-ops/primer.md
ml-ops/269217187ee6	ml-ops	easy	mlops, cuda, stack	What are the four layers of the CUDA stack from top to bottom, and why is version compatibility critical?	"Application (PyTorch/TensorFlow) -> CUDA Toolkit (nvcc, cuBLAS, cuDNN) -> CUDA Driver (nvidia.ko kernel module) -> GPU Hardware. Compatibility is strict: toolkit version must match the application build, driver must be >= toolkit version, driver must support the GPU generation, and kernel version must be compatible with the driver.\n\nRemember: MLOps = DevOps for ML. Training, versioning, deployment, monitoring. ""CI/CD for models.""\n\nAnalogy: The CUDA stack is like a tower — app at the top, hardware at the bottom. Each layer must be compatible with its neighbors or the whole tower falls."	training/library/topics/ai-ml-ops/primer.md
ml-ops/0a74389eca45	ml-ops	easy	mlops, gpu, scheduling	How does GPU scheduling work in Kubernetes by default?	"Kubernetes doesn't natively understand GPUs — you need the NVIDIA device plugin (a DaemonSet). GPU scheduling is all-or-nothing: if a pod requests 1 GPU, it gets exclusive access to that entire GPU. No sharing, no overcommit. Verify GPUs with kubectl describe node | grep nvidia.com/gpu.\n\nRemember: MLOps = DevOps for ML. Training, versioning, deployment, monitoring. ""CI/CD for models.""\n\nGotcha: GPUs cannot be overcommitted like CPU. If you request 1 GPU, you get the whole GPU. Plan node sizing accordingly."	training/library/topics/ai-ml-ops/primer.md
ml-ops/905e59aeaa2a	ml-ops	medium	mlops, gpu, sharing	What is the difference between GPU time-slicing and MIG (Multi-Instance GPU)?	"Time-slicing shares a GPU by rapidly switching between workloads — no memory isolation, all workloads share the same VRAM, high OOM risk. Good for development, bad for production. MIG (available on A100/H100) provides hardware-level partitioning with true memory isolation — one workload can't OOM another. MIG is ideal for inference serving where each model needs a predictable memory slice.\n\nRemember: MLOps = DevOps for ML. Training, versioning, deployment, monitoring. ""CI/CD for models.""\n\nName origin: MIG = Multi-Instance GPU. Introduced with NVIDIA A100 (Ampere architecture, 2020)."	training/library/topics/ai-ml-ops/primer.md
ml-ops/393da03d1354	ml-ops	medium	mlops, model-serving	What is vLLM and how do you configure its Kubernetes deployment for production?	"vLLM is an LLM inference server. Key deployment considerations: set gpu-memory-utilization (e.g., 0.90), mount a PVC for model cache (avoid re-downloading 140GB models on every restart), set initialDelaySeconds on readiness probes to 120+ seconds (models take minutes to load), mount /dev/shm as emptyDir with medium: Memory for PyTorch DataLoader.\n\nRemember: MLOps = DevOps for ML. Training, versioning, deployment, monitoring. ""CI/CD for models.""\n\nName origin: vLLM stands for ""virtual Large Language Model"" — it uses PagedAttention to efficiently manage GPU memory for LLM inference."	training/library/topics/ai-ml-ops/primer.md
ml-ops/26ddfba43c8f	ml-ops	medium	mlops, gpu, oom	What are the five common causes of GPU memory (VRAM) OOM errors and their fixes?	"(1) Batch size too large — reduce batch_size. (2) Model doesn't fit in GPU memory — use model parallelism, quantization, or bigger GPU. (3) Memory fragmentation in PyTorch — set PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True. (4) Memory leak in training loop — ensure tensors are detached/deleted. (5) Multiple users sharing GPU via time-slicing — use MIG for memory isolation.\n\nRemember: MLOps = DevOps for ML. Training, versioning, deployment, monitoring. ""CI/CD for models.""\n\nDebug clue: `nvidia-smi` shows per-GPU memory usage. `torch.cuda.memory_summary()` in Python shows PyTorch\'s allocation breakdown."	training/library/topics/ai-ml-ops/primer.md
ml-ops/46e9e4261011	ml-ops	medium	mlops, storage	What are the three categories of ML storage needs and appropriate solutions for each?	"(1) Model weights (read-heavy, large files): NFS/NAS, S3 with local cache, or ReadWriteMany PVCs. (2) Training data (read-heavy, massive): S3/GCS with streaming, Lustre/GPFS for HPC, or NFS with SSD cache. (3) Checkpoints (write-heavy during training): local NVMe for speed, PVC for persistence, or S3 with periodic sync for durability.\n\nRemember: MLOps = DevOps for ML. Training, versioning, deployment, monitoring. ""CI/CD for models.""\n\nRemember: ""Models = read-heavy, Training data = massive reads, Checkpoints = write-heavy."" Each needs different storage characteristics."	training/library/topics/ai-ml-ops/primer.md
ml-ops/9927ba552216	ml-ops	hard	mlops, pitfalls, devshm	Why is mounting /dev/shm critical for PyTorch training jobs in Kubernetes, and what happens if you forget?	"PyTorch DataLoader uses shared memory for multiprocess data loading. Default /dev/shm in Kubernetes is only 64MB. A training job with 8 data workers will crash because it exceeds this limit. Fix: mount an emptyDir with medium: Memory at /dev/shm with an appropriate sizeLimit (e.g., 16Gi).\n\nRemember: MLOps = DevOps for ML. Training, versioning, deployment, monitoring. ""CI/CD for models.""\n\nGotcha: The default 64MB /dev/shm in Kubernetes is a silent killer for ML workloads. Always mount an emptyDir with `medium: Memory`."	training/library/topics/ai-ml-ops/primer.md
ml-ops/46f344091bd0	ml-ops	hard	mlops, gpu, taints	Why should GPU nodes be tainted in Kubernetes, and what happens without taints?	"Without taints, Kubernetes will schedule CPU-only pods on expensive GPU nodes, consuming CPU and memory that GPU workloads need for data loading. A GPU node with 4x A100s and 64 CPU cores might only run 4 pods (one per GPU), and the CPU/RAM exists to feed the GPUs. Taint GPU nodes and tolerate only GPU workloads to prevent resource waste on $10K-$40K/card hardware.\n\nRemember: MLOps = DevOps for ML. Training, versioning, deployment, monitoring. ""CI/CD for models.""\n\nRemember: ""Taint + tolerate = GPU reservation."" Without taints, cheap CPU pods consume expensive GPU node resources."	training/library/topics/ai-ml-ops/primer.md
ml-ops/6074f0be164f	ml-ops	hard	mlops, gpu, alerts	What are the six critical alerts to configure for a GPU cluster and their severity levels?	"GPU memory > 90% VRAM (Warning), GPU temperature > 83C sustained (Warning — thermal throttling reduces performance 20-40%), GPU utilization < 10% for 30 min (Info — wasting money), XID errors detected (Critical — hardware errors), CUDA OOM pod restart (High), GPU not detected / nvidia-smi fails (Critical). Use DCGM exporter Prometheus metrics for alerting.\n\nRemember: MLOps = DevOps for ML. Training, versioning, deployment, monitoring. ""CI/CD for models.""\n\nRemember: ""DCGM exporter = GPU Prometheus metrics."" It exposes nvidia_gpu_* metrics for Grafana dashboards and alerting."	training/library/topics/ai-ml-ops/primer.md

<!-- wiki:related:start -->
---

## Wiki Navigation

### Related Content

- Datascience Flashcards *(CLI)* (flashcard_deck, L1) — AI/ML Infrastructure Ops
- [The Ops of AI/ML Workloads](../../../../library/topics/ai-ml-ops/index.md) (Topic Pack, L2) — AI/ML Infrastructure Ops

<!-- wiki:related:end -->
