Azure Virtual Machines - Street-Level Ops¶
Baseline CLI Deployment¶
rg=rg-sre-lab
location=eastus
vm=vm-sre-01
az group create \
--name "$rg" \
--location "$location"
az network vnet create \
--resource-group "$rg" \
--name vnet-sre \
--address-prefixes 10.40.0.0/16 \
--subnet-name snet-compute \
--subnet-prefixes 10.40.10.0/24
az vm create \
--resource-group "$rg" \
--name "$vm" \
--image Canonical:0001-com-ubuntu-minimal-jammy:minimal-22_04-lts-gen2:latest \
--size Standard_D2s_v5 \
--vnet-name vnet-sre \
--subnet snet-compute \
--admin-username azureuser \
--assign-identity \
--generate-ssh-keys \
--public-ip-address ""
The official CLI quickstart demonstrates assigning a managed identity during creation. (Microsoft Learn: Create a Linux VM with Azure CLI)
In production, the VM should normally be private-only. Administrative access usually comes through one of these patterns:
- Azure Bastion.
- Site-to-site VPN or ExpressRoute.
- A hardened jump host.
- Entra-based SSH login plus Azure RBAC.
- Just-in-time access workflows.
Direct public SSH/RDP with a broad NSG rule is a lab pattern, not a production baseline.
Golden-Image and Bootstrap Split¶
A stable fleet usually separates:
- Image-time configuration: OS hardening, packages, agents, kernel modules, base users, trusted CAs.
- Boot-time configuration: Environment-specific endpoints, secrets references, cluster membership, service discovery, and one-time registration.
Common tools:
- Packer or Azure VM Image Builder.
- Azure Compute Gallery for versioned, replicated images.
cloud-initfor Linux first boot.- Ansible for convergent post-provisioning configuration.
- VM extensions for platform-managed components such as Azure Monitor Agent, antimalware, custom script, or Entra login.
Do not turn cloud-init into a 45-minute configuration-management system. Long bootstraps increase deployment failure rates, autoscale latency, and replacement risk.
Scale-Set Pattern¶
A typical stateless service stack is:
- VMSS across zones.
- Standard Load Balancer or Application Gateway.
- Health extension or application health probe.
- Autoscale rules based on Azure Monitor metrics.
- Rolling image upgrade policy.
- Automatic instance repairs.
- External state in managed databases/storage.
Example skeleton:
az vmss create \
--resource-group "$rg" \
--name vmss-api \
--orchestration-mode Flexible \
--image Canonical:0001-com-ubuntu-minimal-jammy:minimal-22_04-lts-gen2:latest \
--vm-sku Standard_D2s_v5 \
--instance-count 3 \
--zones 1 2 3 \
--vnet-name vnet-sre \
--subnet snet-compute \
--admin-username azureuser \
--generate-ssh-keys
Verify the exact CLI flags for the installed Azure CLI version — VMSS commands have changed as Flexible orchestration evolved.
Patching and Fleet Operations¶
Azure has several overlapping update mechanisms:
- Automatic VM guest patching.
- Azure Update Manager.
- Image replacement/reimage in a VMSS.
- OS-native tools such as
dnf,apt, WSUS, or configuration management.
For immutable fleets, rebuild the image and replace instances. For long-lived enterprise VMs, use maintenance windows, update assessments, staged rings, and explicit reboot handling. Mixing image replacement and in-place configuration drift without a declared ownership model creates "works until reimage" failures.
Backup, DR, and Recovery¶
- Azure Backup protects VM disks through Recovery Services vault or newer vault patterns depending on workload.
- Azure Site Recovery replicates VMs for disaster recovery and orchestrated failover.
- Disk snapshots are crash-consistent unless application quiescence is coordinated.
- Zone redundancy is not regional DR.
- A backup existing in the same subscription and region is not automatically sufficient against subscription compromise, ransomware, or regional failure.
Observability¶
At minimum, wire VMs to:
- Azure platform metrics.
- Azure Activity Log.
- Azure Monitor Agent using Data Collection Rules.
- Log Analytics workspace.
- VM Insights where its cost and dependency model are justified.
- Resource Health and Service Health alerts.
Host metrics do not include complete guest OS visibility. CPU percentage may be available as a platform metric, but filesystem usage, process health, application logs, and many memory metrics require guest collection.