Anti-Primer: Azure Virtual Machines¶
Everything that can go wrong, will — and in this story, it does.
The Setup¶
A team migrates a stateless API from on-prem to a VM Scale Set behind a Standard Load Balancer. The deadline is tight, and the engineer building it has strong EC2 experience but is new to Azure's resource model.
The Timeline¶
Hour 0: "Stopped" Confusion¶
Believing the VM is stopped and no longer billing, the engineer shuts it down from inside the guest OS instead of running az vm deallocate.
Footgun #1: Stopping inside the guest is not the same as deallocating — the VM remains allocated and billed, and the team doesn't notice the wasted spend for weeks.
Hour 1: Golden Image Built on latest¶
The Packer image pipeline references the Marketplace image tag latest instead of a pinned Azure Compute Gallery version.
Footgun #2:
latestimage references are not reproducible — three weeks later a routine rebuild pulls a newer base image with a different OpenSSL version and the deployment silently changes behavior.
Hour 2: One Availability Set, Called "HA"¶
To save time, the team puts everything in a single availability set and calls the design "highly available."
Footgun #3: Availability Set and Availability Zone are not interchangeable — an availability set only spreads risk within one datacenter. A single power-distribution failure takes down every fault domain in that datacenter at once.
Hour 4: The Scale Set That Can't Change Its Mind¶
Midway through building the VMSS, the team realizes they need per-instance VM management APIs for a maintenance script, but the scale set was created with Uniform orchestration.
Footgun #4: VMSS orchestration mode is immutable — the entire scale set has to be deleted and recreated with Flexible orchestration, losing the week of tuning already invested.
Hour 6: The Extension That Wedges Everything¶
A custom-script extension references an internal package repository that isn't reachable from the new subnet's egress path.
Footgun #5: Extensions can wedge provisioning — every new instance in the scale set gets stuck in a failed provisioning state, even though the OS itself booted fine. Autoscale can't add capacity during a traffic spike because every new instance fails the same way.
Hour 10: Temp Disk Holds the Cache — and the Logs¶
Under pressure to unblock the extension issue, an engineer redirects the application's debug logs to the temporary disk "just for this investigation."
Footgun #6: Temporary disks are disposable — a routine resize operation to fix a capacity problem wipes the temp disk, and the exact logs needed to diagnose the extension failure are gone.
Hour 14: The Cleanup That Wasn't¶
Finally stable, the team deletes the failed test VMs from the earlier troubleshooting — but only through the portal's quick-delete, without checking the "also delete" boxes for disks and NICs.
Footgun #7: Deleting a VM can leave billable resources — a dozen orphaned managed disks and NICs sit in the resource group, quietly costing money until the next billing review catches it.
The Aftermath¶
Nothing here was one catastrophic mistake — it's seven small, individually survivable decisions that compound. Each one made sense in isolation, under time pressure, to someone reasoning from EC2 instincts in an Azure resource model that doesn't quite match.
What Should Have Happened¶
- Use
az vm deallocateexplicitly, and alert on VMs leftStopped(not deallocated) for more than a few minutes. - Pin Azure Compute Gallery image versions in every pipeline; promote versions deliberately.
- Design across Availability Zones for anything called "highly available," and use availability sets only where zone-level resilience genuinely isn't available or needed.
- Decide Flexible vs. Uniform orchestration deliberately before creating a scale set — it's a one-way door.
- Treat extensions as deployment dependencies with their own network/DNS/repo requirements, and test them in the actual target subnet before rollout.
- Never write anything you can't afford to lose to a temporary disk.
- Automate orphan-resource detection (Resource Graph queries or Azure Policy) instead of relying on manual delete-option checkboxes.