π Mission Debrief: Canary Weight Imbalance¶
What Happened¶
Your canary deployment had a 50/50 traffic split (5 stable pods, 5 canary pods) instead of the intended 90/10 split.
This means half your users were exposed to the new, untested canary version - defeating the entire purpose of canary deployments!
Canary deployments should expose only a small percentage of users to the new version for safe testing.
How Kubernetes Behaved¶
Service load balancing (Kubernetes default):
Service selector: app=myapp (matches both stable and canary)
β
Endpoints: All pods with app=myapp label
β
Load balancer distributes evenly across ALL endpoints
β
With 5 stable + 5 canary = 10 total pods
Each pod gets ~10% of traffic
β
stable pods: 5 Γ 10% = 50% traffic β
canary pods: 5 Γ 10% = 50% traffic β
Broken configuration:
Stable (v1): 5 pods βββ
ββββΆ Service βββΆ 50% v1, 50% v2 β
Canary (v2): 5 pods βββ
Result: Half your users are guinea pigs!
Fixed configuration:
Stable (v1): 9 pods βββ
ββββΆ Service βββΆ 90% v1, 10% v2 β
Canary (v2): 1 pod βββ
Result: Only 10% of users test new version
The Correct Mental Model¶
Canary Deployment Strategy¶
Concept: Gradually roll out new version to a small subset of users, monitor for issues, then progressively increase traffic.
βββββββββββββββββββββββββββββββββββββββββββββββ
β PHASE 1: Stable Only β
β ββββββββββββββββ β
β β Service βββββββββΆββββββββββββββββ β
β β (selector: β β Stable Pods β β
β β app=myapp) β β (v1.0) β β
β ββββββββββββββββ β 10 replicas β β
β ββββββββββββββββ β
β β
β All users β v1.0 β
βββββββββββββββββββββββββββββββββββββββββββββββ
βββββββββββββββββββββββββββββββββββββββββββββββ
β PHASE 2: Canary Introduction (10%) β
β ββββββββββββββββ ββββββββββββββββ β
β β Service ββββββ¬βββΆβ Stable Pods β β
β β (selector: β β β (v1.0) β β
β β app=myapp) β β β 9 replicas β β
β ββββββββββββββββ β ββββββββββββββββ β
β β β
β β ββββββββββββββββ β
β ββββΆβ Canary Pods β β
β β (v2.0) β β
β β 1 replica β β
β ββββββββββββββββ β
β β
β 90% users β v1.0 β
β 10% users β v2.0 (testing) β
βββββββββββββββββββββββββββββββββββββββββββββββ
βββββββββββββββββββββββββββββββββββββββββββββββ
β PHASE 3: Increase Canary (50%) β
β ββββββββββββββββ ββββββββββββββββ β
β β Service ββββββ¬βββΆβ Stable Pods β β
β β β β β (v1.0) β β
β β β β β 5 replicas β β
β ββββββββββββββββ β ββββββββββββββββ β
β β β
β β ββββββββββββββββ β
β ββββΆβ Canary Pods β β
β β (v2.0) β β
β β 5 replicas β β
β ββββββββββββββββ β
β β
β 50% users β v1.0 β
β 50% users β v2.0 β
β (No errors? Increase more!) β
βββββββββββββββββββββββββββββββββββββββββββββββ
βββββββββββββββββββββββββββββββββββββββββββββββ
β PHASE 4: Full Canary (100%) β
β ββββββββββββββββ ββββββββββββββββ β
β β Service ββββββ¬βXββ Stable Pods β β
β β β β β (v1.0) β β
β β β β β 0 replicas β β
β ββββββββββββββββ β ββββββββββββββββ β
β β β
β β ββββββββββββββββ β
β ββββΆβ Canary Pods β β
β β (v2.0) β β
β β 10 replicas β β
β ββββββββββββββββ β
β β
β 100% users β v2.0 β
β
β Delete old stable deployment β
βββββββββββββββββββββββββββββββββββββββββββββββ
Traffic Split Calculation¶
Formula:
Examples:
| Stable Pods | Canary Pods | Total | Canary % | Use Case |
|---|---|---|---|---|
| 9 | 1 | 10 | 10% | Initial canary |
| 8 | 2 | 10 | 20% | Expand testing |
| 5 | 5 | 10 | 50% | Half and half |
| 2 | 8 | 10 | 80% | Nearly full |
| 0 | 10 | 10 | 100% | Complete rollout |
| 95 | 5 | 100 | 5% | Large-scale 5% canary |
Common canary percentages:
10% canary: 9 stable, 1 canary
20% canary: 4 stable, 1 canary
25% canary: 3 stable, 1 canary
50% canary: 1 stable, 1 canary
Canary vs Other Strategies¶
| Strategy | Traffic Split | Resource Usage | Rollback Speed | Risk Level |
|---|---|---|---|---|
| Canary | Gradual (10β50β100%) | 1-2x (during transition) | Fast (scale down) | Low (limited exposure) |
| Blue-Green | Instant (0β100%) | 2x (both versions) | Instant (switch selector) | Medium (all at once) |
| Rolling | Gradual (pod by pod) | 1x + surge | Medium (rollback) | Medium (gradual) |
| Recreate | Instant (downtime) | 1x | Slow (redeploy) | High (no testing) |
Real-World Incident Example¶
Company: E-learning platform (2M students) Impact: 8-hour performance degradation affecting 1M students Cost: $1.2M in refunds + massive support load
What happened:
The team developed a new version (v2.5) with a "performance optimization" - they changed the database query pattern. They planned a canary rollout to test it safely.
Intended plan: 10% canary (test with 200K students)
Actual configuration:
# Intended configuration (but NOT what was deployed)
# stable: 90 replicas
# canary: 10 replicas
# Actual deployed configuration:
apiVersion: apps/v1
kind: Deployment
metadata:
name: app-stable
spec:
replicas: 50 # β WRONG
---
apiVersion: apps/v1
kind: Deployment
metadata:
name: app-canary
spec:
replicas: 50 # β WRONG - should be ~5 for 10%
Someone misread the config and deployed 50/50 instead of 90/10!
The cascade (Monday, 8 AM - peak hours):
08:00 - Deploy canary (v2.5)
08:00 - Intended: 10% of traffic to canary
08:00 - Actual: 50% of traffic to canary (50/50 split) β
08:05 - 1M students (50%) now using v2.5
08:10 - Dashboard shows "error rate normal" (because they monitor v1 and v2 separately)
08:15 - Support tickets start flooding in:
"Assignments loading slowly"
"Videos won't load"
"Getting timeout errors"
08:20 - Check metrics: v2.5 has 10x higher latency than v1
08:20 - Realize: "Optimization" broke performance!
08:25 - Decision: Rollback canary
08:26 - Scale canary to 0: kubectl scale deployment app-canary --replicas=0
08:28 - All traffic back to v1
08:30 - Performance returns to normal
08:30 - But damage done: 1M students affected for 30 minutes
During prime homework submission time!
Why it was catastrophic: - Peak hours - 8 AM, students doing homework before school - 50% exposure - Should have been 10%, instead 1M students affected - Misleading metrics - Monitoring didn't aggregate or highlight the split - Performance regression - The "optimization" actually made things worse - Support overload - 5,000 tickets in 30 minutes
What should have happened:
With correct 10% canary:
08:00 - Deploy canary (10% traffic)
08:05 - 200K students on v2.5 (not 1M)
08:10 - Notice latency issue
08:15 - Scale canary to 0
08:16 - Only 200K affected, 1.8M never saw the issue
08:20 - Fix identified and patched
Root cause analysis:
-
Human error in configuration:
-
No validation of traffic split:
-
Insufficient monitoring - No alert for "canary traffic > 15%"
The fix implemented:
-
Automated traffic validation:
#!/bin/bash # canary-deploy.sh STABLE_REPLICAS=$1 CANARY_REPLICAS=$2 TARGET_PCT=$3 ACTUAL_PCT=$((CANARY_REPLICAS * 100 / (STABLE_REPLICAS + CANARY_REPLICAS))) if [ $ACTUAL_PCT -gt $((TARGET_PCT + 5)) ]; then echo "ERROR: Canary will be $ACTUAL_PCT%, target is $TARGET_PCT%" exit 1 fi kubectl scale deployment app-stable --replicas=$STABLE_REPLICAS kubectl scale deployment app-canary --replicas=$CANARY_REPLICAS -
Pre-deployment checks:
-
Real-time monitoring:
Lessons learned: 1. Validate traffic splits before deploying 2. Automate canary percentage calculations - don't rely on mental math 3. Monitor actual traffic distribution - alert if canary gets too much 4. Start very small - 5% or even 1% for risky changes 5. Have rollback automation - instant scale to 0
Commands You Mastered¶
# Scale deployments for canary
kubectl scale deployment app-stable --replicas=9 -n <namespace>
kubectl scale deployment app-canary --replicas=1 -n <namespace>
# Check replica counts
kubectl get deployments -n <namespace>
# Calculate actual traffic split
STABLE=$(kubectl get deployment app-stable -n <namespace> -o jsonpath='{.status.readyReplicas}')
CANARY=$(kubectl get deployment app-canary -n <namespace> -o jsonpath='{.status.readyReplicas}')
TOTAL=$((STABLE + CANARY))
CANARY_PCT=$((CANARY * 100 / TOTAL))
echo "Canary traffic: $CANARY_PCT%"
# Check service endpoints (see all pods)
kubectl get endpoints app-service -n <namespace>
# Test traffic distribution (sampling)
for i in {1..100}; do
kubectl run -it --rm test-$i --image=busybox --restart=Never -n <namespace> -- wget -q -O- app-service
done | grep -c "Canary"
# Quick rollback (scale canary to 0)
kubectl scale deployment app-canary --replicas=0 -n <namespace>
# Progressive rollout (increase canary)
kubectl scale deployment app-canary --replicas=2 -n <namespace> # 20%
# Monitor...
kubectl scale deployment app-canary --replicas=5 -n <namespace> # 50%
# Monitor...
kubectl scale deployment app-canary --replicas=10 -n <namespace> # 100%
kubectl scale deployment app-stable --replicas=0 -n <namespace> # Remove old
Best Practices for Canary Deployments¶
β DO:¶
-
Start small (5-10% canary):
-
Validate traffic split:
-
Monitor canary metrics separately:
-
Use gradual rollout:
-
Have automatic rollback:
β DON'T:¶
-
Don't start with 50/50 split:
-
Don't ignore canary metrics:
-
Don't rush through stages:
-
Don't forget rollback plan:
Canary Deployment Patterns¶
Pattern 1: Progressive Rollout¶
#!/bin/bash
# progressive-canary.sh
TOTAL_PODS=10
STAGES=(1 2 5 10) # 10%, 20%, 50%, 100%
for CANARY in "${STAGES[@]}"; do
STABLE=$((TOTAL_PODS - CANARY))
CANARY_PCT=$((CANARY * 100 / TOTAL_PODS))
echo "Deploying canary at $CANARY_PCT% ($CANARY pods)"
kubectl scale deployment app-stable --replicas=$STABLE
kubectl scale deployment app-canary --replicas=$CANARY
echo "Monitoring for 5 minutes..."
sleep 300
# Check error rate
ERROR_RATE=$(check_canary_errors)
if [ $ERROR_RATE -gt 1 ]; then
echo "ERROR RATE TOO HIGH! Rolling back..."
kubectl scale deployment app-canary --replicas=0
kubectl scale deployment app-stable --replicas=$TOTAL_PODS
exit 1
fi
done
echo "Canary successful! Cleaning up stable..."
kubectl delete deployment app-stable
Pattern 2: A/B Testing Canary¶
# Use different selectors for A and B testing
apiVersion: v1
kind: Service
metadata:
name: app-variant-a
spec:
selector:
app: myapp
variant: a # User segment A
---
apiVersion: v1
kind: Service
metadata:
name: app-variant-b
spec:
selector:
app: myapp
variant: b # User segment B (canary)
Pattern 3: Geographic Canary¶
# Deploy canary to specific region first
apiVersion: apps/v1
kind: Deployment
metadata:
name: app-canary-us-west
spec:
template:
spec:
nodeSelector:
region: us-west
# Canary version
---
# Stable in all other regions
apiVersion: apps/v1
kind: Deployment
metadata:
name: app-stable-global
spec:
# Stable version everywhere else
What's Next?¶
You've learned how to configure canary deployments with proper traffic splits for safe testing.
Next level: StatefulSet vs Deployment! You'll learn why using Deployments for stateful applications can cause data loss.
Key takeaway: Canary deployments must have correct replica ratios to limit user exposure. Always validate your traffic split before deploying!