π― Level 28 Debrief: Service Endpoints & Readiness Probes¶
Congratulations! You've mastered one of the most critical concepts in production Kubernetes: readiness probes and service endpoint management. This seemingly simple configuration detail has prevented countless outages and saved millions in revenue.
π What You Just Fixed¶
The Problem¶
Your service was routing traffic to pods before they were ready to handle requests, causing: - 500 errors during pod initialization - Failed requests hitting pods still loading data - Inconsistent behavior as some pods worked while others didn't - Poor user experience with intermittent failures
The Root Cause¶
# β BROKEN: No readiness probe
apiVersion: v1
kind: Pod
metadata:
name: web-app-1
spec:
containers:
- name: web
image: nginx:1.21
# Missing readinessProbe!
# Pod is added to endpoints IMMEDIATELY
Why this fails: 1. Pod starts and gets an IP address 2. Kubernetes immediately adds pod to service endpoints 3. Service starts routing traffic to the pod 4. Pod is still initializing (loading config, warming cache, etc.) 5. Requests fail with 500 errors or timeouts
The Solution¶
# β
FIXED: Readiness probe configured
apiVersion: v1
kind: Pod
metadata:
name: web-app-1
spec:
containers:
- name: web
image: nginx:1.21
readinessProbe:
httpGet:
path: /
port: 8080
initialDelaySeconds: 5
periodSeconds: 5
failureThreshold: 1
How this works:
1. Pod starts but is marked "Not Ready"
2. After 5 seconds (initialDelaySeconds), Kubernetes checks / endpoint
3. Every 5 seconds (periodSeconds), the probe runs again
4. If the probe succeeds, pod is marked "Ready"
5. Only then is the pod added to service endpoints
6. Traffic flows only to ready pods
π Deep Dive: Readiness vs Liveness Probes¶
Kubernetes has three types of health checks, each serving a different purpose:
1. Readiness Probe (What You Just Used)¶
Purpose: Determine if a pod is ready to receive traffic
When it fails: - Pod is removed from service endpoints - Pod continues running - No traffic is sent to the pod - Kubernetes keeps checking until pod recovers
Use cases: - Application initialization (loading config, warming cache) - External dependency checks (database connection, API availability) - Temporary overload (too many connections, high CPU) - Graceful degradation (shed load during issues)
Example:
readinessProbe:
httpGet:
path: /healthz/ready
port: 8080
initialDelaySeconds: 10
periodSeconds: 5
timeoutSeconds: 1
failureThreshold: 3
2. Liveness Probe¶
Purpose: Determine if a pod is alive and healthy
When it fails: - Pod is restarted (killed and recreated) - Drastic action for deadlocked or hung processes - Traffic may be disrupted during restart
Use cases: - Deadlock detection (app is running but not responding) - Memory leaks (app degraded beyond recovery) - Fatal errors (app in unrecoverable state)
Example:
livenessProbe:
httpGet:
path: /healthz/live
port: 8080
initialDelaySeconds: 30
periodSeconds: 10
timeoutSeconds: 1
failureThreshold: 3
β οΈ Warning: Be careful with liveness probes! Too aggressive settings cause restart loops.
3. Startup Probe (Kubernetes 1.18+)¶
Purpose: Handle slow-starting applications
When it fails: - Pod is restarted (like liveness probe) - Disables liveness probe until startup succeeds - Prevents premature restarts during long initialization
Use cases: - Legacy applications with long startup times - Applications with unpredictable initialization duration - Prevent liveness probe from killing slow-starting pods
Example:
startupProbe:
httpGet:
path: /healthz/startup
port: 8080
initialDelaySeconds: 0
periodSeconds: 10
failureThreshold: 30 # 30 * 10 = 300 seconds max startup time
Comparison Table¶
| Feature | Readiness Probe | Liveness Probe | Startup Probe |
|---|---|---|---|
| Action on Failure | Remove from endpoints | Restart pod | Restart pod |
| Traffic Impact | No traffic sent | Traffic disrupted | Traffic disrupted |
| Use Case | Temporary issues | Fatal errors | Slow startup |
| Aggressiveness | Can be frequent | Must be conservative | One-time check |
| Recovery | Automatic | Requires restart | Requires restart |
π Real-World Horror Story: The $1.2M Endpoint Incident¶
Company: StreamVideo Inc. (Video streaming platform) Date: November 2022 Duration: 45 minutes Impact: $1.2M in lost revenue + 250,000 angry users
The Setup¶
StreamVideo ran a massive video transcoding service on Kubernetes: - 500 pods handling video uploads and transcoding - Each pod needed to load a 2GB ML model on startup - Average startup time: 60-90 seconds - Black Friday promotion driving 10x normal traffic
The Mistake¶
The deployment looked perfect:
apiVersion: apps/v1
kind: Deployment
metadata:
name: transcoder
spec:
replicas: 500
template:
spec:
containers:
- name: transcoder
image: streamvideo/transcoder:v2.3
# β NO READINESS PROBE!
livenessProbe:
httpGet:
path: /health
port: 8080
initialDelaySeconds: 120
periodSeconds: 30
What was missing: Readiness probe! Only liveness probe was configured.
The Incident Timeline¶
11:47 AM - The Rollout Begins - New deployment rolls out with improved ML model - Rolling update strategy: 25% pods at a time - First 125 pods are terminated, new pods start
11:48 AM - Traffic Hits Unprepared Pods - New pods get IP addresses and are added to service endpoints immediately - Load balancer starts sending 25% of traffic to new pods - Pods are still loading the 2GB ML model (45 seconds into 90-second startup) - Users get 500 Internal Server Error responses
11:49 AM - Cascade Begins - 25% of all uploads failing - Users retry, doubling the load - Existing healthy pods overwhelmed with retries - Some healthy pods start becoming slow due to overload
11:51 AM - Full Impact - Second wave: 125 more pods replaced - Now 50% of pods are unready but receiving traffic - 50% failure rate across entire platform - Social media explodes with complaints - #StreamVideoDown trending on Twitter
11:55 AM - The "Fix" Makes It Worse - On-call engineer sees high CPU on healthy pods - Scales deployment from 500 to 700 pods (thinking it's a capacity issue) - 200 new pods start, all receiving traffic immediately - Failure rate jumps to 65% - Now trending #1 on Twitter
12:03 PM - The Realization - Senior engineer joins war room - Checks service endpoints: sees unready pods receiving traffic - Realizes: no readiness probe configured!
12:05 PM - Emergency Fix
# Emergency patch applied
kubectl patch deployment transcoder -p '{
"spec": {
"template": {
"spec": {
"containers": [{
"name": "transcoder",
"readinessProbe": {
"httpGet": {
"path": "/ready",
"port": 8080
},
"initialDelaySeconds": 100,
"periodSeconds": 10
}
}]
}
}
}
}'
12:08 PM - Rollout Continues - New pods now wait 100 seconds before being marked ready - Only ready pods added to endpoints - Failure rate drops to 15%
12:15 PM - Manual Intervention - Engineer manually removes all unready pods from endpoints:
# Get all unready pods
kubectl get pods -l app=transcoder -o json | \
jq -r '.items[] | select(.status.conditions[] | select(.type=="Ready" and .status=="False")) | .metadata.name' | \
while read pod; do
# Force delete to speed up replacement
kubectl delete pod $pod --grace-period=0 --force
done
12:32 PM - Full Recovery - All pods healthy and ready - Traffic normalized - Incident declared resolved
The Damage¶
- $1.2M in lost revenue (45 minutes during Black Friday peak)
- 250,000 users affected (couldn't upload videos)
- 12,000 support tickets generated
- #1 trending on Twitter for wrong reasons
- 15% customer churn in following week
The Root Cause Analysis¶
Immediate cause: - Missing readiness probe allowed traffic to unready pods
Contributing factors: 1. No pre-production testing: Startup time never measured in staging 2. Aggressive rollout: 25% at a time too fast for slow-starting pods 3. Panic scaling: Scaling up made problem worse 4. Inadequate monitoring: No alerting on 5xx error rates 5. No runbook: Team didn't know how to diagnose endpoint issues
The fix that should have been there:
apiVersion: apps/v1
kind: Deployment
metadata:
name: transcoder
spec:
replicas: 500
strategy:
rollingUpdate:
maxUnavailable: 10% # β
Slower rollout
maxSurge: 10%
template:
spec:
containers:
- name: transcoder
image: streamvideo/transcoder:v2.3
# β
Startup probe for slow initialization
startupProbe:
httpGet:
path: /healthz/startup
port: 8080
initialDelaySeconds: 30
periodSeconds: 10
failureThreshold: 12 # 30s + (10s * 12) = 150s max startup
# β
Readiness probe to control traffic
readinessProbe:
httpGet:
path: /healthz/ready
port: 8080
initialDelaySeconds: 5
periodSeconds: 5
successThreshold: 1
failureThreshold: 3
# β
Liveness probe to detect deadlocks
livenessProbe:
httpGet:
path: /healthz/live
port: 8080
initialDelaySeconds: 150
periodSeconds: 30
failureThreshold: 3
Lessons Learned¶
- Always use readiness probes - Never assume pods are ready immediately
- Measure startup time - Know how long your app takes to initialize
- Test rollouts - Simulate deployments in staging with realistic load
- Monitor endpoints - Alert on unready pods receiving traffic
- Slow rollouts - Use conservative maxUnavailable/maxSurge settings
- Separate health endpoints - Different paths for startup/readiness/liveness
- Document procedures - Have runbooks for common issues
π¬ Understanding Service Endpoint Management¶
How Endpoints Work¶
When you create a Service, Kubernetes automatically creates an Endpoints object:
# Service definition
apiVersion: v1
kind: Service
metadata:
name: web-backend
spec:
selector:
app: web
ports:
- port: 80
targetPort: 8080
Kubernetes creates:
# Automatically created Endpoints object
apiVersion: v1
kind: Endpoints
metadata:
name: web-backend # Same name as service
subsets:
- addresses:
- ip: 10.244.1.5 # Pod web-app-1 (READY)
- ip: 10.244.2.8 # Pod web-app-2 (READY)
ports:
- port: 8080
The Endpoint Controller Loop¶
Kubernetes runs a continuous control loop:
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β ENDPOINT CONTROLLER β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β
βΌ
ββββββββββββββββββββββββββββββββββββββββββββββ
β 1. Watch all Pods matching service β
β selector (app=web) β
ββββββββββββββββββββββββββββββββββββββββββββββ
β
βΌ
ββββββββββββββββββββββββββββββββββββββββββββββ
β 2. For each Pod, check Ready condition β
β - Look at status.conditions[] β
β - type: Ready, status: True/False β
ββββββββββββββββββββββββββββββββββββββββββββββ
β
βΌ
ββββββββββββββββββββββββββββββββββββββββββββββ
β 3. Update Endpoints object β
β - Ready pods β addresses[] β
β - Not Ready pods β notReadyAddresses[] β
ββββββββββββββββββββββββββββββββββββββββββββββ
β
βΌ
ββββββββββββββββββββββββββββββββββββββββββββββ
β 4. kube-proxy watches Endpoints β
β - Updates iptables/ipvs rules β
β - Traffic only to addresses[] β
ββββββββββββββββββββββββββββββββββββββββββββββ
Pod Ready Condition¶
The "Ready" condition is determined by:
# Without readiness probe - ALWAYS ready once running
status:
conditions:
- type: Ready
status: "True" # β Immediately true!
# With readiness probe - ready only after probe succeeds
status:
conditions:
- type: Ready
status: "False" # Initially false
# ... time passes, probe runs ...
- type: Ready
status: "True" # β
True after probe succeeds
Viewing Endpoint Status¶
# See which pods are in endpoints
kubectl get endpoints web-backend -o yaml
# Watch endpoint changes in real-time
kubectl get endpoints web-backend -w
# See detailed pod ready status
kubectl get pods -o custom-columns=\
NAME:.metadata.name,\
READY:.status.conditions[?(@.type==\"Ready\")].status,\
IP:.status.podIP
π― Probe Configuration Best Practices¶
1. Choose the Right Probe Type¶
HTTP GET Probe (Most Common)
readinessProbe:
httpGet:
path: /healthz/ready
port: 8080
httpHeaders:
- name: Custom-Header
value: Awesome
initialDelaySeconds: 5
periodSeconds: 10
TCP Socket Probe
- β Good for databases, TCP services - β Simple, low overhead - β Only checks if port is open (not actual health)Exec Probe
readinessProbe:
exec:
command:
- /bin/sh
- -c
- "pg_isready -U postgres"
initialDelaySeconds: 5
periodSeconds: 10
2. Tune Timing Parameters¶
readinessProbe:
httpGet:
path: /ready
port: 8080
initialDelaySeconds: 10 # Wait before first probe
periodSeconds: 5 # How often to probe
timeoutSeconds: 1 # Probe timeout
successThreshold: 1 # Successes needed to mark ready
failureThreshold: 3 # Failures needed to mark not ready
Guidelines:
- initialDelaySeconds: Set to 80% of typical startup time
- periodSeconds: 5-10 seconds for readiness, 10-30 for liveness
- timeoutSeconds: Conservative (1-3 seconds)
- failureThreshold: 3 for readiness (quick removal), 3+ for liveness (avoid restarts)
3. Implement Smart Health Endpoints¶
Bad Health Endpoint:
Good Readiness Endpoint:
@app.route('/healthz/ready')
def ready():
# Check database
try:
db.execute("SELECT 1")
except:
return "Database not ready", 503
# Check cache
if not cache.ping():
return "Cache not ready", 503
# Check critical dependencies
if not critical_service_available():
return "Dependencies not ready", 503
# Check if too overloaded
if get_active_connections() > 1000:
return "Too busy", 503
return "OK", 200 # β
Comprehensive check
Good Liveness Endpoint:
@app.route('/healthz/live')
def live():
# Only check if app is alive (not hung/deadlocked)
# Don't check dependencies - those are readiness concerns
if is_deadlocked():
return "Deadlock detected", 500
if memory_usage() > 95:
return "Out of memory", 500
return "OK", 200 # β
Only critical checks
4. Common Patterns¶
Pattern 1: Gradual Readiness
# Warm up gradually
startup_time = time.time()
warmup_duration = 30 # 30 seconds
@app.route('/healthz/ready')
def ready():
elapsed = time.time() - startup_time
# Phase 1: Initial startup (0-10s)
if elapsed < 10:
return "Still starting", 503
# Phase 2: Warm cache (10-30s)
if elapsed < warmup_duration:
if cache_warmup_progress() < 100:
return "Warming cache", 503
# Phase 3: Ready for production traffic
return "OK", 200
Pattern 2: Dependency Checks
@app.route('/healthz/ready')
def ready():
dependencies = {
'database': check_database(),
'cache': check_cache(),
'api': check_external_api(),
}
failed = [k for k, v in dependencies.items() if not v]
if failed:
return f"Not ready: {', '.join(failed)}", 503
return "OK", 200
Pattern 3: Circuit Breaker Integration
circuit_breaker = CircuitBreaker()
@app.route('/healthz/ready')
def ready():
# If circuit is open, don't accept traffic
if circuit_breaker.is_open():
return "Circuit breaker open", 503
return "OK", 200
π οΈ Debugging Endpoint Issues¶
Symptom 1: "Service returns 500 errors intermittently"¶
Diagnosis:
# Check if all pods are ready
kubectl get pods -l app=web
# Check endpoint status
kubectl get endpoints web-backend
# See which pods are in/out of endpoints
kubectl describe endpoints web-backend
Common causes: - Pods without readiness probes - Readiness probe too aggressive (timing out) - Application not responding to probe path
Fix:
# Check probe configuration
kubectl get pod web-app-1 -o yaml | grep -A 10 readinessProbe
# Check probe logs
kubectl logs web-app-1 | grep -i ready
# Manually test probe endpoint
kubectl exec web-app-1 -- curl localhost:8080/healthz/ready
Symptom 2: "Pod never becomes ready"¶
Diagnosis:
# Check pod events
kubectl describe pod web-app-1
# Check readiness probe failures
kubectl get pod web-app-1 -o yaml | grep -A 20 "conditions:"
# Test probe manually
kubectl exec web-app-1 -- wget -O- http://localhost:8080/healthz/ready
Common causes: - Wrong probe path or port - Application not listening on expected port - Probe timeout too short - Probe starts before app is ready (initialDelaySeconds too small)
Symptom 3: "Endpoints keep changing"¶
Diagnosis:
# Watch endpoints in real-time
kubectl get endpoints web-backend -w
# Check pod restarts
kubectl get pods -l app=web -o custom-columns=\
NAME:.metadata.name,\
RESTARTS:.status.containerStatuses[0].restartCount
# Check for failing probes
kubectl describe pods -l app=web | grep -i "readiness probe failed"
Common causes: - Readiness probe flapping (intermittent failures) - Resource constraints causing slowdowns - Network issues affecting probe
π Key Takeaways¶
Must Remember¶
- Readiness probes control traffic - Pods without readiness probes receive traffic immediately
- Liveness probes restart pods - Use conservatively to avoid restart loops
- Startup probes for slow apps - Handle long initialization times
- Endpoints = Ready pods - Only ready pods receive traffic
- Test your probes - Verify they work correctly before production
Production Checklist¶
Before deploying to production:
- Readiness probe configured on all pods
- Liveness probe configured (if needed)
- Startup probe configured (for slow-starting apps)
- Probe timings tested under load
- Health endpoints implemented with real checks
- Monitoring configured for probe failures
- Alerts set up for endpoint changes
- Rollout strategy conservative (slow maxUnavailable)
- Runbook created for probe troubleshooting
- Load tested with probe configurations
Common Mistakes to Avoid¶
β No readiness probe - Pods receive traffic immediately β Liveness = Readiness - Using same endpoint for both (restart on temp issues) β Aggressive timings - Too short periods cause flapping β Fake health checks - Always returning 200 OK β Ignoring startup time - initialDelaySeconds too small β No dependency checks - Not verifying database, cache, etc. β Probe endpoint errors - Health endpoint crashes the app
π What's Next?¶
You've mastered readiness probes and endpoint management. Next, you'll tackle:
Level 29: LoadBalancer vs NodePort - Understanding service types and cloud provider integration Level 30: Headless Services - StatefulSet DNS and direct pod access
Further Learning¶
- Kubernetes Documentation: Configure Liveness, Readiness and Startup Probes
- Best Practices: Health Checks in Production
- Deep Dive: Endpoint Slices (for large-scale services)
π Achievement Unlocked!¶
Endpoint Master - You can now: - β Configure readiness probes to control traffic flow - β Understand the difference between readiness, liveness, and startup probes - β Implement smart health check endpoints - β Debug service endpoint issues - β Prevent production outages from unready pods
Remember: A simple readiness probe can prevent million-dollar outages. Never deploy a pod without one!
"In production, traffic should only flow to pods that are ready. Readiness probes are not optionalβthey're essential." - Kubernetes SRE Handbook