Lesson 17 / 25
Startup, Liveness and Readiness Probes
Different questions, different actions.
Restart versus remove from traffic
Readiness probes ask "can this pod serve traffic now?"; failing removes the pod from Service endpoints without restarting it (useful during warm-up or temporary dependency issues). Liveness probes ask "is this container stuck?"; failing restarts the container. Startup probes protect slow-starting apps: liveness and readiness checks wait until the startup probe succeeds. Keep liveness checks simple and local (not dependent on databases), or a shared outage will cause restart storms.
Stay healthy, scale with load
Probes tell Kubernetes about health; autoscalers and disruption budgets keep capacity right.
How long before a probe acts? run
I ran this with Python 3. It is a simplified model of Kubernetes behaviour for learning, not the real controller code. With these settings, a startup probe allows up to 150 seconds, liveness restarts a container after about 30 seconds of failures, and readiness removes a pod from traffic after about 10 seconds.
probes = {
"startupProbe": {"periodSeconds": 5, "failureThreshold": 30},
"livenessProbe": {"periodSeconds": 10, "failureThreshold": 3},
"readinessProbe": {"periodSeconds": 5, "failureThreshold": 2},
}
for name, p in probes.items():
print(f"{name:<15} acts after about {p['periodSeconds'] * p['failureThreshold']:>3} s of consecutive failures")
print("startup budget lets a slow app take up to 150 s to start before liveness checks begin")
print("liveness failure -> container restarted; readiness failure -> removed from Service endpoints, not restarted")
Output:
startupProbe acts after about 150 s of consecutive failures livenessProbe acts after about 30 s of consecutive failures readinessProbe acts after about 10 s of consecutive failures startup budget lets a slow app take up to 150 s to start before liveness checks begin liveness failure -> container restarted; readiness failure -> removed from Service endpoints, not restarted
Probe configuration
Not applied to a live cluster in this course; check field names against the API reference for your version.
startupProbe:
httpGet: {path: /healthz, port: 8080}
periodSeconds: 5
failureThreshold: 30 # up to 150 s to start
livenessProbe:
httpGet: {path: /healthz, port: 8080} # cheap, local check
periodSeconds: 10
failureThreshold: 3
readinessProbe:
httpGet: {path: /ready, port: 8080} # may check dependencies
periodSeconds: 5
failureThreshold: 2Quick check: What happens when a readiness probe fails?
- The node is drained
- The container is restarted
- The pod is removed from Service endpoints but not restarted
- The Deployment is deleted
Answer
The pod is removed from Service endpoints but not restarted — Readiness controls traffic; liveness controls restarts.