# Startup, Liveness and Readiness Probes — Kubernetes

Source: https://www.skillbyai.com/en/kubernetes/h-probes

> Different questions, different actions.

## Restart versus remove from traffic

**Readiness** probes ask "can this pod serve traffic now?"; failing removes the pod from Service endpoints without restarting it (useful during warm-up or temporary dependency issues). **Liveness** probes ask "is this container stuck?"; failing restarts the container. **Startup** probes protect slow-starting apps: liveness and readiness checks wait until the startup probe succeeds. Keep liveness checks simple and local (not dependent on databases), or a shared outage will cause restart storms.

## Stay healthy, scale with load

Probes tell Kubernetes about health; autoscalers and disruption budgets keep capacity right.

![Three ideas: probes, Horizontal Pod Autoscaling, disruption budgets.](assets/figures/kubernetes/section-6-map.svg) — Figure 6.1 — Probes, autoscaling and disruption budgets.

## How long before a probe acts? run

I ran this with Python 3. It is a simplified model of Kubernetes behaviour for learning, not the real controller code. With these settings, a startup probe allows up to 150 seconds, liveness restarts a container after about 30 seconds of failures, and readiness removes a pod from traffic after about 10 seconds.

```python
probes = {
    "startupProbe":   {"periodSeconds": 5, "failureThreshold": 30},
    "livenessProbe":  {"periodSeconds": 10, "failureThreshold": 3},
    "readinessProbe": {"periodSeconds": 5, "failureThreshold": 2},
}
for name, p in probes.items():
    print(f"{name:<15} acts after about {p['periodSeconds'] * p['failureThreshold']:>3} s of consecutive failures")
print("startup budget lets a slow app take up to 150 s to start before liveness checks begin")
print("liveness failure -> container restarted; readiness failure -> removed from Service endpoints, not restarted")
```

Output:

```
startupProbe    acts after about 150 s of consecutive failures
livenessProbe   acts after about  30 s of consecutive failures
readinessProbe  acts after about  10 s of consecutive failures
startup budget lets a slow app take up to 150 s to start before liveness checks begin
liveness failure -> container restarted; readiness failure -> removed from Service endpoints, not restarted
```

## Probe configuration

Not applied to a live cluster in this course; check field names against the API reference for your version.

```yaml
startupProbe:
  httpGet: {path: /healthz, port: 8080}
  periodSeconds: 5
  failureThreshold: 30        # up to 150 s to start
livenessProbe:
  httpGet: {path: /healthz, port: 8080}   # cheap, local check
  periodSeconds: 10
  failureThreshold: 3
readinessProbe:
  httpGet: {path: /ready, port: 8080}     # may check dependencies
  periodSeconds: 5
  failureThreshold: 2
```

**Quiz:** What happens when a readiness probe fails?

- [ ] The node is drained
- [ ] The container is restarted
- [x] The pod is removed from Service endpoints but not restarted
- [ ] The Deployment is deleted

*Answer:* The pod is removed from Service endpoints but not restarted. Readiness controls traffic; liveness controls restarts.
