Implementing Health Checks
1. Configuring Liveness Probes
Restart container if probe fails. Catches deadlocks where process is alive but stuck.
livenessProbe:
httpGet: { path: /healthz, port: 8080 }
initialDelaySeconds: 15
periodSeconds: 10
failureThreshold: 3
| Field | Default |
| initialDelaySeconds | 0 |
| periodSeconds | 10 |
| timeoutSeconds | 1 |
| failureThreshold | 3 |
| successThreshold | 1 (always for liveness) |
2. Configuring Readiness Probes
Remove pod from Service endpoints when failing. No restart; lets pod warm up or shed load.
readinessProbe:
httpGet: { path: /ready, port: 8080 }
periodSeconds: 5
failureThreshold: 3
successThreshold: 1
| Use Case | Pattern |
| Cache warmup | Return 503 until cache loaded |
| DB connection | Probe fails until DB pool ready |
| Graceful shutdown | Set unready before SIGTERM |
3. Configuring Startup Probes
Disables liveness/readiness until startup succeeds. Lets slow-starting apps avoid premature restarts.
startupProbe:
httpGet: { path: /healthz, port: 8080 }
failureThreshold: 30 # 30 * 10s = 5 min max startup
periodSeconds: 10
| Property | Note |
| Runs once | Until first success; then liveness/readiness take over |
| Best for | JVMs, legacy apps with long init |
4. Using HTTP Probes
| Field | Use |
| path | HTTP path |
| port | Number or named port |
| scheme | HTTP (default) or HTTPS |
| httpHeaders | Custom headers (e.g. Host) |
| Success codes | 200-399 |
5. Using TCP Probes
readinessProbe:
tcpSocket: { port: 5432 }
periodSeconds: 5
| Use | Caveat |
| DB/queue connectivity | Only checks TCP accept; doesn't validate app health |
6. Using Exec Probes
livenessProbe:
exec: { command: ["pg_isready","-U","app"] }
periodSeconds: 10
| Result | Outcome |
| exit 0 | Healthy |
| non-zero | Unhealthy |
| timeout | Unhealthy (kills process) |
Warning: Exec probes fork a process every period; expensive at scale.
7. Using gRPC Probes
livenessProbe:
grpc: { port: 9000, service: "" } # empty = overall server
| Requirement | Detail |
| App support | Implement grpc.health.v1.Health |
| Stable | Since v1.27 |
8. Setting Initial Delay
| Probe | Typical initialDelaySeconds |
| Fast service | 5 |
| JVM/Node app | 15-30 |
| Slow startup | Use startupProbe instead |
9. Configuring Period Seconds
| Probe Type | Typical periodSeconds |
| Liveness | 10-30 (avoid CPU overhead) |
| Readiness | 5-10 (fast feedback for LB) |
| Startup | 5-10 with high failureThreshold |
10. Setting Failure Threshold
| Probe | Effect of Reaching Threshold |
| Liveness | Restart container |
| Readiness | Remove from Service endpoints |
| Startup | Restart container (init failed) |
11. Setting Success Threshold
| Probe | Allowed Values |
| Liveness / Startup | Must be 1 |
| Readiness | ≥1; require N consecutive successes after failure |
12. Configuring Timeout Seconds
| Property | Note |
| Default | 1s (often too tight) |
| Recommendation | ≥3s for HTTP; ≥5s for exec |
| Tip | Probe timeout should be < periodSeconds |
Probe Cheat Sheet
- Liveness = "restart me when broken" — use sparingly
- Readiness = "stop sending traffic" — essential for rolling updates
- Startup = "wait before judging" — for slow boots
- Never probe external dependencies in liveness (cascading restarts)