Implementing Observability Practices
1. Understanding Three Pillars
| Pillar | Question Answered | Tool |
| Metrics | What is happening (aggregate)? | Prometheus, Datadog |
| Logs | What happened (per event)? | Loki, ELK |
| Traces | Where did time go (per request)? | Jaeger, Tempo |
| + Profiles | Why is this code slow? | Pyroscope, Parca |
2. Implementing Health Check Endpoints
Example: Spring Boot Actuator
GET /actuator/health
{
"status": "UP",
"components": {
"db": { "status": "UP" },
"redis": { "status": "UP" },
"diskSpace": { "status": "UP", "details": { "free": "12GB" } }
}
}
3. Implementing Readiness and Liveness Probes
| Probe | Action on Fail | Best Practice |
| Liveness | Restart container | Shallow check (process responsive) |
| Readiness | Remove from Service endpoints | Includes critical dependencies |
| Startup | Delay liveness during boot | For slow-starting apps |
Warning: Liveness probes that check downstream dependencies cause cascading restarts during outages.
| Type | Tool |
| CPU sampling | async-profiler, perf, pprof |
| Allocation | JFR, async-profiler --alloc |
| Lock contention | JFR, async-profiler --lock |
| Continuous | Pyroscope, Parca, Datadog Continuous Profiler |
5. Understanding Black-Box vs White-Box Monitoring
Black-Box
- External probe; no app knowledge
- Synthetic checks (Pingdom)
- Detects user-visible failures
White-Box
- Internal metrics, traces, logs
- Predicts failures
- Diagnoses root cause
6. Implementing Synthetic Monitoring
| Aspect | Detail |
| Purpose | Continuous probe critical user journeys |
| Tools | Datadog Synthetics, Pingdom, Grafana Cloud, k6 cloud |
| Frequency | 1-5 min per check |
| Multi-region | Detect localized outages |
7. Implementing Real User Monitoring (RUM)
| Metric | Detail |
| Core Web Vitals | LCP, FID/INP, CLS |
| Page load | FCP, TTFB |
| JS errors | window.onerror, unhandledrejection |
| Tools | Datadog RUM, NR Browser, Sentry, SpeedCurve |
8. Implementing Error Tracking and Alerting
| Feature | Detail |
| Group similar errors | By stack trace fingerprint |
| Release tracking | Assign to deploy / commit |
| Source maps | De-minify JS stacks |
| Tools | Sentry, Rollbar, Bugsnag |
9. Understanding Observability-Driven Development
| Practice | Detail |
| Instrument first | Add metrics/traces with new code |
| High-cardinality events | Honeycomb-style wide events |
| Test in production | Feature flags + monitoring |
| Debug unknown unknowns | Slice by any dimension |
10. Implementing Debugging in Production
| Tool | Use |
| Dynamic logging | Toggle DEBUG without restart |
| Tracing slow requests | Tail-based sampling captures outliers |
| Live debugger | Lightrun, Rookout (snapshot at code line) |
| Heap / thread dump | jcmd, jstack |
| eBPF tracing | bpftrace one-liners |