Rubric: L-OBSERVABILITY
Lens: L-OBSERVABILITY ยท Density: D-MED ยท Axes: OBS, OPS, RCV
Criteria
- Can a stranger answer: is it healthy, what failed, what to do next?
- Structured logs vs printf soup; correlation IDs.
- Metrics/traces presence and cardinality sanity.
- Honesty of status endpoints (no false green).
- Runbooks / on-call affordances in-repo.
Citations
| Work | Point |
|---|---|
| Google SRE Book โ Monitoring Distributed Systems | Four golden signals (adapt to monolith/CLI) |
| OpenTelemetry conceptual docs | Traces/metrics/logs correlation |
| ISO/IEC 25010 | Operability-related qualities |
Pros / cons
| Pros | Cons |
|---|---|
| Catches false-healthy dashboards | Easy to demand enterprise APM for a library โ calibrate |
| Status honesty is high leverage | Log volume โ observability |
False green
If status claims OK while a subsystem timed out or was skipped, severity โฅ high under OBS (and CMP if advertised).