When healthy CPU hides a dying dependency
Why cluster reliability reports should treat dependency lag as a first-class signal, not a footnote under infrastructure metrics.

Why cluster reliability reports should treat dependency lag as a first-class signal, not a footnote under infrastructure metrics.

Host charts are comforting. They move slowly, they look scientific, and they often have nothing to say about the outage your customers just felt.
In several Cluster Pulse Point assessments we have opened a reliability review to find CPU, memory, and disk all polite while the application cluster was drowning in synchronized waits on a single dependency. The application analytics that mattered were queue age, retry amplification, and the shape of timeouts — not the temperature of the boxes.
Ask which signals would have predicted the last three incidents before customer impact. If the honest answer is “none of the green charts,” your reliability report should say so. Sponsors can live with bad news; they struggle with surprises wrapped in healthy averages.
Our reports separate capacity headroom from dependency honesty. A cluster can have spare CPU and still be one slow certificate renewal away from a regional stumble. Naming that gap is part of cluster reliability reporting, even when it makes the infrastructure slides look incomplete.