Field notes
Short pieces on application analytics choices and the craft of writing cluster reliability reports.
Short pieces on application analytics choices and the craft of writing cluster reliability reports.
Notes from engagements and roundtables. Nothing here replaces a formal assessment; these pieces share vocabulary and judgement calls we see repeatedly.
Failover exercises produce some of the richest material for cluster reliability reporting — if you capture more than a pass/fail tick.
Read the note
A simple consequence-first ranking for actions that fall out of application analytics — so the backlog does not swallow the only fixes that matter.
Read the noteHow we run draft reviews so technical contacts can correct facts without sanding consequence out of a reliability report.
Read the note
Why cluster reliability reports should treat dependency lag as a first-class signal, not a footnote under infrastructure metrics.
Read the note