Field note

Ranking reliability actions after the report

18 June 2026

A simple consequence-first ranking for actions that fall out of application analytics — so the backlog does not swallow the only fixes that matter.

Every assessment ends with more ideas than a team can finish. Without ranking, the safest items win: new dashboards, renamed alerts, a wiki page nobody will open during an incident.

We rank actions by consequence if left undone, then by effort only as a tie-breaker. An ugly one-line change that stops synchronized retries outranks a beautiful observability project that explains the next outage more eloquently.

Three buckets we use in reports

  1. Stop the bleeding — changes that remove a known amplification path
  2. See earlier — signals that would have predicted the last real incident
  3. Hardening — architecture moves that need a programme, not a sprint

Application analytics feed the first two buckets most often. The third belongs in the report so sponsors understand the horizon — not so the team pretends a quarterly redesign is this week’s chore.

After we leave

The clarification window is for misunderstandings in the text, not for reopening the ranking because a pet project fell into bucket three. If priorities truly change, commission a briefing; do not quietly edit the PDF.


Back to field notes