Evidence

Client stories

Evidence from engagements where cluster reliability reporting had to survive messy access and conflicting memories.

Testimonials

“They spent three days in our staging and production notes before writing a word. The reliability report named the retry storm we had been blaming on the load balancer — awkward to read, but it ended a month of circular arguments.”

— Mira K., platform lead, Melbourne payments team

“Useful assessment, though the first draft assumed more access to our runbooks than we could grant. Once we clarified that, the final brief was fair and specific.”

— Tom R., SRE manager, Brisbane logistics

“The pattern review showed that three ‘different’ outages shared the same cold-cache path after deploys. We changed the release checklist the same week.”

— Anika S., engineering manager, Adelaide retail stack

“Helen’s reading with our risk committee was calm. No theatre, just the consequence of leaving the secondary cluster on a neglected certificate path.”

— David L., head of operations, Sydney insurer

Extended story: payments cluster, Melbourne

A payments platform team asked for a cluster reliability assessment after four partial outages in a quarter. Monitoring showed healthy CPU; customer support showed otherwise. During evidence gathering we found application analytics pointing to synchronized retries against a dependency that only slowed under settlement-window load.

The draft report was contested — several engineers believed the dependency vendor was solely at fault. The owner reading walked through timestamps side by side with the dependency’s own status page. The final report kept the vendor issue but added the client’s retry amplification as a co-equal reliability risk. Remediation (jittered backoff and a circuit on the settlement path) sat with the client; our role ended at the signed report and a thirty-day clarification window.

Extended story: logistics failover, Brisbane

A logistics group commissioned an incident pattern review after weekend failover drills kept failing for “different” reasons. Pattern clustering showed a single missing health-check on a message consumer that only ran during regional failover. The brief was six pages. The team used it to rewrite one runbook section and to justify a follow-on full assessment later that year.