Writing for the people who hold the pager
Engineering notes, operating research and regulatory analysis from the team building the platform — and occasionally from the customers running it.
Blast radius is the only change metric that matters
Change failure rate tells you how often you were wrong. Blast radius tells you how much it cost when you were. Only one of them is something you can decide in advance.
Why every CMDB drifts, and what to do instead
A configuration database records intent. Production records outcome. The gap is not a data-quality problem to be solved — it is a signal to be measured continuously.
Runbooks that write themselves still need a reviewer
Generated procedures are only as safe as the review step that follows them. What we learned putting a human gate in front of every AI-authored runbook.
Quorum-aware patching for availability groups
A technical walk-through of how wave planning derives sequencing constraints from replication topology, and why naive parallelism loses quorum at scale.
Meeting DORA obligations without a two-year programme
Most of what the regulation asks for is evidence you already generate. The problem is that it lives in six systems and nobody can assemble it on demand.
The real cost of an unattended restart at 2am
We modelled the fully-loaded cost of out-of-hours remediation across 40 enterprise estates. The engineer hours are not the expensive part.
One email a month, no product announcements
A summary of what we published, what we got wrong, and what we learned from customer estates. Unsubscribe in one click.
See it run against your own estate
A proof of concept takes two weeks. We deploy inside your network, discover a scoped part of your estate, and run a real change end to end — with your team holding the approvals.