Closing the detection gap
Microsoft - sovereign cloud
Incidents were reaching us the wrong way round: customers noticed regressions before our monitoring did. Auto-mitigation sat at 0%, so every incident consumed a human, and the humans were being paged by the people they were supposed to be protecting.
Not more alerting - 36 production alert rules designed as a system. CDN and traffic-manager coverage mirrored across two sovereign clouds on a tiered warn/elevated/critical model, each rule routed to an owner and hardened against schema drift so it could not quietly stop firing.
Auto-mitigation reached a peak of ~95%, with auto-resolution rising into a 39-65% band. The band is wide on purpose: full auto-resolution only ever went to failure modes with a known-safe recovery, and the rest still stop at a human. A machine that resolves an incident it does not understand is not reliability, it is a slower outage.


