preloader

· azure cloud devops incident-management resilience microsoft infrastructure

A Single Network Change Took Down Azure West US for Four Hours, and Sentinel Went Down With It

Source: Data Center Dynamics / Azure Status History

At approximately 14:44 UTC on July 23, Microsoft customers connected to the Azure West US region began experiencing intermittent connectivity failures and elevated latency. What started as a regional network issue spread quickly: Microsoft’s status page eventually listed more than twenty affected services, and the outage rippled into Microsoft 365, Teams, Outlook, SharePoint, OneDrive, Copilot, Xbox Live and the Microsoft Store.

A network change, not a single service failure

Microsoft traced the disruption to a recent change in West US network infrastructure and resolved it by rolling that change back, with telemetry showing recovery by 18:47 UTC, roughly four hours after the incident began. Because the fault sat in shared network infrastructure rather than in a single application platform or database engine, the blast radius extended across services that, on paper, share little beyond the same region. Among the casualties: Azure Kubernetes Service, Azure Virtual Desktop, ExpressRoute, Application Gateway, and Microsoft Sentinel, Microsoft’s own cloud security monitoring product.

The part worth pausing on: your SIEM went down too

A four-hour regional outage is disruptive on its own. What makes this one worth a closer look for security teams is that Sentinel was among the affected services. If your organisation relies on Sentinel as its primary detection and alerting layer, a regional network fault does not just take your applications offline, it can simultaneously blind the tooling you would use to notice anything unusual happening during the disruption itself. That is a single point of failure that rarely shows up in an architecture review, because monitoring infrastructure is usually assumed to be the one thing that stays up.

The second major outage in a month

This is not an isolated event. Azure West US 2 was offline for roughly fifteen hours after lightning strikes hit multiple datacentre buildings simultaneously earlier this month, and now West US has had a four-hour network-driven outage of its own. Two significant regional incidents in close succession is a pattern worth factoring into resilience planning, particularly for European organisations running workloads or disaster-recovery targets in US Azure regions as part of a global architecture.

If you want an honest assessment of what actually happens to your monitoring and detection capability during a cloud provider outage, not just your applications, contact Excello Digital. We help teams design multi-region and multi-cloud resilience plans that account for the dependencies that only reveal themselves once something breaks.

These news items are automatically aggregated from industry sources and are not individually reviewed. Any inaccuracies are unintentional — let us know and we'll correct or remove it.

We’ll help you resolve your infrastructure challenges

Our team of experts is ready to help you with your infrastructure challenges. We’ll give you honest and personal treatment. Get in touch to learn more.

Get in touch!