preloader

· aws cloud devops resilience incident-management infrastructure europe

AWS Just Had Its Third Reliability Incident in Three Months, and the Cause Was a Handful of Network Devices

Source: Tech Insider / AWS Service Health

Between 3:55 and 4:15 a.m. PDT on July 24, AWS experienced a connectivity failure affecting its US-West-2 region in Oregon. The disruption cascaded quickly: outage trackers reported Apple Pay, DoorDash, Reddit, Hulu and PlayStation Network all going dark within minutes of each other. AWS’s own postmortem pointed to networking devices responsible for routing traffic from the region to the Seattle metro area, and mitigation brought services back within roughly 80 minutes of the first reports. AWS said there was no indication that customer data was lost during the incident.

A pattern, not a one-off

This is the third notable AWS reliability incident in about three months. A thermal event at a Northern Virginia data centre disrupted services in May, and a separate network disruption hit AWS-linked services in June. Set alongside July’s US-West-2 connectivity failure, the common thread across all three is not a software bug or a bad deployment, it is physical and network infrastructure: cooling systems, routing hardware, the unglamorous plumbing that every application, regardless of how well it is architected, ultimately depends on.

Why a US region outage is a European problem too

It is tempting to read this as a US story, since the affected region and most of the visibly disrupted consumer apps sit outside Europe. But a meaningful share of European businesses run production workloads, disaster recovery targets, or critical SaaS vendor dependencies in US-East or US-West regions as part of a global architecture, and plenty more depend on US-hosted vendors that themselves run on AWS. When a hyperscaler region goes down, the blast radius follows the dependency graph, not the geography of your headquarters. NIS2 and DORA both push regulated European entities toward mapping and testing exactly these kinds of third-party and cloud dependencies, and an incident like this is a reasonable prompt to check whether that mapping actually reflects reality.

Building resilience beyond “AWS is very reliable”

Three incidents in three months does not mean AWS is unreliable in absolute terms, hyperscalers still run at availability levels most organisations could never match on their own. It does mean that treating a single AWS region as sufficient resilience is a decision worth revisiting on purpose rather than by default. That means genuine multi-region or multi-cloud failover for anything customer-facing, regular game-day testing of what actually happens when a region degrades rather than fully fails, and an up-to-date map of which of your critical vendors sit on which cloud region, so an outage like this one does not surface dependencies you did not know you had.

If you want an honest assessment of where a single-region outage would actually hurt your business, or help designing a resilience plan that goes beyond your cloud provider’s own SLA promises, contact Excello Digital. We help European teams build and test the failover plans that only get noticed when they are missing.

These news items are automatically aggregated from industry sources and are not individually reviewed. Any inaccuracies are unintentional — let us know and we'll correct or remove it.

We’ll help you resolve your infrastructure challenges

Our team of experts is ready to help you with your infrastructure challenges. We’ll give you honest and personal treatment. Get in touch to learn more.

Get in touch!