preloader

· · devops cloud google-cloud reliability infrastructure incident-response europe

One Engineer Disconnected a Fiber Cable. A Chunk of Google Cloud Went Down.

Source: The Register

Redundant fiber paths are supposed to be the thing that saves you when one connection fails. They don’t help when a single procedural mistake takes all of them out at once, which is exactly what happened to part of Google Cloud on September 1.

Thirteen minutes to disconnect every redundant path

Google’s design for the affected us-central1-b cluster assumed physically separated routing devices and fiber paths would keep the zone reachable even if one link failed. During planned network-fabric maintenance, a procedural error caused an engineer to sequentially disconnect the fiber paths across all the routing devices in the cluster, one after another, over roughly thirteen minutes. Because the disconnections happened in sequence rather than all at once, there was no single alarm moment to catch and reverse before the isolation was complete.

The result was total: traffic-drop rates hit 100 percent at peak, and virtual machines in the affected zone became unreachable from outside it, though they could still talk to each other internally. The disruption ran from roughly 07:41 to 11:52 Pacific time, over four hours.

Redundancy assumes independent failure. This wasn’t independent.

The uncomfortable lesson here is not that Google’s redundancy model was wrong in theory, it is that redundancy protects against independent failures, not a single human action that touches every redundant path in sequence. Google’s systems did eventually detect the isolation and shift traffic to healthy capacity elsewhere in the region, and technicians physically reseated the disconnected fiber, but recovery still took hours, and the company reportedly paused further maintenance in the region pending an audit.

What this means if your infrastructure runs on someone else’s cloud

Multi-zone and multi-region architecture is what protects you from exactly this kind of single-cluster event, but only if your workloads are actually built to fail over rather than just deployed across zones on paper. If a four-hour outage in one us-central1 zone would have taken your production traffic with it, that is worth finding out before it happens rather than after.

If your team wants a review of how your infrastructure actually behaves when a cloud provider has a bad morning, contact Excello Digital. We help European businesses design and test cloud architecture that survives the outages the provider itself does not see coming.

These news items are automatically aggregated from industry sources and are not individually reviewed. Any inaccuracies are unintentional — let us know and we'll correct or remove it.

We’ll help you resolve your infrastructure challenges

Our team of experts is ready to help you with your infrastructure challenges. We’ll give you honest and personal treatment. Get in touch to learn more.

Get in touch!