preloader

· cloud google-cloud other-cloud outage resilience devops europe infrastructure

A Cooling Failure at a Dutch Google Cloud Facility Knocked Out GCVE, NetApp Volumes and Bare Metal for 15 Hours

Source: The Register / Google Cloud Service Health

A cooling failure at a Google Cloud facility serving the europe-west4 region, based in Eemshaven in the Netherlands, took three specialised services offline for around 15 hours. Google Cloud VMware Engine (GCVE), Google Cloud NetApp Volumes and Bare Metal Solutions were all affected after an electrical fault on the utility grid upstream of the datacentre disrupted the site’s power distribution gear and, in turn, its cooling systems. Rather than let hardware run in an uncontrolled thermal environment, Google proactively shut the affected workloads down to protect customer data and equipment, then worked through the outage until service was restored.

Why a single building could take down three services at once

The detail that makes this incident worth understanding, rather than just noting, is that GCVE, NetApp Volumes and Bare Metal Solutions in europe-west4 run out of a discrete datacentre building separate from the standard Compute Engine capacity that makes up most of the region. Customers who assumed their region-level redundancy protected all their Google Cloud services were exposed to a single building’s power and cooling failure regardless. That is a meaningfully different resilience profile than the one most architecture diagrams assume when they show “europe-west4” as a single resilient unit.

This is not a Google-specific quirk. Every major cloud provider houses certain specialised or legacy-inherited services, VMware-based offerings and bare metal among them, in dedicated facilities that do not share the redundancy characteristics of the provider’s core elastic compute fleet. The practical consequence is that “which region” is not a complete answer to “how resilient is my deployment.” Which specific service, and which physical facility that service actually runs from, matters just as much.

What this means for European workloads

For organisations running VMware-based workloads, NetApp storage or bare metal instances in europe-west4 as part of a broader EU data residency or sovereignty strategy, this outage is a concrete prompt to ask a harder question: does our failover plan actually cover this specific service, or only the parts of our architecture running on standard Compute Engine and GKE? A region-to-region failover plan built around VM and container workloads will not protect a workload that only exists as a GCVE cluster or Bare Metal Solutions node in one physical building.

The fix is not necessarily multi-region duplication of every service, which is expensive and often unnecessary. It starts with an honest audit: which of our production services run on infrastructure that shares fate with a single facility, and does our stated recovery time objective actually match what would happen if that facility went dark for half a day.

If you want a clear-eyed review of where your Google Cloud, Azure or AWS architecture has hidden single points of physical failure, and a realistic plan for closing the gaps that matter most, contact Excello Digital. We help European engineering teams turn resilience assumptions into resilience that has actually been tested.

These news items are automatically aggregated from industry sources and are not individually reviewed. Any inaccuracies are unintentional — let us know and we'll correct or remove it.

We’ll help you resolve your infrastructure challenges

Our team of experts is ready to help you with your infrastructure challenges. We’ll give you honest and personal treatment. Get in touch to learn more.

Get in touch!