preloader

· · devops hetzner cloud outage disaster-recovery europe infrastructure

Hetzner’s Storage Cluster Needed Manual Recovery, a Reminder That Object Storage Is Not a Backup Strategy

Source: Hetzner Status

Object storage is sold on durability numbers with a lot of nines in them. Hetzner’s latest incident shows what happens in the space those numbers leave out.

Two hardware failures in short succession broke the safety net

On August 25, Hetzner’s Object Storage service in its NBG cluster degraded after two hardware failures happened close together. Object storage systems are built to tolerate a single failure by rebuilding redundancy from the surviving copies automatically, but Hetzner reported that this automatic recovery failed for a subset of the affected data once the second failure landed. That forced the company to stop all services for roughly a third of the NBG cluster while engineers manually identified and restored the affected data, leaving buckets in that section inaccessible rather than merely slow. Hetzner targeted completion of the manual restore by August 26 at 11:00 CEST, a window measured in more than a day for infrastructure that customers generally assume is always on.

Redundancy is a mitigation, not a guarantee

This is precisely the low-probability, high-impact scenario that object storage redundancy is designed to absorb, and the fact that it still required manual intervention to prevent data loss is the useful signal here, not the outage itself. Every major cloud provider has had an incident where automated recovery hit its limits and needed a human in the loop, and it will happen again on every platform, Hetzner, AWS, Azure, Google Cloud or otherwise. The businesses that came through this particular incident unaffected are the ones that never treated a single storage cluster, in a single provider, in a single location, as their only copy of anything they could not afford to lose.

The gap this exposes for teams running lean on Hetzner

Hetzner is a popular choice across Europe specifically because it is cheaper and closer to home than the US hyperscalers, and plenty of startups and mid-sized companies run production workloads, backups and customer data on it without a second provider or a genuine offsite copy. That is a defensible cost decision, but only if it is a deliberate one made with eyes open to what “one provider, one region” actually means during an incident like this one. A backup that lives in the same cluster as the data it is backing up is not a backup, it is a second copy of the same single point of failure.

If you are unsure whether your current backup and disaster recovery setup would survive an incident like this one, or want a second opinion on your Hetzner, AWS or Azure architecture, contact Excello Digital. We help European teams build infrastructure that keeps running when a single provider has a bad day.

These news items are automatically aggregated from industry sources and are not individually reviewed. Any inaccuracies are unintentional — let us know and we'll correct or remove it.

We’ll help you resolve your infrastructure challenges

Our team of experts is ready to help you with your infrastructure challenges. We’ll give you honest and personal treatment. Get in touch to learn more.

Get in touch!