Three competing AI companies going down at the same moment looks like coincidence until you check whose infrastructure they were all standing on.
A 90-minute window, one shared point of failure
On the morning of September 3, a routing error inside Microsoft Azure’s East US region began around 7:43am Pacific time. Within the same window, ChatGPT and OpenAI’s Codex tool started showing elevated error rates, Anthropic’s Claude models degraded, and xAI’s Grok went down entirely after what the company described as an outage at its Memphis compute centre. Google’s Gemini, built on Google’s own cloud infrastructure rather than Azure, was unaffected throughout. A fix was implemented by around 8:17am Pacific and monitored from there, with services recovering progressively and reaching full restoration by 12:38pm Pacific, nearly five hours after the first errors appeared.
The root cause traces back to Azure East US, the same regional backbone that happens to underpin the compute for three of the four largest AI chatbots on the market. That is not a coincidence of bad luck, it is a coincidence of shared vendor dependency: when the region that several unrelated companies all rely on develops a fault, the visible result is several unrelated products failing at once, in the same window, for the same underlying reason.
Redundancy on paper is not redundancy in practice
This is the same lesson Google Cloud handed the industry on September 1 with its us-central1-b fiber incident, just from the other side of the stack. It does not matter whether the workload belongs to a hyperscaler’s own flagship AI product or a mid-sized European SaaS company: if your production traffic depends on a single cloud region, a single provider’s control plane, or a single account’s blast radius, a regional incident you had no part in causing becomes your outage too. The AI vendors in this incident had the engineering resources of trillion-dollar companies behind them, and it still took the better part of five hours to fully recover.
What this means for anything you have built on Azure
European businesses have spent the past two years wiring Azure OpenAI, Azure AI Foundry, and Copilot into core workflows, often without asking what happens to that workflow when the specific Azure region behind it has a bad morning. The honest answer, for most deployments, is that the workflow stops, exactly like ChatGPT and Claude did on September 3. Multi-region failover, a fallback provider for anything customer-facing, and a documented plan for what your team does when the AI layer goes dark are the difference between an inconvenience and an incident.
If you want a real assessment of what a regional Azure, AWS or Google Cloud outage would actually do to your production systems, not a theoretical one, contact Excello Digital. We help European teams design cloud architecture that keeps working when a provider’s region does not.
