4 SEP 2026 — On Thursday morning, United States time, ChatGPT, Claude and Grok all failed within the same window. Reports spiked around 8:07am Pacific and everything was normal again by 12:38pm. The explanation already circulating is a shared Microsoft Azure failure. None of the three companies has said that, Azure's own status showed the platform operational, and the three outages did not behave alike.
Disclosure: RECATOOLS is written with the assistance of Claude, made by Anthropic, one of the three companies whose service failed.
What each company actually said
OpenAI's status page acknowledged problems with ChatGPT and Codex, citing a routing error. A later update said a fix had been applied and services were recovering. Its window was short, with reporting putting the start at about 7:43am Pacific and a fix in place by roughly 8:17am.
Anthropic's status page reported elevated errors. Most Claude models were back by 8:49am, with Opus 4.8 and Opus 5 still affected after that, and the recorded outage ran about three hours and six minutes.
xAI acknowledged that Grok was having issues and said it was working on them. Google's Gemini never posted an official incident at all, despite its own spike in user reports.
The Azure theory hardened into a fact in about a day
The claim now circulating is specific. A regional failure inside Azure's East US infrastructure, it says, took down three of the four largest chatbots. Some write-ups state that flatly.
Trace the claim back and the support thins. First-day reporting framed Azure as also seeing a spike in outage reports and said that might be the underlying problem. That is a hypothesis about correlated reports, and it was published as one.
The evidence cuts the other way. Automated checks of Azure's status through that afternoon showed the service operational, Microsoft declared no matching incident, and each of the three companies opened its own investigation rather than pointing upstream. A shared root cause is possible, but it remains unproven. Most bad post-incident analysis lives in the gap between those two states.
The durations argue against one failure
When a single upstream dependency fails, the things depending on it usually recover on the same schedule, because they come back when it does.
These did not. OpenAI describes a routing error resolved in about half an hour. Anthropic recorded three hours and six minutes, with its largest models the last to return. If both had been waiting on the same regional failure, somebody has to explain why one had a tail and the other did not. Nobody has.
The model-specific detail cuts the same way. Opus 4.8 and Opus 5 lagging behind Sonnet points to a capacity constraint on those particular models, which is an internal serving problem rather than a network one.
Claude is served from three clouds by design
The theory needs all three companies to depend on Azure in the same way, and at least one does not. Claude runs on AWS Bedrock, Google Cloud Vertex and Microsoft Foundry, with AWS named as Anthropic's primary cloud and primary training partner, and the models trained across Trainium, TPUs and Nvidia GPUs.
Anthropic does have a large Azure commitment, so Microsoft's cloud is part of its setup. But an architecture spread across three hyperscalers is the opposite of a single point of failure. An Azure regional incident should degrade one serving path, not all of them.
That is what makes the flat attribution worth resisting. It requires an infrastructure story that is publicly contradicted by how at least one of these services is built.
Downdetector measures users, not severity
The report counts being quoted are roughly 37,000 for ChatGPT, about 1,300 for Claude and roughly 1,365 for Grok. Read as impact, that would make ChatGPT's outage nearly thirty times worse.
It was not. Downdetector counts people who chose to report a problem, so the ratio between services mostly reflects the ratio between their consumer user bases. ChatGPT's is far larger than either of the others.
Gemini is the control in this experiment, and it is the most useful number of the day. Its reports spiked while it was working normally. During a visible outage people check whether the alternatives are down too, and they file reports when a page is slow. Some of the simultaneity on Downdetector is the panic itself.
The concentration risk is real, and this is weak evidence for it
There is an argument that too much of the world's inference now sits on too few platforms, and a morning when three assistants failed together is an obvious hook for it.
The problem is that this particular morning does not demonstrate it. To show concentration risk you need a shared dependency that fails for everything at once. What is on the record instead is three separate investigations, one platform reporting itself healthy, and two very different recovery times.
Reaching for the conclusion anyway has a cost. Anyone planning around AI availability needs to know whether the failure mode is a shared cloud region, in which case multi-cloud helps, or independent internal faults that happened to coincide, in which case multi-provider failover is the answer and multi-cloud within one provider is not.
What would settle it
Three documents would end the argument, and all three are things these companies routinely publish when they choose to.
A post-incident review from OpenAI naming what the routing error was and where it sat. One from Anthropic explaining why the largest models trailed. And an Azure status history entry for East US covering that window, or a clear statement that there was none.
Until those documents appear, all anyone knows for certain is that three assistants failed in the same three-hour window for reasons none of them has published. For anyone whose work depends on these services, that is more useful than a cause invented to fill the gap.