The AI DownsideDocumenting AI's downsides

LLMs

ChatGPT, Claude and Grok Went Down at the Same Time — and No One Said Why

Three assistants sold as alternatives failed in the same window on 3 September. The outage exposed how much shared plumbing sits underneath them — and how little any of them will tell you.

Editorial illustration for “ChatGPT, Claude and Grok Went Down at the Same Time — and No One Said Why”.

Open ChatGPT on the afternoon of 3 September and, for a while, you got a blank 404. Switch to Claude, the tool you keep for exactly this moment, and it was throwing errors too. Try Grok as a third option and it told you it had been disconnected. The backup for one AI assistant, for a lot of people, is another AI assistant — and on 3 September the backup was down as well.

This is the failure mode nobody prices in when they sign up for a second or third AI subscription. You imagine the tools as independent, so that a bad day at one company is a shrug and a tab-switch. For a few hours that afternoon, three of the biggest names in the business — OpenAI, Anthropic and xAI — failed inside the same window, and the tab-switch led nowhere.

The outages were real, short, and resolved within hours. That part is ordinary; everything goes down sometimes. The uncomfortable part is the two things that came with it: the companies would not say why, and the fact that they failed together hints at something the marketing never mentions — how much shared plumbing sits underneath tools that are sold to you as alternatives.

The afternoon all three went dark

Start with what is documented. Anthropic’s status page opened an incident it called “Elevated errors for multiple models” at 13:26 UTC and marked it resolved at 16:16 UTC — just under three hours — spanning claude.ai, the API, Claude Code and Cowork, and naming Mythos/Fable 5.1 and Opus 5, 4.8 and 4.6 among the affected models. The incident log is precise about what broke and useless about why: there is no stated cause anywhere in it.

OpenAI’s history tells the same shape of story from the other side. Its own log records “elevated errors across ChatGPT and Codex” that day, resolved, with no explanation of the underlying fault. Users filled in the texture the status page left out: chatgpt.com returning a raw 404 with no page around it, Codex calls to the backend failing outright, and — a detail worth holding onto — the error pages carrying Cloudflare’s fingerprints. One person on a Pro account got through while a Plus account on the same machine got the 404; an incognito window, logged out, loaded fine. Whatever was wrong sat somewhere between the login and the model, not in the model itself.

xAI’s Grok completed the set, showing a “Grok has been disconnected” error on its own status surface at around the same time. And the pile-up was wider than three: users pointed at Downdetector reporting problems with Claude, Grok, ChatGPT and Gemini all at once. It was not a clean, total blackout — more a rolling wobble in which existing sessions often survived, Claude’s smaller Sonnet model kept answering while the larger Opus errored, and the failures overlapped without lining up perfectly. But if you sat down at 2pm UTC to get some work done, the practical experience was that the whole shelf of tools was, briefly, out of reach.

Nobody would say why

The information vacuum was its own event. When three services go down together, the reasonable question is not “is it broken” but “is it the same thing breaking” — and none of the companies answered it. Anthropic named the symptoms; OpenAI named the symptoms; xAI said less. As one commenter put it while refreshing the pages, unhelpful status pages in a crisis are “a tradition at this stage.” The people paying for these tools were left to reverse-engineer an outage from cf-ray codes and gut feeling.

Into that vacuum rushed the theories, and they are instructive precisely because they were guesses. Maybe CoreWeave or AWS or a SpaceX-linked datacentre was down, someone offered. Maybe it was DNS, said another, because “it’s almost always DNS.” Maybe, a third suggested more soberly, one service failing sent a surge of traffic to the next and knocked it over in turn. Someone noticed Cloudflare’s own status page was flagging problems around the same time, then immediately added the honest caveat: “Not sure if they are linked.” Nobody outside the companies could tell. That is the point. When the provider won’t explain a correlated failure, the customer is left unable to distinguish coincidence from a single fault running under all of it.

You didn’t buy three independent AI tools. You bought three front-ends, and on 3 September nobody would tell you how much they share behind the login.

The plumbing they share

Here is the mechanism the launch videos skip. A modern AI assistant is not one company’s stack from your keyboard to the GPU. It is a chain of shared layers: a content-delivery and edge network (Cloudflare fronts a large share of the web, which is why its cf-ray IDs turned up in the 404s), DNS, load balancers, and underneath it all a very small number of clouds — Amazon’s, Microsoft’s, Google’s, specialist GPU hosts such as CoreWeave, and xAI’s own SpaceX-linked capacity. The models compete. The substrate they run on largely does not; it is rented from the same handful of landlords.

Concentration like that has a specific consequence: correlated failure. If two of your tools sit behind the same edge network or in the same cloud region, then a fault in that shared layer does not politely pick one victim. It takes down everything downstream of it, including the “alternative” you were about to switch to. You can hold two subscriptions and still have one point of failure, and — this is the part that ought to bother anyone relying on these tools for real work — you generally have no way to find out in advance whether you do. Providers do not publish which cloud, which region or which CDN sits behind a given product, so “keep a backup” quietly becomes “keep a second thing that might die at the same moment for the same reason.”

This is a different complaint from the one we made when Claude kept going down and Anthropic’s own status page said so. That was about one company’s uptime. This is about the shape of the whole market: a few models, fewer clouds, and a redundancy story that only holds if the layers you can’t see happen not to overlap.

Even the maintenance keeps breaking it

How fragile is that shared substrate? Fragile enough that its own scheduled maintenance keeps knocking it over — and here it is worth being scrupulous, because these next incidents are not the 3 September AI outage and no one has shown that they are connected.

Google Cloud went down twice in a fortnight for the same banal reason. On 20 August, its us-west1 region lost network capacity during, in Google’s words, “scheduled fiber optic maintenance.” Then on 1 September, us-central1-b fell over when, as The Register reported on 4 September, “the inadvertent physical disconnection of network fiber-optic cables during a routine hardware maintenance procedure” took it down. Google’s own post-mortem is almost comic in its precision: “A procedural error meant that the physical maintenance action sequentially unplugged 100% of fiber paths across all devices within 13 minutes,” after which “traffic flow drop rates … reached 100%.” The fix was to walk in and plug the cables back in.

None of the reporting on those outages mentions ChatGPT, Claude or Grok, and it would be exactly the kind of hype-in-reverse we try to avoid to pin the AI failures on a fibre cable in Iowa. The reason the GCP incidents belong in this story is narrower and fairer: they are independent, documented proof that the ground these tools stand on can be pulled out from under them by a maintenance crew following a checklist. When the substrate is that easy to trip over, correlated failures upstairs are not a paranoid theory. They are the base rate.

To be fair: what this isn’t

Several honest caveats, because the gap between promise and product cuts both ways. First, this was not a catastrophe. All three services were restored within hours, plenty of existing sessions kept running, and at Anthropic the lighter Sonnet model answered throughout — a partial degradation, not a wall. Second, there is a real argument against a single shared cause: xAI’s capacity is not OpenAI’s, and one commenter noted that the SpaceX-linked hosting some firms use does not sit under all three, which cuts against a tidy “one datacentre took them all out” narrative. A cascade — one outage driving a traffic surge into the next — may account for more of the overlap than any common backbone, and it was GPT-6 Astra’s launch afternoon, a day that already had paying users hitting walls, so unusual load is a fair part of the picture.

And third, redundancy is genuinely hard. Running frontier models is expensive and concentrated by nature; there are only so many places with the GPUs and power to host them, so some overlap is structural rather than negligent. Expecting every AI company to build on a private, fully independent stack is neither realistic nor, on its own, what we’re asking. What we’re asking for is candour: tell paying users enough about the shared layers that “keep a backup” can mean something.

What you can actually do

Until the providers are more open about what sits behind the login, the defensive moves are yours to make. None of them are exotic:

  • Make your backup genuinely independent. A second chatbot on the same cloud is not redundancy. An open-weight model you can run locally, or a provider you know sits on a different cloud, is. On 3 September, the people least inconvenienced were the ones who could fall back to something running on their own machine.
  • Keep your work outside the tool. Save drafts, code and prompts locally as you go, so an outage costs you access for an hour rather than your afternoon’s output. Treat the assistant as a workspace you’re renting, not a filing cabinet you own — especially since getting your data back out of an AI tool is rarely as easy as it should be.
  • Assume shared plumbing until told otherwise. If two tools matter to you, it’s worth knowing whether they share a cloud or a CDN. The companies rarely volunteer it, but even a rough guess beats discovering the overlap mid-outage.
  • Don’t trust the status page to be fast. On the day, the official pages lagged and stayed vague. A community thread or a service like Downdetector will usually tell you it’s not just you long before the vendor admits it.

The models keep getting better, and none of this is a reason to stop using them. But 3 September was a small, clarifying reminder that the reliability of an AI tool is not a property of its model — it is a property of the stack underneath, most of which you can’t see and none of which you control. Three rivals went down together and not one would say why. The least you can take from it is to stop mistaking a second subscription for a safety net.

Frequently asked questions

What actually happened on 3 September 2026?

Within roughly the same window that afternoon (UTC), three major AI assistants failed. Anthropic’s status page opened an incident titled ‘Elevated errors for multiple models’ at 13:26 UTC and marked it resolved at 16:16 UTC, covering claude.ai, the API, Claude Code and Cowork. OpenAI logged ‘elevated errors across ChatGPT and Codex’; users reported chatgpt.com returning raw 404 pages and Codex calls failing. xAI’s Grok showed a disconnection error on its own status page. Downdetector users also reported problems with Gemini.

Did one company’s outage cause the others?

No source has shown that. None of the three published a clear cause, so any single-root-cause story is speculation. There are reasons to doubt a simple shared-datacentre explanation — for example, xAI’s Grok runs on different capacity from OpenAI — and a plausible partial explanation is a cascade, where one service failing sends a surge of traffic to the others. What is verifiable is that the failures overlapped and that none of the companies explained why.

Isn’t using a different AI tool a good backup when one goes down?

Only if the two tools do not share a failure domain, and you usually cannot tell whether they do. Rival assistants lean on the same small set of cloud providers and the same content-delivery and edge networks. When the shared layer wobbles, the ‘alternative’ you switch to can be affected by the same fault. Real redundancy means an independent path — a different provider on a different cloud, or a local model — not a second front-end to the same infrastructure.

Was the Google Cloud fibre outage the reason?

There is no evidence for that. Google Cloud did have two recent outages caused by fibre-optic cables being disconnected during maintenance — us-west1 on 20 August and us-central1-b on 1 September — and The Register reported the details on 4 September. But none of those reports link the fibre incidents to the 3 September AI outages. We mention them only as separate evidence that the shared substrate these tools sit on is more fragile than the marketing suggests.

What can I do to avoid being stranded next time?

Keep a genuinely independent fallback rather than a second tool that may share the same plumbing: an open-weight model you can run locally, or a provider you know sits on a different cloud. Save anything important outside the tool as you go, so an outage costs you access rather than work. And treat status pages as lagging indicators — on 3 September they were slow and vague — cross-checking a service like Downdetector or a community thread when something feels wrong.

Sources

  1. Anthropic status — incident ‘Elevated errors for multiple models’ (claude.ai, API, Claude Code, Cowork; Mythos/Fable 5.1, Opus 5/4.8/4.6); opened 13:26 UTC, resolved 16:16 UTC, 3 Sep 2026Anthropic
  2. OpenAI status history — ‘Elevated errors across ChatGPT and Codex’, 3 Sep 2026 (resolved; no cause stated)OpenAI
  3. xAI status — Grok service statusxAI
  4. Ask HN: ‘Why were OpenAI, Claude, and Grok simultaneously down?’ — halcdev, links all three status pages (3 Sep 2026, 15:07 UTC)Hacker News
  5. ‘ChatGPT outage – Resolved’ — user reports of raw 404s carrying Cloudflare cf-ray IDs and Codex backend failures (3 Sep 2026)Hacker News
  6. ‘Grok outage’ — leecb: ‘ChatGPT here has been giving me Cloudflare errors, and Cloudflare reports that it is having issues’; harrisoned: ‘Downdetector reports Claude, Grok, ChatGPT and Gemini with issues’ (3 Sep 2026)Hacker News
  7. The Register — ‘Google engineer unplugged fiber and took down a chunk of the G-Cloud’: us-central1-b, 1 Sep 2026, ‘sequentially unplugged 100% of fiber paths across all devices within 13 minutes’ (4 Sep 2026)The Register
  8. Tell HN: ‘Both recent GCP outages caused by fiber optic maintenance’ — quotes Google’s us-west1 (20 Aug) and us-central1-b (1 Sep) incident language (3 Sep 2026)Hacker News

Related grievances

All articles →