Something very strange happened yesterday, and I’m surprised it isn’t a much bigger story.
ChatGPT went down. Claude went down. Grok went down. There were even reports of problems with Gemini.
Not over the course of a bad week. At roughly the same time.
These are competing AI platforms, operated by different companies, with enormous amounts of money invested in keeping them online. Yet on September 3, they experienced overlapping outages significant enough to disrupt ChatGPT, Codex, Claude, Claude Code, Grok and their APIs. Ars Technica called simultaneous interruptions across four major AI systems “practically unheard of.”
And yet, a day later, we still don’t really know why.
OpenAI says a routing error made ChatGPT and Codex unavailable for some users. Anthropic says an infrastructure problem caused its outage. xAI traced Grok’s problem to its Memphis compute center. On the surface, those sound like three separate failures.
Maybe that’s exactly what happened.
But that’s one hell of a coincidence.
The obvious question is whether these supposedly independent AI ecosystems aren’t nearly as independent as most people think.
They may have different models, different interfaces and different logos, but underneath them sits an increasingly complicated web of cloud providers, data centers, networking infrastructure, compute partnerships and shared technology.
That’s the part I think deserves much more attention.
We’ve spent the last few years racing to put AI into everything. Businesses are building workflows around it. Developers are letting agents write and deploy code. Customer service is being automated. Companies are beginning to depend on AI systems to perform tasks that humans used to perform.
I’ve certainly done it myself.
But if ChatGPT, Claude and Grok can effectively disappear together, even briefly, where are the actual points of failure?
Interestingly, Cloudflare publicly said its systems were operating normally, and the major cloud providers weren’t reporting corresponding major incidents at the time. That makes the simultaneous timing more interesting, not less.
Wired is now asking essentially the same question: nobody is saying why these outages happened at almost exactly the same time, and no common cause has been established.
I’m not suggesting some conspiracy. Infrastructure failures happen.
I’m suggesting something much more practical.
We may be building an enormous amount of the world’s future infrastructure on top of dependencies that most of us don’t actually understand.
If three of the largest AI platforms can stumble simultaneously and, 24 hours later, there still isn’t a clear explanation for why the timing overlapped, that’s worth talking about.
A lot more than we currently are.
