Claude, Grok outages point to Memphis facility, while OpenAI cites internal routing error

Claude, Grok outages point to Memphis facility, while OpenAI cites internal routing error

N
News Editor
2026-09-05 06:42:21
ChatGPT, Claude, Grok and some Codex users were hit by service issues on the morning of Sept. 3, but there is still no single official explanation linking all three incidents. The clearest cross-company statement came from SpaceX, which said Grok problems followed an outage at its Memphis compute center and added that it also wanted to apologize to impacted compute partners. That matters because Anthropic is paying SpaceX $1.25 billion a month to use Colossus 1 in Memphis, according to the report. Colossus 1 is described as having more than 220,000 NVIDIA GPUs, including H100, H200 and GB200 units. The report says the contract runs through May 2029, implying about $15 billion a year in payments. OpenAI’s timeline was different. It said ChatGPT and Codex were unavailable from 7:43 a.m. because of an internal routing error and recovered at 8:17 a.m., roughly 80 minutes after the earlier issues at Anthropic and xAI began. Status pages from Microsoft Azure, AWS and Google Cloud reportedly showed no matching incidents that day, while Cloudflare said it saw no major service disruption.

ChatGPT, Claude and Grok ran into problems on the morning of Sept. 3, affecting users of those services as well as some Codex users. So far, no official statement has tied all three incidents together.

The only formal comment that directly connects two of the companies came from SpaceX. In a post on X, it said: 「We are sorry for the issues you may have experienced with Grok following an outage at our Memphis compute center this morning. We’d also like to apologize to our impacted compute partners. All systems have now been restored and are functioning nominally.」

The outage timelines did not fully match

The timing reported by each company was close, but not identical.

  • Anthropic said at 6:23 a.m. that request errors had increased for some Claude models, then marked the issue resolved at 9:16 a.m.
  • Grok was affected from 6:30 a.m. until 10:05 a.m.
  • OpenAI said ChatGPT and Codex became unavailable at 7:43 a.m. because of an internal routing error, with service restored at 8:17 a.m.

Claude and Grok began showing problems within seven minutes of each other. OpenAI’s incident started nearly 80 minutes later, leaving open the question of whether there was a shared cause across the three companies.

Anthropic’s leased capacity is also in Memphis

SpaceX’s reference to impacted compute partners was not just boilerplate, according to the report. SpaceX acquired xAI in an all-stock deal on Feb. 2, and the combined company was valued at $1.25 trillion. After that transaction, xAI became a SpaceX business unit.

On May 6, SpaceX signed a compute agreement with Anthropic and opened Colossus 1 to the company. The report says Anthropic pays SpaceX $1.25 billion a month for that capacity, and the facility in question is Colossus 1 in Memphis.

Colossus 1 is said to house more than 220,000 NVIDIA GPUs, including H100, H200 and next-generation GB200 chips. An xAI announcement said the compute would be used to 「directly increase capacity for Claude Pro and Claude Max subscribers.」 The report also places Colossus 1 in a former Electrolux factory in southwest Memphis at 3231 Riverport Rd.

A contract measured in the tens of billions

According to SpaceX’s IPO filing, Anthropic’s payments continue through May 2029. At $1.25 billion a month, that works out to roughly $15 billion a year.

The report says SpaceX’s annual revenue is about $18 billion, which means Anthropic’s payments alone approach the company’s previous full-year revenue base. Anthropic co-founder and chief compute officer Tom Brown also said on X that starting in June, Anthropic would expand usage to GB200 capacity in Colossus 2.

That left one concrete overlap in the public record: Claude and Grok hit problems within seven minutes, and SpaceX acknowledged an outage at the Memphis facility on the same day.

OpenAI gave the clearest company-specific explanation

OpenAI’s account was more explicit than the others. A Hacker News thread titled 「Why did OpenAI, Claude, and Grok all go down at the same time」 drew heavy discussion, and one comment claimed to come from the incident commander for OpenAI that day.

The comment said: 「I work at OpenAI, I was the incident commander for yesterday’s outage. We had an internal routing error in our infrastructure that impacted some products. It was unrelated to the Astra launch. We don’t comment on other companies’ incidents.」 The statement also pushed back on speculation that a new model release was behind the disruption.

The report adds that OpenAI does not have a compute partnership with SpaceX, and Elon Musk is still in litigation with OpenAI. Combined with the 80-minute gap before OpenAI’s own incident began, the company’s outage looks more like a separate problem that happened on the same morning.

No confirmed evidence pointing to a cloud provider failure

There is also no public basis, at least for now, for blaming a third-party cloud provider. The report says status pages for Microsoft Azure, Amazon Web Services and Google Cloud showed no related incidents that day. Cloudflare also said there was 「no major service outage.」

Google Gemini did receive scattered reports of abnormal behavior on the same day, but Google did not confirm any large-scale failure.

Shared-cause theories remain unproven

Online speculation ranged from the possibility of a nonpublic network monitoring feature shared by major AI firms to claims involving government surveillance, including references to past National Security Agency monitoring programs. The report says those theories lack evidence, and notes that collecting traffic records alone would not usually cause customer-facing services to fail at the same time.

A more plausible explanation raised in the report is cascading overload. If one model service goes offline, enterprise and individual users may quickly shift demand to other providers. A sudden spike in traffic could then overwhelm the next platform taking that load. The report notes that AI companies are constrained by GPU supply and data center capacity, which may leave little spare room to absorb all of a rival’s traffic on short notice.

For now, there is still no single official account that fully links the Sept. 3 disruptions across all three companies.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
700

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.