At around 9:30 AM ET, xAI's Grok became the first of the three to show visible problems. Users attempting to run a prompt on X received a message saying, "This model is overloaded right now. Please try again shortly or pick a different model." The wording pointed to a capacity issue rather than a complete network failure, and it affected the Android app, the iOS app, and the web at the same time. Roughly ninety minutes later, OpenAI's status page said there were "elevated errors across ChatGPT and Codex." Anthropic's Claude chatbot, Claude Code, and Claude API similarly began experiencing issues around the same time. According to CJ Avilla, a technical staff member at Anthropic, an "infrastructure issue" caused a partial outage across its services. Anthropic was the first company to confirm a resolution, with service restored at about 12:15 PM ET.
Key facts at a glance
- Grok started having problems at about 9:30 AM ET on Android, iOS, and the web, with an error saying the model was overloaded.
- ChatGPT and Codex experienced elevated errors around 11:00 AM ET.
- ChatGPT's outage affected logins, file uploads, voice mode, search, deep research, image generation, and other features.
- Claude, Claude Code, and the Claude API went down at roughly the same time due to an infrastructure issue.
- xAI linked its Grok outage to a problem at its Memphis data center.
- Anthropic confirmed recovery by 12:15 PM ET, while OpenAI and xAI also restored service shortly afterward.
What happened with ChatGPT and Codex
ChatGPT has evolved far beyond a simple chatbot. It now serves as a productivity platform for millions of people who rely on it for writing, coding, research, image creation, voice conversations, and automated workflows. OpenAI also operates Codex, an AI coding tool designed to help developers write and debug code inside their existing environments. When OpenAI said that both ChatGPT and Codex were experiencing elevated errors, it became clear that the outage was not limited to a single interface or feature.
Users trying to access ChatGPT encountered error messages as they loaded the site. Many could not log in at all, which meant they could not reach their previous conversations or use their custom instructions. File uploads failed for users who wanted to analyze documents. Voice mode did not properly respond, breaking the hands-free conversational experience that many people use on mobile devices. Search and deep research tools also went down, preventing users from getting real-time information. Image generation, a popular feature inside ChatGPT, was likewise disrupted. The timing was particularly awkward because OpenAI had been teasing the launch of Astra, a new AI model. While there was no indication that the outage was connected to the Astra announcement, product launches often place additional strain on infrastructure and draw more attention to service uptime.
ChatGPT's importance has grown significantly since its public debut in late 2022. The product helped start the current generative AI boom and has become one of the most visited AI applications on the internet. For many individuals and companies, it is not just a toy but a primary work tool. Writers use it to brainstorm and edit. Programmers use it to explain bugs and suggest fixes. Students use it for tutoring and study guides. Marketers use it to draft campaign copy. When ChatGPT goes down, even for an hour, the interruption is felt across industries and time zones. This is why OpenAI maintains a status page and usually reports incidents quickly. Still, the company did not offer an immediate public explanation for the root cause of Thursday's failure.
Claude and Anthropic's infrastructure issue
Anthropic's Claude family of products has become one of the leading alternatives to ChatGPT, especially among enterprise customers who want strong safety, long-context understanding, and more cautious AI behavior. On Thursday, Anthropic said that Claude, Claude Code, and the Claude API all experienced a partial outage. Claude Code is particularly important to software developers because it lets them use Claude inside their terminal environment to generate code, review pull requests, fix test failures, and automate repetitive tasks. An outage in Claude Code can stop teammates who have built their daily workflows around the assistant.
Claude's API is also used by many companies to power customer service tools, document summarization applications, legal research systems, and backend automation. When an API goes down, downstream applications can break even if the user interface of those applications appears normal. This makes AI outages a supply chain problem, not just an inconvenience for individual users. Anthropic technical staff member CJ Avilla said that an "infrastructure issue" caused the outage. That kind of language typically points to an internal problem involving networks, databases, virtual machines, or cloud services rather than a deliberate maintenance event. Anthropic acted quickly, and by 12:15 PM ET the company said that the problem had been resolved. Even after the fix, some users may have continued to see errors for a short period as traffic was rebalanced.
Grok and the Memphis data center connection
Grok is xAI's artificial intelligence assistant, deeply integrated with Elon Musk's X platform. It also exists as a standalone app for mobile devices and a web product. On Thursday, Grok stopped generating responses for many users across every interface. The error message users saw described the model as overloaded, suggesting that the systems responsible for running Grok's machine-learning workloads had hit capacity. xAI later revealed that an outage at its Memphis data center was responsible for the issue. The Memphis facility has been central to xAI's expansion strategy, serving as a massive cluster of computing hardware designed to train and run some of the company's most advanced models.
Data centers are the backbone of generative AI. Training and running large language models requires thousands of specialized processors, high-power connections, liquid cooling systems, and sophisticated networking. If any one of those elements fails, the model-serving stack can quickly stop accepting requests. In Grok's case, the issue appeared to be tied to the data center that processes many of the system's queries. Problems at data centers have become more visible as AI adoption has grown. Companies like xAI have raced to build huge facilities at remarkable speed, and these projects can face power grid limitations, cooling challenges, and hardware reliability issues. When a data center outage occurs, it is not just a local issue; it can take down services for users around the world.
Why did all three go down at once?
With three major AI providers failing in the same time window, it was natural for users to wonder whether the outages shared a common cause. As of Thursday afternoon, there was no official evidence that the events were linked. OpenAI did not explain the root problem behind ChatGPT's elevated errors. Anthropic called its issue an infrastructure failure. xAI identified its Memphis facility as the source of Grok's outage. Those three explanations could easily be three independent incidents that simply happened on the same morning. However, observers pointed out that AI companies rely on many of the same third-party services, including cloud providers, internet service providers, domain name systems, content delivery networks, and authentication platforms. A failure in one of those shared layers could theoretically cause multiple unrelated-looking applications to fail at roughly the same time.
There were also other possibilities. Large-scale AI platforms compete for power and network capacity, especially in regions where data centers are concentrated. Grid disturbances, network fiber cuts, or extreme weather could affect more than one company. Yet none of those explanations were confirmed. The more likely answer, based on the available evidence, is that this was a coincidence. ChatGPT, Anthropic, and xAI each run highly complex distributed systems with many potential points of failure. On any given day, there is a nonzero chance that one of these companies will experience an incident. The likelihood that two of them will suffer outages on the same day is already low, but not impossible. Three on the same day is unusual, but it does happen, especially in an industry that is expanding infrastructure as quickly as the AI sector.
What the outage tells us about AI dependency
The simultaneous failures of ChatGPT, Grok, and Claude revealed just how reliant users have become on these tools. In the past, a web service outage was often an inconvenience. But for many people now, losing access to an AI assistant means losing part of their writing process, coding workflow, research pipeline, or customer service operation. Professionals who depend on AI to draft emails, summarize meetings, translate documents, or generate images can find themselves stuck when the service disappears. This is why outages tend to inspire a flood of social media posts, support tickets, and status page checks. On Thursday, the fact that all three major assistants were unavailable forced some users to fall back on traditional search engines and manual work, a reminder of how quickly the AI era has become normal.
The outage also highlights the growing challenge of reliability in AI. Building and maintaining large language models at scale is very different from running a traditional website. Each user request can use a significant amount of computational power, especially for complex reasoning tasks. Popular moments, viral trends, product launches, and new feature announcements can create traffic spikes that overwhelm capacity. Companies like OpenAI, Anthropic, and xAI invest heavily in redundancy, but they also face practical limits. Hardware fails, networks degrade, database clusters become overloaded, and software configuration changes can have unexpected side effects. With customers increasingly asking for enterprise-grade service level agreements, AI providers will need to invest in fault-tolerant architectures that isolate failures and reduce downtime.
There is also a deeper lesson about resilience. An organization that relies on one AI vendor is vulnerable when that vendor has an outage. An organization that relies on three AI vendors is less exposed if only one goes down, but as Thursday proved, even that strategy can provide less protection than expected when multiple services fail at once. Some companies have started to build local models or open-source alternatives into their workflows as a fallback. Others are designing internal applications with multiple model providers so traffic can be rerouted automatically. For individual users, the options are more limited, but the outage served as a useful reminder that free or cheap AI tools are not guaranteed to be available every minute.
AI is still a young industry, and outages remain a fact of life for providers that push the limits of scale. Thursday's event was not destructive in the way that a security breach or data loss incident might be. No reports emerged suggesting that user data was compromised, and all three companies seemed to recover without permanent damage. But the synchronized nature of the disruption offered a rare glimpse into the architecture of the modern AI ecosystem. These powerful services are built on enormous networks of data centers, processors, and software layers. When those systems work, they feel almost magical. When they fail, they remind everyone that the magic depends on thousands of physical machines that must stay online, powered, cooled, and properly connected. The next major outage is probably not far away, and it may again reveal how concentrated the AI world has become.
Source: The Verge News