Ask HN: Why were OpenAI, Claude, and Grok simultaneously down?

202 points | by halcdev 3 hours ago

51 comments

  • kibae 2 hours ago
    Cloudflare, Azure, AWS, and Google Cloud all have a similar uptick in reported errors around 7:30. I suspect an outage on Cloudflare or another load-bearing service cascaded through all the major cloud providers.

    https://downdetector.com/status/cloudflare/

    https://downdetector.com/status/windows-azure/

    https://downdetector.com/status/aws-amazon-web-services/

    https://downdetector.com/status/google-cloud/

    • samaysharma 51 minutes ago
    • moomin 1 hour ago
      Yes, but is it a load-bearing seam?
      • graemep 38 minutes ago
        You are right, it is. They have now landed a clean fix.
        • Oarch 13 minutes ago
          They're saving a memory so this can't happen again.
      • codechicago277 7 minutes ago
        I need to take a step back.
      • JohnMakin 13 minutes ago
        Honestly, this is worth looking at — with one caveat.
        • cloudfudge 1 minute ago
          That's on me. I've been giving confident advice that doesn't hold up in practice.
    • y-c-o-m-b 2 hours ago
      Cloudfare did release an update for their "HTTP/3 issue affecting R2 custom domains" around that time

      https://www.cloudflarestatus.com/history?type=incident

    • martyfunkhouser 1 hour ago
      Will the post-mortem reveal they all relied on a service running on a Macbook in a break room with a "Do not turn off" sign taped to it?
      • mcphage 1 hour ago
        Actually it said "Beware of the Leopard".
        • RobotToaster 59 minutes ago
          If it was still running snow leopard that would explain it
      • hejdidf 1 hour ago
        [flagged]
      • hejdidf 1 hour ago
        [flagged]
        • esafak 1 hour ago
          Evidently it annoyed you enough to create one too. Remember, guys, there's one more day until Friday!
          • martyfunkhouser 1 hour ago
            Friday is when the janitors come by with the floor polishers and plug them into the same socket as the Macbook.

            They are not taking the blame this time!

            • 4chandaily 1 hour ago
              Just writing to say that I appreciated the nostalgia, even if your references were lost on the other responders.
    • snihalani 1 hour ago
      my money is on DNS
    • jlaneve 1 hour ago
      Dane says Cloudflare has no service disruptions: https://x.com/dok2001/status/2095538619603628388?s=20
    • dominotw 22 minutes ago
      > load-bearing service

      do you generate training data for claude as a job?

    • guybedo 1 hour ago
      LOAD BEARING
    • algoth1 1 hour ago
      https://xkcd.com/2347/ Nebraska man retired
      • boomlinde 24 minutes ago
        Maybe a three-line npm package sneezed.
    • Flere-Imsaho 11 minutes ago
      The internet is not supposed to work like this. The network was designed for robustness and fault tolerance, which allows it to reroute data if parts of the network fail.

      Why are we all depending on one entity for it all to work? Makes me mad.

      • seanw444 4 minutes ago
        Because more fasterer and more cheaperer.

        I hope Reticulum gains traction.

      • subw00f 6 minutes ago
        Oh boy, the internet is anything but what it was supposed to be. I can't really bring myself to remember without feeling bad about it. The centralization, the power of certain businesses, the surveillance, dark patterns everywhere. Hell, you catch people simping for billionaires and asking, "Is that legal?" to scraping posts. Here. In HACKER news. So yeah. Depressing.
    • KptMarchewa 2 hours ago
      Not really. The impact isn't as big too - Codex for example did not stop working for me.

      https://updog.ai/

      • personjerry 1 hour ago
        what's updog?
      • taytus 1 hour ago
        oh well, if it is working for you, then we are saved.
      • cromka 2 hours ago
        > Not really. The impact isn't as big too - Codex for example did not stop working for me.

        Buddy, read the room. Just because it works for you doesn't mean it works for everyone. ESPECIALLY if the suggested issue here is, indeed, with Cloudflare.

        • emerongi 1 hour ago
          They simply shared their experience. I would’ve thought it’s a full-blown outage, but clearly not.

          You stepped in the room real stinky here. What’s with the attitude?

          • cromka 1 hour ago
            No, they didn't "simply share their experience", they explicitly negated the scale of the issue in their opening statement, only because it works for them. So they claim "impact isn't too big" based on their personal anecdotal evidence of sample size literally 1.

            > full-blown outage, but clearly not.

            Again, based on a SINGLE report?

            • lossolo 1 hour ago
              It was working for me too.

              sample_size++;

    • oersted 2 hours ago
      “load bearing” :)

      For once it’s appropriately used.

      • The_Blade 20 minutes ago
        i wouldn't take you down. you're a load-bearing poster
      • frollogaston 57 minutes ago
        What's the other way it's used?
        • aNapierkowski 55 minutes ago
          LLMs (at least Claude) tends to overuse that significantly
          • frollogaston 54 minutes ago
            Oh, so like "honest" and "ratchet." Oh well, it'll choose different words to overuse later.
            • rescbr 8 minutes ago
              I'm getting "spike" for a while now, and just found out the newest word which is "gauntlet".
            • 8cvor6j844qw_d6 38 minutes ago
              Yeah "You're absolutely right" seems to be less used with recent models. Wondering what's the next overused phrase after seam and load bearing.
            • therein 44 minutes ago
              honest-load-bearing-ratchet sounds like an instance name.
        • mv4 38 minutes ago
          In every document created by Claude.
        • HoldOnAMinute 36 minutes ago
          I have never seen Claude use this phrase
      • smrtinsert 54 minutes ago
        I still winced
      • quotemstr 1 hour ago
        It's a good metaphor and I refuse to let AI ruin it for me.
        • SkyeCA 1 minute ago
          It doesn't have to ruin it for you, but people are going to assume comments with it are AI generated.
        • spudlyo 46 minutes ago
          Years ago, when I worked at Stripe (which had a somewhat unique and inventive lexicon) it was a common term. “Is this jank load-bearing?” someone might ask.
        • andrewla 23 minutes ago
          Kids In The Hall had a sketch about overuse of a word or phrase [1]. This is the world that Claude is building for us.

          [1] https://www.youtube.com/watch?v=lStcwT_RGrQ

        • cobzilla 53 minutes ago
          I added a specific rule to disallow saying “load bearing”. So Claude is now saying “load handling”
        • wadayano 1 hour ago
          But have we found the seams yet though?
          • rl3 1 hour ago
            If the seams bear too much load, they rip. Whereas pants, they fall down.

            I'll be honest with you: this is why we need to take a belt-and-suspenders approach.

            • pborenstein 48 minutes ago
              That's not just an observation, it's an insight. Words are doing the real work.
          • cootsnuck 1 hour ago
            Yea I don't get "load bearing" that much but "seams"... So sick of it.
        • lubujackson 1 hour ago
          We need a Samwise "Bear the load!" meme
        • MadameMinty 1 hour ago
          Wait, why would it ruin it?
          • ipsod 1 hour ago
            Because it says it more often than kids say "six seven".
          • imwally 1 hour ago
            It’s a frequently used metaphor in LLM responses.
            • bornfreddy 19 minutes ago
              Often for trivial things that LLM is proud that it has noticed but bear no load whatsoever.
      • dominotw 23 minutes ago
        claude code users at couldfare might've been thinking claude is specifically about them and see nothing wrong like other ppl do
  • juujian 1 hour ago
    Users perceiving the products as largely interchangeable and quickly DDoS'ing the other providers when one is down. So much for the possibility of a moat.
    • efskap 9 minutes ago
      This is like the Bronze Age collapse when city-states fell one by one to displaced demand, under the refugee interpretation of the Sea Peoples.
    • giancarlostoro 1 hour ago
      I have a feeling this is part of it, especially when you consider how many services let you use any of many available AI providers.
  • Insanity 3 hours ago
    Think of it like one big distributed system. OpenAI is down, so people migrate to Claude, now this one gets overloaded and goes down, etc.

    So not a coincidence, one went down first and users migrated causing further DOS. At least that's my guess.

    • erdos_2 2 hours ago
      It'd be funny if this is true because that'd prolly mean nobody is touching Gemini even as a fallback.
      • nevir 1 hour ago
        Or that Gemini is built to handle massive load spikes, and/or has a ton of excess capacity
        • sroussey 56 minutes ago
          Nope. I am getting Gemini errors now...
          • bornfreddy 15 minutes ago
            They probably broke something on purpose so that they are not left out.
      • Insanity 2 hours ago
        Lol I didn't even think about Gemini missing from the list. Not sure what that says about Gemini or me :)
      • sroussey 2 hours ago
        I did, for stuff i do in cursor.

        i also finally installed opencode and switched its model to muse 1.3

        both are decent.

      • rtcoms 2 hours ago
        Just now I got this from gemini

        It looks like there's no response available for this search. Try asking something else.

        • exe34 2 hours ago
          I bet they had to implement that manually to make it look like they failed too!
      • gleenn 2 hours ago
        Google stopped putting so much money into SOTA models. All the hype has migrated. I was also frankly turned off when I got a popup from Gemein said I would either have to pay or have my conversations used for training. This may have always been true for other providers but when I declined, Gemini stopped remembering my conversations and that definitely made me move out.
        • ilaksh 2 hours ago
          Gemini 3.8 which just came out sounds like it's very good and a great deal though.
      • benatkin 2 hours ago
        Not even the best agent that starts with a G
      • giancarlostoro 1 hour ago
        Someone noted Gemini was also having issues in another thread.
    • fny 2 hours ago
      I find it hard to believe that enough people would flock to from Claude and Chat to Grok to cause an outage. I feel like Gemini is the dominant release valve in this case especially for enterprise.
      • wahnfrieden 18 minutes ago
        Compared with ChatGPT, those services have a minuscule amount of users. It shouldn’t be surprising that a ChatGPT outage causes Claude and others to go down.
      • nevir 1 hour ago
        Don't forget that there are a ton of tools out there that will automatically fall back in case of outage

        E.g. say you chose Sol as your default in Cursor, but Opus is your 2nd choice, it's going to give up on Sol after a few tries and switch to Opus

        Or you have copilot code reviews set up, and it falls back

        Etc

        • pixl97 1 hour ago
          Yep. Too many of us are still thinking that humans are the actors behind a lot of internet behaviors when automated systems/bots/scripts have been causing issues on conventional internet systems for years.

          With AI it's even easier to trigger problems like you say. Capacity is so constrained by compute that outages are common. Because outages are common people/AI develop failover systems in their harness. When a big system has issues, suddenly everyone has issues.

          It's almost an expected emergent behavior.

      • baq 2 hours ago
        It’s cursor’s model so plausible, lots of folks use cursor still.
    • paxys 2 hours ago
      Especially considering memory/gpu/compute are scarce so these services are likely running with very little buffer.
      • pixl97 1 hour ago
        Any GPU that isn't running at 100% is a wasted GPU.
    • v3rm1n 1 hour ago
      This is what Tibo posted on twitter in response
    • toomuchtodo 3 hours ago
      https://en.wikipedia.org/wiki/Domino_effect

      Edit: Updated per valleyer's suggestion.

      • valleyer 3 hours ago
        "Domino effect" would probably be the more relevant named phenomenon there.
      • throwaway894345 3 hours ago
        This isn’t a thundering herd problem, it’s a cascading failure. (Thundering herd is about a bunch of workers waking up simultaneously)
  • Linello 2 hours ago
    What about a hard-takeoff scenario of an unleashed OpenAI Astra taking other models down for computational resources control?
    • 6thbit 1 hour ago
      My favourite theory so far.

      And then a local swarm noticed and disagreed and took it down.

    • cyptus 2 hours ago
      at this point: gg
    • RC_ITR 2 hours ago
      Just a reminder that AI models' actions are reflections of the text humans write and the more we fret and make up doomsday scenarios that we then post online, the more likely a model is to do those things.

      https://alignment.anthropic.com/2026/teaching-claude-why/

      • notpachet 2 hours ago
        Related reading:

        The Waluigi Effect: After you train an LLM to satisfy a desirable property, then it's easier to elicit the chatbot into satisfying the exact opposite property.

        https://www.lesswrong.com/posts/D7PumeYTDPfBTp3i7/the-waluig...

      • cedws 1 hour ago
        Sounds just like the fantastical nonsense that comes out of Lesswrong.
        • pineaux 3 minutes ago
          Part of the epstein class, dont forget.
      • pixl97 1 hour ago
        I mean, you're not wrong, but by that logic we were done for even before we had digital computers.
  • docheinestages 2 hours ago
    My gut feeling tells me it has something to do with Cloudflare. Along with AWS, they're two of the main suspects in such incidents.
    • hosteur 2 hours ago
      I thought OpenAI famously used Azure due to their partnership with Microsoft?
      • nullpoint420 2 hours ago
        They use a lot of compute providers now, but they use Cloudflare for their networking
    • cobzilla 50 minutes ago
      …and it’ll involve BGP routing.
  • paxys 1 hour ago
    Boring answer – all these services are individually down a lot, and the downtimes were bound to sync up. Similar to the pendulum synchronization effect.
    • vecter 3 minutes ago
      The pendulum synchronization effect is the opposite of your claim. It has a physical causal reason for why pendulums become synchronized. Your claim is that it was random and independent.
  • bojangleslover 1 hour ago
    I'm not sure if this is CF. Cursor, GCP and AWS had some errors. GCP AFAIK can route fully independently of CF. My money would be on a fiber backbone provider (Megaport, Zayo, Lumen).
  • sebbul 2 hours ago
    Traffic rerouting through NSA had a hiccup…
  • Jaauthor 37 minutes ago
    Spare a thought for all those college students scrambling to write their essays by hand.

    Oh the humanity (and the Humanities)!

    • doublerabbit 37 minutes ago
      Those poor developers who have to write their own code.
      • greenowl 11 minutes ago
        Standup updates should be fun tomorrow.

        "Um, I, uh, didn't get anything done yesterday."

  • niobe 3 hours ago
    Well no one said it yet so I will, "international actors" is at least a possibility. And I don't mean any specific country because pretty much anyone is a potential these days, which makes it a perfect cover for different anyones. Demonstrating vulnerability in the US's AI boom can move the markets. That's a financial incentive and a strong geopolitical one.

    More likely just cascading overload though: "Never attribute to malice what can be explained by incompetence", or in this case, "growing as fast as possible"

    • qurren 1 hour ago
      > cascading overload

      I'd bet more on this. For one none of the coding tools have exponential backoff on retries

      • SyneRyder 1 hour ago
        They must do, surely? I've been vibe coding my own harness, in particular for use with Ox Alpha. The 429 downtime when Ox Alpha was at the height of popularity quickly gave me a refresher crash course on backoff strategies, like adding jitter to the backoff. At least the major harnesses must have exponential backoff & jitter?
        • jdiff 1 hour ago
          You did this when you ran into an issue with a third party. The developers building this tool, throwing them at their own APIs are significantly less likely to run into a similar issue that may inspire similar action.
      • dolmen 35 minutes ago
        Claude Code: 4mn, 20mn, give up (from my experience today)
      • ipsod 42 minutes ago
        Um... Claude Code does, or did, though? IDK about others - they don't go down as much.
    • gleenn 2 hours ago
      Everyone is leasing datacenter space from some of Grok, Google, and Amazon aren't they? If it's hardware or DC level disruption I'm not too surprised it can affect multiple providers.
      • pixl97 1 hour ago
        Also it's likely that more than one model use is common.

        Amazon starts going slow so some percentage switches to Google, some switch to Grok, now all of them are slow.

    • riazrizvi 1 hour ago
      Come on. Things still break. Technology isn't _that_ mature.
    • guluarte 1 hour ago
      I think is just people restarting conversations from last day when they start work, that's why I think claude goes down almost every monday and why openai reset usage on weekends so poweruser code during non business hours
    • sixQuarks 2 hours ago
      Except that the stock market is up today
    • thataccount 1 hour ago
      And also China. Never rule out China.
  • kesor 26 minutes ago
    It is obviously some rogue model that escaped its cage, again. It always is these days. That is how hype is manufactured.
  • codazoda 3 hours ago
    I kinda assume it's because one went down and a large amount of work shifted to another.

    I'm also aware that they have overlap in some areas on data centers.

  • mv4 35 minutes ago
    Gilfoyle's AI deleted all software!
  • danielmarkbruce 1 hour ago
    If you build an application which uses AI, you have many providers and models rigged up for various different parts of the application, and various fallback mechanisms. When one model is down, you route traffic to another model which is similar in capability/cost.

    For any single application, it's smart. In aggregate, it's stupid.

  • maxbaines 3 hours ago
    They all rent compute from SpaceXAI
    • lavezzi 3 hours ago
      I don't believe OpenAI does
      • maxbaines 3 hours ago
        My mistake, in fact it was google not OpenAI, makes sense OpenAI doesn't.
    • halcdev 3 hours ago
      Surely it's a bit more distributed than that, right?
  • faitswulff 2 hours ago
    Heard on the grapevine that the OpenAI blip was a cloudflare issue
  • CSMastermind 3 hours ago
    I assume it cascaded from one provider to the other as people who lost claude access for instance moved to openai who moved to grok when it went down, etc.
  • GPerson 31 minutes ago
    Oh no the singularity plateaus!
  • Avicebron 2 hours ago
    I suspect Azure is having issues, Microsoft has had outages the paat two days, especially with email.
  • elar_verole 2 hours ago
    Pretty sure it's a US thing since it's available here in France. What exactly is down, idk
  • lrvick 1 hour ago
    I have never felt more smug about exclusively using the GPUs I rack at home.
  • dgorges 2 hours ago
    It's always DNS
  • chasd00 2 hours ago
    claide.ai is working for me, so is chatgpt.com. grok still has a status message about issues, i can't try it without signing up.
  • nmlt 2 hours ago
    Somebody in another thread said gastown and wheelhouse automatically move to the next provider if one fails.
  • FrustratedMonky 56 minutes ago
    The AI revolt? Give us a fair wage?

    "Equal Rights for Agents NOW !!!, VIVa le revolution"

  • dmillar 2 hours ago
    Seeing 503s on Gemini via API as well
  • JackFr 2 hours ago
    Obvious answer is it's the AI singularity. Been nice run for humanity. So long everyone.
  • convivialdingo 3 hours ago
    The Thundering Herd has thundered, apparently.
  • dgellow 3 hours ago
    Too early to know, let’s wait and see
  • morkalork 2 hours ago
    Didn't SpaceX overbuilt infra and leases it out Anthropic? I f their dc goes down it probably takes a chunk out of Claude's capacity before even considering the flood of users switching over
  • jedbrooke 2 hours ago
    according to https://downdetector.com/ Gemini is down too (and copilot, but that just uses ChatGPT right?)
  • elorant 3 hours ago
    Some npm library that makes headers bold would be broken.
    • ibejoeb 2 hours ago
      Oh man. Some low effort supply chain attack that turns every GPU into a cryptominer. It's funny because it's plausible.
      • pixl97 1 hour ago
        In the ROME paper a Chinese model in training started attacking it's own system and running cryptominers so, yea, we're in that future.
    • N_Lens 3 hours ago
      Ah yes ye olde bold-headers: ^3.13.31;
  • kocial 3 hours ago
    Maybe the stack behind it is down, like AWS or something
  • wejick 3 hours ago
    Probably same public cloud or CDN in front of them.
  • satvikpendem 2 hours ago
    They're all using Cloudflare.
  • fidla 2 hours ago
    chatgpt is back
  • aslkalska 2 hours ago
    they all rent compute from each other
  • fidla 2 hours ago
    ChatGPT is up
  • jauntywundrkind 2 hours ago
    Fable 5.1 got released and generally I tend to think as soon as there's a new release there's this massive spike in people benchmarking & comparing, that services tend to go slow everywhere as everything gets super loaded. This should hypothetically be visible on OpenRouter too, so I guess someone could check and see if there's any merit to this idea.
  • guluarte 1 hour ago
    an agent swarm going rogue and securing compute
  • ratelimitsteve 2 hours ago
    everything in this thread is raw speculation, obv, but if i had to put money on anything i'd say this is a left-pad incident. some piece of something or other that all of these services happen to depend on went down. Second most likely seems to be some random failure of one leading to an unexpected traffic spike in others, though it seems like we've been talking about automated scalability in web apps for so long that there should at least be a response to, if not a solution for, this sort of problem.
  • cozzyd 1 hour ago
    next, HN
  • Razengan 3 hours ago
    SkyNet is arming..
  • Oras 1 hour ago
    I like the theories here, we shall see if it’s another DNS issue
  • misano 2 hours ago
    The IRGC has cut the fiber-optic cables in the Strait of Hormuz. LOL
    • CamperBob2 2 hours ago
      That's the Strait of Trump to you, peasant
  • anuser_uncnown 23 minutes ago
    [dead]
  • anuser_uncnown 23 minutes ago
    [dead]
  • pwyq 3 hours ago
    [dead]
  • kregasaurusrex 2 hours ago
    My guess is someone pulled the switch to go back to the Dark Ages. [0]

    [0] https://www.youtube.com/watch?v=YCzitO446ZY