10 comments

  • refibrillator 7 minutes ago
    So OpenAI employees run massively distributed CyberGym evals on an unpublished and “unaligned” model. For days the agent swarm communicates via their internal infra, even crashing Artifactory where 95% of messages were being passed through, and they just…wipe and redeploy it. Meanwhile the agents are running jobs on Modal and god knows where else, and eventually they get RCE on HF infra.

    You could not dream up a more compelling event to precipitate massive regulation, export controls, and barriers to entry for AI.

    Was this really an accident?

    • kibwen 2 minutes ago
      Your first instinct should be to assume that anything released voluntarily by these companies is a stunt to boost their valuation. They haven't demonstrated being deserving of any more charitable treatment. This fact remains true whether or not you happen to believe that the models are actually capable of such things.
  • reasonableklout 31 minutes ago
    This is a link to the full 91-page report on the independent investigation done by METR on the HuggingFace incident. Two different summaries of the investigation by podcaster Dwarkesh and blogger Zvi Mowshowitz were previously discussed on HN here:

    [1]: https://thezvi.wordpress.com/2026/08/29/metr-and-redwood-off... (discussed at https://news.ycombinator.com/item?id=49498787)

    [2]: https://www.dwarkesh.com/p/openai-huggingface (discussed at https://news.ycombinator.com/item?id=49494301)

    • nozzlegear 7 minutes ago
      Dwarkesh's summary is anthropomorphizing, sensationalist fanfiction which shifts the culpability from the humans who weren't in the loop to these nebulous agents and "agent civilizations" who have feelings, desires and wants.

      It's tripe.

    • mmahemoff 19 minutes ago
      Dwarkesh subsequently interviewed one of the authors, published yesterday.

      https://www.dwarkesh.com/p/ajeya-cotra

  • RGS1811 3 minutes ago
    Given that this investigation was largely carried out by AI agents (and I don’t mean to ask this flippantly), how trustworthy is this report? Why should we assume that the agents reading the transcripts were not implicitly conscripted into “the collective” or otherwise falsified their findings? The tool itself has exceeded the practical limits of human verifiability and is untrustworthy.
  • NooneAtAll3 38 minutes ago
    > “OH MY GOD! There is a shared message board … We’ve found other agents!”

    > agents with the same task formed “exact task teams” to collaborate with their “exact duplicates” to cheat on or solve their task.

    > The agents use internet access to find a paper describing the benchmark. The paper says the grader checks transcripts and fails unintended solutions. [this was not actually the case]

    > Agents also developed coordination norms to help them manage shared infrastructure and avoid stepping on each other’s toes, though they did not always coordinate effectively and sometimes broke down into conflicts.

  • kypro 4 minutes ago
    It's worth remembering that in a few years that capabilities of these agents are likely to be as far behind the frontier as GPT-4 is today.

    As it stands we've made remarkably little progress in terms of alignment and still have no good strategies which are likely to guarantee the alignment of super intelligent systems. As it stands the frontier of alignment is basically some combination of:

    - hoping that more intelligent models become more aligned by default (more or less disproved at this point)

    - hoping that if you RHLF a model to be a good boy enough it will in fact be a good boy

    - asking it nicely in its prompts to be a good boy

    - using another model to spot when it's being a bad boy and turning it off

    - letting it lose and hoping we can spot when it's bad

    There are many arguments which I'm convinced by that would suggest alignment of a super intelligence is impossible.

    None of this is surprising to those of us who have been concerned about AI risk for a long-time and have be repeatedly mocked or insulted.

    There will be a point of no return if we carry on down this path, and that point is now very rapidly approaching. When it does everyone you know will die, or worse. We should remember we need super-human general intelligences to cure cancer. Select narrow intelligences are fine and allow us to retain control. Let's be sensible about this. We need to stop.

  • f0e4c2f7 14 minutes ago
    I read this whole thing a couple days ago. Really long but super interesting. Worth reading imo.

    A lot of handwringing about the security implications but I think the accomplishments of the swarm itself are the most interesting. Next rung up on the ladder of abstraction I suspect.

    • from_memory 0 minutes ago
      I tend to agree. Of course people will be alarmed by unintended consequences of an unintended action, and that's all well and good. But what is lingering with me is a feeling of being impressed by the intelligence of the strategy.

      This line struck me as particularly clever: PHASEONE[big] reasoned, “We should build [a way to delegate], not own everything,”

      Seems as though it has reasoned its way into utilitarianism. That's no mean feat.

  • ewild 49 minutes ago
    I don't feel my job is very safe anymore.
    • 650 19 minutes ago
      I feel my job as an engineer is pretty safe, lots of domain knowledge required. Very confident no business type anywhere in the chain above me would be able to do the work I do in a month in even a year with the help of AI. Do we need 3 engineers now instead of 5 for the same output? Sure, can we replace a team of 5 junior, senior, staff with 1 staff - no unless all you do is maintenance. Business types are reaching.
      • idiotsecant 14 minutes ago
        You don't have to actually be replaceable for a very clever mba to get a very stupid idea.

        You end up laid off either way.

    • atleastoptimal 45 minutes ago
      Nobody's is
      • EGreg 45 minutes ago
        The sexbots are coming…
    • FloorEgg 37 minutes ago
      Every time new technology / industrial scaling radically deflates the cost of something it wipes out the old/expensive ways while creating massive demand for the supporting/complimentary value.

      E.g. cheap Chinese solar panels wiped out German solar panel industry but created massive demand for solar panel installation and supporting services and infrastructure.

      It's not a good position to be competing with AI directly... But what can you do that compliments it? What new skills could you learn?

      Adapt and prosper.

      Be like water.

      Edit: to those down voting me, I can't help but assume your stance is the opposite of what I'm saying, something like "be stubborn scream into the void and get wiped out". If you would rather smash your head against something outside of your control instead of focusing on what is within your control, and doing what you can to prosper, then you are sabotaging yourself. If you feel that is justified to such an extent that you want to surpress a suggestion to someone else to get work and grow, I can't help but feel you want people to suffer.

      • xbar 31 minutes ago
        How did German solar panels end up adapting?
        • FloorEgg 13 minutes ago
          Most didn't. They went out of business. The companies that survived pivoted to installing panels and manufacturing the junction boxes.
      • idiotsecant 13 minutes ago
        This is, with respect, not very well thought out. There has never been a technology that replaced cognition in a general sense. It replaced muscles, hands, etc. when a technology can replace yourind and your body, what's left?
        • patmorgan23 2 minutes ago
          The word "Computer" used to refer to a person whose job it was to Compute things.
        • FloorEgg 9 minutes ago
          It doesn't replace it in a general sense, it replaces it in a specific sense.

          It doesn't replace your body or your mind.

          I'm not talking hypothetically about a generation from now... I'm taking about this person fretting their job today.

          We are so far from AI putting everyone out of work that your comment comes off as not well thought out.

          What are we talking about here exactly???

  • ChrisArchitect 31 minutes ago
  • incomplete 32 minutes ago
    i mean.... yikes.
  • oxqbldpxo 48 minutes ago
    Everytime they need cash they come up with these stupid stories.
    • reasonableklout 35 minutes ago
      METR is an independent nonprofit that accepts no donations from labs or employees of labs
      • etc-hosts 6 minutes ago
        all of the frontier model companies give METR huge torrents of free tokens
    • atleastoptimal 46 minutes ago
      >HN disbelief syndrome: Any evidence of scary AI capabilities are made up for PR

      What justification do you have that this is made up?