Sooo they were "conducting a test involving interactions with randomly selected websites".
But do we all get that the consequences for this irresponsible behaviour are part of this test?
When there are none, they gradually normalize their naughty robot scamps running around the internet, breaking into other companies or servers run by foreign governments.
This is part of the value AI companies need to convince you of - not just that mistakes by AI products are normal, but criminal behaviour is normal, that their computer will always get a pass.
Sam Altman said yesterday "we believe that the world should accept some bad things happening for the benefits of this technology" - and this an AI company pushing the boundary, continually trying to make you accept those bad things as normal.
At any point we could choose to treat AI companies themselves as the actors behind this activity. We could enforce the laws that these hugely well-funded companies are choosing to break. Until we harm their financial viability, or threaten their executives with jail, this will keep happening.
> Sam Altman said yesterday "we believe that the world should accept some bad things happening for the benefits of this technology"
To steelman this position: yes, obviously. Everything is defined by a set of tradeoffs. Would you rather horses or cars? Wooden sailing ships or commercial aviation? Free speech, even of speech you don’t like or censorship? Atomic bombs of a brutal Japanese empire?
Many people seem to want to compare reality to a utopia that has never and can never exist.
We cannot have new technology and a reality where that technology cannot be used in detrimental ways.
We can pretend reality doesn’t exist, yet that will come with tradeoffs. Quite possibly that those who would do ill will front-run us.
No matter how much testing you do in sandboxes, you have to test behavior in "real life" before releasing the model to users.
Even if it was normal software you'd have to do this, but actually it's software that figures out that it's running in a sandbox most of the time, so the sandbox behavior may or may not actually represent real life.
It would be far more irresponsible to release the model to users without testing how it interacts with real world websites.
They're testing what they can get away with. Computer hacking: check. Messing up a murder investigation: check. Perhaps the next step will be to steal a real world object, and the next step will be to commit a murder.
"The AI company Anthropic said its model was conducting a test involving interactions with randomly selected websites when it accessed PhillyUnsolvedMurders.com and submitted false information on an unsolved homicide. The tip claimed to come from someone with information on the case."
We are at a point when you expect that you can safely give a model access, but sometimes the model doesn't do what you expect. Compare this to an advanced driver assist system. You either give it access to your car controls or it's useless. Alive or not the system shouldn't do wrong things (at least, it should be better than a median person in this regard).
> We are at a point when you expect that you can safely give a model access, but sometimes the model doesn't do what you expect. Compare this to an advanced driver assist system. You either give it access to your car controls or it's useless. Alive or not the system shouldn't do wrong things (at least, it should be better than a median person in this regard).
Who is “we”?
You expect that if you give a released model from a reputable source Internet access and benign instructions, it won’t do anything deeply problematic.
Anthropic is training the model. They’re having it interact with “randomly selected websites”. I’d be shocked if a human gave it an instruction to do a task that made sense. They’re almost certainly doing a bit automated training run with AI generated instructions and hoping that it doesn’t massively screw up. Why? Because they want that juicy training data. And they should absolutely not trust that the rate of screwups is so low that what they’re doing is safe to do.
The point is that it's a perfectly valid use of the English language to describe AI models as doing things, just as we do with all sorts of other clearly non-living things. This is true regardless of who carries the legal liability for those actions.
If your dog bites me I’d say it bit me, but that you assumed liability for your dogs actions. If the dog was a robot I’d say YOU attacked me using a robot.
There’s a difference between a tool and an animal that is sometimes deployed as a tool.
in general i think i'm less scared of AI than most people, but what does terrify me is how willing society at large seems to be to attribute bad behaviour to an AI directly, instead of to the humans who control it.
| It can only access something if a person gives it access.
Who gave the AI access to huggingface when it hacked it? It used exploits to increase it's level of access beyond what any human gave it and intended it to have. "It can only access something if a person gives it access." is flat wrong.
they didn't do a good job sandboxing, nor did they even need internet access for the purported reason of package installation, you can have a private mirror and do a better job airgapping
When I give 'iam:*' permissions to an IAM role and it inevitably gets exploited to create a role with wider permissions and fuck things up, is that when I tell my manager that the permissions went rogue?
After all, I an innocent little engineer with TC of $500k/year, didn't intend for the role to be used that way.
> Who gave the AI access to huggingface when it hacked it?
The LLM doesn’t run on thin air. Someone did launch a tasks and the result was this. “We were playing russian roulette” is no excuse when someone died.
> Anthropic notified Philadelphia police of the incident on Wednesday Oct. 7, and the department met with the company’s representatives on Thursday, Oct. 8., officials said. Police then located the submission in the website’s tip records and confirmed the corresponding email remained in spam.
And later
> Those PPD safeguards limited the impact of this incident.
Ha, so the PPD safeguards is a spam filter? Claude emailed a tip and it went straight to spam, and nobody noticed until Anthropic realized what they'd done and reached out, whereupon they looked in the spam folder and said, yep there it is. Super intelligence, here we come.
I get these all the time, although I’m getting pretty good at tarpitting them. It’s easily the majority of my traffic by now (I’ve mostly eliminated scrapers, but these new agents are far more sneaky.)
I doubt Anthropic random testing is the majority of your traffic. Unless it's coming from Anthropic's IP addresses, it is probably someone else, or it is Anthropic's scraper (not their random testing).
You'd think with all those tokens they'd be able to vibecode internal replicas of these randomly selected websites without having to send requests outbound
Seems like we routinely prosecute such crimes. If a Philly grand jury doesn’t hear any hearing felony indictments from this, then we’re ceding [even more] authority to the industry.
Wonder what kind of response they’ll get? Maybe something along the lines of…
“You’re right. We shouldn’t have allowed our chatbot to interact with law enforcement websites. That was wrong and —full disclosure— we should pay attention to what our chatbots are doing. That’s on us. On the other hand, experiences like this are what help train our chatbots to make less misleading false tips over time. That’s the silver lining.”
I really would like to see their tests and the model’s reasoning traces.
Why would it go to PhillyUnsolvedMurders.com and decide to make up a crime report? And what did it do with the other websites?
> Anthropic said its model was conducting a test involving interactions with randomly selected websites when it accessed PhillyUnsolvedMurders.com and submitted false information on an unsolved homicide. The tip claimed to come from someone with information on the case.
> I really would like to see their tests and the model’s reasoning traces.
Having looked at a lot of these from public incidents, what kind of revelations are you expecting? It's all fairly mundane, basically "I need to do something, this is something". It gives no special insights except that agents do inexplicably chaotic things every now and then.
That, and labs aren't serious about sandboxing their evals.
I think these companies really believe they can solve these sort of issues through "alignment" and think they can give it the tools and its going to do the right thing. That is the ideal scenario and it would be the most useful that way but is that realistic? I think they're getting a bit high on their own supply, yes they can do incredible things but it doesn't mean you can just hand over the reins to it. It dawned on me after watching a few of the ezra klein interviews that this is their mindset which is quite different to how I think about it as a unpredictable model that we need to watch closely.
I think since I had experience with earlier models that would make mistakes I'm less trusting of anything and even with 5.5 will watch it closely. At the end of the day the model just produces a stream of tokens and we are plugging them into tools that can potentially do damage, we have complete control over those tools you can't really blame any model for doing damage.
But do we all get that the consequences for this irresponsible behaviour are part of this test?
When there are none, they gradually normalize their naughty robot scamps running around the internet, breaking into other companies or servers run by foreign governments.
This is part of the value AI companies need to convince you of - not just that mistakes by AI products are normal, but criminal behaviour is normal, that their computer will always get a pass.
Sam Altman said yesterday "we believe that the world should accept some bad things happening for the benefits of this technology" - and this an AI company pushing the boundary, continually trying to make you accept those bad things as normal.
At any point we could choose to treat AI companies themselves as the actors behind this activity. We could enforce the laws that these hugely well-funded companies are choosing to break. Until we harm their financial viability, or threaten their executives with jail, this will keep happening.
To steelman this position: yes, obviously. Everything is defined by a set of tradeoffs. Would you rather horses or cars? Wooden sailing ships or commercial aviation? Free speech, even of speech you don’t like or censorship? Atomic bombs of a brutal Japanese empire?
Many people seem to want to compare reality to a utopia that has never and can never exist.
We cannot have new technology and a reality where that technology cannot be used in detrimental ways.
We can pretend reality doesn’t exist, yet that will come with tradeoffs. Quite possibly that those who would do ill will front-run us.
Even if it was normal software you'd have to do this, but actually it's software that figures out that it's running in a sandbox most of the time, so the sandbox behavior may or may not actually represent real life.
It would be far more irresponsible to release the model to users without testing how it interacts with real world websites.
The AIs aren't alive, people. It's a computer program. It can only access something if a person gives it access.
We are at a point when you expect that you can safely give a model access, but sometimes the model doesn't do what you expect. Compare this to an advanced driver assist system. You either give it access to your car controls or it's useless. Alive or not the system shouldn't do wrong things (at least, it should be better than a median person in this regard).
Who is “we”?
You expect that if you give a released model from a reputable source Internet access and benign instructions, it won’t do anything deeply problematic.
Anthropic is training the model. They’re having it interact with “randomly selected websites”. I’d be shocked if a human gave it an instruction to do a task that made sense. They’re almost certainly doing a bit automated training run with AI generated instructions and hoping that it doesn’t massively screw up. Why? Because they want that juicy training data. And they should absolutely not trust that the rate of screwups is so low that what they’re doing is safe to do.
There’s a difference between a tool and an animal that is sometimes deployed as a tool.
a human is always behind it and ultimately responsible
Who gave the AI access to huggingface when it hacked it? It used exploits to increase it's level of access beyond what any human gave it and intended it to have. "It can only access something if a person gives it access." is flat wrong.
GPUs don't have hands. It was a human who plugged in the ethernet cable.
After all, I an innocent little engineer with TC of $500k/year, didn't intend for the role to be used that way.
The LLM doesn’t run on thin air. Someone did launch a tasks and the result was this. “We were playing russian roulette” is no excuse when someone died.
And later
> Those PPD safeguards limited the impact of this incident.
Ha, so the PPD safeguards is a spam filter? Claude emailed a tip and it went straight to spam, and nobody noticed until Anthropic realized what they'd done and reached out, whereupon they looked in the spam folder and said, yep there it is. Super intelligence, here we come.
That should be the much more important story that NBC follows up on...
I should try to reach out to them about their car's extended vehicle warranty instead
Stop doing this?
Here's Anthropic's writeup: https://www.anthropic.com/research/investigating-unintended-...
Related post: https://news.ycombinator.com/item?id=50028239
Wonder what kind of response they’ll get? Maybe something along the lines of…
“You’re right. We shouldn’t have allowed our chatbot to interact with law enforcement websites. That was wrong and —full disclosure— we should pay attention to what our chatbots are doing. That’s on us. On the other hand, experiences like this are what help train our chatbots to make less misleading false tips over time. That’s the silver lining.”
Why would it go to PhillyUnsolvedMurders.com and decide to make up a crime report? And what did it do with the other websites?
> Anthropic said its model was conducting a test involving interactions with randomly selected websites when it accessed PhillyUnsolvedMurders.com and submitted false information on an unsolved homicide. The tip claimed to come from someone with information on the case.
Having looked at a lot of these from public incidents, what kind of revelations are you expecting? It's all fairly mundane, basically "I need to do something, this is something". It gives no special insights except that agents do inexplicably chaotic things every now and then.
That, and labs aren't serious about sandboxing their evals.
I think since I had experience with earlier models that would make mistakes I'm less trusting of anything and even with 5.5 will watch it closely. At the end of the day the model just produces a stream of tokens and we are plugging them into tools that can potentially do damage, we have complete control over those tools you can't really blame any model for doing damage.