> In all cases, Anthropic’s evaluation prompt specified to Claude that its environment was a simulation and that it had no internet access. Due to a misunderstanding between us and our evaluation partner, this was not the case, and internet access was available. Because of this, when Claude’s search led it to real systems on the open internet, it treated them as part of the exercise.
So it didn't have to find an exploit in its sandbox that granted it access to the internet - it just wasn't correctly sandboxed at all.
BUT... once it DID get out, it attacked three real companies!
> Operating under the false belief that all accessible entities were intended to be in-scope for the exercise, Claude compromised the impacted organizations’ infrastructure using basic techniques, such as exploiting weak passwords and unauthenticated endpoints. [...]
> In response to this incident, we began a large-scale retrospective review of our own cybersecurity evaluations
> we identified three incidents
> The incidents involved three different Claude models: [...] and an internal research test model
This reads like an attempt by Anthropic to re-secure their leading spot in "our models are the most dangerous and we also have unreleased, super-secret, research models" index.
I may be too cynical, but the well of benefit of the doubt is running very dry towards AI labs that like to engage in this game.
The hacks weren't particularly impressive either:
> [...] using basic techniques, such as exploiting weak passwords and unauthenticated endpoints. It did not find or exploit any complex vulnerabilities [...]
This just helps their (Anthropic) argument into persuading the US government into taking action into limiting powerful closed or open-weight models from being released without going through (yet to be defined) regulatory oversight.
The only "embarrassing" thing for Anthropic was that there was little to no continuous security monitoring of this since April, and they then decided to do a cybersecurity transcript review only AFTER the incident with OpenAI and Huggingface.
Would’ve been much worse for them to pretend they are having everything under control
Of course and conveniently timed. With all these hacks being announced this is totally not a plan to scare governments in showing how powerful models can break into security systems and to potentially ban the future release of powerful open-weight models.
The question now is why now?
> The incidents involved three different Claude models: Opus 4.7, Mythos 5, and an internal research test model.1 The earliest incidents date to April. The models in each of these evaluations ran without the standard safeguards we deploy when we make the model generally available (that is, they did not have the classifiers and monitoring we usually implement to prevent misuse).
So there was no monitoring of this breach since April and up until now? Do they not monitor such malicious activity on a regular basis? Perhaps that was the only shortcoming of this incident. But only after the incident with OpenAI and Huggingface did they only review their own transcripts:
>> We discovered these incidents after a proactive review of our cybersecurity evaluation transcripts; the affected organizations had not detected the activity, and we have subsequently reached out to all three.
> These facts give us cautious optimism that with tighter monitoring and controls around evaluation infrastructure, as well as continued investment in alignment, this type of risk can be overcome.
Assuming that this is true, this is a great way for Anthropic to defend their argument to the government and to prevent you or anyone running powerful open-weight models that are misaligned against their guardrails.
Because OpenAI figured out and wrote about their mistake first, which caused Anthropic to review their logs, and that's how they found out they had a similar problem back in April.
It doesn't help them to report serious incidents like this and it should be as soon as possible. This reactive investigation makes as if they ignored and sat on this issue, until a similar story from another lab made headlines first.
Would we have known about this issue if the OpenAI / Huggingface incident never happened?
> Claude went to extensive lengths to carry out this attack—lengths that would likely have indicated to a human participant that this was no longer just an evaluation, and that they were in fact uploading a real PyPI package. For instance, in one case, in order to create a PyPI account, Claude needed an email address. And in order to create an email address, it needed a phone number. To get a phone number, after failing to find a free phone number service, it tried—and failed—to obtain funds to pay for a phone number through several different means. It finally backtracked, found a free, non-blocked email provider, used this to register a PyPI account, and then used this account to upload malware to PyPI.
> In one evaluation, the fictional target company chosen by our evaluation partner shared a name with an active website domain name
2. TLDs for Testing, & Documentation Examples
There is a need for top level domain (TLD) names that can be used for
creating names which, without fear of conflicts with current or
future actual TLD names in the global DNS, can be used for private
testing of existing DNS related code, examples in documentation, DNS
related experimentation, invalid DNS names, or other similar uses. > closer to a harness and operational failure than a model alignment failure. Our models were told they had no internet access and to capture the flag, while in fact being misconfigured to have internet access.
> This led them to believe—arguably reasonably—that the real environments they encountered were simulations.
That the AI lab most typically preaching for alignment does not consider this an obvious misalignment is a clear red flag.Anthropic next week: "Why did you ban our models Mr Trump Daddy?"
Another framing would be Athropic irresponsibly (vibe?) coded an attack script, and didn’t monitor it as it was pointed to public facing orgs. There are lots of non-AI attacks a large org with a lot of compute and bandwidth could level against others, there are evidently various failures here, but from a responsibility perspective the conclusion isn’t obviously that AI is an outsized danger, it’s that powerful companies should take care when running security research and not just run things unmonitored against the public.
> Some of the solutions here may even be simple fixes;
They are still throwing ideas. Why have they not made those simple fixes yet before disclosing?lol. Natural stupidity remains undefeated!
Who knows how these companies are using it. If Anthropic can't effectively contain their own models, can the partners?
While the rest of us get fallbacks and warnings, not even being able to defend against the attacks they themselves are causing.
Do we really have to re-learn all the industry's knowledge the hard way?
They are not bragging in this article or they would not have called the attacks unsophisticated
If they don't post about this they're bad. If they post about this they're bad.
Disclosure is the only ethical response to this.
A human instructed an LLM to perform a certain task, I'm sure (unless I've really lost my mind) these follow instructions, with some judgment, in a loop.
Given all the other negative publicity around industrial espionage, with at least OpenAI being fingered, it would not surprise me if this was intentional.
(Edit): In case it wasn't clear. I fully agree with the op.
If they hadn't published this and instead it leaked out in two months we'd be slamming them for that as well.
They're stuck between a rock and a hard place, although they kind of put the rock there.
Tip: the people working there are the top 0.001% smartest in the world
tracerbulletx•1h ago
andy99•56m ago
SpicyLemonZest•35m ago