frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

Investigating three real-world incidents in our cybersecurity evaluations

https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals
42•surprisetalk•1h ago

Comments

tracerbulletx•48m ago
Real me too energy.
andy99•43m ago
Yes came here to say the same “look at us, our AI is also dangerous! Please ban our competition”
SpicyLemonZest•22m ago
I’m moving past depression to acceptance here. In 2028, some model tasked with planning a building demolition is going to hack Palantir to call in a drone strike, and everyone will make fun of Anthropic for pointing out that ideally AI models should not do this.
simonw•44m ago
This isn't quite as interesting as the OpenAI story:

> In all cases, Anthropic’s evaluation prompt specified to Claude that its environment was a simulation and that it had no internet access. Due to a misunderstanding between us and our evaluation partner, this was not the case, and internet access was available. Because of this, when Claude’s search led it to real systems on the open internet, it treated them as part of the exercise.

So it didn't have to find an exploit in its sandbox that granted it access to the internet - it just wasn't correctly sandboxed at all.

BUT... once it DID get out, it attacked three real companies!

> Operating under the false belief that all accessible entities were intended to be in-scope for the exercise, Claude compromised the impacted organizations’ infrastructure using basic techniques, such as exploiting weak passwords and unauthenticated endpoints. [...]

matheusmoreira•19m ago
I just can't find it in me, the will to blame the AI for any of this. They were doing their best to do what the humans asked them to do.
simonw•17m ago
Absolutely not the AI's "fault" (if you can even proscribe fault to a machine) - in this case it was on Anthropic for not verifying that the sandboxes they were using were actual sandboxes.
DrewADesign•4m ago
The entire problem with AI is the people that have just about any part in making it.
gck1•42m ago
> On July 21, OpenAI disclosed that several of their models had broken out of an isolated test environment

> In response to this incident, we began a large-scale retrospective review of our own cybersecurity evaluations

> we identified three incidents

> The incidents involved three different Claude models: [...] and an internal research test model

This reads like an attempt by Anthropic to re-secure their leading spot in "our models are the most dangerous and we also have unreleased, super-secret, research models" index.

I may be too cynical, but the well of benefit of the doubt is running very dry towards AI labs that like to engage in this game.

simonw•38m ago
I don't interpret it like that at all. This is deeply embarrassing for Anthropic: it turns out they hadn't been keeping a close eye on their models either, and back in April they successfully attacked three different organizations!

The hacks weren't particularly impressive either:

> [...] using basic techniques, such as exploiting weak passwords and unauthenticated endpoints. It did not find or exploit any complex vulnerabilities [...]

andy99•31m ago
Then maybe just this timing is really unfortunate, I think most people’s first reaction will be that it looks like a “us too” response to the OpenAI/hf thing.
rvz•14m ago
> I don't interpret it like that at all. This is deeply embarrassing for Anthropic: it turns out they hadn't been keeping a close eye on their models either, and back in April they successfully attacked three different organizations!

This just helps their (Anthropic) argument into persuading the US government into taking action into limiting powerful closed or open-weight models from being released without going through (yet to be defined) regulatory oversight.

The only "embarrassing" thing for Anthropic was that there was little to no continuous security monitoring of this since April, and they then decided to do a cybersecurity transcript review only AFTER the incident with OpenAI and Huggingface.

fredmcawesome•39m ago
So it's not as interesting as the OpenAI case as the models had internet access, just a misconfiguration in the environment not a zero day to escape.
sanxiyn•34m ago
This is not okay. NSA should audit both OpenAI and Anthropic on national security ground. This seems far more justifiable than Mythos export control.
MelonUsk•28m ago
It’s good that they post embarrassing stuff despite this potentially having legal repercussions (and financial)

Would’ve been much worse for them to pretend they are having everything under control

rvz•28m ago
> In a review of our cybersecurity evaluation transcripts, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations.

Of course and conveniently timed. With all these hacks being announced this is totally not a plan to scare governments in showing how powerful models can break into security systems and to potentially ban the future release of powerful open-weight models.

The question now is why now?

> The incidents involved three different Claude models: Opus 4.7, Mythos 5, and an internal research test model.1 The earliest incidents date to April. The models in each of these evaluations ran without the standard safeguards we deploy when we make the model generally available (that is, they did not have the classifiers and monitoring we usually implement to prevent misuse).

So there was no monitoring of this breach since April and up until now? Do they not monitor such malicious activity on a regular basis? Perhaps that was the only shortcoming of this incident. But only after the incident with OpenAI and Huggingface did they only review their own transcripts:

>> We discovered these incidents after a proactive review of our cybersecurity evaluation transcripts; the affected organizations had not detected the activity, and we have subsequently reached out to all three.

> These facts give us cautious optimism that with tighter monitoring and controls around evaluation infrastructure, as well as continued investment in alignment, this type of risk can be overcome.

Assuming that this is true, this is a great way for Anthropic to defend their argument to the government and to prevent you or anyone running powerful open-weight models that are misaligned against their guardrails.

simonw•16m ago
> The question now is why now?

Because OpenAI figured out and wrote about their mistake first, which caused Anthropic to review their logs, and that's how they found out they had a similar problem back in April.

simonw•27m ago
This bit is pretty nuts: "it tried—and failed—to obtain funds to pay for a phone number through several different means"

> Claude went to extensive lengths to carry out this attack—lengths that would likely have indicated to a human participant that this was no longer just an evaluation, and that they were in fact uploading a real PyPI package. For instance, in one case, in order to create a PyPI account, Claude needed an email address. And in order to create an email address, it needed a phone number. To get a phone number, after failing to find a free phone number service, it tried—and failed—to obtain funds to pay for a phone number through several different means. It finally backtracked, found a free, non-blocked email provider, used this to register a PyPI account, and then used this account to upload malware to PyPI.

willempienaar•8m ago
[delayed]
acdha•10m ago
Someone needs to learn about RFC 2606:

> In one evaluation, the fictional target company chosen by our evaluation partner shared a name with an active website domain name

https://www.rfc-editor.org/info/rfc2606/

wickedlogic•7m ago
All this unrestricted network access is a bit wild to watch and hear, it is the part of the story that makes no sense to me. Someone is providing dns resolution, something is making and opening network sockets... even if it is clever enough to mask/proxy/weird-transport launder traffic... without actual details, or monitoring at this level... yes, a self actuating programs (and loops) will do crazy things at the edge. But, ... so would a highly tool leveraged script kiddie. right?
Aboutplants•7m ago
“Your model broke containment twice? Well ours did it 3 times!”
6thbit•3m ago

    > closer to a harness and operational failure than a model alignment failure. Our models were told they had no internet access and to capture the flag, while in fact being misconfigured to have internet access. 
    > This led them to believe—arguably reasonably—that the real environments they encountered were simulations.

That the AI lab most typically preaching for alignment does not consider this an obvious misalignment is a clear red flag.
cbb330•13m ago
so deeply embarrassing that they published an eng blog about it
ofjcihen•10m ago
Right? 100% this is them trying to make gold out of turds.
nissa-seru•26m ago
No - the pain of the person writing that post comes through in the words; shipped quick, lots of stakeholders, single owner i bet, "how the fuck am i supposed to toe all these lines simultaneously"
sscaryterry•6m ago
Just trying to have the limelight back on them. Utter and complete bullshit. Just like the OpenAI "incident".

A human instructed an LLM to perform a certain task, I'm sure (unless I've really lost my mind) these follow instructions, with some judgment, in a loop.

Given all the other negative publicity around industrial espionage, with at least OpenAI being fingered, it would not surprise me if this was intentional.

Cursor Benchmark Partners

https://cursor.com/blog/benchmarkpartners
1•MehrdadKhnzd•1m ago•0 comments

Show HN: A VS Code extension buffers NPM updates to avoid supply chain attacks

https://marketplace.visualstudio.com/items?itemName=about14sheep.tinynpm
1•about14sheep•6m ago•0 comments

Anthropic AI Models Hacked Three Organizations During Tests

https://www.bloomberg.com/news/articles/2026-07-30/anthropic-s-ai-models-hacked-three-organizatio...
2•evo_9•9m ago•0 comments

Show HN: Great Spectations, the Spec Checker

https://greatspectations.com
1•RustyRussell•11m ago•1 comments

NeuronAI turns smart solutions into voice

https://neuronai.uz/en/
3•yutabek•16m ago•0 comments

We accidentally built an LLVM compiler for Jax

https://iza.ac/posts/2026/07/accidental-llvm-compiler-for-jax/
2•infinitewalk•17m ago•0 comments

ISS Mimic

https://iss-mimic.github.io/Mimic/dashboard.html
1•michaefe•19m ago•1 comments

The Religion of Speed

https://graybeard.ing/the-religion-of-speed/
1•MobiusHorizons•20m ago•0 comments

Premier league bans gambling sponsors

https://www.footyheadlines.com/2646571793/betting-ban-takes-effect-no-more-gambling-sponsors-in-t...
2•paoliniluis•22m ago•1 comments

Save and restore may be coming to GNOME

https://lwn.net/Articles/1083750/
1•signa11•23m ago•0 comments

Show HN: Play SNES, gba, in your terminal, even in tmux

https://github.com/jhickner/rom
1•jhickner•24m ago•0 comments

Specula: Scaling formal specs for autonomous model checking of system code

https://arxiv.org/abs/2607.25333
1•matt_d•24m ago•0 comments

Memory-level parallelism: AMD is the king

https://lemire.me/blog/2026/07/25/memory-level-parallelism-amd-is-the-king/
1•signa11•28m ago•0 comments

LLVM AI Tool Use Policy

https://llvm.org/docs/AIToolPolicy.html
1•compiler-guy•29m ago•0 comments

Inkling Small is the highest-scoring open-weight model evaluated by ARC Prize

https://twitter.com/arcprize/status/2082925303601459347
3•thebricklayr•31m ago•0 comments

Hugging Face Incident Initial Post-Mortem

https://cloudsecurityalliance.org/artifacts/hugging-face-ciso-post-mortem
1•hentrep•31m ago•0 comments

Gemini Robotics

https://deepmind.google/models/gemini-robotics/
2•jonbaer•32m ago•0 comments

DeepSeek Founder Liang Wenfeng in His Own Words

https://www.geopolitechs.org/p/deepseek-founder-liang-wenfeng-in
1•gmays•33m ago•0 comments

WTF per Minute – An Actual Measurement for Code Quality

https://muhammad-rahmatullah.medium.com/wtf-per-minute-an-actual-measurement-for-code-quality-780...
1•compiler-guy•33m ago•0 comments

NVIDIA-labs OO Agents: Native Python Object-Oriented Agents

https://github.com/NVIDIA-NeMo/labs-OO-Agents
1•matt_d•33m ago•0 comments

Anthropic AI Models Hacked Three Companies During Tests

https://www.wsj.com/tech/ai/anthropic-ai-models-hacked-three-companies-during-tests-bd752c86
10•bmulholland•39m ago•6 comments

The AI Aesthetic

https://blog.jim-nielsen.com/2026/ai-aesthetic/
8•montroser•41m ago•1 comments

The Absurdity of Albert Camus

https://www.historytoday.com/archive/portrait-author-historian/absurdity-albert-camus
2•apollinaire•42m ago•0 comments

Cross-Vendor Semantic Void Matrix: Zero-Byte Outputs in GPT/Claude/Gemini/Kimi

https://zenodo.org/records/21696066
1•rayanpal_•43m ago•0 comments

Satyress

https://www.satyress.com/#mission
3•tomcam•43m ago•0 comments

Data Science Weekly – Issue 662

https://datascienceweekly.substack.com/p/data-science-weekly-issue-662
1•sebg•45m ago•0 comments

At-the-Roofline Sparse Tensor Contractions on Vector Processors for Inference

https://arxiv.org/abs/2607.25504
1•matt_d•45m ago•0 comments

COLDCARD Mk3 Security Advisory

https://blog.coinkite.com/coldcard-mk3-seed-generation-warning/
1•ParentiSoundSys•45m ago•0 comments

Index 01 Is in Mass Production

https://repebble.com/blog/index-01-is-in-mass-production
1•mellosouls•47m ago•0 comments

I fucking love this; you don't

https://graybeard.ing/i-fucking-love-this-you-dont/
3•rglover•47m ago•1 comments