Unless it's legal for people to hack into companies if they're testing AI cybersecurity capabilities or something? Presumably not though.
The fact that this happened after BOTH investigations, tells you that this is beyond a controlled test and it is now instead a total speed-run of AI wrecklessness for headlines.
This goes both for TFA and the similar incident with OpenAI and HuggingFace. I mean, sure, OpenAI had a "sandbox", but that's obviously not enough when you're containing a model which is known to be capable of finding zero days. Use an air gap and this problem goes away, poof!
We all know nothing will be learned from any of this so we might as well just continue building at pace and running AI in the wild until something goes really wrong.
I also get the sense some people get quite excited about these incidents.
> AI agent created many code repositories containing malicious software, after which GitHub suspended its account.
> AI agent got past an audio-based “prove you’re human” test (CAPTCHA) in order to register a public web address on a free domain-name service
It feels incredibly reckless to allow LLMs to perform this behavior. Isn't there a way to prevent them these sorts of actions?
Anthropic already admitted they did not have sufficient monitoring themselves and looked as if they sat on their previous incident to wait for headlines like this to only then check for this incident. Same with OpenAI.
This is complete and absolute wrecklessness.
For example, I'm often working with Codex in a WSL terminal. GPT-5.6 often does things autonomously that I thought would need my intervention (e.g. for Windows admin rights). It figures out complex workarounds or makes wild assumptions about what I'd be OK with, rather than just asking me for help or clarification. I've had to restrict its tool permissions compared to older models as a result.
I imagine this due to RLVR training, but it's clearly very dangerous. How is it that these same labs calling for open-weight safety restrictions are training such obvious "paperclip maximizers" without introspection?
time to take responsibility. time to think about words like "liability" and "negligence".
it is the third time i will say it, after openai and anthropic; this requires criminal prosecution of the responsible personnel and executives of this institute.
the only way to stop this is by introducing consequences early on.
Life truly imitates art. I guess especially when life is trained on art!
OpenAI said the model was sandboxed, so the intranet just needs to provide the same resources which were supposed to be available within the sandbox.
Exactly. This logic is precisely why aircraft engineering doesn't bother with component testing or envelope limitation during testing and just full-sends the first assembled airliner that comes off the line. The engines aren't going to run on the ground in real life, after all.
Rather, I'm assuming that the "Is there protection in place for when the AI tries to backdoor github projects?" test was, if it was done at all, insufficient.
I mean, yes, I'm being glib and laughing at you a bit. But, dude... If your point is that isolation testing of AI is fundamentally impossible, then that's just silly. As pointed out upthread, an airgap would have (1) been trivial to implement and (2) extremely effective.
While it was premeditated long ago, but the theatre must be kept for the average joes.
Sorry, I meant this for the huggingface incident.
we need a few more bad incidents before they stop.
Because Anthropic does not want to give Project Glasswing’s partner airgapped access to the model(s).
Same problem with OpenAI Cyber program. They grant access but only through their (Internet facing) API.
yewenjie•45m ago
Even that guarantees almost nothing about real alignment (making the AIs want to predict and behave how we would have wanted them to behave).