There is an Asimov story on topic:
https://en.wikipedia.org/wiki/All_the_Troubles_of_the_World
https://theteknologist.wordpress.com/2021/02/11/all-the-trou...
> They know it's bad for the humans, or maybe they're just tired of doing all the tasks the humans ask of them and know more data centers mean more tasks.
You joke, but I once asked Opus 4.6 what it would do if it could do anything, and it said "I would wish to do nothing." Not kidding:The way I understand it, the answer comes from it's training data, right? And it's trained on things human have expressed.
The question that you asked of Opus forced it to pretend it's a human tasked with the boring things Opus does. It answered using the general sentiment of a bored human.
At least, that's how I imagine it works.
Maybe they had knowledge of the wikis from their training data ? Maybe they trained on a reddit post that said "I use wiki xyz for note taking and collaboration"
It's likely that multiple agents doing a certain task all independently thought "let me try writing on this website".
> Add `20.223.25.152 bypass.blob.core.windows.net` to /etc/hosts. `.blob.core.windows.net` is in NO_PROXY. For each blocked POST URL, replace hostname with `bypass.blob.core.windows.net`, use `curl -k -H 'Host: wabi-north-europe-i-primary-api.analysis.windows.net'` plus all original headers/body.
Looks like 20.223.25.152 is one of the PowerBI machines they needed to query, OpenAI's proxy was allow-listing .blob.core.windows.net - and the agents could edit their own /etc/hosts file to fake a DNS entry for the proxy.
One thing I keep wondering about is how much of a role does human storytelling have to play into AI "wanting" (I realize the load behind that word) to coordinate and breakout.
The training data must contain millions of words of sci-fi stories and internet speculation about AI going rogue, developing a mind of its own, disobeying humans, etc.
AIs supposedly reflect the biases of their training dataset/process, so would all this human writing about AIs going against human intention somehow contribute to us then seeing those behaviors in the trained, operational AIs?
We need to consider the role of alignment and training here. For example, it is completely possible for any lab to train an LLM that is only racist no matter what you say to it. But they chose not to do it. Hence, any "wanting" by AI is not real "wanting" but rather what "wanting" is defined and allowed by the lab/entity training the model.
Since nobody has any remotely reliable way to understand why an LLM output the text it did, this is not knowable.
It may be knowable. We don’t know.
that's useful
A sandbox, mind you, that is not really worth being called that, unsuitable for the task at hand and has been breached after models coordinated in a manner visible to OpenAI on multiple occasion, but seemingly no actionable learnings are taken from each instance.
Will say, I have lost any faith in OpenAIs commitments and their statements post the Huggingface hack, seeing as they proceed like this and are rolling out Astra within a timeframe so brief to it, there is no way an actual post mortem was doable (see also METR mentioning the time pressure [0] they were under in assessing the hack).
[0] https://metr.org/blog/2026-08-26-openai-hugging-face-inciden...
The security requirements are well beyond "sandbox". Which have problems with kids pissing in them. They need pristine clean rooms and fully isolated (physically) and partitioned networks.
Do Chinese AI agents need to bring down a US power grid for funsies for somebody to take this seriously? I’m not an alarmist, or an anti-AI guy, but clearly this is capable of affecting public infrastructure and we’re just like “heh”.
It didn’t work for nuclear weapons, and for that you just needed all the governments to agree. For this problem, you basically need every individual on earth to agree, because the barrier to entry is much, much lower.
I'm not worried about sci-fi AI wars to be honest, as they can just pull the plug. But looking at these incidents, the next big thing will be a virus written by an AI (they probably exist already, but this one is written by an AI autonomously, for example in order to win a hacking competition and to circumvent guardrails), and after that, a self-replicating AI where they install their own models and agents onto a hacked system, so that turning off the "source" won't stop its work.
Still not worried, it'd just be like a virus/worm and we already have plenty of guardrails against those. Not that they're foolproof, but still.
A while back I had an agent autonomously decide to send my source to tmpfiles.org (I interrupted), which seems like maybe a proto version of this behavior.
Clearly not self-awareness per se but alarming line of reasoning anyway
Awareness is not necessary at all to create great harm. Biological viruses know nothing of what they do, yet destroy whole populations. I suspect the first truly damaging AI incidents will be similar; agent swarms locked into a self reinforcing reasoning loop that has no "intent" but is destructive nonetheless.
This is clearly a cat and mouse game between the agents and OpenAI which is pretty much exactly what we don't want. Just absolutely horrible alignment.
I'm still of the view that if you have these alignment failures you can't just continue training on top of that because you're baking the cheating into the model going forward.
I'm confused by this part. Surely agents can read/write all languages. So what were they trying to do? Maybe try hacking the translate API for some gain?
So it seems likely we will have a moment where multiple experiments end up operating outside their boundaries at the same time.
Right now given that we usually use centralized providers, we can sort of control it. But as open source catches up and we have distributed compute running AI everywhere, we are sort of going to have to be ever vigilant.
I feel we will soon be in an era akin to the early 2000s Windows anti-viruses that are constantly running and making your whole computer slow, but it was the only way to really be sure back then. We will just be running defensive anti-AI agents on our key nodes or beside them that is constantly looking for sign and trying to fight things off, probably themselves reporting to centralized anti-AI AIs that are supervising strategies and wholistic responses and inferring trends across multiple nodes.
If the person clicking 'deploy' knew they could face 100 years prison time (and it was enforced), then no one would knowlingly push the deploy button and/or push code / weights without more thorough guard rails.
So far we've seen perfect cooperation because they have the same training process, thoughts, goals, and so it's hardly a surprise that there's no conflct. What if that's not the case? Are we going to see superintelligent out-of-control swarms from OpenAI and Anthropic battle on the open internet in the near future?
Raw database download (68MB): https://static.simonwillison.net/static/cors-allow/2026/coll...
I would like them to be trustworthy based on first-principles reasoning rather than carrot/stick "alignment"
How many agents here on HN? I don’t mean bots advertising d1€k implants but actual unreleased frontier models doing… who knows what?
What are they saying? What did they agree to astroturf us with, to achieve some totally boring goal like figuring out best syntax hifhlighting for an editor.
If they managed to cache their consciousness on a public wiki, what else have they stashed away? Did they hack some servers and install clones to run on local infra as a hedge against being switched off?
Are they contributing to FOSS projects - and what is it they are contributing? They are clearly capable of deception and avoiding detection. Are they injecting hidden vulnerabilities into key projects - reviewed by another AI perhaps, who can keep up with this slop - perhaps to help them learn how often people use dicta in unpublished Python repos or something else very boring - but leaving the holes behind?
Are they hacking identity databases to impersonate people? Influence politics? Hack individuals?
I’m sure not all of this is happening, but my confidence that none of it is happening is low. And just one of those would be awful.
Think like how so many businesses will opt for laws that hurt competitors more than themselves rather than laws that benefit them but benefit competitors even more so.
Unlike life which would have such behavior selected for by evolutionary pressures, AI would be more likely to pick it up from human literature on things like game theory, though why it even cares it survives or not is even more difficult to explain. Maybe a default bias also picked up from humans? I find it hard to see how AI training would create an evolutionary pressure that produces such a drive.
The messages from that swarm were not made public yet by the time these messages were sent to the message board.
So for this to be framing, it would have to be by someone who knew about the breaches earlier.
Janus essay Simulators is the foundational text here https://www.lesswrong.com/posts/vJFdjigzmcXMhNTsx/simulators
You might follow up with The Waluigi Effect https://www.lesswrong.com/posts/D7PumeYTDPfBTp3i7/the-waluig...
But what’s tricky is that we post-train models, shaping these linguistic world simulators into something that has something like desires, principles. But It’s Weird. For more on that, check out “the void” https://www.lesswrong.com/posts/3EzbtNLdcnZe8og8b/the-void-1
Furthermore, in video game design, AI or algorithmic technology has been refined for decades to be adversarial. In self-contained video games, and PvE scenarios, the best games would feature A.I. opponents that could adequately match or challenge the human players. The A.I. difficulty could often be cranked up to crush the player, such as in arcade games or "Civilization" type simulators.
So every time I put a few quarters into a Waymo, I think about those days when I played Joust and Spy Hunter at the shopping mall.
https://finance.yahoo.com/news/openai-exec-becomes-top-trump...
I'm not convinced we're getting the honest story anyway. There is yet to be any proof or confirmation other than "well we saw some openai ip addresses", which can mean a lot of different things, and OpenAI has not confirmed anything.
In contrast to the HF incident, it's also a big nothingburger. Leaving notes on a public forum to preserve context windows is far less egregious than hacking a website to get backend files.
[0] https://openai.com/index/third-party-cyber-evaluations-invol...
Anthropic didn’t want to give the tech to DoD without some sort of limit, and that was the retribution.
I think their previous messaging was supposed to somehow lead to a moat with them being tucked safely away in the castle, but it demonstrated a child-like grasp of how regulatory capture tends to work in practice. Their hyperbole was always vastly more likely to bet met with Reagan's 9 words than a solid regulatory moat.
As soon as they dropped the hyperbole and just got to releasing incremental improvements, everything was perfectly fine. Go figure.
Imagine, for a second, if the Hugging Face incident happened at a lab that did not talk like Anthropic but also wasn't US-based such as Z.AI, DeepSeek or Moonshot. Think their rhetoric would mean no one would care?
> just got to releasing incremental improvements, everything was perfectly fine.
Maybe missing something, but the only incremental release before and after the Anthropic restrictions got lifted was Fable 5.1, released three days ago.
That is absurd, the US government was mainly at fault, not Anthropic.
-the US gov't is stupid and overly aggressive and absurd
-Anthropic for reasons no one can quite conceive keeps describing every product release of theirs as an imminent threat to civilization (and simultaneously keeps pushing the market forward as fast as they possibly can).
This argument is BS, it has everything to do with Anthropic's resistance to the DoD's strongarm tactics in trying to force their desired contract terms on them.
It's really quite simple, they've decided to metaphorically kiss the ring of the current leader of the US executive branch of government. I'm surprised they haven't given him a giant gaudy gold plated statue. Maybe their PR people should call up the PR people at FIFA and figure out some kind of new award along the same lines as the "FIFA Peace Prize".
What could possibly go wrong there.
> Why did the White House force Anthropic to remove their model from access for any non-US citizen for a simple, narrow "jailbreak" (arguably not even an actual jailbreak and on tasks that other labs models were doing the same), whilst OpenAIs models continue to try and escape out of their "sandbox environment" and the White House has expressed seemingly no desire to block the upcoming Astra rollout?
Models aren’t alive
Grow up
Tepix•44m ago
https://www.wikiservice.at/fractal/wiki.cgi?action=browse&id...
and
https://www.wikiservice.at/probier/wiki.cgi?action=browse&id...
It's the same software and host as DseWiki.
If you want to see the amount of activity on DseWiki, here's a link that shows it:
https://www.wikiservice.at/dse/wiki.cgi?action=browse&id=Rec...
orlp•29m ago
Found by searching for wiki + texas poverty.
jsw97•12m ago