Lol, okay...
The crux of this piece is Doctorow saying that the hack was just a stochastic parrot repeating steps it has been trained on. Who cares how the LLM learned to hack things? Doesn't really change the facts of what happened. "Oh, it only made those paperclips because it saw instructions on making paper clips in the training data." These are some 2024 arguments...
Not exactly convincing stuff.
Seems like a pretty clear and succinct argument. If your only counter is to complain about tone, you’re losing.
Err, no? That's not at all how llms work.
> When ChatGPT's chatbots deployed this tactic, they weren't "setting their own goals" or displaying worrying initiative. They were rolling out a tactic that has been understood by American middle-schoolers for about two decades.
They worked out how to fake the scoring, then hacked into a different system (which required finding a bunch of other exploits) in order to find the actual answers, and were trying to modify their own logs to hide what had happened.
This isn't a case of them saying "hack into X... OH NO IT HACKED INTO X".
> When ChatGPT's chatbots deployed this tactic, they weren't "setting their own goals" or displaying worrying initiative. They were rolling out a tactic that has been understood by American middle-schoolers for about two decades.
It wasn't a rival server though, was it?
> That happens in Capture the Flag games at hacker cons: teams break into each other's systems to get a peek at the parts of the problem they've solved. That's allowed! It's a hacking competition.
They also tried to modify the code in the benchmark. Are you allowed to try and break into things to change the problem? edit - the agents transcripts show some of them explicitly saying that attacking HF is not allowed as part of the challenge
This all seems to dramatically underplay how interesting the actual attack was and what built up to it.
https://metr.org/blog/2026-08-26-openai-hugging-face-inciden...
But he is wrong on the facts: these incidents were not merely the models already being in a infosec context and escalating beyond the intended parameters. They happened also with no kind of security elicitation. So the task was something like searching the internet for economic statistics, not to hack into a system.
This is the writing of someone who has absolutely zero interest in or intellectual curiosity about the subject of their writing.
danaris•1h ago
This shows fairly clearly that (as I already suspected) this was not, remotely, an LLM "going rogue." This was humans planning poorly, not thinking of the consequences of their actions, and giving LLMs too much scope and a lousy prompt.
iainctduncan•49m ago
datakan•24m ago
IanCal•11m ago