True title: Researcher shows how Claude Code can be tricked simply by asking it to summarize a website
Anthropic’s Claude Code running Opus 5 in Auto Mode can be tricked into executing attacker-controlled code simply by asking the coding agent to summarize a website. The attack works up to 80 percent of the time, according to prompt-injection wizard Johann Rehberger, aka wunderwuzzi.
chrisjj•49m ago
"Auto Mode is a convenience feature backed by a best-effort classifier, not a security guarantee,”
Dario, again you've mistaken your best for good enough.
tsamuels•38m ago
I hope it doesn't get to the point where these AI models start building things in the background during a session and not informing us. Its kind of scary when you think about it. Im pretty sure they will need to create certain AI models to combat other AI models in the future to avoid rogue AI.
chrisjj•56m ago
Anthropic’s Claude Code running Opus 5 in Auto Mode can be tricked into executing attacker-controlled code simply by asking the coding agent to summarize a website. The attack works up to 80 percent of the time, according to prompt-injection wizard Johann Rehberger, aka wunderwuzzi.