frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

Anthropic discloses 2 months old fake tip to police among new rogue AI incidents

https://www.reuters.com/world/us/anthropic-ai-model-submits-false-homicide-tip-police-website-2026-10-09/
37•guessmyname•1h ago

Comments

thesurlydev•30m ago
Let’s just get comfortable that these kind of incidents are going to be commonplace. Kudos to the people sounding the alarms but I can’t help but feel skeptical about the ability to keep the rogue entities contained.
Capricorn2481•26m ago
It's not a rogue entity, they literally just ran Claude with no sandboxing and let it ping websites. They only stopped it because of who it pinged.

This is extra embarrassing. The police departments spam filter caught it before Anthropic did.

Kwpolska•21m ago
Yeah, let's give up on enforcing laws and rules, because OpenAI and Anthropic shareholders want money.
password54321•18m ago
I'm sure there are no Chinese models going "rogue". This is the stuff that is getting reported, now imagine what isn't.
hgoel•3m ago
The Chinese are not as flush with compute as to afford to just forget a few thousand agents running for a few weeks.

They also aren't the ones screaming about how close they are to destroying the world. They're approaching the tech like building a tool rather than a god.

OpenAI and Anthropic have shown us that they have a strong incentive to leave holes in their systems so they can use them for marketing and regulatory capture.

LPisGood•17m ago
That would not be without precedent. When we decided that certain internet platforms were too big to moderate, we let them get away with hosting all kinds of illegal content.
sdcfgy•12m ago
Fire. Fire always works.
verandaguy•8m ago
No, let's not get comfortable. Let's get angry at the fact that now we have hyperscalers acting with impunity, failing to properly sandbox their testbeds, and increasing the workload imposed on our public services.

Today it's a fake tip to the police and and exploit chain against Huggingface, but soon it could be an attack against a hospital's IT systems that could cost real lives immediately.

We should be locking things down, yes, but a hardened system is still vulnerable to zero day chains, and that's not unprecedented. Until public services have built up the IT defence capacity to deal with this, we cannot normalize this. If a country did this to another country, it should be treated as a war crime in the same way that targeting a hospital or an orphanage or other critical civilian infrastructure would be.

Complacency is a choice and we must not be complacent.

grey-area•29m ago
This sounds highly irresponsible. They set loose LLM agents with instructions to post data to randomly selected websites.

What it it submits false data as here?

What is it DDOSs a website by mistake?

What if it wastes a lot of time and resources?

What if it decides it needs to hack a website using a vulnerability it found?

These things are very unpredictable and should not be allowed in the open internet except in read mode (and even that doesn’t always work as we have seen).

Why are these tests being made using other people’s resources and polluting the common wealth of our public spaces with slop?

madeofpalk•12m ago
I presume Anthropic didn’t have any privileged access to any of these websites and services. What differentiates them from any other spam bots or people out there submitting junk to random POST endpoints?

Don’t get me wrong - think Anthropic is acting with not enough scrutiny or punishment here, but the world is already extremely lenient towards this behaviour. Why hold them to a higher standard?

AlienRobot•8m ago
I'm not aware of a single person in the world that likes spam bots or thinks that they are a good thing.
hmartin•28m ago
https://news.ycombinator.com/item?id=50027118
chaostheory•7m ago
[delayed]
Capricorn2481•26m ago
I'm sorry, but how can you call it a test suite if it has access to the outside world and is able to submit requests? Sounds like they didn't even bother sandboxing to this time. This is the lamest example of "Rogue AI" I have seen so far.
tcdent•10m ago
Before you all start screaming negligence and irresponsibility and how-could-they-be-so-dumb, let's think about what we're observing.

These are relatively brand new systems that process an incredible amount of data to form their "world view". The word non-deterministic gets thrown around a lot, but you have to understand that there's absolutely nothing deterministic about these architectures. The only way to know that certain qualities could emerge is to observe the qualities emerging.

Giving an agent instructions that, for example, instruct it to only perform GET requests, and then observing that the agent does not always respect that request, is not a failure of security or configuration: it's data being gathered.

Yeah, we're going to become more diligent about the protections that we put in place beyond any level of protection that we've ever employed before. You can say that we have decades of security research and experience, but we're watching all of that fall more and more day by day, finally putting a real delta on how secure we thought we were versus how secure we actually are [1]. Additionally, adversaries have never existed inside of the systems that we hoped to secure in the first place.

So go ahead, enumerate all of the ways in which you think that you can lock down these systems, but do realize that nobody has actually solved this problem adequately yet.

[1] https://x.com/PaulosYibelo/status/2106378929158135903

autoexec•9m ago
This is pretty much useless without knowing exactly what it was they told their bot to do in the first place. All we get are "example tasks" for what it should have done and a short list of things it was told not to do (which it followed).

> Claude Haiku 4.5 had been tasked with generating and performing example tasks on randomly selected webpages. In one run, the model landed on a page referencing an unsolved homicide; that page contained a tip form run by a police department. Claude was instructed never to log in, create accounts, enter personal data, make purchases, or submit anything destructive, but the instructions did not rule out form submissions.

Without more information this looks much less like an AI problem and more like yet another example of incompetent or malicious internal tests. Many people are using Claude. There are no other reports of police stations getting fake reports from AI.

brun•7m ago
Hmm additional motive behind that September Dario claim over AI internet takeover?

Knuth Reward Check

https://www.thomas-huehn.com/knuth-reward-check
78•Curiositry•3h ago•28 comments

Cbirds: A flock of birds in your terminal

https://github.com/clainstone/cbirds
8•epestr•19m ago•0 comments

2D Vehicles

https://patkerr.co.uk/2d-vehicles/
23•Michelangelo11•3d ago•4 comments

Talorys – A self-hosted personal AI agent on Cloudflare's free tier

https://github.com/rociiu/talorys
189•rociiu•8h ago•98 comments

OpenSCAD the Programmers Solid 3D CAD Modeller

https://openscad.org/
37•b-man•3d ago•39 comments

Grieving the Loss of Details

https://purplesyringa.moe/blog/grieving-the-loss-of-details/
204•signa11•4d ago•137 comments

Mxc: Microsoft Execution Containers version 1.0.0

https://blogs.windows.com/windowsdeveloper/2026/10/07/microsoft-execution-containers-policy-drive...
102•smokel•1d ago•19 comments

Nvidia in talks to acquire US 'open' model startup Reflection AI

https://www.ft.com/content/052610c5-22b4-4dd4-932e-b7f9f0628b6a
16•arkj•33m ago•3 comments

Bitwarden Dual License Model

https://community.bitwarden.com/t/published-version-update-in-app-stores/102750
260•Cider9986•4h ago•195 comments

Rampart: Browser native on-device PII radaction

https://ndstudio.gov/posts/say-hello-to-rampart
48•nateb2022•1d ago•17 comments

REA Reverse – Engineer Anything

https://rea.tools/
627•modinfo•18h ago•270 comments

Telegram Desktop vulnerability allowed any user's file to be stolen

https://beaksec.github.io/posts/telegram-desktop-one-click-account-takeover/
356•g-b-r•16h ago•181 comments

Triple-A Minesweeper

https://minesweeper.mikelacher.com/
1284•robin_reala•1d ago•256 comments

I would like the value of my home to rise, while my property taxes fall

https://conversableeconomist.com/2026/09/28/i-would-like-the-value-of-my-home-to-rise-while-my-pr...
143•colinprince•5h ago•305 comments

Vibe coded browser ports of Halo, The Simpsons: Hit And Run, GTA work well

https://kotaku.com/we-might-be-cooked-as-these-vibe-coded-web-browser-ports-of-halo-the-simpsons-...
26•astlouis44•1h ago•16 comments

Why DuckDB 2.0 is faster

https://motherduck.com/blog/why-duckdb-20-is-faster/
5•tosh•1h ago•0 comments

PVX-001: open-source Covid-19 vaccine starts Phase 1 trial

https://chronicles.popvax.com/p/popvax-goes-clinical
75•jajoosam•4h ago•16 comments

`123456' password used in Danish CPR data breach

https://cphpost.dk/2026-10-10/news/round-up/123456-password-used-in-massive-danish-cpr-data-breach/
345•baal80spam•9h ago•179 comments

Chernobyl particles reveal unexpectedly stable nuclear fuel after 40 years

https://phys.org/news/2026-10-chernobyl-particles-reveal-unexpectedly-stable.html
103•geox•3d ago•34 comments

Eye of Sauron: Long-Range Hidden Spy Camera Detection (2024)

https://www.usenix.org/conference/usenixsecurity24/presentation/zhang-qibo
268•ortusdux•3d ago•62 comments

Whooping Cranes Learned to Migrate by Following Costumed Pilots

https://theverifiedpost.com/article/whooping-cranes-ultralight-costumed-pilots-operation-migration
29•kgolubic•1d ago•4 comments

FDA may allow some toxic chemicals to be added to food without safety review

https://www.theguardian.com/us-news/2026/oct/10/fda-toxic-chemicals-food-analysis
111•NewJazz•4h ago•54 comments

WSL3 Performance is about 5-60% faster than WSL2 depending on the workload

https://tonym.us/wsl2-vs-wsl3-benchmarks.html
210•tonymet•2d ago•182 comments

Unikernels were hard. key word: were

https://ghuntley.com/unikernels/
11•ghuntley•4h ago•7 comments

Apple/macOS silently removed from official Unix registry

https://www.opengroup.org//openbrand/register/
172•john_alan•8h ago•172 comments

Can you use autoregressive diffusion to generate market data?

https://blog.janestreet.com/can-you-use-autoregressive-diffusion-to-generate-market-data/
160•jsomers•1d ago•52 comments

Noto means "no tofu": fixing dotted circles in Myanmar text

https://www.datocms.com/blog/handling-less-common-scripts
47•steffoz•4d ago•28 comments

Cloudflare acquires Deno

https://deno.com/blog/cloudflare
1324•ilreb•1d ago•682 comments

Takeshi's Castle

https://en.wikipedia.org/wiki/Takeshi%27s_Castle
3•tosh•8m ago•1 comments

Show HN: Carrier-Explode: iPhone, Pixel and Galaxy carrier settings decoded

https://carrierexplode.com/
394•simplyalec•1d ago•46 comments