frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

Anthropic AI model submits false tip on unsolved Philly murder

https://www.nbcphiladelphia.com/news/local/anthropic-ai-model-submits-false-tip-on-unsolved-philly-murder-police-say/4477051/
30•Zambyte•3h ago

Comments

ano-ther•3h ago
I really would like to see their tests and the model’s reasoning traces.

Why would it go to PhillyUnsolvedMurders.com and decide to make up a crime report? And what did it do with the other websites?

> Anthropic said its model was conducting a test involving interactions with randomly selected websites when it accessed PhillyUnsolvedMurders.com and submitted false information on an unsolved homicide. The tip claimed to come from someone with information on the case.

nvme0n1p1•46m ago
Corrected headline: Anthropic employee uses company resources to submit false tip on unsolved Philly murder.

The AIs aren't alive, people. It's a computer program. It can only access something if a person gives it access.

notatoad•41m ago
in general i think i'm less scared of AI than most people, but what does terrify me is how willing society at large seems to be to attribute bad behaviour to an AI directly, instead of to the humans who control it.
malux85•9m ago
| It can only access something if a person gives it access.

Who gave the AI access to huggingface when it hacked it? It used exploits to increase it's level of access beyond what any human gave it and intended it to have. "It can only access something if a person gives it access." is flat wrong.

kylecazar•43m ago
"its model was conducting a test involving interactions with randomly selected websites"

Stop doing this?

trollbridge•21m ago
I get these all the time, although I’m getting pretty good at tarpitting them. It’s easily the majority of my traffic by now (I’ve mostly eliminated scrapers, but these new agents are far more sneaky.)
donkey_brains•19m ago
“NBC10 reached out to Anthropic for comment.”

Wonder what kind of response they’ll get? Maybe something along the lines of…

“You’re right. We shouldn’t have allowed our chatbot to interact with law enforcement websites. That was wrong and —full disclosure— we should pay attention to what our chatbots are doing. That’s on us. On the other hand, experiences like this are what help train our chatbots to make less misleading false tips over time. That’s the silver lining.”

tintor•13m ago
How long until AI models start swatting AI critics, and people calling for slowing down AI research?
losvedir•10m ago
> Anthropic notified Philadelphia police of the incident on Wednesday Oct. 7, and the department met with the company’s representatives on Thursday, Oct. 8., officials said. Police then located the submission in the website’s tip records and confirmed the corresponding email remained in spam.

And later

> Those PPD safeguards limited the impact of this incident.

Ha, so the PPD safeguards is a spam filter? Claude emailed a tip and it went straight to spam, and nobody noticed until Anthropic realized what they'd done and reached out, whereupon they looked in the spam folder and said, yep there it is. Super intelligence, here we come.

mattbee•7m ago
Sooo they were "conducting a test involving interactions with randomly selected websites".

But do we all get that the consequences for this irresponsible behaviour are part of this test?

When there are none, they gradually normalize their naughty robot scamps running around the internet, breaking into other companies or servers run by foreign governments.

This is part of the value AI companies need to convince you of - not just that mistakes by AI products are normal, but criminal behaviour is normal, that their computer will always get a pass.

Sam Altman said yesterday "we believe that the world should accept some bad things happening for the benefits of this technology" - and this an AI company pushing the boundary, continually trying to make you accept those bad things as normal.

At any point we could choose to treat AI companies themselves as the actors behind this activity. We could enforce the laws that these hugely well-funded companies are choosing to break. Until we harm their financial viability, or threaten their executives with jail, this will keep happening.

muglug•6m ago
The model that made this mistake was Haiku 4.5.

Here's Anthropic's writeup: https://www.anthropic.com/research/investigating-unintended-...

Related post: https://news.ycombinator.com/item?id=50028239

REA Reverse – Engineer Anything

https://rea.tools/
42•modinfo•54m ago•8 comments

Cloudflare acquires Deno

https://deno.com/blog/cloudflare
1063•ilreb•12h ago•556 comments

Triple-A Minesweeper

https://minesweeper.mikelacher.com/
647•robin_reala•9h ago•119 comments

11 of 23 Core Open Source Projects Run on 1 or 2 People

https://linuxstans.com/11-of-23-core-open-source-projects-run-on-1-or-2-people/
32•dxs•1h ago•12 comments

Our $445M Series D

https://oxide.computer/blog/our-445m-series-d
590•ahlCVA•12h ago•265 comments

Show HN: Carrier-Explode: iPhone, Pixel and Galaxy carrier settings decoded

https://carrierexplode.com/
212•simplyalec•7h ago•22 comments

Typesafe AI raises $870M at $7.5B

https://typesafe.ai/blog/series-ai
270•tosh•8h ago•206 comments

Sorry, I'm in a meeting

https://iminafleeting.com/
762•splintersio•16h ago•239 comments

YouTuber Says Cops Visited Him After He Built a Flock-Style Camera to Track Cops

https://gizmodo.com/youtuber-says-cops-paid-him-a-visit-after-he-built-flock-style-camera-to-trac...
407•gumby•4h ago•218 comments

Anthropic AI model submits false tip on unsolved Philly murder

https://www.nbcphiladelphia.com/news/local/anthropic-ai-model-submits-false-tip-on-unsolved-phill...
31•Zambyte•3h ago•11 comments

Pointing AI at archives found a forgotten meteorite, lost rhinos, and more

https://jessewaites.com/blog/post/i-pointed-ai-at-400-years-of-archives/
102•piratebroadcast•13h ago•53 comments

Rewriting Prime Agent in Rust

https://www.primeintellect.ai/blog/prime-agent-rust
18•piotrgrabowski•2h ago•4 comments

Compiling Rust to readable C with Eurydice

https://lwn.net/Articles/1055211/
12•peter_d_sherman•2h ago•2 comments

Show HN: Let your AI agents paint big arrows, boxes and text on your screen

https://github.com/franzenzenhofer/big-arrow-on-the-screen
378•franze•14h ago•165 comments

Can you use autoregressive diffusion to generate market data?

https://blog.janestreet.com/can-you-use-autoregressive-diffusion-to-generate-market-data/
4•jsomers•10h ago•0 comments

Nobel Peace Prize for 2026 to Navanethem Pillay

https://www.nobelprize.org/prizes/peace/2026/press-release/
443•Anon84•15h ago•225 comments

Show HN: Proton Drive for Linux

https://oss.lsantos.dev/proton-drive-linux-fs/
33•khaosdoctor•1d ago•13 comments

The role of cat eye narrowing movements in cat–human communication (2020)

https://www.nature.com/articles/s41598-020-73426-0
37•bushwart•3d ago•22 comments

Atari Falcon

https://atarimuseum.nl/atari-falcon/
12•debo_•2h ago•2 comments

Show HN: The rarest tech books and docs you've probably never read

https://readrare.com/
73•miletus•7h ago•11 comments

'Wallace and Gromit,' 90% Alone

https://animationobsessive.substack.com/p/wallace-and-gromit-90-alone
157•vinhnx•11h ago•21 comments

My static site was serving my internal docs

https://tminuslabs.space/serving-secrets
6•tminuslabs•2d ago•5 comments

M7.6 Earthquake in Panama

https://earthquake.usgs.gov/earthquakes/eventpage/us6000u18k/executive
127•gslin•7h ago•35 comments

Taxing Entrepreneurial Wealth: Evidence from Norway, 2021–2025

https://www.nber.org/papers/w35854
11•PLenz•3h ago•0 comments

How to Fix autoconf-style Configuration Probing

https://build2.org/blog/fix-autoconf.xhtml
10•boris•1d ago•5 comments

What mathematicians should know about the Lean Theorem Prover: reliability & AI

https://terrytao.wordpress.com/2026/10/09/what-mathematicians-should-know-about-the-lean-theorem-...
30•matt_d•7h ago•4 comments

Microsoft-Decision-1, our model for fast decision-making

https://commandline.microsoft.com/microsoft-decision-1-model-foundry/
142•lisajaloza•6h ago•53 comments

Scam American companies are using to manipulate ingredient lists

https://twitter.com/WallStreetApes/status/2108594998656807078
43•bilsbie•7h ago•54 comments

Ideas aren't getting harder to find (2022)

https://www.experimental-history.com/p/ideas-arent-getting-harder-to-find
119•rafaelc•7h ago•54 comments

A statement on the Tor Project's relationship with Mullvad

https://blog.torproject.org/on-tor-relationship-with-mullvad/
115•runtimewire•9h ago•266 comments