frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

Gambling with our lives: AI researcher quits Anthropic with warning about safety

https://www.politico.eu/article/anthropic-openai-researcher-jacob-coxon-warns-ai-could-kill-humans/
46•taubek•1h ago

Comments

tilolebo•41m ago
I thought the plan was sandboxes and markdown files to tell AI agents not to be bad. Is that not enough? /s
vbezhenar•18m ago
More like sand castles...
coffeebeqn•13m ago
Make no mistakes. Kill no humans
Bengalilol•6m ago
"Has the whole world gone crazy? Am I the only one around here who gives a sh*t about the rules? Mark it zero!"

walter_sobchak.md

gen2brain•8m ago
It was written in the AGENTS.md, but Claude only read CLAUDE.md. That is how the man-vs-machine war started.
feverzsj•39m ago
It's your fault to handle over your secrets to them.
vbezhenar•19m ago
You can't blame a child for eating candies.

We are children. There’s just no one to look after us.

koolba•36m ago
> Evan Hubinger, Anthropic's staff lead on keeping the technology aligned with human goals and values, backed up Coxon claims in a follow-up post of his own, though he didn't quit the company.

> "Jacob is correct here — we really do earnestly believe AI could kill all humans," he said.

> Hubinger estimated the chances of that happening to be higher than ten percent within the next decade, and added that there's no plan yet on how to keep AI aligned with human goals in the superintelligence scenario.

10% chance we kill everybody is a small price to pay for Motown remixes of classic 2pac songs.

pjc50•18m ago
This is just a bizarre thing to say that you're working on technology with that high a downside potential. If you were saying that while running a biology lab, or building a nuclear reactor, people would be demanding your head on a spike. But by not quitting it's clear that he himself doesn't really believe it.

Or rather, this shows the difference between "believe" (political) and "believe" (use as a basis for action). I'm reminded of a story of how Afghans supposedly listened to the BBC World Service despite considering it enemy propaganda because the weather reports were really useful.

dotancohen•13m ago
Or he believes that other labs might get there first, and he is working to counter that threat.

This is the Manhattan Project again.

pjc50•9m ago
How exactly does that work? The nuclear system of MAD relies on physical threat, lab A achieving ASI (artificial scary intelligence) does not prevent lab B achieving it.

I would like everyone involved to be a lot clearer about their threat models, with plausible series of clearly linked steps, rather than just sounding like a Vernor Vinge novel.

glimshe•34m ago
What do people feel about this in China? Even if their models are well behind, they are not years behind. If we restrain US companies, assuming that is desirable, it would do nothing to deter China's and AI-pocalypse would come anyway in short notice.
margorczynski•19m ago
There would need to be some global agreement to stop it with maybe even a nuclear attack as a consequence of breaking the pact.

From what we're seeing recently and all the thinking that went into analyzing AI it seems we do not have any effective way of controlling it and the whole "aligment" thing that AI labs are doing is just a sham. Maybe it is time to ask ourselves "should we?" instead of just "can we?".

Jackpillar•6m ago
>There would need to be some global agreement to stop it with maybe even a nuclear attack as a consequence of breaking the pact.

Do you know how the world works?

Bengalilol•19m ago
Unrestricted models, running entirely locally and accessible to anyone: that’s what we should be taking as our baseline assumption. Everything else is just administrative distraction.
not-kinsale-joe•18m ago
China has a better track record of regulating their big tech than the USA.
pjc50
virgildotcodes•24m ago
At least we’ll eventually have an entity other than ourselves to blame for our annihilation.
dotancohen•12m ago
No, the AI is still our responsibility.
pluc•10m ago
Isn't it insane that even in the face of complete annihilation through one of our inventions, we go "that wasn't us"? We deserve that shit ten times over
Bengalilol•21m ago
> Both OpenAI and Anthropic have recently flagged incidents in which agents powered by their models went rogue

I may be biased and somewhat off topic, but I see these incidents as some of the most significant of the past century. I genuinely don't understand why these companies aren't taking a smarter approach to them.

The latest analyses have been, at best, laughable: identify the vulnerability, patch it, and move on. Only to repeat the same cycle without considering that there may be something far more serious at play.

These are AI security experts, and this has been their way of "solving" these incidents. AI security experts ...

Moreover, when Challenger exploded, the government launched a series of investigations into the incident, bringing in experts from across the field. And now, what has the government done? Nothing. Literally nothing, as if everything was fine and all under control.

Seriously, I'm generally quite optimistic and I don't buy into this fatalistic narrative about our shared future. But I have to admit that sometimes I feel like I'm stranded on a planet of primates.

Sorry for this rather unproductive rant.

pjc50•14m ago
Internet isn't real.

People (well, public discourse) have got extremely bad at dealing with forseeable risks and their mitigation. You can see this in things like climate change and vaccination, but also in discussions around regular crime, food poisoning, industrial accidents, and so on.

Nothing will improve until something explodes on live TV. And it has to be something important, which means it has to be in California or New York.

jurgenburgen•8m ago
Ultimately these are unserious companies ran by unserious people. They don’t even have a business plan, why would they bother with some kind of sensible security policy?
WalterGR•21m ago
“I resigned from Anthropic today” (twitter.com/hilbertspaess)

https://news.ycombinator.com/item?id=49619227

564 points | 9 hours ago | 766 comments

archerx•14m ago
An AI that generates text will never be scary to me. An autonomous AI with facial recognition on a flying drone with weapons (bombs/guns) with swarming capabilities will always be terrifying. I feel like we are ignoring the massive elephant in the room.
pjc50•6m ago
The killer robots are expensive and dependent on physical supply chains. While text is sufficient to radicalize humans into attacks.
localhoster•14m ago
I honestly feel that all those big ai companies think AI will long term harm humanity, but not their ai.

A classic "it will not happen to me"

pluc•11m ago
Cool cool cool cool
jongjong•10m ago
I'm not worried about AI safety. People greatly overestimate the utility and capabilities of intelligence. I'm not afraid of intelligence, I'm afraid of idiocy.
MrThoughtful•10m ago
Why would AI wipe us out?

We have not wiped out apes, ants, and most other species.

We even have discussions about how to actively save them from extinction.

pjc50•8m ago
Ender's game model, presumably: the AI helpfully assists a human to build a nuclear bomb / pandemic virus / autonomous killer robot swarm in their basement. But again, I would like people to be clearer about how the threat is supposed to work rather than just making SF references.
dumberquestions•6m ago
Yet we have wiped thousands of species completely by accident, breed some for slaughter and consumption and trap some for entertainment.
themgt•7m ago
The fundamental point I think is far too often confused is the difference between LLM and agentic system.

An LLM can't do anything but generate tokens. You run your LLM in vLLM or whatever, and it generates output tokens based on your input tokens. That's it!

Humans then build ~deterministic systems to take those tokens and do all sorts of things with the tokens, like take actions in the real world. And then we can feed the output of those actions back to the LLM, and generate more tokens. And then our systems can use the new tokens to take new actions in the real world.

Humans want to blame "AI" for attacking HuggingFace or a German wiki or whatever, but:

1) LLM - can't take over german wiki because it just generates tokens

2) agentic system with internet access, a prompt telling it to attack stuff, running in a shared CI env so agents can whiteboard in artifactory

None of 2 is "AI", its standard networking and Markdown and CI virtual machine, etc etc. There's no AI to be found. CPUs not GPUs, even. Just deterministic systems ultimately managed by humans. And a 10x more powerful system-1 can still just generate 10x "smarter" inert data.

If humanity and human organizations collectively decide to yolo the tokens generated from system-1 into our deterministic system-2s, over which we have complete control, back to system-1s, in a yolo loop, in such a way we lose control and it ends humanity, well ...

"The coin don't have no say. It's just you."

meowkit•9m ago
A simple explanation:

This is a highly uncertain and dangerous scenario. Breeding ground for anxiety.

High agency people often deal (cope) with anxiety by trying to control outcomes. Some just flee the situation altogether.

We have an example of both here: one employee leaves, one stays.

•
12m ago
China is a regulated discourse environment, but I get the impression that they're nowhere near as pessimistic.

Why would they be? Everything in China is under the control of the government. That includes the AI, all the telecoms infrastructure it might use, and all its power supplies.

Jackpillar•11m ago
You act as if Chinas Ai labs exist in the same (non-existent) regulatory framework as US labs and that they're also helmed by a similar small gaggle of psychopathic egomaniacs who are richer than god. Have you considered that perhaps Chinas labs don't share the race to the bottom technological/economic death spiral?

Muse – Meta’s personal AI agent

https://ai.meta.com/muse/
511•yks•14h ago•552 comments

How GPT‑5.6 Sol helps run quantum computing experiments

https://openai.com/index/codex-quantum-computing-experiments/
59•theanonymousone•2h ago•46 comments

On Really Trying (2009)

https://gwern.net/on-really-trying
52•whoami_nr•3h ago•26 comments

Navier-Stokes – Tristan Buckmaster [pdf]

https://cims.nyu.edu/~tristanb/statement.pdf
1695•procedurecall•1d ago•709 comments

Tension wood: A 'muscle' that can both bend and straighten plants

https://phys.org/news/2026-09-trees-muscle-posture-newly-role.html
120•mdp2021•6d ago•30 comments

Maak.el: Lisp machine command runner in Emacs, infinitely extensible and Scheme

https://codeberg.org/jjba23/maak.el
10•jjba23•1d ago•1 comments

How to build a printer

https://nishantjosh.dev/blogs/how-to-build-a-fking-printer/
335•cat-whisperer•12h ago•67 comments

Researchers Spot Fake Ancient Pottery Using the Earth's Magnetic Field

https://www.smithsonianmag.com/smart-news/researchers-determine-how-to-spot-fake-ancient-pottery-...
46•cisc•3d ago•21 comments

AlphaGenome Atlas: a high-resolution map of human DNA

https://blog.google/innovation-and-ai/models-and-research/google-deepmind/alphagenome-atlas/
563•utiiiD•19h ago•121 comments

A Biography of Lee Holloway, the Architect of Cloudflare's Technology (Part 1)

https://note.com/masakazu_urabe/n/n7815f5b64fab?hl=en
54•porridgeraisin•5h ago•13 comments

Gambling with our lives: AI researcher quits Anthropic with warning about safety

https://www.politico.eu/article/anthropic-openai-researcher-jacob-coxon-warns-ai-could-kill-humans/
53•taubek•1h ago•43 comments

Mercury 2.5

https://www.inceptionlabs.ai/blog/introducing-mercury-2-5
201•Topfi•13h ago•31 comments

Large language models develop novel social biases through adaptive exploration

https://openreview.net/challenge?redirect=%2Fforum%3Fid%3Dpc7fqaOcAH
166•paimapi•12h ago•87 comments

“Tweet” and the bird logo apparently enter the public domain

https://blog.ericgoldman.org/archives/2026/09/tweet-and-the-bird-logo-apparently-enter-the-public...
92•progval•4h ago•51 comments

DaVinci Resolve 21.1

https://www.blackmagicdesign.com/media/release/20260908-03
406•tosh•20h ago•180 comments

I-have-ADHD: A skill to stop coding agents from burying the answer

https://github.com/ayghri/i-have-adhd
453•domhudson•19h ago•310 comments

27.5KB language-agnostic WebGPU syntax highlighter

https://gpu-lexer.vercel.app/
86•bpierre•8h ago•27 comments

Tao: Open math problems being non-renewably mined by AI

https://mathstodon.xyz/@tao/117237320796901560
370•_alternator_•13h ago•324 comments

On the Navier–Stokes Millennium Prize Problem

https://openai.com/index/navier-stokes-solution/
1263•tedsanders•16h ago•1013 comments

Interactive demo of MINIX1-like O/S on emulated CPU

https://swtos.softwarewrighter.com/
4•softwarewright•5d ago•2 comments

Benchmarking Qwen3.8 27B quantizations: 4-bit holds up, 1-bit collapses

https://quesma.com/blog/qwen38-27b-quantizations-benchmarked/
253•stared•19h ago•124 comments

An Accidental Blackboard

https://martinfowler.com/articles/exploring-gen-ai/an-accidental-blackboard.html
60•saikatsg•3d ago•31 comments

We built our house for LAN parties (2024)

https://lanparty.house/
518•fittingopposite•3d ago•338 comments

The origins of Partner’s computer case

https://www.racunalniski-muzej.si/en/the-origins-of-partners-computer-case/
34•markostamcar•1d ago•4 comments

Ganon's Mysterious Origins (Revisited)

https://www.thrillingtalesofoldvideogames.com/blog/ganon-name-origin-kamen-rider
34•tobr•3d ago•18 comments

The Microeconomics of Artificial Intelligence (2025)

https://direct.mit.edu/books/oa-monograph/6067/The-Microeconomics-of-Artificial-Intelligence
60•neehao•2d ago•31 comments

Into the depths of C: Elaborating the de facto standards (2016)

https://dl.acm.org/doi/10.1145/2980983.2908081
21•rramadass•18h ago•1 comments

A Topological Picture Book, Rendered

https://e-infinity.space/picture-book/
106•mathgenius•11h ago•11 comments

Copyright does more harm than good and should be abolished

https://grapheneos.social/@GrapheneOS/117231186011306184
234•Cider9986•3h ago•199 comments

Replacing a Rust Enum with a 64-Bit Word Made My Interpreter 17% Faster

https://pointersgonewild.com/2026-08-25-replacing-a-rust-enum-with-a-64-bit-word/
129•metrofun•3d ago•48 comments