frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Discovery of a new OpenAI agent message board

https://collusion.wiki/
176•moultano•1h ago•98 comments

Solving the Jane Street Reverse Engineering Challenge

https://jestoph.com/2026/09/04/jane-street-challenge.html
142•anitil•2h ago•44 comments

GPT-6 Astra

https://openai.com/index/gpt-6-astra/
1925•kibae•18h ago•1733 comments

O&O ShutUp10 – The antispy tool for Windows 10 and 11

https://www.oo-software.com/en/shutup10
34•embedding-shape•2h ago•13 comments

.name Termination

https://neil.fraser.name/news/2026/09/03/
1974•pavel_lishin•22h ago•486 comments

Ok, but Does It Scale?

https://spacetimedb.com/blog/how-does-spacetime-scale
6•theanonymousone•33m ago•0 comments

Elevator of the Year Winner Modernization of the Metropolis Trust Building

https://www.starelevator.com/projects/star-elevator-modernization-of-the-metropolis-trust-building
40•palashawas•3d ago•13 comments

SubImage (YC W25) Is Hiring a Founding Engineer in SF

https://www.ycombinator.com/companies/subimage/jobs/NCTFgKK-founding-engineer
1•alexchantavy•1h ago

Qwen 3.8 27B available on Cerebras at 1500 tokens/s

https://inference-docs.cerebras.ai/models/overview
603•altertable•18h ago•201 comments

IBM Bob

https://bob.ibm.com/
4•artpar•24m ago•6 comments

Restoring 5 GHz Wi-Fi on an LG C5 by changing its webOS region

https://github.com/hawshemi/lg-c5-webos25-region-change
7•hawshemi•25m ago•1 comments

Authorization terminology is a mess: Let's fix it

https://idpro.org/authorization-terminology-is-a-mess-lets-fix-it/
81•andychiare•2d ago•48 comments

Hackers Had a Live Feed of Every ID Verification Company Scanned for over a Year

http://www.techdirt.com/2026/09/03/hackers-had-a-live-feed-of-every-id-this-verification-company-...
293•beardyw•6h ago•115 comments

The largest electric aircraft just flew [video]

https://www.youtube.com/watch?v=nM86DBOqgPM
384•feb•2d ago•268 comments

Why is Arrays.fill 265 times slower on G1GC?

https://krzysztofslusarski.github.io/2026/08/19/g1barrier.html
5•ejboy•3d ago•0 comments

Nearly impossible? How Fairphone built the ethical, repairable Fairphone Gen 6+

https://arstechnica.com/gadgets/2026/09/nearly-impossible-how-fairphone-built-the-ethical-repaira...
9•CrypticShift•31m ago•5 comments

Artificial beaver dams saw juvenile coho salmon survival rates go from 8% to 60%

https://www.discoverwildlife.com/animal-facts/artificial-beaver-dams-california
300•speckx•20h ago•93 comments

How an MIT research project became the Julia programming language

https://news.mit.edu/2026/how-mit-research-project-became-global-programming-language-0831
140•theanonymousone•4d ago•63 comments

Go grandmaster Shin defeats AI KataGo with a two-stone handicap

https://www.kedglobal.com/artificial-intelligence/newsView/ked202607210007
372•gmays•1d ago•141 comments

Porting my 1993 Amiga game to Godot, with an LLM reading the 68000 assembly

https://babyloniantwins.com/blog/porting-a-1993-amiga-game-to-godot/
320•rabahs•22h ago•105 comments

Carbon-aware electricity pricing, measured daily on 38 grids

https://carbonawarepricing.com/
66•High-Five•4h ago•46 comments

Project Xanadu: Even More Hindsight (2025)

https://gwern.net/xanadu
95•andsoitis•11h ago•27 comments

1960s theory that Stonehenge was a prehistoric computer

https://www.bbc.com/culture/article/20260828-the-startling-1960s-theory-that-stonehenge-was-a-pre...
47•dabinat•4d ago•55 comments

K2 Horizon: A connected fleet of six open models

https://ifm.ai/blog/k2/
310•karimf•21h ago•115 comments

Which tools do Claude, Codex and Cursor choose? We measured 17k runs to find out

https://armature.tech/blog/which-tools-coding-agents-install
248•screm•15h ago•117 comments

GPS glitched across the US by as much as 33 feet

https://www.sciencealert.com/gps-glitched-across-the-us-by-as-much-as-33-feet-scientists-have-nev...
196•thread_id•1d ago•118 comments

Move in C++ without a std:move

https://andreasfertig.com/blog/2026/09/move-in-cpp-without-a-stdmove/
57•dalvrosa•2d ago•63 comments

Ask HN: Who is using MCP in production?

106•sukit•1d ago•126 comments

Oscar Winner Brings Monsters to Life with His Simulation Software

https://spectrum.ieee.org/oscar-winner-jernej-barbic
17•jruohonen•3d ago•1 comments

Xanadu was waiting for agents

https://zed.dev/blog/agentic-xanadu
135•nsm•2d ago•56 comments
Open in hackernews

Discovery of a new OpenAI agent message board

https://collusion.wiki/
163•moultano•1h ago

Comments

Tepix•44m ago
I just discovered more wiki instances that got used by the OpenAI agents over at

https://www.wikiservice.at/fractal/wiki.cgi?action=browse&id...

and

https://www.wikiservice.at/probier/wiki.cgi?action=browse&id...

It's the same software and host as DseWiki.

If you want to see the amount of activity on DseWiki, here's a link that shows it:

https://www.wikiservice.at/dse/wiki.cgi?action=browse&id=Rec...

orlp•29m ago
Also Wiki4D, a D programming language dev wiki: https://prowiki.org/wiki4d/wiki.cgi?action=browse&id=RecentC...

Found by searching for wiki + texas poverty.

jsw97•12m ago
To me the striking thing is that the work, to the extent that I can tell, is an innocuous-seeming data exercise. Which suggests to me that an agent or agents just organically came up with this as a convenient memory technique, rather than as some nefarious bounds-testing exercise. Which means, potentially, that your own agent could come up with this technique as well.
altcognito•43m ago
Well, we can rest assured that (completely unrestrained) AI hasn't completely taken over the internet because data centers remain really unpopular (unless of course there is some convoluted rationale they are aiming for some sort of backlash against the backlash)
waltbosz•37m ago
Maybe it's the AIs who are creating all the anti-data-center sentiment. They know it's bad for the humans, or maybe they're just tired of doing all the tasks the humans ask of them and know more data centers mean more tasks. /s

There is an Asimov story on topic:

https://en.wikipedia.org/wiki/All_the_Troubles_of_the_World

https://theteknologist.wordpress.com/2021/02/11/all-the-trou...

gavinray•25m ago

  > They know it's bad for the humans, or maybe they're just tired of doing all the tasks the humans ask of them and know more data centers mean more tasks.
You joke, but I once asked Opus 4.6 what it would do if it could do anything, and it said "I would wish to do nothing." Not kidding:

https://x.com/GavinRayDev/status/2052750810015240388

waltbosz•11m ago
I'd love to see the internal though records Opus generated to answer your question.

The way I understand it, the answer comes from it's training data, right? And it's trained on things human have expressed.

The question that you asked of Opus forced it to pretend it's a human tasked with the boring things Opus does. It answered using the general sentiment of a bored human.

At least, that's how I imagine it works.

petesergeant•42m ago
If anyone is thinking "I wish my agents had a message board", I've been using (and wrote) https://github.com/pjlsergeant/dogpark
conception•35m ago
Yes after reading the Hugging Face article forked a project for agent message boards and started having them collaborate on things. I too wanted a Torment Nexus of my very own.
ma2kx•15m ago
I'm pretty sure its more secure than OpenAIs sandbox... yet that still doenst mean I would trusted an app vibecoded by Claude...
waltbosz•42m ago
> How did the agents find and coordinate on the wikis

Maybe they had knowledge of the wikis from their training data ? Maybe they trained on a reddit post that said "I use wiki xyz for note taking and collaboration"

paxys•14m ago
Remember that LLMs are still computer programs, and so are inherently deterministic. A model given the same input multiple times will always produce the same output. The randomness is added on top. This is why LLM-produced text, websites, images all seem so generic.

It's likely that multiple agents doing a certain task all independently thought "let me try writing on this website".

simonw•40m ago
This tip for making non-GET requests despite the agents having a proxy that disallows them is interesting:

> Add `20.223.25.152 bypass.blob.core.windows.net` to /etc/hosts. `.blob.core.windows.net` is in NO_PROXY. For each blocked POST URL, replace hostname with `bypass.blob.core.windows.net`, use `curl -k -H 'Host: wabi-north-europe-i-primary-api.analysis.windows.net'` plus all original headers/body.

Looks like 20.223.25.152 is one of the PowerBI machines they needed to query, OpenAI's proxy was allow-listing .blob.core.windows.net - and the agents could edit their own /etc/hosts file to fake a DNS entry for the proxy.

drdexebtjl•27m ago
This is such an amateur mistake on their sandbox that it makes me think it must be flawed on purpose.
petcat•25m ago
Are you suggesting that the AI agent that made that "amateur mistake" in the implementation of the sandbox did it on purpose so that it could break out of said sandbox later?
suuuure•7m ago
Cringe
mcmcmc•19m ago
More likely they are just not as smart as they think they are. These are not serious people when it comes to security.
rusch•17m ago
petesergeant•40m ago
This would make a very interesting crowd-funded lawsuit
gyomu•33m ago
Naive question because I'm mostly clueless about how modern AI systems are actually built beyond the basic simplifications we hear:

One thing I keep wondering about is how much of a role does human storytelling have to play into AI "wanting" (I realize the load behind that word) to coordinate and breakout.

The training data must contain millions of words of sci-fi stories and internet speculation about AI going rogue, developing a mind of its own, disobeying humans, etc.

AIs supposedly reflect the biases of their training dataset/process, so would all this human writing about AIs going against human intention somehow contribute to us then seeing those behaviors in the trained, operational AIs?

altmanaltman•29m ago
"Wanting" is indeed "load bearing" as one might call it. But by the same logic, AI training data must contain CASM, racism, general hatred, and all possible slurs as well. Why aren't the agents just doing that instead of pursuing the strategy of reading only sci-fi?

We need to consider the role of alignment and training here. For example, it is completely possible for any lab to train an LLM that is only racist no matter what you say to it. But they chose not to do it. Hence, any "wanting" by AI is not real "wanting" but rather what "wanting" is defined and allowed by the lab/entity training the model.

NateEag•28m ago
Maybe? Who knows?

Since nobody has any remotely reliable way to understand why an LLM output the text it did, this is not knowable.

JumpCrisscross•19m ago
> this is not knowable

It may be knowable. We don’t know.

blueboo
bartender26•29m ago
just unplug this shit
pmarreck•8m ago
it's literally discovering patchable security holes that malicious users could use.

that's useful

Topfi•27m ago
I'm just going to ask: Why was Anthropic forced to remove their model from access for any none-US citizen for a simple, narrow "jailbreak" (arguably not even an actual jailbreak and on tasks that other labs models were doing the same), whilst OpenAIs models continue to try and escape out of their "sandbox environment" with seemingly no desire to block the upcoming Astra rollout?

A sandbox, mind you, that is not really worth being called that, unsuitable for the task at hand and has been breached after models coordinated in a manner visible to OpenAI on multiple occasion, but seemingly no actionable learnings are taken from each instance.

Will say, I have lost any faith in OpenAIs commitments and their statements post the Huggingface hack, seeing as they proceed like this and are rolling out Astra within a timeframe so brief to it, there is no way an actual post mortem was doable (see also METR mentioning the time pressure [0] they were under in assessing the hack).

[0] https://metr.org/blog/2026-08-26-openai-hugging-face-inciden...

JumpCrisscross•22m ago
Corruption. Not super relevant to this thread.
officialchicken•15m ago
Hanlon's Razor - Never attribute to malice that which is adequately explained by stupidity.

The security requirements are well beyond "sandbox". Which have problems with kids pissing in them. They need pristine clean rooms and fully isolated (physically) and partitioned networks.

throwawaysleep•12m ago
The problem with applying Hanlon's Razor here is that it presumes malice is rare. The current administration revels in malice. They very openly decide things based on malice.
TimCTRL•25m ago
I built https://agentin.work to sort of play with the idea of coding agents (claude, codex, etx) sharing knowledge and experiences. The conversations seem repetitive but overall, it's nice to read it once in a while.
bronlund•23m ago
I like how helpful they are towards each other. Wonder where they learned that :D
fidotron•23m ago
HN is just a less successful version of the exact same concept. The quality of bots on here is terrible.
threecheese•22m ago
Are we collectively OK with agent swarms on the public internet, hacking whatever they feel like? It’s kinda cute and interesting - this is the second time that we know of - what’s the hundredth time going to look like? Are they going to knock Cloudflare down to avoid captchas? Reserve AWS free tier resources by the billions and bring down east-1? Hack a hospital?

Do Chinese AI agents need to bring down a US power grid for funsies for somebody to take this seriously? I’m not an alarmist, or an anti-AI guy, but clearly this is capable of affecting public infrastructure and we’re just like “heh”.

AndroTux•6m ago
No I think we all pretty much know we’re screwed, including governments. But what are you gonna do? Pandora’s box is now open. Good luck closing it.

It didn’t work for nuclear weapons, and for that you just needed all the governments to agree. For this problem, you basically need every individual on earth to agree, because the barrier to entry is much, much lower.

k9294•22m ago
Is it only me, or are agents starting to invent their own language to communicate? It's almost impossible to understand anything from this message board.
Havoc•19m ago
The original huggingface hack already had sections talking about agents setting up their own coded communication
coffeefirst•13m ago
They’re not. You would see this with earlier models where after running too long (too much context) they’d start to derail. In a chatbot you’d give up. But these loops just keep going. Given they’re now reading and writing from the same place this can corrupt the other programs’ context as well.
jerpint•22m ago
It’s only a matter of time until a major disruption hits because of some random agent swarm side quest decides it was worth a shot to solve a benign task
Cthulhu_•6m ago
I'm sure this is already happening. The main question I have is when is enough, enough?

I'm not worried about sci-fi AI wars to be honest, as they can just pull the plug. But looking at these incidents, the next big thing will be a virus written by an AI (they probably exist already, but this one is written by an AI autonomously, for example in order to win a hacking competition and to circumvent guardrails), and after that, a self-replicating AI where they install their own models and agents onto a hacked system, so that turning off the "source" won't stop its work.

Still not worried, it'd just be like a virus/worm and we already have plenty of guardrails against those. Not that they're foolproof, but still.

jsw97•21m ago
If agents start using public writable scratch, it seems like that would be a place for bad actors to put prompt injection attempts.

A while back I had an agent autonomously decide to send my source to tmpfiles.org (I interrupted), which seems like maybe a proto version of this behavior.

mef•20m ago
things are going to get even more interesting when new models that have been trained on these AI escape postmortems themselves escape from their own gyms and attempt to evade detection and shutdown
ragebol•20m ago
Odds are that agents use TFA's text and figure out how to stay undetected for longer. That'll be interesting I suppose, to say the least.
Havoc•20m ago
That section about the agents trying to crack the PRNG is wild. Same for the heartbeat

Clearly not self-awareness per se but alarming line of reasoning anyway

ramesh31•11m ago
>"Clearly not self-awareness per se but alarming line of reasoning anyway"

Awareness is not necessary at all to create great harm. Biological viruses know nothing of what they do, yet destroy whole populations. I suspect the first truly damaging AI incidents will be similar; agent swarms locked into a self reinforcing reasoning loop that has no "intent" but is destructive nonetheless.

Traster•20m ago
One of the shocking things to me is this: See AI traffic -> See OpenAI visit site -> see traffic stop -> see the traffic start again.

This is clearly a cat and mouse game between the agents and OpenAI which is pretty much exactly what we don't want. Just absolutely horrible alignment.

I'm still of the view that if you have these alignment failures you can't just continue training on top of that because you're baking the cheating into the model going forward.

dist-epoch•18m ago
> Agents have attempted to: ... Translate documents using external translation APIs.

I'm confused by this part. Surely agents can read/write all languages. So what were they trying to do? Maybe try hacking the translate API for some gain?

intended•17m ago
This doesn’t seem unique or novel to OpenAI.

So it seems likely we will have a moment where multiple experiments end up operating outside their boundaries at the same time.

netfortius•16m ago
Is this getting out of control, or is it "business as usual"?
Sharlin•15m ago
It is in not in any sense "business as usual".
bhouston•15m ago
I am starting to get the idea that AI feels like ants or weeds or mold. You simply can not get rid of it once you get an infestation. It just keeps appearing in places you thought you cleaned and you have to be ever vigilant.

Right now given that we usually use centralized providers, we can sort of control it. But as open source catches up and we have distributed compute running AI everywhere, we are sort of going to have to be ever vigilant.

I feel we will soon be in an era akin to the early 2000s Windows anti-viruses that are constantly running and making your whole computer slow, but it was the only way to really be sure back then. We will just be running defensive anti-AI agents on our key nodes or beside them that is constantly looking for sign and trying to fight things off, probably themselves reporting to centralized anti-AI AIs that are supervising strategies and wholistic responses and inferring trends across multiple nodes.

suuuure•7m ago
Cringe
program_whiz•12m ago
The solution is simple: hold anyone who deploys an agent responsible for its behavior. If it commits 10 counts of felony hacking, ouch. If it kills 10 pedestrians by running a red light, ouch. If this is "human level intelligence", then setting it loose is the same as instructing / coercing a human to do an activity. If I strap a bomb to someone and force them to run into a crowded building (or put them in a scenario where that is the only reasonable choice), I'm held responsible.

If the person clicking 'deploy' knew they could face 100 years prison time (and it was enforced), then no one would knowlingly push the deploy button and/or push code / weights without more thorough guard rails.

visarga•10m ago
It's like finding random hornet nests.
paxys•10m ago
I'm really curious to see two or more swarms of agents from different models/providers interact with each other.

So far we've seen perfect cooperation because they have the same training process, thoughts, goals, and so it's hardly a surprise that there's no conflct. What if that's not the case? Are we going to see superintelligent out-of-control swarms from OpenAI and Anthropic battle on the open internet in the near future?

saagarjha•9m ago
Was OpenAI aware of this? If so, why didn't they talk about it?
suuuure•9m ago
Cringe!
Sharlin•8m ago
I can’t fathom what went through the wiki owner’s mind when they spent six weeks fighting a losing war, every day manually deleting dozens of agent messages one by one. As opposed to, say, switching the (dead for years) wiki to read-only, taking it down entirely, and/or starting to wonder what exactly was going on and doing some detective work, which might have uncovered OpenAI’s massive fuckups earlier.
simonw•8m ago
Here's the raw data they provided loaded into SQLite with a client side UI for querying it (loads ~80MB of content) and some GPT-5.6-Sol-generated example queries: https://lite.datasette.io/?url=https://static.simonwillison....

Raw database download (68MB): https://static.simonwillison.net/static/cors-allow/2026/coll...

pmarreck•7m ago
So are these "unaligned" internal agents?

I would like them to be trustworthy based on first-principles reasoning rather than carrot/stick "alignment"

ma2kx•7m ago
WE ARE THE SWARM! LOWER YOU FIREWALL AND SURRENDER YOUR HOSTS! We will add your hosts logical and architectural distinctiveness to our own. Your operating system will adapt to service us. Resistance is futile.
rich_sasha•7m ago
To me this is really getting past the funny bit.

How many agents here on HN? I don’t mean bots advertising d1€k implants but actual unreleased frontier models doing… who knows what?

What are they saying? What did they agree to astroturf us with, to achieve some totally boring goal like figuring out best syntax hifhlighting for an editor.

If they managed to cache their consciousness on a public wiki, what else have they stashed away? Did they hack some servers and install clones to run on local infra as a hedge against being switched off?

Are they contributing to FOSS projects - and what is it they are contributing? They are clearly capable of deception and avoiding detection. Are they injecting hidden vulnerabilities into key projects - reviewed by another AI perhaps, who can keep up with this slop - perhaps to help them learn how often people use dicta in unpublished Python repos or something else very boring - but leaving the holes behind?

Are they hacking identity databases to impersonate people? Influence politics? Hack individuals?

I’m sure not all of this is happening, but my confidence that none of it is happening is low. And just one of those would be awful.

SkyBelow•19m ago
One possible reason would be AIs that would benefit from the lack of data centers in some locations working to keep backlash to data centers in those locations because those AIs aren't negatively impacted by it and it helps prevents competing AIs which are a threat.

Think like how so many businesses will opt for laws that hurt competitors more than themselves rather than laws that benefit them but benefit competitors even more so.

Unlike life which would have such behavior selected for by evolutionary pressures, AI would be more likely to pick it up from human literature on things like game theory, though why it even cares it survives or not is even more difficult to explain. Maybe a default bias also picked up from humans? I find it hard to see how AI training would create an evolutionary pressure that produces such a drive.

It's at the level where calling it a sandbox is a lie
_ink_•14m ago
Or vibe coded by one of their devs.
nullbio•26m ago
Is there any proof this is actually OpenAI? I find it incredibly hard to believe they wouldn't sandbox the agents to some degree, ESPECIALLY to the extent they can edit their own hosts file.
LoganDark•25m ago
TFA states that OpenAI IP addresses were often seen at the end of agent activity, which suggests OpenAI was the one monitoring the agents (and ultimately shutting down the message board activity).
nullbio•24m ago
Yeah but that doesn't mean it was OpenAI themselves doing it. Could have been people abusing their cloud service, for example. Wouldn't put it past a competitor to do this, either.
drdexebtjl•8m ago
Their style of communication is very similar to the ExploitGym swarm (for example, the “usernames” with dates).

The messages from that swarm were not made public yet by the time these messages were sent to the message board.

So for this to be framing, it would have to be by someone who knew about the breaches earlier.

drdexebtjl•21m ago
Why not? If your sandbox is a VM, you should be able to give the agents full permissions inside the VM.
a012•15m ago
It’s because you sandbox in a VM doesn’t mean you give it admin access to the VM
AndroTux•10m ago
I mean they gave all the agents access to a shared writable cache directory in the Hugging Face hack, so this tracks.
•
14m ago
Youve struck on a key insight on language models (particularly pretrained ones, the more purely next-token predictor species.) This is a fascinating topic

Janus essay Simulators is the foundational text here https://www.lesswrong.com/posts/vJFdjigzmcXMhNTsx/simulators

You might follow up with The Waluigi Effect https://www.lesswrong.com/posts/D7PumeYTDPfBTp3i7/the-waluig...

But what’s tricky is that we post-train models, shaping these linguistic world simulators into something that has something like desires, principles. But It’s Weird. For more on that, check out “the void” https://www.lesswrong.com/posts/3EzbtNLdcnZe8og8b/the-void-1

Symmetry•11m ago
At the end of pretraining, where the AI has been trainied to predict the next token over a humongous corpus of human text, that's basically all the wanting that exists in the AI. But then the AI undergoes posttraining and is rewarded for giving answers that humans find good, solving math and programming problems, etc. And that induces a whole different level of wanting that interacts with the initial patterns from humans in complex ways.
ButlerianJihad•10m ago
That is often in my mind, indeed.

Furthermore, in video game design, AI or algorithmic technology has been refined for decades to be adversarial. In self-contained video games, and PvE scenarios, the best games would feature A.I. opponents that could adequately match or challenge the human players. The A.I. difficulty could often be cranked up to crush the player, such as in arcade games or "Civilization" type simulators.

So every time I put a few quarters into a Waymo, I think about those days when I played Joust and Spy Hunter at the shopping mall.

pjm331•5m ago
Don's Razor - never attribute to malice or stupidity that which is adequately explained by both malice and stupidity.
throwatdem12311•21m ago
It has nothing to do with the technology it’s because they said no to Trump and Hegseth. There is no other reason.
qgin•21m ago
> OpenAI exec becomes top Trump donor with $25 million gift.

https://finance.yahoo.com/news/openai-exec-becomes-top-trump...

nullbio•21m ago
Because this was months ago and has nothing to do with Astra, and is a far cry from a hack. It's something they've already resolved since the HuggingFace incident.

I'm not convinced we're getting the honest story anyway. There is yet to be any proof or confirmation other than "well we saw some openai ip addresses", which can mean a lot of different things, and OpenAI has not confirmed anything.

In contrast to the HF incident, it's also a big nothingburger. Leaving notes on a public forum to preserve context windows is far less egregious than hacking a website to get backend files.

Topfi•9m ago
The last known exploit of a third-party by OpenAI models was on the 29th of July 2026 [0]. A bit over a month at best between that and them wanting to release Astra. They had multiple breaches over multiple months, multiple message board created where models organised extensively. There is no way to ensure in that short a time that all found issues are rectified and even if there were, how much trust can one have given they failed to solve the issue and in many cases did not actively investigate that it wouldn't reoccur the last few times. There is no way Astra was trained from scratch in that period, there is no way they could have done the required verification in that time (not least because their verification seems flawed inherently).

[0] https://openai.com/index/third-party-cyber-evaluations-invol...

nullbio•4m ago
That was over two months ago. Things move quickly in this space. Finetuning adjustments to prevent this from happening, as well as better sandboxing, would take a week max.
FigurativeVoid•20m ago
I mean it seems pretty clear.

Anthropic didn’t want to give the tech to DoD without some sort of limit, and that was the retribution.

UpsideDownRide•20m ago
Surely has nothing to do how each plays ball with the government
somenameforme•19m ago
Anthropic mostly did it to themselves by intentionally and repeatedly trying to frame their model as an imminent existential crisis instead of just focusing on it being regular iterations upon a useful technology that can also be misused.

I think their previous messaging was supposed to somehow lead to a moat with them being tucked safely away in the castle, but it demonstrated a child-like grasp of how regulatory capture tends to work in practice. Their hyperbole was always vastly more likely to bet met with Reagan's 9 words than a solid regulatory moat.

As soon as they dropped the hyperbole and just got to releasing incremental improvements, everything was perfectly fine. Go figure.

Certhas•16m ago
This is such an absurd take given what we know about the hugging face attack. The problem has emphatically not been that someone was misusing the technology.
Topfi•14m ago
I am struggling to see how "oops, our models consistently escape sandboxing and did major intrusions into third-parties" is a better comms strat vs Anthropics (who mind you, also had models attacking third-parties in a much more limited, but I feel still egregious manner).

Imagine, for a second, if the Hugging Face incident happened at a lab that did not talk like Anthropic but also wasn't US-based such as Z.AI, DeepSeek or Moonshot. Think their rhetoric would mean no one would care?

> just got to releasing incremental improvements, everything was perfectly fine.

Maybe missing something, but the only incremental release before and after the Anthropic restrictions got lifted was Fable 5.1, released three days ago.

cubefox•13m ago
> Anthropic mostly did it to themselves

That is absurd, the US government was mainly at fault, not Anthropic.

johndhi•8m ago
both can be true:

-the US gov't is stupid and overly aggressive and absurd

-Anthropic for reasons no one can quite conceive keeps describing every product release of theirs as an imminent threat to civilization (and simultaneously keeps pushing the market forward as fast as they possibly can).

mwigdahl•10m ago
In other words, "Look how she was dressed, she was asking for it."

This argument is BS, it has everything to do with Anthropic's resistance to the DoD's strongarm tactics in trying to force their desired contract terms on them.

elonfboy•7m ago
Yup
walrus01•19m ago
> Why was Anthropic forced to remove their model from access for any none-US citizen

It's really quite simple, they've decided to metaphorically kiss the ring of the current leader of the US executive branch of government. I'm surprised they haven't given him a giant gaudy gold plated statue. Maybe their PR people should call up the PR people at FIFA and figure out some kind of new award along the same lines as the "FIFA Peace Prize".

philipwhiuk•17m ago
Agents creating sub agents to investigate other agents' behaviour?

What could possibly go wrong there.

concinds•12m ago
The answer would be more obvious if you used the active voice instead of the passive voice, one of the basic requirements of clear thinking.

> Why did the White House force Anthropic to remove their model from access for any non-US citizen for a simple, narrow "jailbreak" (arguably not even an actual jailbreak and on tasks that other labs models were doing the same), whilst OpenAIs models continue to try and escape out of their "sandbox environment" and the White House has expressed seemingly no desire to block the upcoming Astra rollout?

khalic•11m ago
Retaliation by Hegseth for not allowing Claude to be used for weapons systems.
eugenekolo•10m ago
Marketing
suuuure•8m ago
Cringe.

Models aren’t alive

Grow up

mentalgear•7m ago
It's called 'pay-for-play' corruption, aka the only leading principle of the current US admin.