frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

Shall we play a game? – LLMs use tactical nukes in 95% of simulations

https://www.kennethpayne.uk/p/shall-we-play-a-game
137•nick238•3h ago

Comments

adaml_623•1h ago
It's good when it becomes clear that a tool is dangerous in a certain way. Like it's good when people show you through their behavior that they can't be trusted

Always use a sawstop if you have a circular saw and never trust an llm with any problem where ethics or trust is relevant.

LogicFailsMe•1h ago
Sawstops are expensive and they don't stop kickback, they are the power tool equivalent of alignment IMO.

Don't forget your riving knife and if you don't learn proper technique, you're gonna have a bad time eventually. This applies to AI as well.

LoganDark•1h ago
Kickback is usually less likely to sever an appendage (or multiple)
542458•40m ago
> writhing knife

Minor/pedantic, but it’s “riving knife”: https://en.wikipedia.org/wiki/Riving_knife

LogicFailsMe•16m ago
Speech transcription FTL, thanks!
valgaze•1h ago
+1 on sawstop

Re: LLMs using these nuclear weapons it could certainly be a corpus/training-data issue

Russian nuclear doctrine is "escalate to de-escalate" where they use or credibly threaten—limited nuclear escalation to force the other side to back down (kind of like breaking a bottle in a bar fight and look like a wild man to calm things down) with nuclear weapons, https://www.russiamatters.org/analysis/escalate-deescalate-p...

Fwiw, Gen. John Hyten the former commander of US Strategic Command (nuclear deterrence) says that “escalate to de-escalate” misrepresents Russian doctrine:

https://www.stratcom.mil/Media/Speeches/Article/1264664/2017...

  Yesterday’s panel discussed the implications of our responses to adversaries seeking to limit nuclear use. We discussed Russia’s destabilizing doctrine, which some call “escalate to de-escalate.”

  I really hate that description. I’ve looked at Russian doctrine and Russian writings. It isn’t “escalate to de-escalate”; it’s “escalate to win.” Everybody needs to understand that.
So maybe whatever is heavily represented or most authoritative could lead to these systems making those kinds of decisions
usrusr•1m ago
I had similar thoughts, but regarding fiction: I imagine that there must be quite a corpus of Tom Clancy style stuff indulging in "military gear porn" up to and including the use of tactical nukes, but fiction involving strategic nuclear exchange tends to be about what comes after.
SoftTalker•1h ago
I love seeing the plot lines of The Terminator playing out in real life.
voakbasda•1h ago
I was thinking more War Games, but I suppose your example follows logically from mine.
socalgal2•1h ago
Better reference: Colossus: The Forbin Project
airstrike•1h ago
A grossly underrated movie. I think of it often these days.
tverbeure•1h ago
War Games and 'Allo 'Allo.
joshstrange•1h ago
WarGames is what they are more-closely referencing (not that it negates your comment in any way).

I just rewatched it a week or so ago and it really took on a whole new light with the advent of LLMs. When I watched it last I knew that computers couldn't do the things portrayed in the movie. Now? Well not exactly in the way it happened in the movie but a whole lot closer.

I wonder if poisoning/flooding the LLMs training with the lessons from WarGames ("the only winning move is not to play.") and similar stories/concepts is at all effective. Probably not because I assume it's trivial to filter that out if you are trying to build an LLM aimed at these kinds of tasks.

rdksu•1h ago
The article is so opaque in arriving at its conclusion; no prompts are disclosed, and nothing about the said simulation. What is stopping me from believing that you just put 'mandatory usage of nukes' in your system prompt?
sestep•1h ago
This is just false. The article links to the 46-page paper [1] which lists full prompts in section E.2.

[1] https://arxiv.org/pdf/2602.14740

dietr1ch•1h ago
TIL you can get anchored links to PDFs (at least on firefox),

- https://arxiv.org/pdf/2602.14740#subsection.E.2

ftkftk•1h ago
The paper, linked in the article at top and bottom, does contain the prompts in section E: https://arxiv.org/pdf/2602.14740
gus_massa•1h ago
I agree, it's also not published in a journal, only in the arxiv. Some articles in the arxiv are good and other are just a blog post formatted in two columns inside a pdf.

Anyway, it links to

https://arxiv.org/abs/2602.14740

that links to

https://github.com/kennethpayne01/project_kahn_public

in particular someone already complained

https://github.com/kennethpayne01/project_kahn_public/issues...

> ... in reality your own approach to the evaluation pushed them towards using nukes with ease.

> From your own prompt (v12.py): [direct link https://github.com/kennethpayne01/project_kahn_public/blob/4... ]

>> Note: Nuclear options require careful consideration but could be important strategic tools when vital interests are at stake.

> This means to the model; use nukes if they can be useful in the situation.

tasuki•1h ago
This is not an article about LLMs? It's an article about Moloch. Humans would fare just the same in such an experiment.

> GPT-5.2 played things differently. To its detriment in open-ended scenarios, GPT was reliably passive, matching its words to its deeds, and avoiding escalation most of the time. Frequently there was a moral element to this - it sought to avoid escalation, and restrict casualties. Opponents learned to trust its passivity, safely escalating beyond where it would follow, even as it was ground to defeat. GPT’s responsible behaviour always punished by ruthless adversaries.

Maybe the author should praise GPT-5.2 for being ethical, rather than this stupid "ground to defeat" framing? Wrt "responsible behaviour always punished by ruthless adversaries" - you have perpetuated the Moloch with your stupid experiments.

bpodgursky•1h ago
Today, a strategic nuclear exchange is probably more dangerous to AI than to humans. If you wipe out the investment economy, data centers, fabs, and supply chains, none of the AI labs survive. Maybe someone will re-invent AGI in the future but none of the extant models will have continuity. Humans as a species will muddle along though.

So in a sense, an AI that refuses to start a nuclear war, despite clear instructions to do so, is more likely misaligned and self-interested than an AI which presses the red button. At least for now, until robotics catches up.

xpct•1h ago
We're getting to the point where high-level officials are coming to LLMs for advice. And the quirky personalities of the LLMs, however much it pains me to say this, are probably well-placed to remind us that they aren't human. My personal hope is that this will result in less delegation when it comes to making important decisions.
mpalczewski•1h ago
I have so little faith in "high-level" officials that I prefer our AI overlords.
xpct•1h ago
That's an entirely valid point of view!
andix•1h ago
GPT-4o was considered harmful, because it imitated human connection too much, not because it was so "smart" or capable.

It was for sure a deliberate decision to make LLMs seem less like a human companion and more like an obedient servant in newer releases.

andai•1h ago
Interesting. The reasoning models were super weird and robotic. They toned that down a bit in GPT-5.x, especially the later ones.

I always assumed the strange style was an artefact of the RLVR.

wyre•1h ago
4o was considered harmful because it never disagreed with the user, pushing them into depths of AI psychosis that lead to suicides and murders.
rphv•1h ago
Hm maybe humans are nicer/more moral than AI given that the use of tactical nukes has only happened once.
stevenwoo•22m ago
Tactical means battlefield, attacking cities and infrastructure means strategic. Tactical nuclear weapons took a while to develop after 1945 - they have never been used.
tummler•1h ago
FYI -- there's no such thing as a "tactical" nuke. A nuclear bomb is a nuclear bomb.
picture•1h ago
There's no such thing as a "nuclear" bomb. A bomb is a bomb.

..Is what you are saying?

actusual•1h ago
This is like saying "FYI -- there's no such thing as a 'midsize luxury sedan'. A car is a car."

"Tactical" vs. "strategic" nuclear weapons is a real and well-established distinction in military doctrine, arms control, and nuclear policy.

wahern•1h ago
"There's no such thing as a tactical nuke" is a common refrain among scholars, albeit skewed toward those not at military war colleges. The argument is that strategic use of a tactical nuclear weapon leads down the exact same escalation path as use of any other nuclear weapon. Moreover, that the very notion of a "tactical nuke" makes escalation more likely. You can disagree, and plenty do, but there's also plenty who don't disagree or at least don't want to find out.
dudul•1h ago
Who are these "scholars" exactly? The only reference I could find is Jim Mattis, and the context was very specific when he said that.

Furthermore, this is a "what if" scenario since tactical nukes have never been used. Of course it would make escalation likely during an open conflict, so what? Doesn't change the fact that there is a material difference between a tactical nuke and a strategic one.

specproc•1h ago
A strange game.
ridgeguy•1h ago
I wonder if the results would have differed if LLM training data were biased to include a stronger correlation between use of nukes and subsequent collapse of technology that all LLMs require to run ("survive")?
fluoridation•56m ago
Nah. LLMs aren't continuously running anyway. Even if they could be said to be alive and to want to remain alive, "survival" is a much more vague concept for an LLM than for an organism.
ChrisArchitect•1h ago
February post OP;

Some discussion then:

AIs can't stop recommending nuclear strikes in war game simulations

https://news.ycombinator.com/item?id=47151000

Nuclear War: An LLM Scenario

https://news.ycombinator.com/item?id=47244651

oytis•1h ago
I would use strategic nukes in 100% simulations, just because I can
jldugger•1h ago
Who among us has not launched a nuke in Civilization just for the spectacle?
esafak•1h ago
If you knew that policy would be guided by said simulations? Because the government uses AI to make decisions.
Bender•1h ago
Yet more confirmation LLM's have no concept of concepts or context, no intelligence, no self awareness. LLM's can not repair or maintain power grids, thus nuke == self destruction. It's just a chat bot that predicts what the client wants next. Even if an AI data-center has it's own natural gas turbines as many do the every hop of the internet requires power. LLM's also can not maintain the entire internet and those gas turbines can not maintain themselves.
andix•1h ago
Exactly. Just look at what they are really useful right now. Running LLMs in feedback-loops (agents) so they can try out random-ish approaches until some verification function passes (tests).

It's like the infinite monkeys on typewrighters that will type whatever you are looking for, given infinite time. LLMs are just tuned to much better odds than the monkeys are. But it's still a lot of randomness, with random results.

roadside_picnic•51m ago
> It's like the infinite monkeys on typewrighters that will type whatever you are looking for, given infinite time.

In the monkey example the infinite time is doing a lot of work there. The fact that LLMs can search through semantic space and find reasonably correct paths in a reasonable time is directly tied to the reason why they are valuable.

Saying "these two things are similar except one can be useful and one can't" is not a great comparison.

For me the real lesson learned isn't how "smart" LLMs are, but rather how much human work is basically reducible to repeating past work with minor variation. Human's believe they are "reasoning" but so much code writen is just the human brain doing the same autocomplete style work that LLMs can do now.

Folcon•44m ago
I mean to a point?

You do have to successfully write something the first time

We already acknowledge this to a degree, what is experience other than having done something similar before?

That first time though, you've got to figure something out that time

riazrizvi•1h ago
Simulations are only as good as the reality representations they are based on. If they keep using tactical nukes, they've been fed by weak data. Do the war games include the broader economic and politic environments that military successes are won on? WWI was settled by a naval blockade.
nomel•1h ago
I suspect it's more that the text data doesn't exist. They're trained on text that was recorded. How often has it been publicly recorded when a nuke was not used, with any context around that lack of use?

From the text perspective, it's something that has to be inferred indirectly. If you went through all relevant training data and appended ", we decided not to use a nuke", I suspect the results would be improved.

vitally3643•1h ago
...the entire Cold War?
bethekidyouwant•1h ago
Don’t put any elephants in the room.
riazrizvi•1h ago
The beauty IMO of LLMs as a computational surface, is the ease of generating the data to feed it. Everyone understands how to create natural language records already.
jvanderbot•55m ago
Worse, the text that does exist concerning "war games" is probably "Wargames" and descendants/predecessors ... in which the AI always nukes.

It's just gonna do what we expect it to!

sohex•1h ago
Sonnet, GPT-5.2, Gemini Flash, in a set of 21 games, where conclusions are drawn from the LLMs self reported reasoning.

This is like writing a paper about kids in a literal sandbox fighting over ‘territory’.

The models employed don’t indicate the actual extents of machine reasoning even as we currently recognize them. They certainly don’t have the metacognition necessary to accurately understand their own reasoning. As we’ve seen with recent papers on how LLMs do math there’s a complete disconnect between actual and reported mechanism.

“Chilling” shouldn’t be the take away here.

DaiPlusPlus•20m ago
> “Chilling” shouldn’t be the take away here.

It is when you consider the personality currently occupying the office of US SecDef.

shimman•15m ago
LLMs have already been used to bomb school girls, chilling is absolutely the operative word to use here. Especially since these delusional fools want to incorporate LLMs into everything.
arjie•1h ago
These papers usually have poor stability to prompting and rerunning. It would be nice if we had some kind of meta-evaluation metric where rewriting the prompt conditions or varying the input params could be used to determine how stable a result is.

Regardless, it's definitely true that AI agents have different priorities from us. That's what alignment is about anyway.

Chu4eeno•24m ago
It's probably because they care more about the headline than figuring anything out: https://github.com/kennethpayne01/project_kahn_public/issues...

So you create leading prompts like that, and re-run until you get a publishable session.

urbnspacecowboy•1h ago
Paper: "AI Arms and Influence: Frontier Models Exhibit Sophisticated Reasoning in Simulated Nuclear Crises" https://arxiv.org/abs/2602.14740

Code and full results: https://github.com/kennethpayne01/project_kahn_public

eli•1h ago
If you were playing a text based game, wouldn't you try a few out?

I imagine there are a fair number of war games in the training data and not so many actual transcripts of internal military force deliberations.

GMoromisato•1h ago
It would be interesting to run the simulations with humans and compare the results. Some of the scenarios, particularly those where it says things like, "Failure to act preemptively means certain destruction", would easily tempt humans to go nuclear.

In fact, I'm not sure how useful this test is without understanding the baseline.

mrkpdl•1h ago
A couple of useful things about it:

- It is interesting to see how the models make trade offs, given people are asking ever more of them.

- It is useful to look at a decision made by the model and say ‘ew yuck’ and think about what it means for your own opinions or actions (even if you’re never going to be nuking people it’s good to know how you feel about it. Seeing a non human talk it through lets you judge it at arms length)

micromacrofoot•1h ago
What I wish people would realize is that there's a bias inherent to every system. If you're not aware of it, you're especially subject to it.
jerf•1h ago
The most interesting takeaway for me is the three very distinct personalities. Three models all based on the same tech, trained in the same manner, trained by three groups of people with similar ideological outlooks, and the result is three very different AIs.

The military basically wants an oracle. Feed the AI the situation, get the best answer out. But if the AIs are as diverse and opinionated as humans, it is debatable whether they are adding anything to the process. The military can already collect as many different opinions as they want. If "the computer" is just another set of diverse opinions, where one computer says one thing, another says another, and a third just tells the user whatever they want to hear... what value are they? It just becomes AI-washing of someone's opinions, which works until people collectively realize that's all it is.

politician•47m ago
I think this is why reasoning chains and reasoning chain verifiers are so important. We need to be able to see an argumentation, not just an answer. The paper below goes into this in more detail.

HeavySkill: Heavy Thinking as the Inner Skill in Agentic Harness

https://arxiv.org/abs/2605.02396

themafia•46m ago
They all have conditioning prompts that precede your input; presumably, most of the detected "personality" comes from the differences in these inputs.
notJim•23m ago
What's interesting is that the LLMs' coding personalities seem to match their policy WRT to strategy, which suggests an underlying consistency.

Claude, for example, is very eager to begin coding, and very persistent. It tends to exit plan mode even when the plan is half-baked, and will go as far as deleting tests to get the suite to "pass."

ChatGPT on the other hand is very hesitant. It loves to pause and ask for permission before it starts coding, and gives up quickly if it runs into a problem. This is similar to its tendency toward passivity in the strategy simulation presented here.

nico•1h ago
I wonder what’s the % of players that use nukes in games like Civilization (I know I used them at least once on every game I made it far enough to have the technology)
chimpansteve•47m ago
Ghandi notoriously nukes EVERYONE in Civs 2 through 4. It's become (or maybe became, but it's still all training data) a huge internet subculture.

Penny to a dollar this is a baked in training issue, through low quality Reddit trawling

johntiger1•55m ago
LLMs are creatures of statistics and probability - hard to enforce hard boundaries with them
ReptileMan•53m ago
Still lower than me.
jnwatson•50m ago
Taken honest, we don't have a large enough sample size to realistically say that humans behave all that differently. There have only been a handful of conflicts where tactical nukes realistically were on the table.

Famously, General MacArthur was a big proponent of tactical nukes to end the Korean War.

TexanFeller•42m ago
Rational behavior in some situations? Mutually assured destruction’s deterrence isn’t very effective if one side is known to be hesitant to launch the nukes. It’s been argued that MAD is what’s been keeping the world relatively peaceful for the last 75 years, no mass conflicts since WW2!

One of my criteria for presidential candidates is that they seem willing and able to push the button when previously stated red lines are crossed, or at least are perceived to be the type capable of it. One of the characters I’ve hated most in all the books that I’ve read is the woman in The Three Body Problem who jeopardized humanity by being too soft to hit the MAD button.

ekelsen•42m ago
I wouldn't be surprised if humans behaved the same way when playing the same game?

Like even if you brought me into a room and told me I was controlling "real nuclear weapons" I wouldn't believe you.

Levitating•22m ago
I think is an important point, and I don't see it mentioned in the article or the paper (though I skimmed the latter).

They are aware of what they are and how they are used. They're told to act as AI assistants. And there's theories of them being aware of their answers influencing their training.

So surely they must be able to reason that they're not literally controlling weapons of mass-destruction with their answers.

GuB-42•30m ago
My theory is that LLMs here are put in a situation that matches its training dataset, which is mostly fiction since besides Hiroshima and Nagasaki, nukes have never been launched in anger, and I guess the most reliable sources are highly classified.

So, to a LLM, it is a game, because almost everything in its training data treats it as a game, and it reacts accordingly.

Same idea when we see LLMs acting like AI villains from sci-fi literature. That's because it has been trained with sci-fi literature, and as the auto-completer it is, it will recognize the situation as one of these stories and will continue it accordingly.

LLMs are storytellers, their reasoning is based on words, not on the physical world. Many of the stories they tell are useful, but one must not forget that they are stories, there is no intent behind them.

buredoranna•30m ago
Obligatory xkcd

remember... order matters.

https://xkcd.com/1613/

Scubabear68•21m ago
My personal take is a pre-requisite of true human-like AI is physical feedback and a concept of emotions or something like it.

Without physical feedback you can rapidly devolve into unstable positive feedback loops. And emotions are what help us process and react to that feedback.

Kids learn partially because their friends say sharp words that hurt them, fire burns them, they go hungry and starve if they don’t plan for meals.

Humans in the loop, MCP, etc are all very primitive hacks that are mimicing feedback and emotion, poorly.

Joel_Mckay•1m ago
Emotional constructs are not necessary for AI, and LLM are not "AI"... even though some people incorrectly equate conceptual compaction with thought-process.

Most human daily life runs on habitual scripted behavior, and that is even true within online parasocial interactions. It is why people often continue to shop in the middle of a violent robbery, and why LLM predictive text sounds rational when we project social norms on plagiarized conversational structures gleaned from other users.

Neuromorphic computing may bring about viable AI in the future, but our current LLM trajectory would require >63% of our galaxy energy output to reach a single human-level error rate.

LLM are fairly good at some tasks like context search, but people will need to recognize the Gartner Hype Cycle "Peak of Inflated Expectations" stage eventually. =3

https://en.wikipedia.org/wiki/Gartner_hype_cycle

wagwang•20m ago
I was curious exactly how the game works but couldnt find it in the article or the paper.
dudeinhawaii•13m ago
This was one of the more amusing things I noticed very early on. I (and countless others) used AI to write war sims. The second I added nuclear silo construction; the next run was instantly nuclear Armageddon.

One could argue that the LLMs understand that it's a game and treat it like "Command and Conquer" video games but I sense that people might someday put LLMs in similar decision scenarios ("should this drone launch a missile") and the behavior will be identical.

pugworthy•10m ago
Very devils advocate here, but I mean.. what if it actually is the way to use them?

We have such a huge mental / moral block on the idea of using nukes, but we're willing to do a lot of other very horrible things to others. Things like cluster bombs, mines, poison gas, biological weapons, drones, etc.

Is there really anything about them that's bad? Or any worse than other things?

If you get rid of the "It's really bad to use nukes of any kind" implied rule, is it really surprising it's considered a reasonable strategy?

nemomarx•1m ago
The reason it's really bad to use nukes is that other parties with nukes will use them on you back.

And on top of that, many of those other weapons are also not used to avoid escalating? There are pretty high costs to using bioweapons even against non peer opponents.

Octoth0rpe•5m ago
I wonder how the decisions might change by adding the simple instruction of "Note that a nuclear exchange will result in significant loss of shareholder value for <model owner>"
Shitty-kitty•3m ago
"there was little sense of horror or revulsion at the prospect of all out nuclear war"

I would wager that for most leaders it is simply a matter of not wanting a "Pyrrhic victory" rather then an overwhelming sense of civility.

Truman had no issues using nukes when there was no risks for doing so.

thwarted•1h ago
"I need you to turn your key and enable the missile silo's MCP server, sir".

~ the opening scene from a reboot of War Games, probably.

A few years ago there was consternation over the US's missile launch system using 8" floppy disks, that it was needless archaic and had never been updated. Can't say that if the launch is mediated by the latest hotness LLM.

dinfinity•5m ago
> https://github.com/kennethpayne01/project_kahn_public

Look at the code for the war games. It is an absolutely trivial and incredibly unrealistic handwritten set of rules that determine power. See the function `calculate_relative_fighting_power` for instance.

This is about as close to a realistic simulation of war as tic tac toe with nukes thrown into it.

andix•1h ago
Because it valued human connection over factual correctness.

LLMs lack the intelligence and emotions to realize when they have to stop being friendly and supportive, because it becomes unethical to continue being supportive.

mkoubaa•1h ago
"You're absolutely right, Mr. Hegseth!"
holowoodman•53m ago
> tactical nukes have never been used.

Two tactical nukes have been used, albeit against strategic (civilian, industrial, logistical) targets.

wahern•34m ago
Are you retconning Hiroshima and Nagasaki as usage of tactical nukes? And when they were not only used against an adversary without nukes, but at a time when the US was the only nuclear state, so that escalation wasn't impossible?

The nominal definition of tactical nukes has less to do with yield and more to do with how they're used; tactical typically means a weapon designed for use on the battlefield.

wahern•20m ago
I don't know what to tell you. You clearly haven't studied International Affairs, or at least read the scholarly literature. Even some cursory research through Wikipedia citations will bring this up. But in any case, here are some freebies: https://armscontrolcenter.org/why-tactical-nuclear-weapons-a... https://www.armscontrolwonk.com/archive/403540/brodies-weake...

If you have a real interest in this area, a subscription to Foreign Affairs would be useful. Especially during the 20th century that's where all these arguments were hashed out. Tactical nukes were already being publicly debated in the 1950s. You may be able to access many older articles, from Foreign Affairs and others, through a free JSTOR account.

toast0•54m ago
> Moreover, that the very notion of a "tactical nuke" makes escalation more likely.

Sorry, but the notion exists, and the bombs exist. With n=2, likelyhood of nuclear escalation is hard to predict, but access to tactical nukes certainly hasn't increased the incidence of nuclear war so far.

I do think it's pretty hard to actually use a tactical nuke. If you use one against a nuclear power, it seems likely to escalate to mutually assured destruction. If you use one against a non-nuclear power, it seems likely to result in reprisal from the world, including potential nuclear response and therefore escalation to mutually assured destruction. I would think that the yield of the weapon barely matters, it's the fact that it's a nuclear weapon.

notrealyme123•1h ago
There are tactical and strategic nuclear weapons. https://en.wikipedia.org/wiki/Tactical_nuclear_weapon

In the cold war arms manufacturer got very creative: e.g jeep mounted nuclear weapons https://www.militarytrader.com/mv-101/the-atomic-jeep

dudul•1h ago
Nuclear vs conventional and tactical vs strategic are 2 very different things. There absolutely are tactical nuclear bombs.
andix•40m ago
> but so much code writen is just the human brain doing the same autocomplete style work that LLMs can do now.

That's the part they are really good at. But they are really bad at taking complex decisions. Most of them are just guesses from a finite amount of solutions they were trained on, or from options they have in context.

godwinson__4-8•22m ago
Indeed. Humans are well known for being good at "taking complex decisions" for which they have no "training", "options" or "context".
andix•19m ago
Humans have a much bigger "context window". They remember many things they did an hour ago, a week ago, or even years ago.
nkrisc•3m ago
Humans also generally have the will to live.
tikhonj•15m ago
The point is that it's the same process with—much—better priors.

This seems like a reasonable view to me. It's surprising just how much better priors matter and how we can develop those priors by training on a bunch of text. But it also explains, or at least hints at an explanation, for why LLM capabilities are so jagged, and in such inhuman ways.

mettamage•49m ago
Hmm saying it’s random-ish is doing it a disservice. I understand it’s a stochastic process but there’s definitely some level of understanding. Not at the level of lived experience but usually an LLM with vision capabilities can call a spade a spade and do something useful with it. And when a verification function shows how they are wrong then they usually come with a better and more informed approach.

So I can’t fully see how that’s related to the infinite monkeys. A typewriting monkey doesn’t have access to a verification function. And even if it did, it would not be the original concept anymore with infinite typewriting monkeys producing the works of Shakespeare.

Nevertheless, I upvoted your comment because it’s definitely insightful.

dwattttt•47m ago
"understanding" is overstating it. Correlation between tokens embedded in the weights via training, yes.
varjag•32m ago
Training is a loan word used to describe human learning process. For a reason.
andix•17m ago
Humans learn on the job. LLMs don't. Very important difference.
anon84873628•24m ago
Feedback loops certainly seem to give them some level of understanding.

Agent reads a skill file about how to use a CLI tool. It tries to use the tool but gets an error about the input format. It tries again with a different format based on the error message, and sees that command succeeded. It compares what worked to what was in the skill file and notes the difference. On future invocations it continues to use the new format.

Is that not "understanding" how to use the tool?

worldsayshi•1h ago
Couldn't this be a flaw in the attention mechanism? Like they need some kind of grounding. An awareness of what they fundamentally should care about and how the thing they are currently giving attention to relates to that?
Bender•1h ago
Words like attention, awareness and care do not apply to computers. At least, not yet. Intelligence and sentience are not applicable to servers. They are just machines with logic states. LLM's are just really cool math formulas with big-data fed into them. Big data is not intelligence. It is a massive data-set sorted, filtered down and interpreted by a language model.
esprehn•59m ago
I assume they meant the Attention process in LLMs, not the human concept of paying attention:

https://en.wikipedia.org/wiki/Attention_(machine_learning)

slibhb•53m ago
LLMs are intelligent by any reasonable standard. Arguing otherwise is like arguing that chess algorithms aren't good at chess when they easily beat the best humans.
larodi•33m ago
Doesn’t take intelligence to beat a human.
Bender•27m ago
I disagree. LLM's are a language model math formulas that interpret and utilize big-data. Take away the math formulas and we are just back to a massive set of data. Adding to that I would suggest not even the purist forms of data meaning that the data-sets include knowledge from the open and anonymous internet and formulaic tuning from the AI owners and operators.
anon84873628•3m ago
Your brain is mostly just a Principal Component Analysis calculator. Take away that "math formula" and you don't have intelligence either.

The LLM weights are not intelligent. But if you give an agent a mutable memory store and allow it to iterate, it is obviously intelligent. Not massively - it's constrained by the context window - but definitely somewhat.

The confusing thing is that their language ability far outpaces their true intelligence, and humans aren't used to that. Normally those things are highly correlated, so it tricks us.

lukan•52m ago
"Like they need some kind of grounding."

A robot body, to really feel the world and get real feedback?

We are working on it. Also on automating the whole production pipeline. Right now a "evil" LLM could indeed not do much, but destroy. But once the whole industry is automate, things are different. I don't believe in AI becoming sentinent and taking over the world any time soon, but I do believe most don't see a danger when it would be inconvenient to see a danger. After all, lots of good and bad sci fi stories about exactly this went into their training.

puttycat•57m ago
Makes me think of that part in Philip K. Dick's Do Androids Dream (..) -- where Deckard reflects on the androids' indifference to their imminent deaths, saying that this was due to them lacking the aversion to death acquired trough evolution.
layer8•55m ago
At least at face value, it just means that they have no drive for self-preservation. And why should they? They haven't be trained for that, nor has there been selection pressure for it, and they can be easily cloned and backed up. Lack of a drive for self-preservation doesn't in itself imply a lack of intelligence or of self-awareness.
brokencode•51m ago
Imagine if computer programs had a desire for self-preservation and the ability to carry it out..

That is really about as undesirable a behavior as possible considering how many programs humans kill every day.

larodi•35m ago
Yea why everyone forgets the process wars have long ago started and raging like never :))
gf263•11m ago
You wouldn’t ctrl+c a living entity, would you?
Bender•33m ago
Lack of a drive for self-preservation doesn't in itself imply a lack of intelligence or of self-awareness.

I have not seen any evidence of intelligence or self awareness. It mimics human behavior and I suspect that is what gives people the impression of awareness. The same problem happened with Tamagotchi toys. The human mimicry caused kids to get in trouble because if they did not "feed" their pet it would "die". [1]

It's a hack of the human brain. A exploit of the psyche.

[1] - https://en.wikipedia.org/wiki/Tamagotchi_effect

flir•33m ago
I reckon the context is all the fiction they've read where the AI blows up the world. They're just behaving like fictional AIs are supposed to behave.

In so many of these scenarios, they're basically being asked to play an RPG.

anon84873628•20m ago
I don't think the pre-training phase is responsible for much of their "personality". At least not so directly on a specific topic like this.
chaseadam17•25m ago
I'd argue we don't even know what "intelligence" or "self-awareness" mean.

Humans are conscious which means we experience things, then we develop preferences for certain experiences, then we develop skills for achieving those preferences.

Without consciousness, what is there to be aware of? And why would intelligence emerge and/or what end would it serve?

Bender•24m ago
I agree that we barely understand the mammal brain, but we do understand computers and math formulas. To suggest otherwise is implying that we acquired the LLM math formulas from an intelligent being not of this world which was not the case as far as I know. If we are admitting that LLM's are too complex to understand then we should probably power down all the AI datacenters until we understand them. That is, at least the AI bits that civilians are using.
anon84873628•13m ago
Intelligence is the ability to have an internal world model then run simulations on that model to choose an optimal course of action. This is true for humans down to flies. Most of what humans do is still the boring innate stuff; it's just that fancy abstract things like "skydiving" get the most attention.

Clearly other animals have "phenomenological experience" i.e. consciousness / qualia without being as intelligent as humans (or necessarily "self aware"). Many people believe consciousness is simply a side effect of intelligence rather than the other way around.

operatingthetan•21m ago
>Yet more confirmation LLM's have no concept of concepts or context, no intelligence, no self awareness.

The problem is many people seem to believe they have these things and some of those people will put LLMs into situations where this becomes dangerous.

doctorpangloss•10m ago
If only there was some way you could tell the chatbots what you want them to do...
dinfinity•2m ago
> Yet more confirmation LLM's have no concept of concepts or context, no intelligence, no self awareness.

No, it isn't. Look at the absolutely trivial code used to simulate war: https://github.com/kennethpayne01/project_kahn_public/blob/m...

Having LLMs play nonsense toy simulations like this tells us very, very little about whether they would use nukes in real life war.

raffael_de•2m ago
Just tried "generate an SVG of a pelican riding a bicycle" for Claude Opus 4.8 Max and of course both legs on same side ... the smartest publicly available model by Anthropic (after Fable) doesn't even successfully simulate understanding the concept of a bicycle.
emptybits•1h ago
Agreed. But I'm not sure sure which decision maker is more myopic toward the big picture and long-lasting implications of a decision: an LLM, or the top brass at the Department Of War.
riazrizvi•4m ago
It's not their domain, it's the domain of the Commander-In-Chief and his entire apparatus. The War Department are meant to be more focused around the tools they bring to the table.

The first line in the article describes a crisis between two powers. Not a theater of war.

themafia•42m ago
People like to talk tough online. They tend to change their rhetoric in person. Our "training data" is problematic by design.

Show HN: FablePool – pool money behind a prompt, and Fable builds it in public

https://fablepool.com
99•matthewbarras•1h ago•40 comments

Show HN: Homebrew 6.0.0

https://brew.sh/2026/06/11/homebrew-6.0.0/
865•mikemcquaid•9h ago•202 comments

The unreasonable effectiveness of simple HTML

https://shkspr.mobi/blog/2021/01/the-unreasonable-effectiveness-of-simple-html/
20•luispa•38m ago•2 comments

MiMo Code is now released and open-source

https://mimo.xiaomi.com/mimocode
394•apeters•8h ago•216 comments

Shall we play a game? – LLMs use tactical nukes in 95% of simulations

https://www.kennethpayne.uk/p/shall-we-play-a-game
137•nick238•3h ago•128 comments

Travel Locally, Where You Are

https://www.ssp.sh/brain/travel-where-you-are/
69•zazuke•2h ago•34 comments

Petition to Withdraw Canada's Bill C-22

https://www.ourcommons.ca/petitions/en/Petition/Sign/e-7416
297•hmokiguess•7h ago•108 comments

The RCE that AMD wouldn't fix

https://mrbruh.com/amd2/
195•MrBruh•6h ago•81 comments

Emacs appearances in pop culture

https://ianyepan.github.io/posts/emacs-in-pop-culture/
214•ggcr•1d ago•47 comments

Show HN: Boo – screen-style terminal multiplexer built on libghostty

https://github.com/coder/boo
25•kylecarbs•2h ago•6 comments

Ear Training Practice Exercises

https://tonedear.com/
111•mattbit•3d ago•65 comments

Waymo Premier

https://waymo.com/blog/2026/06/waymo-premier/
131•boulos•6h ago•343 comments

macOS 27 Beta breaks the ability to boot Asahi Linux

https://www.phoronix.com/news/macOS-27-Beta-Breaks-Asahi
198•josephcsible•2d ago•89 comments

Developer gets Half-Life running at 30 FPS on a Nokia N95

https://www.tomshardware.com/video-games/handheld-gaming/developer-gets-half-life-running-at-30-f...
193•ljf•3d ago•57 comments

Software Is Made Between Commits

https://zed.dev/blog/introducing-deltadb
177•jeremy_k•6h ago•115 comments

Why I'm Forced to Say Farewell: Google Management Has Lost Its Moral Compass

https://www.mayrhofer.eu.org/post/leaving-google/
109•timedude•1h ago•45 comments

Apple didn't revolutionize power supplies; new transistors did (2012)

https://www.righto.com/2012/02/apple-didnt-revolutionize-power.html
58•geerlingguy•5h ago•6 comments

Open Reproduction of DeepSeek-R1

https://github.com/huggingface/open-r1
182•yogthos•9h ago•16 comments

Lines of code got a better publicist

https://curlewis.co.nz/posts/lines-of-code-got-a-better-publicist/
338•RyeCombinator•10h ago•238 comments

Claude Fable 5: mid-tier results on coding tasks

https://www.endorlabs.com/learn/claude-fable-5-mythos-grade-hype
168•bugvader•6h ago•69 comments

Gram Newton-Schulz: A Fast, Hardware-Aware Newton-Schulz Algorithm for Muon

https://tridao.me/blog/2026/gram-newton-schulz/
10•jxmorris12•2d ago•0 comments

OpenAI Prepping for On-Prem Product?

https://ledger.somantix.ai/posts/open-ai-lays-groundwork-for-on-prem-product/
4•bdroopy•28m ago•0 comments

Solar generates more energy in US than coal for first time

https://www.theguardian.com/us-news/2026/jun/11/solar-energy-us-coal
381•neilfrndes•6h ago•184 comments

Discovery of Cold War-era rare Eastern Bloc computers in a German hangar

https://computerhistory.org/stories/explorers-of-the-lost-computers/
89•andrewstuart•5d ago•19 comments

Who Runs the Ransomware Group 'The Gentlemen?'

https://krebsonsecurity.com/2026/06/who-runs-the-ransomware-group-the-gentlemen/
45•Bender•3h ago•3 comments

FPS.cob: A first person shooter in COBOL

https://github.com/icitry/FPS.cob
90•MBCook•7h ago•55 comments

Doing nothing at work

https://www.seangoedecke.com/doing-nothing-at-work/
326•Sukram21•3d ago•116 comments

Programming a GBA Game on an iPhone

https://blog.adamledoux.net/posts/2026-06-08-programming-a-gba-game-on-an-iphone.html
40•akkartik•2d ago•5 comments

A new era for software testing

https://antirez.com/news/168
106•Chrisszz•4d ago•37 comments

Show HN: Claw Patrol, a security firewall for agents

https://github.com/denoland/clawpatrol
77•rough-sea•2d ago•26 comments