frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Position: LLMs Can't Jump

https://openreview.net/challenge?redirect=%2Fforum%3Fid%3DklU4737opt
63•theanonymousone•1h ago•30 comments

Three Six Mafia – Data about "6/6/6 dating" (2024)

https://divingintheshallowend.com/three-six-mafia/
48•embedding-shape•34m ago•19 comments

Civilian plane crash in New Mexico tied to military GPS blocking

https://www.wired.com/story/a-civilian-plane-crashed-in-new-mexico-was-the-militarys-tech-to-blame/
112•dzdt•1h ago•39 comments

Stateless MCP has recaptured my interest

https://simonwillison.net/2026/Jul/31/stateless-mcp/
282•tosh•4d ago•147 comments

“Gravity is worth asking about”

https://unsung.aresluna.org/gravity-is-worth-asking-about/
140•nozzlegear•6d ago•83 comments

Helsinki Hacker News Meetup

https://calpaterson.com/helsinki-hn.html
127•calpaterson•3h ago•76 comments

Scaling NumPy on Free-Threaded Python

https://labs.quansight.org/blog/scaling-numpy-on-free-threaded-python
25•ngoldbaum•5d ago•5 comments

Pi's Minimalism Is Its Advantage

https://earendil.com/posts/pi-autoresearch-and-databricks/
428•luispa•14h ago•203 comments

Mistral's Shieldstral: 3B open-weights model for multimodal moderation

https://mistral.ai/news/shieldstral/
447•riadsila•20h ago•114 comments

Birduino: A card-triggered audio player for [learning] the birds

https://hannahilea.com/blog/birduino/
15•austinallegro•5d ago•2 comments

Show HN: Simple algorithm and color space to generate diverse skin tones

https://toneyalexander.github.io/inclusive-color-space/
558•automatoney•21h ago•96 comments

Could psilocybin be the key to treating anorexia?

https://www.scientificamerican.com/article/psilocybin-could-kick-start-anorexia-recovery-early-re...
12•tzury•1h ago•3 comments

Zero-Mem: Zero-Token Memory Operations for LLM Agents

https://arxiv.org/abs/2607.29377
64•theanonymousone•8h ago•10 comments

Bubble Memory

https://en.wikipedia.org/wiki/Bubble_memory
19•jacquesm•2d ago•2 comments

The Pneumatics of Hero of Alexandria (1851)

https://www.thehopkinthomasproject.com/TheHopkinThomasProject/TimeLine/Wales/Steam/URochesterColl...
38•gregsadetsky•4d ago•3 comments

IP and DNS Leaks in WebKit Affecting Proxy Browsers and iCloud Private Relay

https://mysk.blog/2026/08/04/webkit-proxy-icloud-private-relay-ip-leak/
135•lapcat•13h ago•22 comments

Show HN: Maple-Preview – Ternary 20B MoE running at 120 tok/s on a iPhone

https://deepgrove.ai/maple-preview
145•edwardbzhang•16h ago•43 comments

In Memory of My Wife, Elise Cawley, with Thanks for 36 Wonderful Years

https://writings.stephenwolfram.com/2026/08/in-memory-of-my-wife-elise-cawley-1961-2026-with-than...
1466•jdcampolargo•17h ago•84 comments

DuckDB – Data power tools for your laptop, now in Clojure (2023)

https://techascent.com/blog/just-ducking-around.html
125•sourdecor•14h ago•18 comments

Bugtraq is back

https://lists.securityfocus.com/hyperkitty/list/bugtraq@securityfocus.com/thread/CHKLXLA7SJEWLDFH...
60•bashtoni•12h ago•20 comments

Polling shows a growing political reckoning is coming for data centers

https://www.politico.com/news/2026/07/21/poll-data-centers-democrats-moratorium-01001799
13•mapping365•1h ago•36 comments

An SLM trained on $8 ESP32-S3

https://github.com/Carloscodix/qapla
45•pavelai•8h ago•15 comments

Rio-vt and librio: Rio's terminal engine, now embeddable

https://rioterm.com/blog/2026/07/27/rio-vt-and-librio
44•vinhnx•1w ago•14 comments

We finally learned to center a div, then browsers added sidebars

https://seg6.space/posts/center-div/
155•seg6•14h ago•125 comments

Eight Myths on Software Engineering and GenAI

https://queue.acm.org/detail.cfm?id=3807963
240•tchalla•12h ago•200 comments

Why is it all in the kernel?

https://lawrencecpaulson.github.io//2026/07/30/Collatz.html
59•ibobev•4d ago•25 comments

There Will Come Soft Rains (1950) [pdf]

https://users.wpi.edu/~zrbutzke/Docs/BradburyStories(1).pdf
397•pmg101•1d ago•407 comments

AI fuels more than half of cybercrime in Africa as scams surge – Interpol

https://www.africanews.com/2026/08/04/ai-fuels-more-than-half-of-cybercrime-in-africa-as-digital-...
267•bookofjoe•14h ago•207 comments

Video2NAND – Abusing video codecs for great computational power

https://sharedobject.blog/posts/vp8-combinatorial-logic/
74•firer•2d ago•16 comments

libexpat now funded by the City of Munich for up to 6 months

https://blog.hartwork.org/posts/libexpat-city-of-munich-open-source-sabbatical/
303•spyc•13h ago•66 comments
Open in hackernews

Position: LLMs Can't Jump

https://openreview.net/challenge?redirect=%2Fforum%3Fid%3DklU4737opt
58•theanonymousone•1h ago

Comments

jvanderbot•52m ago
Came for: "A computer once beat me at chess, but it was no match for me at kick boxing."

TFA was actually about leaps of intuition, sadly.

One of the experiments I've heard proposed around here is to somehow create an LLM from all text up to 1980 or 1990 and see if it can get back to making itself.

Normal_gaussian•42m ago
The curious case here is how much of a description do we give it of itself? That would almost certainly dominate success rates.

My feeling is that a prompt would have to provide a vague description of a program that meaningfully passes something like a Turing test, an API to conform to, an expectation of novel construction (no 'ifs all the way down'), and then a requirement to search broadly and pursue promising ideas and not get hung up on the philosophy. Anything more precise feels like it would corrupt the test, but as it is that description feels doomed to loop before even trying the interesting parts.

ModernMech•35m ago
I wonder if we could just tell it to invent itself without any description and see if it can I introspect enough through its own interface to figure out what it is.
elar_verole•28m ago
I think this can't work because an LLM needs too much data, and before the internet there probably just wasn't enough to get close to what we have now
jvanderbot•23m ago
Even simpler: Can GPT-2 anticipate and build Gwen/Deepseek? I think the answer is almost trivially "no", so I wonder what changed?
throwaway314155•28m ago
Article was plenty interesting to me.
nativeit•51m ago
Previous: https://news.ycombinator.com/item?id=49162791
nativeit•50m ago
Also https://news.ycombinator.com/item?id=49136070
HelloUsername•20m ago
You posted this twice now. Why not just edit your original comment?
nativeit•47m ago
Also:

https://news.ycombinator.com/item?id=49136070 https://news.ycombinator.com/item?id=49096837 https://news.ycombinator.com/item?id=46890333 https://news.ycombinator.com/item?id=46870562

All of these are titled “LLMs Can’t Jump”

conartist6•46m ago
This tracks for me as someone making keeps of intuition in little-explored areas.

I just don't see any of the LLM users around at all. Clearly some force is guiding them all away from thinking any of the "leap of faith" thoughts that I am thinking.

sobiolite•43m ago
The theory is that creative leaps in theoretical physics require a grounding in sensory experience, but the obvious counter-argument is that humans can make creative leaps in abstract fields without such sensory grounding. They do address this at the end, saying

"In abstract domains such as Mathematics or Computer Science, the Sense Experience (E) may be grounded in high-dimensional topology or have other goals such as generality or minimality."

But if such sense experience is possible in abstract domains via some high-dimensional topology, why could a sufficiently advanced LLM not develop an equivalent high-dimensional topology for domains like physics and use it to make creative leaps?

Spacecosmonaut•35m ago
Isnt the sensory grounding even in abstract cases some (limited) intuition that simulates in a mental world model?
bob001•30m ago
That's an interesting analogy. My gut sense is that theoretical mathematics requires a high level of intelligence versus more grounded domains. That may imply that deficiency in grounding can be made up for with intelligence and basically reverse engineering the gaps in grounding from first principles/limited grounding. The ultimate question would then be what is the tradeoffs between grounding and raw intelligence for the same outcome.
roenxi•21m ago
> ...but the obvious counter-argument is that humans can make creative leaps in abstract fields without such sensory grounding...

But we have no idea at how good humans are at that. Given the appalling failures of humans to handle even basic statistical situations like identifying that the same thing happens over and over, it might be that they are hilariously bad at creative leaps in abstract fields, it is just we have had nothing better available to measure against. We've spent about as long as decision theory existed trying to convince people to use it instead of flailing. Limited success, usually in exceptional cases.

And the paper seems a bit dodgy, we have models created with sensory data available. No reason a LLM can't be trained on more sensory data than a human can accumulate in one lifetime. There is a lot of visual data on YouTube.

Mistletoe•40m ago
I feel like you could just add some noise or randomness to the LLM and start approximating the leaps that the human mind uses to solve and understand unrelated things. Maybe that’s naive, it’s just coming from my organic computer in my skull.
Der_Einzige•35m ago
Hahaha you just derived temperature from first principles.

Turns out temperature is pretty bad too, you can find ways to sample from deeper in the distribution without distorting it. Great example is XTC (exclude top choices), In a few weeks/months it'll also have a proper scholarly paper with peer review.

GodelNumbering•25m ago
I have been writing a 'paper' [1] on an adjacent topic for months now. At some point, I decided to make it an empirical paper vs position paper. I am still chasing the experiments (when I get some free time waiting for agentic loops)

For this paper specifically, after reading the abstract [2], I felt almost certain that the author would have used Judea Pearl's ladder of causation (https://web.cs.ucla.edu/~kaoru/3-layer-causal-hierarchy.pdf) but they did not. Would have probably been a better argument to make.

[1] paper in quotes because it may never get published (it is over 20 pages atm). the core argument is that lack of native adjacency resolution makes problems harder and sample inefficient, not impossible

[2] "Using Einstein’s formulation of General Relativity as a case study, we demonstrate that LLMs are structurally incapable of creating new foundational axioms, particularly when observational data is scarce. "

Also, the claim that 'LLMs are structurally incapable of creating new foundational axioms' is provably false depending on where you place 'fundamental'.

bob1029•21m ago
I think this is more of a function of the harness and the environment than the LLM. I've seen some LLM interactions over complex environments like Godot and Unity that challenges the notion that there is no "jumping" going on at all.

An LLM in isolation from its environment might as well be a brain in a vat in some dark cave. You need an external environment to sample from and act upon to make forward progress.

altmanaltman•9m ago
Why does the cave need to be dark if its just a brain in a vat?
defgeneric•16m ago
Worth reposting a follow-up tweet from the author Tom Zahavy [1] after this made the rounds on X/Twitter recently:

> A few reflections on my "LLMs Can’t Jump" paper:

> My position paper recently got some traction here, so I wanted to share a few thoughts and clarify a few things.

> First things first: some people are framing this as "DeepMind is throwing cold water on AI for science" or claiming the paper argues LLMs can never make real scientific discoveries. This is NOT the case.

> This is a personal position paper, not the company's view on AI for science. This is also not my position. As a core contributor to AlphaProof (the first AI system to win an IMO medal), I know firsthand that my colleagues at DeepMind, other frontier labs, and academia have made amazing discoveries with LLMs and will continue to do so. This paper is NOT an "LLMs are a dead end" kind of thing.

> Rather, the paper is the result of a deep dive I took to study the invention of General Relativity. I wanted to explore what it would take for a modern AI system to make that exact kind of jump. Specifically, I focused on the equivalence principle—a key axiom that Einstein formulated through thought experiments grounded in his physical intuition. I was trying to figure out what it would take to give modern AI systems that sort of thinking.

> Giving AI this specific capability isn't necessarily the most urgent thing to do next. It is very likely that improving our current recipes will lead to many exciting discoveries in the near future. In fact, that is what I am personally working on these days (sorry to disappoint you!). It is also quite possible that I am wrong, and that simply scaling our current systems will lead to new inventions in physics and elsewhere.

> Nevertheless, this was my position last winter when I wrote the paper, and I'm sticking to it. I think that there are a few interesting ideas to explore in this space which could influence the next generation of AI systems. I was very lucky to receive a lot of interesting feedback about this position—thank you for all the messages!

[1] https://x.com/TZahavy/status/2082401499628376180

kfarr•14m ago
Best comment on this from 6 months ago: https://news.ycombinator.com/item?id=46870575
reliablereason•9m ago
Clearly LLMs cant do leaps of intuition since their "intuition" is locked after training ends.

The only way a LLM can come up with new ideas if the "idea" appeared as a generalisation durring training or if it was achieved using reason in chain of thought.

quantum_mcts•9m ago
The popular retelling of how Einstein created Special Relativity to "Resolve the contradictions of Michelson-Morly experiments" is very reductive to the history of the question. The epitome is the quote from the paper:

> From the two postulates, Einstein derived the Lorentz trans- formation ...

If Einstein derived them, who is "Lorentz"?

The groundwork for Special Relativity was the study of electrodynamics and symmetries of Maxwell equations. The Einsteins paper was literally called "On the Electrodynamics of Moving Bodies" and never cites Michelson and Morley.

rsfern•8m ago
I found this paper really thought provoking, but I think the conclusion of “world models are the solution” leaves something to be desired. People are already equipping agentic systems with physical simulation tools and exploring action-conditioned world models. This is cool because you can change the rules of the simulation and observe what happens, but it doesn’t address the core question of what to change the rules to, or even what the goal should be in the first place.
brainless•7m ago
I have a weird thought experiment: If you give a GPT-2/3 level LLM tools to search the internet - any document, can it build bigger, better LLMs?

You may think this is not a good test because an older (or say a smaller) LLM can study from the knowledge on the Internet and build. But we are like that - we can access the Universe through our senses.

Can we ever produce anything that is beyond this Universe? I think an LLM that is lacking in knowledge can build more complex systems as long as it can access more data.

SkyBelow•9m ago
If there is enough cross over between real world knowledge engrams and abstract knowledge engram, would this allow for the jump?

One interesting (albeit sad) area which might be related are humans who are never raised with a first language. They seem to never developer abstract reasoning and even seem to lose the ability to develop it later in life. This might indicate there is some 'real world senses' -> 'direct language' -> 'indirect language' -> 'abstract abduction' hierarchy that develops, perhaps related to more real world abductions as a necessary side chain to developing abstract ones.

One of the obvious problems with this is just how difficult we find it to study intelligence purely in humans. We are measure a LLMs by a yardstick that is already known broken, but maybe this is still the right path.