frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

AliExpress runs silent WebAudio fingerprinting that breaks Bluetooth multipoint

https://blog.laserphile.com/2026/08/aliexpress-webpage-keeping-multipoint.html
460•emctech•5h ago•150 comments

Malicious Rust crate Arrayref runs a build-time payload

https://safedep.io/arrayref-proc-macro1-rust-build-time-malware/
201•abhisek•2h ago•146 comments

Show HN: I trained a 125M model to autocomplete piano on-device

https://simedw.com/2026/08/20/midi-autocomplete/
240•simedw•3h ago•61 comments

Clean up Claude 5's token vomit with a separate LLM

https://github.com/zachahn/vomit
14•Bluestein•19m ago•1 comments

HTML Can Do That

https://chrisburnell.com/html-can-do-that/
104•encyclopedism•1d ago•16 comments

DiffusionGemma Technical Report

https://arxiv.org/abs/2608.00146
55•gmays•2h ago•9 comments

Hacking with Claude on a $27 Smart Watch

https://www.mikekasberg.com/blog/2026/08/19/hacking-with-claude-on-a-27-smart-watch.html
31•speckx•1h ago•15 comments

I like 'em thick: an apology to my English teachers

https://www.experimental-history.com/p/i-like-em-thick
16•Ariarule•1d ago•2 comments

Every Model Cheats

https://dreadnode.io/research/every-model-cheats-prompt-level-mitigation-of-cheating-on-offensive...
19•vga805•1h ago•6 comments

CIA funding helped keep NeXT afloat in the 80s

https://www.wsj.com/tech/steve-jobs-apple-next-cia-161b65f9?st=NWWds1&reflink=desktopwebshare_per...
52•EwanG•15h ago•15 comments

Launch HN: Vendo (YC S26) – Let users build features on top of your product

https://github.com/runvendo/vendo
5•yousefh409•15m ago•0 comments

An elliptic curve of rank ≥ 30

https://elliptic-rank.icarm.cloud/curve/273
22•robinhouston•1h ago•8 comments

Mojo is now open source

https://www.modular.com/blog/mojo-open-source
193•visheshdembla•1d ago•51 comments

Proof of Human (YC S23) Is Hiring a Member of Technical Staff

https://www.ycombinator.com/companies/proof-of-human/jobs/ZTZHEbb-member-of-technical-staff
1•timshell•3h ago

Xorg-Server 26.0.99.901

https://lists.x.org/archives/xorg-announce/2026-August/003741.html
30•st_goliath•2h ago•3 comments

Windows brings out the Rorschach test in everyone (2003)

https://devblogs.microsoft.com/oldnewthing/20030825-00/?p=42803
302•luu•9h ago•106 comments

Git at any scale

https://cursor.com/blog/git-at-any-scale
94•meetpateltech•1d ago•7 comments

Why the Ocean Cleanup hasn't solved the plastic pollution crisis

https://therevelator.org/why-ocean-cleanup-has-not-solved-plastic-pollution/
41•sohkamyung•2h ago•34 comments

Bun 1.4

https://bun.com/blog/bun-v1.4
98•meetpateltech•1h ago•35 comments

Double-double: 31 digits of precision without leaving the FPU

https://marekfiser.com/blog/double-double-arithmetic/
7•iliketrains•3d ago•0 comments

Nearly 1,400 live streams from Japan

https://tomarigi.me/
39•pajop•2d ago•8 comments

Theory of Fluids Enters the 21st Century

https://www.quantamagazine.org/theory-of-fluids-enters-the-21st-century-20260817/
36•librasteve•2d ago•2 comments

Show HN: Open-source Stripe Connect alternative

https://zoneless.com
16•tinyprojects•1h ago•5 comments

Show HN: Check if any of the $656M in unclaimed royalties at The MLC is yours

https://pub.doub.ly/
11•knaught•1h ago•2 comments

An American Mosaic (interactive map of ancestry census data)

https://www.nytimes.com/interactive/2026/07/01/us/america-ancestry-census-data-map.html
21•cckolon•3d ago•4 comments

Stop Anthropomorphizing Intermediate Tokens as Reasoning/Thinking Traces

https://arxiv.org/abs/2504.09762
99•nunodonato•1d ago•22 comments

Stwipe Acquires OpenWouter

https://stwipe.com/
138•eatonphil•1h ago•18 comments

Seeing beyond BMI: Estimating cardiometabolic risk with smartphone imagery

https://research.google/blog/seeing-beyond-bmi-estimating-cardiometabolic-risk-with-smartphone-im...
42•leanderjanssen•4h ago•19 comments

CI jobs artifacts should not be difficult

https://deadsimpleci.sparrowhub.io/doc/job-artifacts
14•melezhik•2h ago•10 comments

Risk Engineering

https://risk-engineering.org/
55•throwaw12•5h ago•7 comments
Open in hackernews

Stop Anthropomorphizing Intermediate Tokens as Reasoning/Thinking Traces

https://arxiv.org/abs/2504.09762
99•nunodonato•1d ago

Comments

florianherrengt•23h ago
> While a human may say “aha” to indicate exactly a sudden internal state change, this interpretation is unwarranted for models which do not have any such internal state, and which on the next forward pass will only differ from the pre-aha pass by the inclusion of that single token in their context. Interpreting the “aha” moment as meaningful exemplifies the long-neglected assumption about long CoT models – the false idea that derivational traces are semantically meaningful, either in resemblance to algorithm traces or to human reasoning.

This paper addresses something that has always bothered me about LLMs. You read their reasoning, see something like “Wait, that’s wrong” and then watch them make the exact mistake they just identified.

Jeff_Brown•22h ago
By itself, "aha" carries no insight, but the insight is probably stated immediately after it. In that case the aha is semantically useful, by identifying the insight it is near.
paimapi•20h ago
it's a rhetorical heuristic that a writer should know to use when directing a reader to a declarative that they want them to pay attention to, usually because it's a non-obvious or roundabout insight

when utilized by AI, it's a probabilistic output and it's variable whether or not that rhetorical trick is useful. it also pushes a non-skeptical reader to focus too much on the following text or even to believe that they, themselves, derived some insight. this is effectively a kind of persuasive sophistry which is not helpful - adding rules around it prevents people from deluding themselves with AI

abitmoa•17h ago
It amounts to noise overall, but it has further unwanted and potentially misleading 'properties'. I think it's rather sobering to see how much bandwidth is still being wasted.
ghostpepper•18m ago
Did not read the paper so apologies if this is covered but isn't it possible that there is some recognizable semantic pattern in the training data where an "aha" is often followed by a subtle semantic shift that proves closer to the original premise in some critical way, and by emitting the "aha" token the model causes itself to produce such a subtle semantic shift that pushes the subsequent reasoning closer to the desired response?
wizzwizz4•20h ago
> but the insight is probably stated immediately after it.

If the intermediate tokens represent reasoning or thought, you would expect "aha" to occur after the thoughts that led to the realisation, including the thoughts encoding the explanation: they don't have any other state. There is no reason to draw the conclusion you've drawn. Furthermore, what LLMs are doing isn't thought.

deaton•21m ago
It really isn't useful though, unless it is a summary. At best it is a semantic trick to tell the next iteration to come up with something smart.
Terr_•23h ago
I've been calling them film noir internal monologues, within the documents being generated by the LLM which happen to look like movie scripts.

In other words, it isn't qualitatively different from character dialogue. "Keep cheese on your pizza by using glue" is the same problem regardless of whether the script calls for the character to speak it out-loud or not.

clhodapp•22h ago
Seems like they are closer to scratch than reasoning... Generating some scratch to draw from helps make it easier to compute the real answer.
cyanydeez•20h ago
I assume theyre searching the local gradient to see if theres a better descent before proceeding.
c0_0p_•12h ago
I don't think there's anything like that going on. They just word vomit into a secondary area, and then there is an internal prompt that says "clean this up and summarize for the user".
eigenspace•8h ago
LLMs dont do gradient descent to generate tokens.

They are trained by gradient descent, but inference doesnt involve it.

forgotTheLast•2h ago
That's my personal theory too. The model is stuffing its own context with vaguely related tokens, which helps the attention heads retrieve the right tokens.
basedpolymer•20h ago
The anthropomorphization of LLMs should be discouraged as much as possible. It perpetuates bad practices and encourages the use of these bots for tasks they are not intended for (particularly as chatbots).

Thinking traces should be treated as black boxes. There is no point in reading them. Only the LLMs’ conclusions are relevant. This is particularly true of Opus 5, which employs reasoning that seems highly questionable but very often reaches excellent conclusions (compared to its peers)

FloorEgg•18h ago
Sometimes I monitor thinking traces for misunderstandings (missing context / bad assumptions). If it's going to go off on a ~20 min task and I can catch it's going in the wrong direction in the first minute I save a lot of tokens and wasted time. I don't monitor the whole thing, mostly just the first bit to see if there was a gap or misalignment in intention.

As an aside, anthropomorphization has nothing to do with my motivations.

qarl2•3h ago
> The anthropomorphization of LLMs should be discouraged as much as possible.

And yet, they have extensive human-like behavior. If you treat them nicely or encourage them, they perform better.

Ignoring that human-like behavior is wrong headed.

andai•33m ago
A while back I made an "OpenClaw in 50 lines" by just wrapping Claude Code in a Telegram bot.

I asked it for the weather. "I don't know that. I'm just a programmer."

I added "believe in yourself, you can do anything" to sysprompt, suddenly it had the confidence to Google the weather...

thaanpaa
porridgeraisin•19h ago
Related:

Poster side dialogue and Q&A about this work at ICML.

https://news.ycombinator.com/item?id=49277303

smugtrain•11h ago
Strong dislike for papers that tell me what to do in the title, especially when even the paper admits a loose correlation of the intermediate tokens compared to solution correctness. My solutions work and they speak for themselves.
fabsalvadori•6m ago
There is a useful engineering consequence here beyond terminology.

If intermediate tokens are not a faithful representation of the computation, then they are a pretty bad audit artifact too. We probably shouldn't be trying to make the model's internal narration more interpretable., but rather the computation around it more reproducible.

Record the actual inputs, model/version/configuration, tool observations and outputs, then make the execution replayable enough that differences between runs can be isolated.

In other words, don't ask the model to explain what it thought, and instead make the system able to show what actually happened.

•
33m ago
That's not a consequence of an LLM. It's a consequence of the training data. In fact, I would argue that the latest models aren't nearly as sensitive to the tone of input anymore. It's an issue that has been addressed by better curating training data.
fedpost•30m ago
I think you're ending that train of thought too early. Why does this occur?

Well... We can hypothesize that these things are largely trained on internet dialogue so there's probably some correlation between threads where people are not flaming each other and the quality of the replies. They're just statistical engines so anything you can do to raise the odds of a helpful next token...

I'm essentially just making shit up here, maybe it's right, maybe it isn't, but rather than saying "it's human and we should treat it so" we're trying to get to the ground truth of how it works.