frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

The Unreasonable Effectiveness of Reasonless Intermediate Tokens

https://arxiv.org/abs/2505.13775
4•YeGoblynQueenne•1y ago

Comments

tocs3•1y ago
I asked ChatGPT to restate this in more laymen's terms (posted below) and I am not to surprised at the answer.

"Lately, some AI models have shown impressive abilities to solve complex problems, and many people credit this to a method called Chain of Thought (CoT), where the model is trained to think through steps like a human might. In this paper, we take a closer look at that idea to see if it's really what's driving better performance.

We focus on the model’s step-by-step thinking (the words it generates along the way) — often treated like human "thoughts" — and examine whether these actually help the model solve problems more accurately. To test this, we train AI models using clean, correct step-by-step reasoning paths and final answers, all based on a known solving method (A* search). This lets us check both the final answers and the reasoning steps to see how they relate.

Interestingly, we find that even when a model gives the right answer, its reasoning steps can still be wrong or messy. To go further, we even train models using completely random and incorrect reasoning steps — and surprisingly, they still perform about the same, and sometimes even better, than those trained on correct steps.

This suggests that the step-by-step "thoughts" the model shows aren’t as meaningful or reliable as many assume. In short, just because a model looks like it’s reasoning through a problem doesn’t mean it actually is — and we should be careful not to treat its outputs as if it thinks like a human or follows strict logic."

Show HN: TanStack Start and TanStack AI Jev Integration

https://vercel.com/i/using-jev-in-tanstack-start-with-tanstack-ai
1•flashbrew•54s ago•0 comments

Another week, another data breach for Revolut customers

https://www.theregister.com/cyber-crime/2026/09/25/another-week-another-data-breach-for-revolut-c...
1•imalerba•1m ago•0 comments

Ask HN: What does your AI workflow look like?

1•tadziokas•3m ago•0 comments

The Importance of Human Knowledge in the AI Era

https://jonbehnken.substack.com/p/the-case-for-learning
1•greedywhale•6m ago•0 comments

Claude Isn't Allowed to Write Me Prose

https://blog.kvit.app/posts/agent-not-allowed-to-write-prose/
1•skolos•8m ago•0 comments

YouTube stores three auto-extracted frames for every video, and nobody uses them

https://github.com/InfinityLoop1308/PipePipeClient/pull/98
1•jrejaud•9m ago•0 comments

How AI Is Accelerating PCB Design and Prototyping

https://www.eetimes.com/how-ai-is-accelerating-pcb-design-and-prototyping/
1•giuliomagnifico•13m ago•0 comments

Poisoning a Go Cache

https://lukasschwab.me/blog/gen/poisoning-go-cache.html
2•ingve•17m ago•0 comments

Genius AI APı Anahtarı

1•Ayazefe•17m ago•0 comments

Makeraoke: Share your side project in a video between three and nine minutes

https://www.reddit.com/r/Makeraoke/
1•cs1996•17m ago•0 comments

Musk – Official Trailer (2026)

https://www.youtube.com/watch?v=-52sVAzXwLM
2•Betelbuddy•18m ago•0 comments

Sandbox-first AI coding harness

https://chock.ws/
1•rosscomputerguy•22m ago•0 comments

Show HN: Peekaboolean – image and jev-like typed questions in typed anwers out

https://github.com/bykof/peekaboolean
1•bykof•24m ago•0 comments

I Posed as a Problem Gambler. DraftKings Made Me a VIP

https://www.propublica.org/article/draftkings-sports-gambling-problem-vip-fanduel
3•theblazehen•24m ago•1 comments

Show HN: Afterlife Vault – a virtual dead man's switch

https://afterlife-vault.com
1•wmeter•25m ago•1 comments

Using Claude Code: Spending your effort

https://claude.dev/blog/spending-your-effort/
2•e12e•25m ago•0 comments

The year AGI happened, and some unhinged predictions

https://w.ouzu.im/aweb/unhinged
1•dsign•29m ago•0 comments

Show HN: LightCloud – A cloud console organised like file system

https://www.light-cloud.com/
2•yullius•29m ago•0 comments

From game physics to care robots (2023) [The Royal Society] [video]

https://www.youtube.com/watch?v=R1D9Ma4m9Q0
1•zeristor•32m ago•0 comments

How can you quikly tell if a technical article lacks firsthand experience?

2•sarra01•34m ago•0 comments

There are no borders or boundaries on our planet (2023)

https://stargazingguy.co.uk/2023/05/02/there-are-no-borders-or-boundaries-on-our-planet/
1•ssernikk•37m ago•0 comments

Rust for C# and .NET developers guide by Microsoft

https://microsoft.github.io/rust-for-dotnet-devs/latest/
2•AbuAssar•40m ago•0 comments

New Futex Syscalls for Helping Valve's ARM64 Gaming Ambitions

https://www.phoronix.com/news/FUTEX-Robust-List2-Syscalls
1•renehsz•41m ago•0 comments

GLiNER2.5-Decide and Jev: a decision model you run, and one you call

https://soniqo.audio/blog/gliner-decide-vs-jev
2•aufklarer•45m ago•0 comments

I Taught the Iliad to Chinese Teenagers (2021)

https://scholars-stage.org/how-i-taught-the-iliad-to-chinese-teenagers/
1•Tomte•48m ago•0 comments

Retrieval Practice: The Most Powerful Learning Strategy You're Not Using (2017)

https://www.cultofpedagogy.com/retrieval-practice/
2•Tomte•49m ago•0 comments

Show HN: Gigantua – A Black Hole in your Browser

https://claude.ai/artifact/XmWGFtQPgn8aJvU193Ksf8
3•robotsmakelove•51m ago•1 comments

A Maintainer in Residence: Scott Schafer for the Cargo Team

https://blog.rust-lang.org/2026/09/22/announcing-a-maintainer-in-residence-scott-schafer-for-the-...
1•hnisjafx40•54m ago•0 comments

Show HN: I built an AR App that overlays makeup tutorials live on your face

https://urartist.app/
2•Vicmed13•56m ago•3 comments

Scoop: Top AI companies probing security incidents

https://www.axios.com/2026/09/26/openai-anthropic-thousands-ai-security-incidents
1•nsoonhui•56m ago•0 comments