frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

The Unreasonable Effectiveness of Reasonless Intermediate Tokens

https://arxiv.org/abs/2505.13775
4•YeGoblynQueenne•1y ago

Comments

tocs3•1y ago
I asked ChatGPT to restate this in more laymen's terms (posted below) and I am not to surprised at the answer.

"Lately, some AI models have shown impressive abilities to solve complex problems, and many people credit this to a method called Chain of Thought (CoT), where the model is trained to think through steps like a human might. In this paper, we take a closer look at that idea to see if it's really what's driving better performance.

We focus on the model’s step-by-step thinking (the words it generates along the way) — often treated like human "thoughts" — and examine whether these actually help the model solve problems more accurately. To test this, we train AI models using clean, correct step-by-step reasoning paths and final answers, all based on a known solving method (A* search). This lets us check both the final answers and the reasoning steps to see how they relate.

Interestingly, we find that even when a model gives the right answer, its reasoning steps can still be wrong or messy. To go further, we even train models using completely random and incorrect reasoning steps — and surprisingly, they still perform about the same, and sometimes even better, than those trained on correct steps.

This suggests that the step-by-step "thoughts" the model shows aren’t as meaningful or reliable as many assume. In short, just because a model looks like it’s reasoning through a problem doesn’t mean it actually is — and we should be careful not to treat its outputs as if it thinks like a human or follows strict logic."

Waterfox Version 6.7.4 Released

https://www.waterfox.com/releases/6.7.4/
1•tmtvl•39s ago•0 comments

When the Debugger Lies

https://danielmangum.com/posts/when-the-debugger-lies/
1•hasheddan•2m ago•0 comments

Show HN: Online Google Maps Scraper

https://gmapscrawl.com
1•qwikhost•2m ago•0 comments

Anthropic Is Suing Meta

https://twitter.com/bunjavascript/status/2102630092451217782
1•TiredOfLife•2m ago•0 comments

F5 patches BIG-IP APM zero-day flaw exploited in RCE attacks

https://www.bleepingcomputer.com/news/security/f5-warns-of-big-ip-apm-remote-code-execution-zero-...
1•sbulaev•12m ago•0 comments

OpenGym, A self-hosted gym and body-weight tracker (AGPL)

https://github.com/DuarteSantos8/openGym
1•duarteSantos•14m ago•0 comments

Z80 REPL

https://abagames.github.io/z80-repl/index.html
2•adunk•15m ago•0 comments

The Download: why AI's latest breakthroughs and fears may be more hype than rea

https://www.technologyreview.com/2026/09/22/1144910/the-download-dont-believe-ai-hype/
5•joozio•18m ago•0 comments

Memristive Singular Value Decomposition

https://www.nature.com/articles/s41467-026-76272-2
1•bookofjoe•19m ago•0 comments

Blood Transfusion: Jehovah's Witnesses revise medical guidelines

https://gazettengr.com/blood-transfusion-jehovahs-witnesses-revise-medical-guidelines/
1•nephihaha•20m ago•0 comments

Rust Allocator API Stabilized

https://github.com/rust-lang/rust/pull/156882
1•Georgelemental•21m ago•0 comments

Llama.cpp Under the Hood

https://www.cppdepend.com/blog/llama-cpp-under-the-hood/
1•cppfirst•23m ago•0 comments

The darker side of being a doctor

https://drericlevi.pages.dev/the-darker-side-of-being-a-doctor/
28•Danhale93•25m ago•2 comments

Show HN: I made my own scripting language for my game engine

https://github.com/ArcadeMakerSources/ArcadeMaker/tree/master
1•themondayguy•27m ago•0 comments

Ask HN: Do you think we'll ever have local models of Fable level?

2•Nair0•27m ago•0 comments

Anthropic-linked CVEs pile up, attackers mostly shrug

https://www.theregister.com/security/2026/09/21/anthropic-linked-cves-pile-up-attackers-mostly-sh...
1•vismit2000•27m ago•0 comments

Code Genome Project

https://codegenomeproject.org/
3•mihau•28m ago•0 comments

Solar Aquagrid

https://www.solaraquagrid.com
1•blondie9x•30m ago•0 comments

Solitaire Alone Together making solitaire a little social

https://eieio.games/blog/solitaire-alone-together/
1•Bluestein•31m ago•0 comments

Ask HN: Jev - Anyone built anything useful for day to day work with Jev?

1•sankalpdomore•32m ago•0 comments

Apple Intelligence uses up to 30GB+ on macOS 27

https://www.macrumors.com/2026/09/23/apple-intelligence-30gb-some-macs-macos-27/
1•tosh•34m ago•1 comments

Show HN: A game where a bell goes techno

https://www.technobell.run/?r=hn
1•dejv_cz1•35m ago•1 comments

Manycore Processor

https://en.wikipedia.org/wiki/Manycore_processor
3•peter_d_sherman•38m ago•1 comments

2015 Office of Personnel Management data breach

https://en.wikipedia.org/wiki/2015_Office_of_Personnel_Management_data_breach
1•chistev•38m ago•0 comments

Ask HN: How do you handle sensitive data in distributed systems?

1•Securelytixdev•42m ago•0 comments

The Booker Prize 2026

https://thebookerprizes.com/the-booker-library/prize-years/2026
1•downbad_•42m ago•0 comments

GPT-6 Astra has made a major breakthrough in the Goldbach Conjecture

https://twitter.com/captain_sude/status/2100277626120163664
4•nsoonhui•42m ago•0 comments

Better prompt caching for GPT-6

https://openai.com/index/better-prompt-caching-for-gpt-6/
1•theanonymousone•44m ago•0 comments

A thread of early explorations with Claude Opus 5.5

https://twitter.com/claudeai/status/2102471866635919731
1•tzury•46m ago•0 comments

Next.js 16.3.6 fixes critical ImageResponse RCE (CVE-2026-94545)

https://nextjs.org/blog/nextjs-security-update-september-22-2026
2•joshcsimmons•46m ago•1 comments