frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

The Unreasonable Effectiveness of Reasonless Intermediate Tokens

https://arxiv.org/abs/2505.13775
4•YeGoblynQueenne•1y ago

Comments

tocs3•1y ago
I asked ChatGPT to restate this in more laymen's terms (posted below) and I am not to surprised at the answer.

"Lately, some AI models have shown impressive abilities to solve complex problems, and many people credit this to a method called Chain of Thought (CoT), where the model is trained to think through steps like a human might. In this paper, we take a closer look at that idea to see if it's really what's driving better performance.

We focus on the model’s step-by-step thinking (the words it generates along the way) — often treated like human "thoughts" — and examine whether these actually help the model solve problems more accurately. To test this, we train AI models using clean, correct step-by-step reasoning paths and final answers, all based on a known solving method (A* search). This lets us check both the final answers and the reasoning steps to see how they relate.

Interestingly, we find that even when a model gives the right answer, its reasoning steps can still be wrong or messy. To go further, we even train models using completely random and incorrect reasoning steps — and surprisingly, they still perform about the same, and sometimes even better, than those trained on correct steps.

This suggests that the step-by-step "thoughts" the model shows aren’t as meaningful or reliable as many assume. In short, just because a model looks like it’s reasoning through a problem doesn’t mean it actually is — and we should be careful not to treat its outputs as if it thinks like a human or follows strict logic."

The High Ponytail That Made Lydia Corbett Picasso's Muse

https://www.wsj.com/real-estate/lydia-corbett-picasso-muse-37f2e088
1•petethomas•6m ago•0 comments

The Computer, De-Invented

https://daylightcomputer.com
1•signa11•6m ago•0 comments

Calm Technologies That Excite Me

https://abhi.now/blog/calm-technologies/
1•signa11•8m ago•0 comments

Show HN: Mcpgen – Turn OpenAPI/Postman Specs into Python MCP Servers

https://github.com/JnanaSrota/mcpgen
1•princea747•9m ago•0 comments

Why care about programming languages

https://ebellani.github.io/blog/2026/why-care-about-programming-languages/
1•signa11•10m ago•0 comments

How good is your AI Gateway?

https://www.highflame.com/blog/the-three-moments-your-ai-gateway-can-ruin/
1•sharathr•12m ago•2 comments

Local AI that finds sensitive files on your Mac before attackers do

https://www.vaultsort.com/guardian
1•VaultSort•13m ago•0 comments

AI chatbots can be as effective as humans at emotional support, sometimes better

https://www.manchester.ac.uk/about/news/ai-chatbots-can-be-as-effective-as-humans-at-emotional-su...
1•giuliomagnifico•14m ago•0 comments

SteamForge – Modern PC Gaming Achievement Tracker (No Account Required)

https://steamforge.app/
1•Gengar•14m ago•0 comments

India's solar push idles factories unable to shake reliance on China

https://www.reuters.com/business/energy/indias-solar-push-idles-factories-unable-shake-reliance-c...
2•JumpCrisscross•20m ago•0 comments

Was the Hannibal Directive activated on October 7, 2023?

https://medium.com/freedomofthought/was-the-hannibal-directive-activated-on-october-7-2023-52b42f...
2•raynchad•21m ago•0 comments

Stop pretending billionaires built the future

https://www.elysian.press/p/billionaires-didnt-create-innovation
3•aard•21m ago•0 comments

Warsh Confuses Traders, Makes Them Guess Next Week's Rate Move

https://www.bloomberg.com/news/articles/2026-07-22/warsh-leaves-bond-traders-in-the-dark-on-next-...
1•petethomas•21m ago•0 comments

Show HN: I built my wife an ad-free news brief that fact-checks and flags bias

https://beamwire.ai/
1•malammar•24m ago•1 comments

ProkitQ – Free tools for entrepreneurs with no signup required

https://prokitq.com/tools
2•isurumahesh•27m ago•1 comments

A searchable index of nonfiction book prizes

https://book-prize-index.vercel.app
1•saikatsg•29m ago•0 comments

How to Write a Quine

https://czterycztery.pl/slowo/quine-EN.html
1•mci•31m ago•0 comments

Startup founders urge Trump not to shut off Chinese open weight AI

https://www.politico.com/news/2026/07/22/startup-founders-urge-trump-not-to-shut-off-chinese-open...
3•tosh•35m ago•0 comments

Show HN: Trangram – a free browser-based vector editor (illustrator + timeline)

https://www.trangram.com
2•trangram•37m ago•0 comments

AI, hackers, and prediction markets threaten the indie web

https://stephenfollows.com/p/what-just-happened-to-thenumberscom-should-worry-us-all
1•alexnew•37m ago•0 comments

Fractal by Plasma AI

https://www.plasma.ai/research/fractal
1•handfuloflight•37m ago•0 comments

Hiring Senior Front End Developer

1•birthdayn•38m ago•1 comments

Pip 26.2: –only-deps solves 16 years of app deployment hacks

https://jamesoclaire.com/2026/07/23/pip-26-2-only-deps-solves-16-years-of-app-deployment-hacks/
1•ddxv•39m ago•0 comments

A gas-rich outback station could host a $28B off-grid AI data centre

https://thenextweb.com/news/australia-28bn-data-centre-project
1•langfo•40m ago•0 comments

How Agile Are You Really?

https://rethinkingsoftware.substack.com/p/how-agile-are-you-really
1•aard•42m ago•0 comments

Show HN: KalmDown – a phone lock with a passcode even you don't know

https://www.kalmdown.app/
1•dmsehuang•43m ago•0 comments

Travis Kalanick's Atoms raises $1.7B

https://atoms.co/unfinished-business
2•buildingrobots•49m ago•1 comments

Show HN: Free tool that reads a recruiter email and writes your reply

https://www.offerflowai.com/tools/recruiter-reply
1•csongorczezar•53m ago•0 comments

Using sed to make indexes for books (long)

https://www.pement.org/sed/make_indexes.txt
1•TMWNN•54m ago•0 comments

Show HN: RSVP Reader specially for PDF books

https://catb1t.github.io/seshat/
1•cab1t•57m ago•0 comments