frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

The Unreasonable Effectiveness of Reasonless Intermediate Tokens

https://arxiv.org/abs/2505.13775
4•YeGoblynQueenne•1y ago

Comments

tocs3•1y ago
I asked ChatGPT to restate this in more laymen's terms (posted below) and I am not to surprised at the answer.

"Lately, some AI models have shown impressive abilities to solve complex problems, and many people credit this to a method called Chain of Thought (CoT), where the model is trained to think through steps like a human might. In this paper, we take a closer look at that idea to see if it's really what's driving better performance.

We focus on the model’s step-by-step thinking (the words it generates along the way) — often treated like human "thoughts" — and examine whether these actually help the model solve problems more accurately. To test this, we train AI models using clean, correct step-by-step reasoning paths and final answers, all based on a known solving method (A* search). This lets us check both the final answers and the reasoning steps to see how they relate.

Interestingly, we find that even when a model gives the right answer, its reasoning steps can still be wrong or messy. To go further, we even train models using completely random and incorrect reasoning steps — and surprisingly, they still perform about the same, and sometimes even better, than those trained on correct steps.

This suggests that the step-by-step "thoughts" the model shows aren’t as meaningful or reliable as many assume. In short, just because a model looks like it’s reasoning through a problem doesn’t mean it actually is — and we should be careful not to treat its outputs as if it thinks like a human or follows strict logic."

Made by Google 2026

https://blog.google/products-and-platforms/devices/pixel/made-by-google-2026/
1•saikatsg•24s ago•0 comments

Did poop enable the evolution of complex animals?

https://arstechnica.com/science/2026/08/feces-fueled-a-flurry-of-evolution-during-the-cambrian-st...
1•jnord•3m ago•0 comments

DeepMind Models

https://deepmind.google/models/
1•nimsarajay•4m ago•0 comments

How do I self-host a good-looking AI generated website

1•pcblues•4m ago•0 comments

How much memory is needed to render a TrueType font to bitmap

https://nothings.org/gamedev/font_rendering_malloc.txt
1•anitil•6m ago•1 comments

The Strongest El Niño Ever Forecast and the Hunger It Will Leave Behind

https://www.4hunger.org/p/scorched-harvest-the-strongest-el
2•fredski42•8m ago•0 comments

Show HN: LinkShip – Turn any file into a trackable, shareable link

https://linkship.net
1•sunpy•10m ago•0 comments

Show HN: Find the jobs that fit your resume in 60 seconds

https://www.seekerscore.com
3•not_wowinter14•15m ago•0 comments

ChatGPT Desktop (Codex Desktop) for Linux

https://openai.com/codex/
5•allanrbo•19m ago•0 comments

Size and Lifespan

https://xkcd.com/3283/
4•dgudkov•20m ago•0 comments

Vibe Coding Interview

http://funcall.blogspot.com/2026/08/vibe-coding-interview.html
5•zombot•21m ago•2 comments

Ratings Firm Accused of Grade Inflation Vouched for $40B of Insurer Debt

https://www.wsj.com/finance/ratings-firm-accused-of-grade-inflation-vouched-for-40-billion-of-ins...
5•petethomas•30m ago•0 comments

Patterns and problems in emerging multiagent systems

https://www.anthropic.com/research/multiagent-systems
4•ledoge•32m ago•0 comments

BrowserMesh – isolated Playwright sessions for MCP clients

https://github.com/scrollDynasty/multi-agent-browser-mcp
4•scroll11•35m ago•0 comments

Celld: Self-hosted, distributed Durable Objects

https://celld.dev/
5•godisdad•36m ago•0 comments

CFTC Advisory on Self-Certification of Incentive Programs for Prediction Markets

https://www.cftc.gov/PressRoom/PressReleases/9282-26
5•petethomas•37m ago•0 comments

Living with Depression: The Part I Never Say Out Loud

https://medium.com/freedomofthought/living-with-depression-the-part-i-never-say-out-loud-3e9c218b...
5•raynchad•37m ago•0 comments

What Kind of Particularism?

https://medium.com/freedomofthought/what-kind-of-particularism-489c113a8a9c
5•raynchad•38m ago•0 comments

Lossless codec for AI agent messages – 36% fewer tokens, overhead counted

https://github.com/reh8n/a2acompress
4•reh8n•40m ago•0 comments

Who invented the great numerical algorithms? [pdf]

https://math.pku.edu.cn/teachers/litj/notes/numer_anal/inventorstalk.pdf
4•nill0•40m ago•0 comments

My First Post

4•VimalRB•41m ago•0 comments

Show HN: CoreTrace, a visual 16-bit CPU simulator

https://coretrace.srianjaneyam.me/
5•mr-anjaneyam•54m ago•2 comments

AI to Kids by Letting Them Tweak Local Chatbots [video]

https://www.youtube.com/watch?v=fNkt41q7eqs
5•dbreunig•59m ago•0 comments

Rsync 3.5 Released as "Extraordinary" Update to Fix 33 Security Issues

https://www.phoronix.com/news/Rsync-3.5
4•Bender•59m ago•0 comments

Long Story Short: Daily word game

https://gizmodo.com/long-story-short-wants-to-be-the-wordle-of-editing-bad-sentences-2000797842
3•akashwadhwani35•1h ago•0 comments

General Dynamics Failed to Build Artillery Shells for the Army

https://www.propublica.org/article/general-dynamics-artillery-factory-failed
5•colinprince•1h ago•0 comments

Is Bending Spoons Good for Startups?

https://s-1.vercel.app/posts/bending-spoons-a-capital-cycle-play/
4•Altaba•1h ago•0 comments

Show HN: UTC Time - live clock, ISO 8601, Unix timestamp

https://utctime.app/
7•nadermx•1h ago•0 comments

DeepSeek V4 Pro 0813: Intelligence, Performance and Price Analysis

https://artificialanalysis.ai/models/deepseek-v4-pro
4•theanonymousone•1h ago•0 comments

Install Manticore Search with one command

https://medium.com/@s_nikolaev/install-manticore-search-with-one-command-8ab5ba07a06b
3•snikolaev•1h ago•0 comments