frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

The Unreasonable Effectiveness of Reasonless Intermediate Tokens

https://arxiv.org/abs/2505.13775
4•YeGoblynQueenne•1y ago

Comments

tocs3•1y ago
I asked ChatGPT to restate this in more laymen's terms (posted below) and I am not to surprised at the answer.

"Lately, some AI models have shown impressive abilities to solve complex problems, and many people credit this to a method called Chain of Thought (CoT), where the model is trained to think through steps like a human might. In this paper, we take a closer look at that idea to see if it's really what's driving better performance.

We focus on the model’s step-by-step thinking (the words it generates along the way) — often treated like human "thoughts" — and examine whether these actually help the model solve problems more accurately. To test this, we train AI models using clean, correct step-by-step reasoning paths and final answers, all based on a known solving method (A* search). This lets us check both the final answers and the reasoning steps to see how they relate.

Interestingly, we find that even when a model gives the right answer, its reasoning steps can still be wrong or messy. To go further, we even train models using completely random and incorrect reasoning steps — and surprisingly, they still perform about the same, and sometimes even better, than those trained on correct steps.

This suggests that the step-by-step "thoughts" the model shows aren’t as meaningful or reliable as many assume. In short, just because a model looks like it’s reasoning through a problem doesn’t mean it actually is — and we should be careful not to treat its outputs as if it thinks like a human or follows strict logic."

Writing with AI Is Stupid

https://lambdaland.org/posts/2026-08-07-ai-writing-stupid/
1•frizlab•37s ago•0 comments

Better Batteries

https://matklad.github.io/2026/08/20/better-batteries.html
1•olexsmir•3m ago•0 comments

MOQ Video Streaming for Robots, Drones, and Embedded Devices

https://www.red5.net/blog/moq-video-streaming-for-robots-drones-and-embedded-devices/
1•mondainx•14m ago•0 comments

Claude Opus 4.6 returned nothing 900/900 times. Should agents retry?

https://zenodo.org/records/21696066
4•aiagenttester•18m ago•1 comments

Why to Start a Startup in a Bad Economy (2008)

https://paulgraham.com/badeconomy.html
1•tosh•19m ago•0 comments

Stop Clutching Your FPV Drones

https://foxandlion.pub/analysis/stop-clutching-your-fpv-drones
4•ewfeber•21m ago•0 comments

El Niño set to be 'strongest in living memory', says Met Office

https://www.bbc.co.uk/weather/articles/c3ekg93vjz9o
2•iamben•21m ago•0 comments

ISO 24495-1:2023 Plain language

https://www.iso.org/standard/78907.html
2•Bluestein•22m ago•0 comments

Show HN: Meg – I processed 259M nodes on my laptop using 6.91MB RAM

https://github.com/blueray313164-a11y/Quantum-Drive-Solver
1•blueray313164•26m ago•0 comments

See Through Walls with WiFi

https://github.com/ruvnet/RuView
1•chunkyslink•28m ago•1 comments

The Lost Treasure of Sid Meier's Pirates

https://remapradio.com/articles/the-lost-treasure-of-sid-meiers-pirates/
4•spankibalt•28m ago•0 comments

Going Freestanding

https://antonz.org/going-freestanding/
2•ingve•32m ago•0 comments

Small Models Can Introspect, Too (2025)

https://vgel.me/posts/qwen-introspection/
1•networked•33m ago•0 comments

Reducing C++ template bloat by factoring out the type-dependent portions

https://devblogs.microsoft.com/oldnewthing/20260820-00/?p=112629
1•ingve•34m ago•0 comments

Where to Draw the Line Which decisions should not be handed over to LLMs?

https://queue.acm.org/detail.cfm?id=3834784
1•tmanolatos•38m ago•2 comments

Domain-Driven Design matters more when AI writes your code

https://threedots.tech/post/ddd-and-ai-coding/
1•BerislavLopac•40m ago•0 comments

The LLM rewrote the email after I approved it

https://www.crusbro.com/en/product-architecture.html
1•crusbro•41m ago•0 comments

Updating a side project with AI in 275 commits

https://benhoyt.com/writings/updating-gifty-with-ai/
1•ingve•42m ago•1 comments

Show HN: Plainspeak, make AI agents write like a human

https://sufiyan.cc/plainspeak
1•codersufiyan•43m ago•0 comments

Show HN: I built a TV like website for Indian nostalgic shows

https://www.dabbatv.xyz/
2•codepeddler•48m ago•2 comments

From Magna Carta to microchip – The measurement of time (1981)

https://www.rigb.org/explore-science/explore/video/magna-carta-microchip-measurement-time-1981
1•pillars•49m ago•0 comments

The Download: polycrisis support networks and a hydrogen gold rush

https://www.technologyreview.com/2026/08/20/1142579/the-download-polycrisis-support-networks-unde...
1•joozio•50m ago•0 comments

We Rebuilt the Linux MicroVM Stack on Apple Silicon

https://encore.dev/blog/firecracker-apple-silicon
10•signa11•52m ago•0 comments

Show HN: Real Valued Chess – chess on a continuous board

https://martinylu.com/real-chess
2•martinAsdf•52m ago•2 comments

Video to Prompt

https://www.videotoprompt.dev
1•daisyjin•55m ago•0 comments

RL Policy Churn

https://twitter.com/id_aa_carmack/status/2090514515129520516
1•tosh•59m ago•0 comments

Young Americans increasingly fear AI will take their jobs

https://www.pewresearch.org/short-reads/2026/08/18/young-adults-in-the-us-are-increasingly-wary-o...
4•01-_-•1h ago•0 comments

Mathematically Perfect solar array in Factorio [video]

https://www.youtube.com/watch?v=CDzJ1_p3uVg
1•operatorius•1h ago•0 comments

Pretraining a Mini Kimi K3

https://books.vizuara.ai/book/pretraining-a-mini-k3
1•ngaut•1h ago•0 comments

Trump Threatens 'Tremendous Consequences' for Countries Doing Business with Iran

https://canews24.online/?p=66
2•Edymilson•1h ago•2 comments