frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

The Unreasonable Effectiveness of Reasonless Intermediate Tokens

https://arxiv.org/abs/2505.13775
4•YeGoblynQueenne•1y ago

Comments

tocs3•1y ago
I asked ChatGPT to restate this in more laymen's terms (posted below) and I am not to surprised at the answer.

"Lately, some AI models have shown impressive abilities to solve complex problems, and many people credit this to a method called Chain of Thought (CoT), where the model is trained to think through steps like a human might. In this paper, we take a closer look at that idea to see if it's really what's driving better performance.

We focus on the model’s step-by-step thinking (the words it generates along the way) — often treated like human "thoughts" — and examine whether these actually help the model solve problems more accurately. To test this, we train AI models using clean, correct step-by-step reasoning paths and final answers, all based on a known solving method (A* search). This lets us check both the final answers and the reasoning steps to see how they relate.

Interestingly, we find that even when a model gives the right answer, its reasoning steps can still be wrong or messy. To go further, we even train models using completely random and incorrect reasoning steps — and surprisingly, they still perform about the same, and sometimes even better, than those trained on correct steps.

This suggests that the step-by-step "thoughts" the model shows aren’t as meaningful or reliable as many assume. In short, just because a model looks like it’s reasoning through a problem doesn’t mean it actually is — and we should be careful not to treat its outputs as if it thinks like a human or follows strict logic."

Show HN: Is this photo edited? Client-side image forensics

https://vajba.com/image-forensics/
1•trivsamt•1m ago•0 comments

AI can't mark its own homework

https://www.techradar.com/pro/ai-cant-mark-its-own-homework
1•ilreb•6m ago•1 comments

Rebuilding AUTOMATIC1111 with Gradio Workflow

https://huggingface.co/blog/gradio-workflow-1111
1•vertigoruntime•7m ago•0 comments

AI rent can save humanity

https://nonlineartransform.substack.com/p/how-ai-rent-can-save-humanity
2•program_whiz•7m ago•1 comments

Why and how you should run your own DNS resolver

https://www.immibis.com/blog/own-dns-resolver
1•immibis2•8m ago•0 comments

Mathematicians want proof OpenAI didn't use their work

https://www.theverge.com/ai-artificial-intelligence/993263/where-does-openai-get-mathematics-trai...
6•kevcampb•8m ago•0 comments

I Made a Real Fly Brain Play Pong. It Didn't Learn. That's the Interesting Part

https://jonatasperaza.medium.com/i-made-a-real-fly-brain-play-pong-it-didnt-learn-and-that-s-the-...
1•sebg•11m ago•0 comments

Show HN: Draw a wheel. See the road it rolls on

https://baselashraf.com/roadwright/
1•BaselAshraf81•12m ago•0 comments

A New Chess Benchmark for Language Models

https://chessbench.ai/timeline
2•soltanov•13m ago•0 comments

The iPhone Duo, the Intelligent Personal Hub, Apple Watch Audio Intelligence

https://stratechery.com/2026/the-iphone-duo-the-intelligent-personal-hub-apple-watch-audio-intell...
1•swolpers•13m ago•0 comments

A New Observable

https://observablehq.com/@observablehq/a-new-observable
1•bluemanshoe2•14m ago•0 comments

Practices I Abandoned with Agents: An Ode to Test-Driven Development

https://adamtornhill.substack.com/p/practices-i-abandoned-with-agents
2•nephrenka•17m ago•0 comments

God Help Us, Let's Try to Learn About Mechanistic Interpretability Techniques

https://www.astralcodexten.com/p/god-help-us-lets-try-to-learn-about
1•andyjohnson0•17m ago•0 comments

Arm Neoverse CSS N4 Launched for Next-Gen CPUs and DPUs – ServeTheHome

https://www.servethehome.com/arm-neoverse-css-n4-launched-for-next-gen-cpus-and-dpus/
1•rbanffy•19m ago•0 comments

What happens in practice if the P=NP problem is solved?

https://heather.cafe/posts/what-if-p-is-np-is-solved/
1•CJefferson•19m ago•0 comments

How Would AI Kill Us All? What to Know About the AI Doomsday Debate

https://www.wsj.com/tech/ai/how-would-ai-actually-kill-us-all-what-to-know-about-the-ai-doomsday-...
2•doener•19m ago•0 comments

The Dress

https://en.wikipedia.org/wiki/The_dress
4•thunderbong•20m ago•0 comments

Arista Networks Arbitrary RCE Vulnerability

https://www.arista.com/en/support/advisories-notices/security-advisory/24730-security-advisory-0174
3•joeblubaugh•21m ago•0 comments

The $2k Phone Should Be Free

https://victorantos.com/posts/the-2000-dollar-phone-should-be-free/
3•victorbuilds•22m ago•0 comments

AI agents that QA-test your web app like real users

https://www.sanitycheck.dev/
1•russellvancuren•22m ago•0 comments

LeMap – Learning Enterprise Map

https://datasong.app
1•thallukrish•23m ago•1 comments

Healthcare AI's next test is integration

https://www.technologyreview.com/2026/09/10/1141421/healthcare-ais-next-test-is-integration/
1•joozio•26m ago•0 comments

Ask HN: What is the proper way of preventing AI labs to use your private data

1•ozgung•27m ago•0 comments

Humble Bundle – 24 Linux, Cloud and Agentic AI Books

https://www.humblebundle.com/books/ultimate-linux-cloud-infrastructure-bundle-packt-books
1•kartikeypan•28m ago•0 comments

Static hosting a coding agent can drive end to end

https://novence.ai
1•LincChad•28m ago•0 comments

Yemen's Houthis Seize Strategic Red Sea Port, Officials Say

https://www.nytimes.com/2026/09/10/world/middleeast/yemens-houthis-seize-strategic-red-sea-port-o...
2•ceejayoz•28m ago•0 comments

Search engine may access the user's location without prompting them

https://chromium.googlesource.com/chromium/src.git/+/main/docs/x_geo_header.md
1•kozika•31m ago•0 comments

iPhone Pro Materials, Ranked

https://volatileinputs.com/2026/04/iphone-pro-materials-ranked/
1•simonebrunozzi•35m ago•0 comments

Keyword Research Tools

https://topkeywordresearchtools.com/
1•Keywordstat•36m ago•0 comments

The Einstein test: what happens when AI tries to rediscover relativity?

https://www.nature.com/articles/d41586-026-02804-x
1•bookofjoe•38m ago•0 comments