frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

The Unreasonable Effectiveness of Reasonless Intermediate Tokens

https://arxiv.org/abs/2505.13775
4•YeGoblynQueenne•1y ago

Comments

tocs3•1y ago
I asked ChatGPT to restate this in more laymen's terms (posted below) and I am not to surprised at the answer.

"Lately, some AI models have shown impressive abilities to solve complex problems, and many people credit this to a method called Chain of Thought (CoT), where the model is trained to think through steps like a human might. In this paper, we take a closer look at that idea to see if it's really what's driving better performance.

We focus on the model’s step-by-step thinking (the words it generates along the way) — often treated like human "thoughts" — and examine whether these actually help the model solve problems more accurately. To test this, we train AI models using clean, correct step-by-step reasoning paths and final answers, all based on a known solving method (A* search). This lets us check both the final answers and the reasoning steps to see how they relate.

Interestingly, we find that even when a model gives the right answer, its reasoning steps can still be wrong or messy. To go further, we even train models using completely random and incorrect reasoning steps — and surprisingly, they still perform about the same, and sometimes even better, than those trained on correct steps.

This suggests that the step-by-step "thoughts" the model shows aren’t as meaningful or reliable as many assume. In short, just because a model looks like it’s reasoning through a problem doesn’t mean it actually is — and we should be careful not to treat its outputs as if it thinks like a human or follows strict logic."

Mirror Life Is Theoretical. The Danger Is Anything But

https://www.rand.org/pubs/commentary/2026/08/mirror-life-is-theoretical-the-danger-is-anything-bu...
1•jonbaer•39s ago•0 comments

Ponytail: Lazy Senior Engineer Skill

https://ponytail.dev/
1•zatkin•4m ago•0 comments

Abliteration.ai is making a business out of removing AI guardrails

https://techcrunch.com/2026/09/03/abliteration-ai-is-making-a-business-out-of-removing-ai-guardra...
1•NoRagrets•13m ago•0 comments

Always free and privacy-first image editor – PicFiddle

https://picfiddle.com
1•iplaypc•13m ago•1 comments

New method predicts where earthquakes will strike

https://phys.org/news/2026-09-method-massive-earthquakes.html
1•toomuchtodo•18m ago•1 comments

Don't build your organization around a model provider

https://juanreyero.com/article/ai/model-provider-independence
1•juanre•19m ago•0 comments

Everyone Has a Distaste for the Artificial

https://www.emailstotheworld.com/p/everyone-has-a-distaste-for-the-artificial
1•rsktaker•20m ago•0 comments

Speech2Speech – browser-based STT and TTS for interacting with local LLMs

https://github.com/rhulha/Speech2Speech
1•Curiositry•20m ago•0 comments

Bespoke: A Programming Language for People Who Say Please – Larvitz Blog

https://blog.hofstede.it/bespoke-a-programming-language-for-people-who-say-please/
1•rbanffy•27m ago•0 comments

Music Theory

https://brianberns.github.io/
1•thunderbong•28m ago•0 comments

Army to spend $465M on 'Group 3 killer' anti-drone laser

https://taskandpurpose.com/news/army-locust-laser-drone-killer/
1•bushwart•35m ago•0 comments

ROCm 10.0: A Decade of Open Compute, Built for the Age of Agentic AI

https://rocm.blogs.amd.com/ecosystems-and-partners/rocm-x-blog/README.html
2•metrofun•36m ago•0 comments

AI Is Already Making Us Less Human

https://www.theatlantic.com/ideas/2026/09/open-ai-consciousness-morality/688535/
1•pseudolus•44m ago•2 comments

GPT-6 Astra beats portal [video]

https://www.youtube.com/watch?v=g5u2y0BwRJ0
1•aizk•47m ago•0 comments

Airbnb cut authentication code by 60% with server-driven flows

https://medium.com/airbnb-engineering/flexible-authentication-reimagining-authentication-for-mill...
2•CoderLim110•59m ago•0 comments

Where Temporary Decisions Become Permanent Behavior

https://medium.com/@gurvinder372/article-8-where-temporary-decisions-become-permanent-behavior-0f...
2•gps372•59m ago•0 comments

216M Spy TVs – The LG Smart TV Problem [video]

https://www.youtube.com/watch?v=6IFVTcM28KA
4•treve•1h ago•0 comments

MathKernel: An evidence-aware multi-engine mathematics kernel and MCP server

https://github.com/Staatsgeheim/MathKernel
2•staatsgeheim•1h ago•0 comments

Meditations on Moloch (2014)

https://www.slatestarcodexabridged.com/Meditations-On-Moloch
1•vermilingua•1h ago•0 comments

The Death of Educational Content on YouTube [video]

https://www.youtube.com/watch?v=-Gnrp_caPvo
1•josephcsible•1h ago•1 comments

Prediction of maternal and infant outcomes with a Mother-Child AI agent

https://doi.org/10.1038/s41591-026-04694-y
1•sbulaev•1h ago•0 comments

LED Sports Ticker – A high-performance 60fps stadium-style scoreboard

https://play.google.com/store/apps/details?id=com.tickerled.sports&hl=en_US
1•scrolling•1h ago•0 comments

Hard-Chat – A serverless, RAM-only P2P terminal chat

https://github.com/mrhardlint/Hard-Chat
3•hardlint•1h ago•0 comments

Five dead after Amazon-branded cargo plane overruns runway at Miami airport

https://www.nbcnews.com/news/us-news/amazon-plane-crash-miami-rcna596369
1•r_singh•1h ago•0 comments

App Review Shenanigans

https://kaylees.site//app-review-shenanigans.html
3•latexr•1h ago•1 comments

Four Weeks of a Vegan Diet Alter Signs of Inflammation and Aging

https://www.uniklinik-freiburg.de/en/press/press-releases/detailed-view/6938-vier-wochen-vegane-e...
26•marvinborner•1h ago•22 comments

CrisperWhisper – Speech to Text Model That Transcribes What You Say Verbatim

https://github.com/nyrahealth/CrisperWhisper
2•Curiositry•1h ago•0 comments

The Turpentine Effect (2010)

https://ribbonfarm.com/2010/03/18/the-turpentine-effect/
2•Curiositry•1h ago•0 comments

Pivot to AI safety, I beg you

https://ceselder.substack.com/p/pivot-to-ai-safety-i-beg-you
7•doitLP•1h ago•7 comments

Genius AI Detector

https://geniusaidetector.com/
1•ohjeez•1h ago•0 comments