frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

The Unreasonable Effectiveness of Reasonless Intermediate Tokens

https://arxiv.org/abs/2505.13775
4•YeGoblynQueenne•1y ago

Comments

tocs3•1y ago
I asked ChatGPT to restate this in more laymen's terms (posted below) and I am not to surprised at the answer.

"Lately, some AI models have shown impressive abilities to solve complex problems, and many people credit this to a method called Chain of Thought (CoT), where the model is trained to think through steps like a human might. In this paper, we take a closer look at that idea to see if it's really what's driving better performance.

We focus on the model’s step-by-step thinking (the words it generates along the way) — often treated like human "thoughts" — and examine whether these actually help the model solve problems more accurately. To test this, we train AI models using clean, correct step-by-step reasoning paths and final answers, all based on a known solving method (A* search). This lets us check both the final answers and the reasoning steps to see how they relate.

Interestingly, we find that even when a model gives the right answer, its reasoning steps can still be wrong or messy. To go further, we even train models using completely random and incorrect reasoning steps — and surprisingly, they still perform about the same, and sometimes even better, than those trained on correct steps.

This suggests that the step-by-step "thoughts" the model shows aren’t as meaningful or reliable as many assume. In short, just because a model looks like it’s reasoning through a problem doesn’t mean it actually is — and we should be careful not to treat its outputs as if it thinks like a human or follows strict logic."

Show HN: I made a site where you can explore any football club's trophy cabinet

https://palmares.fyi/
1•naeemnur•3m ago•0 comments

Pareto Front

https://en.wikipedia.org/wiki/Pareto_front
1•binyu•3m ago•0 comments

Flickr iOS App Reverted

https://www.gyford.com/phil/writing/2026/07/28/flickr-ios-app-reverted/
1•speckx•4m ago•0 comments

Does SEO Still Generate Inbound Leads, or Did AI Search Break It?

https://cuescout.com/blog/does-seo-still-generate-inbound-leads
2•weeklyrunner•4m ago•1 comments

Show HN: I built a $8/mo Intercom alternative that integrates with Paddle/Stripe

https://www.userscom.com/
1•ibuildproducts•4m ago•0 comments

Disrupting supply chain attacks on NPM and GitHub Actions

https://github.blog/security/supply-chain-security/disrupting-supply-chain-attacks-on-npm-and-git...
1•nyku•4m ago•0 comments

Foundation Models for Oversight

https://transluce.org/foundation-models-for-oversight
1•paraschopra•5m ago•0 comments

Show HN: I gave Fable a robot body

https://www.talraviv.co/p/i-gave-claude-a-robot-body
1•talsraviv•6m ago•0 comments

Show HN: PassControl – so your AI agents never hold your real API keys

https://passcontrol.vertias.eu
1•vertias3u•6m ago•0 comments

Show HN: A unique social media platform that will replace X, Reddit, Discord:)

https://heahy.com/c/heahychat
1•ignasheahy•7m ago•1 comments

Mark Zuckerberg Disapproves of Centralization of A.I. Power

https://www.nytimes.com/2026/07/28/technology/mark-zuckerberg-meta-ai.html
1•garyrob•8m ago•0 comments

Detecting CSAM Text-to-Image LoRAs from Weights

https://arxiv.org/abs/2607.25750
1•sbulaev•9m ago•0 comments

Deleting Codeberg

https://thanosapollo.org/posts/deleting-codeberg/
2•speckx•9m ago•0 comments

Show HN: Aero Frame – a photo frame for planes flying near you

https://www.aeroframe.app/
1•ipsilo•10m ago•1 comments

Show HN: Dhan360 – Privacy-first portfolio analytics for Indian investors

https://dhan360.in
1•anirudhgoel•10m ago•0 comments

Intel unveiled its iconic Core 2 Duo family 20 years ago – dethroned AMD Athlon

https://www.tomshardware.com/pc-components/cpus/intel-unveiled-its-iconic-core-2-duo-family-20-ye...
1•ksec•11m ago•0 comments

Motherboard VRM thermal testing – budget vs. high-end boards, does it matter?

https://www.tomshardware.com/pc-components/motherboards/motherboard-vrm-thermal-testing-budget-vs...
3•ksec•12m ago•0 comments

How Real-Time Apps Work

https://www.abdulqudus.com/blog/how-real-time-apps-work/
2•jideabdqudus•13m ago•0 comments

Alternative ways to say you are unemployed (from a laid-off Quizlet worker)

https://www.instagram.com/reel/DbWcGXdIfmU/
2•mooreds•14m ago•1 comments

A Backlash Against Anthropic Is Brewing in Silicon Valley

https://www.wsj.com/tech/ai/a-backlash-against-anthropic-is-brewing-in-silicon-valley-3b3ddc80
3•1vuio0pswjnm7•14m ago•1 comments

The Unbearable Slowness of Being: Why do we live at 10 bits/s? (2024)

https://arxiv.org/abs/2408.10234
2•surprisetalk•15m ago•0 comments

Show HN: GPX and FIT File Tools – convert maps, plan routes, edit and more

https://gpxfile.pro/gpx-route-planner
2•jway66•16m ago•0 comments

The First Transatlantic Telegraph Cable Was a Bold, Beautiful Failure

https://spectrum.ieee.org/the-first-transatlantic-telegraph-cable-was-a-bold-beautiful-failure
2•sparsesignal•17m ago•0 comments

Windows Server 2003: The Road to Gold Part One: The Early Years

https://web.archive.org/web/20040202024836/http://www.winsupersite.com/reviews/winserver2k3_gold1...
2•ksec•17m ago•0 comments

AI-found bugs aren't proving any easier to exploit despite the hype

https://www.theregister.com/security/2026/07/28/ai-found-bugs-arent-proving-any-easier-to-exploit...
3•Tomte•18m ago•0 comments

Amazon, Meta, Microsoft face skeptical investors, Google report sparks sell-off

https://www.cnbc.com/2026/07/28/hyperscalers-face-higher-capex-scrutiny-after-alphabet-report-pan...
2•1vuio0pswjnm7•20m ago•0 comments

French musician Kavinsky found dead

https://www.euronews.com/culture/2026/07/29/dj-kavinsky-known-for-his-track-nightcall-found-dead-...
7•bristleworm•21m ago•0 comments

Super Intelligent AI Brain

https://github.com/RATCAR302/Fractal-Pyramidal-Wave-Computer
3•RATCAR302•21m ago•0 comments

Show HN: Experiments with Weather Data

https://strataweather.com/map
3•chrisfreder•23m ago•0 comments

NPM publish-time malware scanning and dual-use metadata

https://github.blog/changelog/2026-07-28-npm-publish-time-malware-scanning-and-dual-use-metadata/
3•Bobaso•25m ago•2 comments