frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

The Unreasonable Effectiveness of Reasonless Intermediate Tokens

https://arxiv.org/abs/2505.13775
4•YeGoblynQueenne•1y ago

Comments

tocs3•1y ago
I asked ChatGPT to restate this in more laymen's terms (posted below) and I am not to surprised at the answer.

"Lately, some AI models have shown impressive abilities to solve complex problems, and many people credit this to a method called Chain of Thought (CoT), where the model is trained to think through steps like a human might. In this paper, we take a closer look at that idea to see if it's really what's driving better performance.

We focus on the model’s step-by-step thinking (the words it generates along the way) — often treated like human "thoughts" — and examine whether these actually help the model solve problems more accurately. To test this, we train AI models using clean, correct step-by-step reasoning paths and final answers, all based on a known solving method (A* search). This lets us check both the final answers and the reasoning steps to see how they relate.

Interestingly, we find that even when a model gives the right answer, its reasoning steps can still be wrong or messy. To go further, we even train models using completely random and incorrect reasoning steps — and surprisingly, they still perform about the same, and sometimes even better, than those trained on correct steps.

This suggests that the step-by-step "thoughts" the model shows aren’t as meaningful or reliable as many assume. In short, just because a model looks like it’s reasoning through a problem doesn’t mean it actually is — and we should be careful not to treat its outputs as if it thinks like a human or follows strict logic."

Jev Based Code Review

https://github.com/egma-ai/jev-code-reviewer
1•namanbhulawat•31s ago•1 comments

Plex: Remote Streaming via Third-Party Apps to Require a License

https://netguide.io/news/en/2026/09/22/plex-remote-streaming-via-third-party-apps-to-require-a-li...
2•sippingabonedry•1m ago•0 comments

Proofmode

https://proofmode.org/
1•F3nd0•3m ago•0 comments

Vulnerability Cve-2026-82958

https://vulnerability.circl.lu/vuln/cve-2026-82958
1•anonyoum•8m ago•0 comments

What your phone is made of

https://anas-sabbar.ca/writing/what-your-phone-is-made-of/
1•Roccan•9m ago•0 comments

Fzakaria/omnibin: Every binary nixpkgs ever shipped, on your PATH

https://github.com/fzakaria/omnibin
1•Tomte•9m ago•0 comments

Finding Bugs

https://matklad.github.io/2026/09/19/finding-bugs.html
1•jamilbk•12m ago•0 comments

What's New in Oracle Solaris 11.4 SRU 95

https://blogs.oracle.com/solaris/whats-new-in-oracle-solaris-11-4-sru-95
2•zdkaster•15m ago•0 comments

Pompeii test case demonstrates Gaussian Splatting's potential

https://www.gim-international.com/news/pompeii-test-case-demonstrates-gaussian-splatting-s-potent...
1•bryanrasmussen•15m ago•0 comments

Every Package Is Installed

https://fzakaria.com/2026/09/24/every-package-is-already-installed
1•ghuntley•19m ago•0 comments

Perch: Semantic Code Linting with Jev

https://github.com/lakeday-org/perch
1•handfuloflight•20m ago•0 comments

Is indirect prompt injection still a big threat as models get more advanced?

https://realarcherl.github.io/posts/prompt_injection/
1•ArcherL•26m ago•1 comments

AI/ML Engineer

https://docs.google.com/forms/d/e/1FAIpQLSdrsnW2CpvYDBRZnS91cL4nXgnvi1mwQSYqF6IhvGPImaLKYw/viewform
1•lets_talk•39m ago•0 comments

We Built Safety into Muse

https://research.meta.ai/blog/security-and-safety-for-ai-agents-our-approach-with-muse
1•AnhTho_FR•41m ago•0 comments

Micron Technology sucessfully demonstrates first 512GB DDR5 RDIMM module

https://investors.micron.com/news/press-release/2026/Micron-Advances-Memory-Innovation-With-the-W...
1•peter_d_sherman•44m ago•0 comments

September's ozone hole biggest in 20 years, warning signs appeared months ago

https://www.rnz.co.nz/news/science-and-technology/1572285/september-s-ozone-hole-biggest-in-20-ye...
2•colinprince•44m ago•0 comments

Collaborative Editing in Wordgard

https://marijnhaverbeke.nl/blog/collaborative-editing-wordgard.html
3•fagnerbrack•48m ago•1 comments

Plan Mode Is Dead

https://www.aymannadeem.com/artificial/intelligence,/developer/tools/2026/09/24/plan-mode-is-dead...
1•jmvldz•50m ago•0 comments

Power your agents: Gemini 3.8 Live with Live Avatar is now generally available

https://cloud.google.com/blog/products/ai-machine-learning/gemini-3-8-live-with-live-avatar-is-no...
1•planteur•54m ago•0 comments

What I Found Out About Organ Donation Shocked Me

https://www.nytimes.com/2026/09/18/opinion/organ-donation-system-ethics-trust.html
2•Alien1Being•54m ago•3 comments

Scalable Subagents

https://github.com/ringlochid/oh-my-subagents
1•ringlochid•55m ago•0 comments

We Built a City of 10k AI Researchers

https://lab.cloud/society/
1•simonpure•1h ago•0 comments

Congestion pricing funds $5M program for overnight truck deliveries in NYC

https://gothamist.com/news/congestion-pricing-funds-5-million-program-for-overnight-truck-deliver...
1•geox•1h ago•0 comments

Show HN: I built 11 Sudoku variants – killer cage generation was the hard part

https://iqiqgame.com/tags/sudoku
1•fang3651•1h ago•1 comments

Side-Channel Leakage from the File-Notification System

https://inoti.fyi/
1•xiaoyu2006•1h ago•1 comments

Learning to Discover Interesting Mathematics

https://arxiv.org/abs/2609.28603
1•E-Reverance•1h ago•0 comments

Research into file-notification attacks on Linux

https://lwn.net/Articles/1096431/
1•xiaoyu2006•1h ago•0 comments

FDA panel endorses first-of-a-kind cancer blood test from Grail

https://apnews.com/article/grail-test-cancer-fda-dna-8fc5480be365b59c450043daa38cad30
1•gmays•1h ago•0 comments

Government considers £11 monthly UK broadband fee to fund BBC TV

https://www.ispreview.co.uk/index.php/2026/09/government-considers-11-monthly-uk-broadband-fee-to...
1•ksec•1h ago•1 comments

Show HN: Today I deployed this Durable Chat application and was impressed

https://durable-chat-template.netly.workers.dev/PghOYV3NHew0fYlhGAUZB
1•DinakarS•1h ago•0 comments