frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

The Unreasonable Effectiveness of Reasonless Intermediate Tokens

https://arxiv.org/abs/2505.13775
4•YeGoblynQueenne•1y ago

Comments

tocs3•1y ago
I asked ChatGPT to restate this in more laymen's terms (posted below) and I am not to surprised at the answer.

"Lately, some AI models have shown impressive abilities to solve complex problems, and many people credit this to a method called Chain of Thought (CoT), where the model is trained to think through steps like a human might. In this paper, we take a closer look at that idea to see if it's really what's driving better performance.

We focus on the model’s step-by-step thinking (the words it generates along the way) — often treated like human "thoughts" — and examine whether these actually help the model solve problems more accurately. To test this, we train AI models using clean, correct step-by-step reasoning paths and final answers, all based on a known solving method (A* search). This lets us check both the final answers and the reasoning steps to see how they relate.

Interestingly, we find that even when a model gives the right answer, its reasoning steps can still be wrong or messy. To go further, we even train models using completely random and incorrect reasoning steps — and surprisingly, they still perform about the same, and sometimes even better, than those trained on correct steps.

This suggests that the step-by-step "thoughts" the model shows aren’t as meaningful or reliable as many assume. In short, just because a model looks like it’s reasoning through a problem doesn’t mean it actually is — and we should be careful not to treat its outputs as if it thinks like a human or follows strict logic."

Wall Street Is Growing Skeptical of the Data Center Boom

https://www.nytimes.com/2026/09/21/business/ai-data-center-ipos.html
1•mikhael•1m ago•0 comments

HHS Seeks Public Input on Electromagnetic Fields and Wireless Radiation

https://www.hhs.gov/press-room/hhs-seeks-public-input-electromagnetic-fields-wireless-radiation.html
2•XnoiVeX•1m ago•0 comments

The Steam Deck Is Being Used to Test Mars Rovers [video]

https://www.youtube.com/watch?v=IJSO-BL3pEU
1•HelloUsername•1m ago•0 comments

The NASA/ESA Mars Sample Return mission has been canceled

https://www.science.org/content/article/nasa-s-mars-sample-return-mission-dead
1•Muhammad523•1m ago•0 comments

Show HN: VernLLM – LLM fallback, no gateway

https://vernllm.dev/
1•Buddo•2m ago•0 comments

B.C. government to sue OpenAI after Tumbler Ridge mass shooting

https://www.cbc.ca/news/canada/british-columbia/bc-government-announce-update-openai-legal-action...
2•dagmx•4m ago•0 comments

Enabling teachers to create learning interactives with generative UI

https://research.google/blog/the-future-of-practice-enabling-teachers-to-create-learning-interact...
1•raahelb•5m ago•0 comments

Is iPhone 18 Pro Max camera really stands out? Keynote was sharper for sure

https://twitter.com/iHarnoorSingh/status/2101129453418295507
1•iharnoor•6m ago•0 comments

V7 gives AI agents institutional memory

https://openai.com/index/v7/
1•rdslw•6m ago•0 comments

Woven Matter v0.2.0 adds tooling for deeper workspace integration

https://github.com/wovenmatter/wovenmatter/releases/tag/v0.2.0
1•trey131•6m ago•0 comments

A scrollable feed of scientific papers

https://papertok.app/feed
1•jslakro•7m ago•0 comments

Show HN: Factlabel: catches AI agents lying about the data they're reporting on

https://github.com/generallymatthew/factlabel
1•generallymatt9•7m ago•0 comments

Bugs that broke driving: Machine Learning edition

https://blog.comma.ai/ml-bugs/
1•LorenDB•8m ago•0 comments

We Are Simulating Humanity

https://www.null0.ai
1•amaarc•9m ago•0 comments

In Search of a Compositional Theory of Self-Stabilization

http://muratbuffalo.blogspot.com/2026/09/in-search-of-compositional-theory-of.html
2•matt_d•11m ago•0 comments

You can use any LLM just like JEV

https://www.reddit.com/r/LocalLLaMA/comments/1wlxpaw/you_can_use_any_llm_just_like_jev/
2•theanonymousone•12m ago•0 comments

Evergarden

https://evergarden.moe/
1•birdculture•14m ago•0 comments

Adaptive Business Engine (ABE): Evidence, Authority and Governance

https://zenodo.org/records/22881485
1•ileuza_maya•14m ago•0 comments

Temporary Flight Restriction over SpaceX's McGregor, Texas, Test Facility

https://notams.aim.faa.gov/notamSearch/createNotamPdf?transactionid=82465898
2•uticus•15m ago•0 comments

RoboHarm: Do Frontier Robot Policies Refuse Unsafe Instructions?

https://robocurve.org/roboharm/
1•msadowski•17m ago•0 comments

An Age of Experimentation [pdf]

https://thomasdullien.github.io/about/slides/An-age-of-experimentation-BlueHat-Asia-2026.pdf
1•porridgeraisin•17m ago•0 comments

Jevmade.com

https://jevmade.com/#jev
1•zenoware•17m ago•0 comments

Can JEV Play Chess

1•dev_marcospimi•18m ago•0 comments

I'm so bad at billiards that I ended up in the hyperbolic plane [video]

https://www.youtube.com/watch?v=kL9BTbIGxLg
2•accrual•22m ago•1 comments

Frontier AI on Your Own Hardware

https://timdettmers.com/2026/09/21/dlab-open-source-week/
1•pretext•22m ago•0 comments

Behind the Rise of Censorship and Distrust in Europe [video]

https://www.youtube.com/watch?v=vL4VvmFYpW0
1•PorciiVorbesc•25m ago•0 comments

JavaScript is enough to build native firmware for microcontrollers

https://geastack.com/one-pager
1•arbayi•25m ago•0 comments

US preparing 'massive' potash deal with Belarus

https://kyivindependent.com/us-preparing-massive-potash-deal-with-belarus-trump-says/
2•consumer451•25m ago•0 comments

Project Terminal – A macOS workspace for terminals and coding agents

https://www.projectterminal.app/
1•soybelli•25m ago•0 comments

Planning with Agents: Divided Worlds, Boundary Objects, and Thicker Interfaces

https://maggieappleton.com/planning-agents
1•azhenley•27m ago•0 comments