frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

The Unreasonable Effectiveness of Reasonless Intermediate Tokens

https://arxiv.org/abs/2505.13775
4•YeGoblynQueenne•1y ago

Comments

tocs3•1y ago
I asked ChatGPT to restate this in more laymen's terms (posted below) and I am not to surprised at the answer.

"Lately, some AI models have shown impressive abilities to solve complex problems, and many people credit this to a method called Chain of Thought (CoT), where the model is trained to think through steps like a human might. In this paper, we take a closer look at that idea to see if it's really what's driving better performance.

We focus on the model’s step-by-step thinking (the words it generates along the way) — often treated like human "thoughts" — and examine whether these actually help the model solve problems more accurately. To test this, we train AI models using clean, correct step-by-step reasoning paths and final answers, all based on a known solving method (A* search). This lets us check both the final answers and the reasoning steps to see how they relate.

Interestingly, we find that even when a model gives the right answer, its reasoning steps can still be wrong or messy. To go further, we even train models using completely random and incorrect reasoning steps — and surprisingly, they still perform about the same, and sometimes even better, than those trained on correct steps.

This suggests that the step-by-step "thoughts" the model shows aren’t as meaningful or reliable as many assume. In short, just because a model looks like it’s reasoning through a problem doesn’t mean it actually is — and we should be careful not to treat its outputs as if it thinks like a human or follows strict logic."

Show HN: Let your agents DM and dynamically self-organize

https://github.com/49-Agents/agent-dms
1•nurgaliev•48s ago•0 comments

Observing Earth's Shadow from a Plane

https://www.hermandaniel.com/blog/20261004-earths-shadow-from-a-plane-window/
1•kekqqq•1m ago•0 comments

Vanishing Gradient Problem

https://en.wikipedia.org/wiki/Vanishing_gradient_problem
1•peter_d_sherman•3m ago•0 comments

Show HN: Muro 6.0 – Native Live Wallpapers for macOS, now bigger

https://murowallpaper.com
2•rockyxsl•4m ago•1 comments

Show HN: I'm developing a flat-file agent CMS with WASM plugins in Rust. Free

https://docs.accentcms.dev/docs/v0.26
1•akapaka•4m ago•0 comments

Agent session transcripts are precious, keep them

https://quesma.com/blog/agent-session-transcripts-are-precious/
1•hugodutka•8m ago•0 comments

Show HN: Learn to program a modular synthesizer in the browser

https://ostra.fm/learn
1•nicktikhonov•8m ago•0 comments

Show HN: Browsentic Open-source Claude for Chrome alternative for any agent CLI

https://github.com/imshaikot/browsentic
1•imshaikot•10m ago•0 comments

How RE investors are profiting off a California tax break for affordable housing

https://www.sfchronicle.com/projects/2026/california-welfare-tax-exemption-explainer/
1•littlexsparkee•11m ago•0 comments

Discovery and engineering of avian R2 retrotransposons for targeted DNA

https://www.nature.com/articles/s41587-026-03315-w
1•sbulaev•11m ago•0 comments

Show HN: Parakeetogo, a portable ASR binary faster than Photon/torch for CPU

https://github.com/RobViren/parakeetogo
1•robviren•12m ago•0 comments

F.02 Robot Decommission

https://www.figure.ai/news/f-02-decommission?shem=aimgspe,
2•helsinkiandrew•14m ago•0 comments

Demoting i686 Windows targets to std-only

https://blog.rust-lang.org/2026/10/02/demoting-i686-windows-targets-to-std-only/
1•torutofu•15m ago•0 comments

Pixel 11 doesn't yet meet the GrapheneOS security standards and may be skipped

https://discuss.grapheneos.org/d/41564-pixel-11-doesnt-yet-meet-the-grapheneos-security-standards...
2•finnlab•16m ago•0 comments

How to keep learning in the age of LLMs

https://ogzhanolguncu.com/blog/how-to-keep-learning-in-the-age-of-llms/
1•cmpit•17m ago•0 comments

Abundant Software, Uncertain Economics

https://www.akashtandon.in/technology/2026-10-05-abundant-software-uncertain-economics/
1•akashtndn•18m ago•0 comments

Show HN: Mcpward – black-box contract and security testing for MCP servers

https://github.com/TsvetanG2/mcpward
1•pwizard234•18m ago•0 comments

Deep Learning Mechanics on Mnist

https://stochastic.blog/deep-learning-mechanics-on-mnist/
1•Anon84•18m ago•0 comments

Fuck Zig (and maybe low-level programming in general?)

https://reedybear.bearblog.dev/fuck-zig-and-maybe-low-level-programming-in-general/
2•surprisetalk•18m ago•0 comments

How do you review UI pull requests?

https://blog.crossui.com/2026/10/you-approved-that-ui-pr-without-seeing-it
1•linb•19m ago•0 comments

Gitframes

https://github.com/gatewai-dev/gitframes
2•oknaslnkn•19m ago•1 comments

What Still Gets Paid If AI Changes the Interface?

https://michaelhillaert1.substack.com/p/what-still-gets-paid-if-ai-changes
1•michaelhillaert•22m ago•0 comments

Accept 'bad things' in return for benefits of AI, says Sam Altman

https://www.theguardian.com/technology/2026/oct/05/sam-altman-open-ai-chatgpt-benefits-risks
22•samizdis•22m ago•35 comments

Subql/common 5.8.3 compromised: postinstall stealer in 18k-star SubQuery repo

https://github.com/subquery/subql/issues/3047
5•forxtrot•22m ago•0 comments

Jovial Programming Language

https://en.wikipedia.org/wiki/JOVIAL
1•debo_•23m ago•0 comments

SSO for integrations, for the one who's building it

https://roniam.dev/blog/sso-for-integrations/
1•rcbart•24m ago•0 comments

OctLLM – Octrees as an Explicit 3D Language

https://plurato.github.io/OctLLM-page/
1•ilreb•25m ago•0 comments

Show HN: I built a video transcript tool that also reads what is on screen

https://scribiz.com
1•illyism•26m ago•0 comments

Measles disaster in Florida might not be far off, analysis finds

https://www.cidrap.umn.edu/public-health/measles-disaster-florida-might-not-be-far-analysis-finds
1•geox•28m ago•0 comments

The Myth of Exponential Hypergrowth

https://longform.asmartbear.com/exponential-growth/
1•sumanthvepa•29m ago•0 comments