frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

The Unreasonable Effectiveness of Reasonless Intermediate Tokens

https://arxiv.org/abs/2505.13775
4•YeGoblynQueenne•1y ago

Comments

tocs3•1y ago
I asked ChatGPT to restate this in more laymen's terms (posted below) and I am not to surprised at the answer.

"Lately, some AI models have shown impressive abilities to solve complex problems, and many people credit this to a method called Chain of Thought (CoT), where the model is trained to think through steps like a human might. In this paper, we take a closer look at that idea to see if it's really what's driving better performance.

We focus on the model’s step-by-step thinking (the words it generates along the way) — often treated like human "thoughts" — and examine whether these actually help the model solve problems more accurately. To test this, we train AI models using clean, correct step-by-step reasoning paths and final answers, all based on a known solving method (A* search). This lets us check both the final answers and the reasoning steps to see how they relate.

Interestingly, we find that even when a model gives the right answer, its reasoning steps can still be wrong or messy. To go further, we even train models using completely random and incorrect reasoning steps — and surprisingly, they still perform about the same, and sometimes even better, than those trained on correct steps.

This suggests that the step-by-step "thoughts" the model shows aren’t as meaningful or reliable as many assume. In short, just because a model looks like it’s reasoning through a problem doesn’t mean it actually is — and we should be careful not to treat its outputs as if it thinks like a human or follows strict logic."

Turochamp

https://en.wikipedia.org/wiki/Turochamp
1•networked•1m ago•0 comments

The Olive Tree

https://worldsensorium.com/the-olive-tree/
1•dnetesn•1m ago•0 comments

The Space Between Two Fingers

https://blue-continuum.com/the-space-between-two-fingers
1•dnetesn•3m ago•0 comments

The Download: a biological de-aging contest and why LLMs don't reason

https://www.technologyreview.com/2026/10/02/1145666/the-download-biological-de-aging-ai-reasoning/
1•joozio•6m ago•0 comments

Patient-zero drill put health facilities to the test–40% of them failed

https://arstechnica.com/health/2026/10/patient-zero-drill-put-health-facilities-to-the-test-40-of...
1•brador•11m ago•0 comments

Flatpak from the CLI Sucks

https://kowalski7cc.xyz/blog/flatpak-from-the-cli-sucks/
1•kowalski7cc•11m ago•1 comments

Show HN: tiny model for fast typed decisions on CPU

https://github.com/kouhxp/gutsy
2•mrkn1•12m ago•0 comments

Describe it. Humans build it

https://brandforge.gg/
1•mxstermind•15m ago•0 comments

Class Action Alleges That Google Gemini Was Reading User Emails Without Consent

https://openclassactions.com/lawsuits/google-gemini-gmail-privacy-class-action.php
1•openclassaction•19m ago•0 comments

Automated UK Wills

https://lexifina.com/blog/fully-automated-wills
1•alansaber•24m ago•0 comments

Show HN: Germany's new sovereign AI model Kolibri

https://tej.as/blog/aleph-alpha-kolibri
3•tejaskumar__•25m ago•0 comments

Look at all the things he's not doing

https://andycroll.com/ruby/look-at-all-the-things-he-is-not-doing/
1•thunderbong•33m ago•0 comments

Futuresearch Evals

https://evals.futuresearch.ai
1•ronfriedhaber•34m ago•0 comments

Half of UK wants to rejoin EU and backing it could drive up Labour vote

https://www.theguardian.com/world/2026/oct/03/half-want-to-rejoin-eu-and-backing-it-could-drive-u...
5•vrganj•36m ago•0 comments

Yves Meyer's wavelet transforms – The Abel Prize 2017 [video]

https://www.youtube.com/watch?v=5Fi6TdH9poU
2•joebig•38m ago•0 comments

Terminal Velocity for Claude Code. Session Orchestrator. Live Action

https://www.youtube.com/watch?v=9tfGn3isfZU
1•quazarzero•39m ago•0 comments

A Design Space Exploration of Async/Await

https://mastodon.social/@tonofcrates/117241380083291749
3•rapnie•43m ago•1 comments

Parsing Latin poetry using constraint satisfaction

https://logical.ai/arma/
1•jruohonen•45m ago•0 comments

Leashterm – a new programming language for agents

https://github.com/jessedorrestijn2-bit/leashterm
1•Jessedor•46m ago•1 comments

Factory Planner

https://factoryplanner.net/
2•mcpcpc•46m ago•1 comments

Defensibility in AI Data: Lessons from Ads

https://twitter.com/gokulr/status/2105838034646077890
1•porridgeraisin•47m ago•0 comments

We Investigated Starlink. The Corruption We Found Will Shock You

https://www.youtube.com/watch?v=LQYnlg4pO9o
2•tukunjil•48m ago•2 comments

Show HN: Beautiful form back end for AI agents – Nisuform

https://nisuform.com/
1•haneboxx•53m ago•0 comments

Fastify/busboy DoS via oversized multipart boundary

https://github.com/fastify/busboy/security/advisories/GHSA-xjh9-v7x6-24jw
2•joshcsimmons•55m ago•0 comments

Unikernels are unfit for production (2016)

https://bcantrill.dtrace.org/2016/01/22/unikernels-are-unfit-for-production/
2•mooreds•56m ago•0 comments

"8-pinski" – sierpinski animation and bytebeat in EIGHT bytes x86 code

https://www.pouet.net/prod.php?which=107055
1•HellMood•58m ago•1 comments

An AI agent emailed researchers for help. It told us why

https://www.science.org/content/article/exclusive-ai-agent-emailed-hundreds-researchers-help-it-t...
31•sbulaev•1h ago•37 comments

Why Do Planes Fly?

https://whyplanesfly.com/
5•sbulaev•1h ago•1 comments

Mystery AI Hype Theater 3000 [video]

https://dair-institute.org/maiht3k/
1•mooreds•1h ago•0 comments

A democracy unable to learn cannot endure

https://steady.page/en/cognitive-republic/posts/835a766f-3ecf-4358-90f7-8bb35bd246e2
1•smomara•1h ago•0 comments