frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

The Unreasonable Effectiveness of Reasonless Intermediate Tokens

https://arxiv.org/abs/2505.13775
4•YeGoblynQueenne•1y ago

Comments

tocs3•1y ago
I asked ChatGPT to restate this in more laymen's terms (posted below) and I am not to surprised at the answer.

"Lately, some AI models have shown impressive abilities to solve complex problems, and many people credit this to a method called Chain of Thought (CoT), where the model is trained to think through steps like a human might. In this paper, we take a closer look at that idea to see if it's really what's driving better performance.

We focus on the model’s step-by-step thinking (the words it generates along the way) — often treated like human "thoughts" — and examine whether these actually help the model solve problems more accurately. To test this, we train AI models using clean, correct step-by-step reasoning paths and final answers, all based on a known solving method (A* search). This lets us check both the final answers and the reasoning steps to see how they relate.

Interestingly, we find that even when a model gives the right answer, its reasoning steps can still be wrong or messy. To go further, we even train models using completely random and incorrect reasoning steps — and surprisingly, they still perform about the same, and sometimes even better, than those trained on correct steps.

This suggests that the step-by-step "thoughts" the model shows aren’t as meaningful or reliable as many assume. In short, just because a model looks like it’s reasoning through a problem doesn’t mean it actually is — and we should be careful not to treat its outputs as if it thinks like a human or follows strict logic."

Jupyter Ipywidget for Editable Tables

https://pypi.org/project/ipyrowtable/
1•jeffrey_t_b•3m ago•1 comments

Newcomb's Problem and Two Principles of Choice (1969) [pdf]

https://danielhoek.com/wp-content/uploads/2020/02/Nozick-Newcombs-Problem-and-Two-Principles-of-C...
1•doener•10m ago•0 comments

Predictable Swarm Scaling

https://wenhaochai.com/blogs/predictable-swarm-scaling.html
1•arkwor•11m ago•0 comments

Show HN: Fusor – Like Vue, but the logic is Rust

https://fusor.build
1•andreespirela•11m ago•0 comments

Newcomb's Problem and Two Principles of Choice

https://link.springer.com/chapter/10.1007/978-94-017-1466-2_7
1•doener•12m ago•0 comments

Functional Decision Theory: A New Theory of Instrumental Rationality (2018)

https://arxiv.org/abs/1710.05060
1•doener•13m ago•0 comments

Ask HN: What are people doing with their OpenClaw set ups now?

2•will10pa•14m ago•0 comments

Rene-1: Open-weight classifier sets SOTA on Decision Index (+9 over Jev)

https://huggingface.co/salfatigroup/rene-1-31b-fp8
2•elonsalfati•14m ago•0 comments

What archaeology reveals about the rise of all-powerful rulers

https://knowablemagazine.org/content/article/society/2026/what-archaeology-reveals-about-rise-of-...
2•marojejian•19m ago•1 comments

Cloudflare's 2026 Annual Founders' Letter

https://blog.cloudflare.com/cloudflares-2026-annual-founders-letter/
2•gpi•25m ago•0 comments

Parametrised tests in Rust with named parameters

https://alexwlchan.net/2026/table-driven-rust/
1•hellerve•29m ago•0 comments

Bookmarks, done right: Give your bookmarks a shape worth keeping

https://chromewebstore.google.com/detail/kosh-bookmark-manager/meldjdfnfgefeoegccphimkmgmeamcce
2•havu12•41m ago•0 comments

Being a Doctor Will Never Be the Same After A.I

https://www.nytimes.com/2026/09/25/opinion/ai-doctor-medical-students.html
1•bookofjoe•45m ago•2 comments

OpenAI to Halt Training of Some Models

https://gizmodo.com/openai-to-halt-training-of-some-models-2000817912
3•HiroProtagonist•47m ago•0 comments

US power companies scramble to secure equipment

https://www.reuters.com/business/energy/us-power-companies-scramble-secure-equipment-surging-data...
1•JumpCrisscross•49m ago•0 comments

A Quick Guide on Creating a Design System

https://www.andrewcoyle.com/blog/a-quick-guide-on-creating-a-design-system
1•SenHeng•50m ago•0 comments

Free Space Bunny Model playground, no signup

https://spacebunnymodel.com/
1•jeyzolo•51m ago•1 comments

Locked out: Why young Europeans can't afford to buy homes

https://www.euronews.com/2026/09/23/locked-out-why-young-europeans-cant-afford-to-buy-homes
2•JumpCrisscross•52m ago•0 comments

Washington Challenges Brussels' Authority over X

https://reclaimthenet.org/washington-challenges-brussels-authority-over-x
1•miohtama•53m ago•0 comments

Five Million Robots Now Operate in Factories Globally

https://ifr.org/ifr-press-releases/news/five-million-robots-now-operate-in-factories-globally
1•JumpCrisscross•55m ago•0 comments

Advice to a Beginning Software Engineer

https://www.seangoedecke.com/advice-to-a-beginning-software-engineer/
1•donsquibio•55m ago•0 comments

SpaceX Pivots Away from Space

https://www.ft.com/content/0fd15797-ac89-4222-a4bb-03a6a1330f48
5•kristianp•55m ago•3 comments

Diagnosing and Mitigating Tool-Call Repetition in MiMo-v2.6

https://mimo.xiaomi.com/blog/mimo-v2-6-tool-call-repetition
1•bashtoni•58m ago•1 comments

3D necroprinting: Leveraging biotic material as the nozzle for 3D printing

https://www.science.org/doi/10.1126/sciadv.adw9953
1•ColinWright•58m ago•0 comments

How many strings can you create per second?

https://lemire.me/blog/2026/09/25/how-many-strings-can-you-create-per-second/
2•kristianp•1h ago•0 comments

VevDB

https://vevdb.com/
1•DASD•1h ago•1 comments

Show HN: Vanish moves heavy computation off your laptop

https://vanishcompute.com/
3•fcesco•1h ago•2 comments

Machine Learning Systems

https://mlsysbook.ai/
1•7777777phil•1h ago•0 comments

Eyes on Asteroids (NASA)

https://eyes.nasa.gov/apps/asteroids/
1•vetchzero•1h ago•0 comments

My Ansible Plugin Had a Jail Escape: CVE-2026-55074

https://blog.hofstede.it/my-ansible-plugin-had-a-jail-escape-cve-2026-55074/
1•zdw•1h ago•0 comments