frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

The Unreasonable Effectiveness of Reasonless Intermediate Tokens

https://arxiv.org/abs/2505.13775
4•YeGoblynQueenne•1y ago

Comments

tocs3•1y ago
I asked ChatGPT to restate this in more laymen's terms (posted below) and I am not to surprised at the answer.

"Lately, some AI models have shown impressive abilities to solve complex problems, and many people credit this to a method called Chain of Thought (CoT), where the model is trained to think through steps like a human might. In this paper, we take a closer look at that idea to see if it's really what's driving better performance.

We focus on the model’s step-by-step thinking (the words it generates along the way) — often treated like human "thoughts" — and examine whether these actually help the model solve problems more accurately. To test this, we train AI models using clean, correct step-by-step reasoning paths and final answers, all based on a known solving method (A* search). This lets us check both the final answers and the reasoning steps to see how they relate.

Interestingly, we find that even when a model gives the right answer, its reasoning steps can still be wrong or messy. To go further, we even train models using completely random and incorrect reasoning steps — and surprisingly, they still perform about the same, and sometimes even better, than those trained on correct steps.

This suggests that the step-by-step "thoughts" the model shows aren’t as meaningful or reliable as many assume. In short, just because a model looks like it’s reasoning through a problem doesn’t mean it actually is — and we should be careful not to treat its outputs as if it thinks like a human or follows strict logic."

TORI 2.0 – A Math Framework for Cosmic AI Loops and Data Decay

https://zenodo.org/records/23262675
1•SamarthNarsipur•1m ago•0 comments

OpenBSD on a Raspberry Pi 4

https://www.mtsapv.com/rpi4obsd/
1•turtleyacht•2m ago•0 comments

Nectarine Demoscene Radio

https://scenestream.net/demovibes/
1•caminanteblanco•3m ago•0 comments

Microsoft and Anthropic play invoice tennis with startup's $17,600 Claude bill

https://www.theregister.com/paas-and-iaas/2026/10/09/microsoft-and-anthropic-play-invoice-tennis-...
1•mundial2k2•4m ago•0 comments

Design for Agent GUIs

https://lexifina.com/blog/agent-guis
1•theanonymousone•8m ago•0 comments

OLM to PST Converter Tool

https://www.perfectdatasolutions.com/en/olm/olm-to-pst-converter.html
1•tieanderson•8m ago•0 comments

An open-source distraction blocker for Windows

https://github.com/maakbaas/bullseye
1•maakbaas•8m ago•0 comments

Prompt Injection Detection and Defense Tools for Enterprise AI Agents

https://www.arcade.dev/blog/prompt-injection-detection-tools/
1•manveerc•9m ago•0 comments

Shape-sensing sheet digitally tracks its own movement as it bends and twists

https://news.mit.edu/2026/shape-sensing-sheet-digitally-tracks-movement-bends-twists-1008
1•bookofjoe•12m ago•0 comments

Apache Karaf – The Modulith Runtime

https://karaf.apache.org/
1•locknitpicker•15m ago•0 comments

The Thing (Listening Device)

https://en.wikipedia.org/wiki/The_Thing_(listening_device)
2•red369•16m ago•0 comments

AI borrowing slows as investors grow wary of debt binge

https://www.ft.com/content/9c13d40e-d2b2-45d5-921c-a68cd4b308f9
2•1vuio0pswjnm7•16m ago•2 comments

The AI Torture Chamber

https://jesse.id/blog/the-ai-torture-chamber
1•jesse_dot_id•17m ago•0 comments

Why Do We Have Baby Teeth and Adult Teeth?

https://www.the-scientist.com/why-do-we-have-baby-teeth-and-adult-teeth-73746
1•Anon84•17m ago•0 comments

Autism in boys, girls looks genetically similar but early, late diagnoses don't

https://www.psypost.org/autism-in-boys-and-girls-looks-genetically-similar-but-early-and-late-dia...
1•theanonymousone•22m ago•0 comments

The Nature of the Swarm

https://hallerite.com/posts/on-the-nature-of-the-swarm/
1•hallerite•25m ago•0 comments

Microsoft Has $80B of AI Chips It Can't Plug in [video]

https://www.youtube.com/watch?v=es4FfRU8saQ
1•xbmcuser•25m ago•0 comments

Turing: Opposition from the Intellectuals

1•daly•27m ago•0 comments

Ariodb – a proxy that checks and can undo what AI agents do to your DB

https://github.com/roozjalali/ariodb
2•roozbehjalali•29m ago•0 comments

Advancing mathematics research with AI-driven formal proof search

https://www.science.org/doi/10.1126/science.aej2213
1•01-_-•30m ago•0 comments

The very bad ad for the very first laptop

https://buttondown.com/suchbadtechads/archive/hx-20-rubber-ducky/
2•rfarley04•33m ago•0 comments

Cited but Not Consulted: A Counterfactual Audit of Legal Chain-of-Thought

https://arxiv.org/abs/2610.12361
1•sbulaev•33m ago•0 comments

LLMs Aren't Inevitable

https://deadsimpletech.com/blog/llms-arent-inevitable
11•rizsyed1•36m ago•2 comments

Vega: A decision model that rolls a ball down a hill

https://www.nandakishorm.com/writing/vega
1•yarapavan•38m ago•0 comments

AI model submitted false tip about unsolved murder, Philadelphia police say

https://6abc.com/post/anthropic-ai-model-submitted-false-tip-unsolved-murder-philadelphia-police-...
1•geox•39m ago•0 comments

If LLMs Can Decide Without Fine-Tuning, Do We Still Need Models Like Jev?

https://itsodeleo.github.io/posts/local-decision-workflow/
1•LeoisNotAI•42m ago•0 comments

Custom Agents Are Antipatterns

https://simianwords.bearblog.dev/custom-agents-are-antipatterns/
1•simianwords•43m ago•0 comments

Show HN: AntiPDF – PDF as it should have been all along

https://antisoftworks.com/
1•civnode•44m ago•0 comments

The Beam Book, Second Edition

https://happihacking.com/resources/the-beam-book/
1•Tomte•44m ago•0 comments

Sitecopy.ai

1•Arizonabusiness•45m ago•0 comments