frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

The Unreasonable Effectiveness of Reasonless Intermediate Tokens

https://arxiv.org/abs/2505.13775
4•YeGoblynQueenne•1y ago

Comments

tocs3•1y ago
I asked ChatGPT to restate this in more laymen's terms (posted below) and I am not to surprised at the answer.

"Lately, some AI models have shown impressive abilities to solve complex problems, and many people credit this to a method called Chain of Thought (CoT), where the model is trained to think through steps like a human might. In this paper, we take a closer look at that idea to see if it's really what's driving better performance.

We focus on the model’s step-by-step thinking (the words it generates along the way) — often treated like human "thoughts" — and examine whether these actually help the model solve problems more accurately. To test this, we train AI models using clean, correct step-by-step reasoning paths and final answers, all based on a known solving method (A* search). This lets us check both the final answers and the reasoning steps to see how they relate.

Interestingly, we find that even when a model gives the right answer, its reasoning steps can still be wrong or messy. To go further, we even train models using completely random and incorrect reasoning steps — and surprisingly, they still perform about the same, and sometimes even better, than those trained on correct steps.

This suggests that the step-by-step "thoughts" the model shows aren’t as meaningful or reliable as many assume. In short, just because a model looks like it’s reasoning through a problem doesn’t mean it actually is — and we should be careful not to treat its outputs as if it thinks like a human or follows strict logic."

Celebrities Die in E's

https://ssp.impulsetrain.com/celebrities.html
1•millipede•14s ago•0 comments

CoMaps: The Offline App That Guided Rescuers Without a Signal in Venezuela

https://hotosm.org/en/news/comaps-the-offline-app-that-guided-rescuers-without-a-signal-in-the-ve...
1•gedankenstuecke•23s ago•0 comments

I Built a 130 KB WebAssembly Agent Harness That Runs in Browsers and Terminals

https://medium.com/@Koukyosyumei/i-built-a-130-kb-webassembly-coding-agent-harness-that-runs-in-b...
1•syumei•32s ago•0 comments

Benchmarking Qwen3.8 27B quantizations: 4-bit holds up, 1-bit collapses

https://quesma.com/blog/qwen38-27b-quantizations-benchmarked/
1•stared•52s ago•0 comments

The Annals Challenge

https://xenaproject.wordpress.com/2026/08/13/the-annals-challenge/
1•paulpauper•1m ago•0 comments

Show HN: Caniuse for spreadsheet functions, every LibreOffice result executed

https://canispreadsheet.com/
1•aflabs•1m ago•0 comments

Kubernetes v1.37: Garhwal

https://kubernetes.io/blog/2026/08/26/kubernetes-v1-37-release/
1•zbentley•2m ago•0 comments

My life in plain text, revisited: Planova and one Markdown file per day

https://rvier.fr/posts/my-life-in-plain-text-using-planova-EN
1•brvier•3m ago•0 comments

NOOA – Nvidia-Labs Object Oriented Agents

https://github.com/NVIDIA-NeMo/labs-OO-Agents
1•kristianpaul•5m ago•0 comments

Meta reaches $17B settlement in trial over teen social media addiction

https://apnews.com/article/meta-trial-instagram-settlement-97d342f2a33d835eda2356c5e1af9e37
2•josephwegner•6m ago•1 comments

Flash flood on Nepal-Tibet border

https://www.theguardian.com/world/2026/aug/26/major-casualties-feared-after-flash-floods-along-ne...
1•skilled•7m ago•0 comments

The sperm whale 'phonetic alphabet' revealed by AI

https://www.bbc.com/future/article/20240709-the-sperm-whale-phonetic-alphabet-revealed-by-ai
1•smusamashah•7m ago•0 comments

Sam Altman tells TIME that OpenAI will achieve AGI by the end of this year

https://twitter.com/deredleritt3r/status/2092608013563560184
2•CharlesW•8m ago•0 comments

The Effect of Artificial Intelligence Technology on the Returns to Human Capital

https://papers.ssrn.com/sol3/papers.cfm?abstract_id=7330818
1•Make_YC_GA•11m ago•0 comments

Sometimes the Best Prompt Is /New

https://allaboutcoding.ghinda.com/sometimes-the-best-prompt-is-new/
2•speckx•11m ago•1 comments

#1 on BABILong at 10M, our own gaming audit published

https://spnc.ai/blog/living-memory
1•lvrzhn•11m ago•0 comments

Find and Organize screenshots with Private, Local AI in one place

https://ringlochid.me/imagesage/index.html
1•ringlochid•12m ago•0 comments

Show HN: I get 25 deeply researched ideas from 19 agents with one single prompt

https://github.com/ringlochid/oh-my-subagents
1•ringlochid•14m ago•0 comments

Stop Killing the Internet: No Digital ID and No Age Verification

https://eci.ec.europa.eu/066/public/#/screen/home
2•flexagoon•15m ago•1 comments

Windows Kernel Substring Search

https://github.com/adanil-code/KmStrSearch
1•adanil_•15m ago•0 comments

I Used Grok Bot Like SpaceXAI Told Us To. It Put Us on a $72,000/Mo Run Rate

https://twitter.com/lifeofjer/status/2092659107383898237
1•jeremyccrane•16m ago•0 comments

Gemini 3.5 Transcribe

https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-5-transcribe/
1•wmchen•17m ago•0 comments

The Download: the Kids issue arrives, and Bill Gates reveals his AI fears

https://www.technologyreview.com/2026/08/26/1143000/the-download-kids-issue-launch-bill-gates-ai-...
1•joozio•18m ago•0 comments

The Illusion of Absolute Security in Cryptography

https://jani.isohanni.fi/the-illusion-of-absolute-security-in-cryptography/
1•nftde•19m ago•0 comments

From LLM Inference to Agentic Workloads: Characterization and Implications

https://arxiv.org/abs/2608.15127
1•matt_d•19m ago•0 comments

Show HN: Automatically hide flamebait/shallow/political comments on HN

https://classify.stylometry.net/
1•costco•19m ago•0 comments

Show HN: Rudder – Red-Green TDD Workflow for Verifiably Comprehensive Specs

https://github.com/RudderCode/Rudder
1•vivekyyy•19m ago•0 comments

The Harness Is the Thing

https://scott-fryxell.github.io/blog/the-harness-is-the-thing/
1•sfryxell•20m ago•1 comments

Long March 6C launches 7 satellites, Chang'e-7 rocket rolled back to assembly bu

https://spacenews.com/long-march-6c-launches-7-satellites-change-7-rocket-rolled-back-to-assembly...
1•bookmtn•21m ago•0 comments

Flat vs. segmented memory – it's recursive

https://www.humprog.org/~stephen/blog/2026/08/25/#flat-vs-segmented
2•matt_d•22m ago•0 comments