frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

The Unreasonable Effectiveness of Reasonless Intermediate Tokens

https://arxiv.org/abs/2505.13775
4•YeGoblynQueenne•1y ago

Comments

tocs3•1y ago
I asked ChatGPT to restate this in more laymen's terms (posted below) and I am not to surprised at the answer.

"Lately, some AI models have shown impressive abilities to solve complex problems, and many people credit this to a method called Chain of Thought (CoT), where the model is trained to think through steps like a human might. In this paper, we take a closer look at that idea to see if it's really what's driving better performance.

We focus on the model’s step-by-step thinking (the words it generates along the way) — often treated like human "thoughts" — and examine whether these actually help the model solve problems more accurately. To test this, we train AI models using clean, correct step-by-step reasoning paths and final answers, all based on a known solving method (A* search). This lets us check both the final answers and the reasoning steps to see how they relate.

Interestingly, we find that even when a model gives the right answer, its reasoning steps can still be wrong or messy. To go further, we even train models using completely random and incorrect reasoning steps — and surprisingly, they still perform about the same, and sometimes even better, than those trained on correct steps.

This suggests that the step-by-step "thoughts" the model shows aren’t as meaningful or reliable as many assume. In short, just because a model looks like it’s reasoning through a problem doesn’t mean it actually is — and we should be careful not to treat its outputs as if it thinks like a human or follows strict logic."

The Eternal Complement

https://openai.com/index/the-eternal-complement/
1•gmays•46s ago•0 comments

Local AI is real now. And its blowing my mind

https://hive.technology/lab-notes/local-ai-is-real/
1•tylermhall•2m ago•0 comments

Show HN: Git extension for controlling Git worktree mess

https://getzit.org/
1•ironfootnz•3m ago•0 comments

Claude Isn't Allowed to Write Me Prose

https://blog.kvit.app/posts/agent-not-allowed-to-write-prose/
2•skolos•19m ago•0 comments

The art of defusing a second world war bomb

https://www.theguardian.com/news/ng-interactive/2026/oct/06/it-could-knock-a-whole-street-down-th...
2•sandebert•21m ago•0 comments

Ichneumon eumerus

https://en.wikipedia.org/wiki/Ichneumon_eumerus
2•GaryBluto•24m ago•0 comments

Show HN: A browser-only checker for text hiding under PDF black boxes

https://camlabs.ai/redaction-checker/
1•camlabsai•25m ago•0 comments

Revisiting Soundness for Occurrence Typing, Semantically

https://arxiv.org/abs/2609.16299
1•azhenley•25m ago•0 comments

Show HN: EasyToSTL – Convert 2D images and text to watertight, print-ready STLs

https://easytostl.com/
1•dawn3727•28m ago•0 comments

Show HN: ImgKit – 62 image tools that run in the browser

https://imgkit.xyz
2•jaydippatel83•31m ago•0 comments

Jev isn't a better judge than Claude

https://deeplearningguy.github.io/blog/jev-vs-claude-judge/
2•deepbuilder•32m ago•0 comments

Show HN: NanoMuse – An open-source AI agent for your phone and computer

https://github.com/nano-muse/nanoMuse
1•ilreb•34m ago•0 comments

Show HN: Vim-like, hyper-efficient, LLM power tool – 20-80k tokens not 350k+

https://easiest.ai/
1•skhameneh•35m ago•0 comments

Multiplayer AI coding workspace for teams and agents

https://www.tryfridaywork.com
1•arunjdass•36m ago•0 comments

Biology, Buddhism and AI: the Platonic space model and the rogue agent swarm

https://okossa.com/biology-buddhism-and-ai-ff213a852eb4
1•okwe•38m ago•0 comments

An Open Challenge to GitHub Users: Test Das Architecture with Meta Muse Sentinel

https://zenodo.org/records/23201404
1•sangamdas•40m ago•0 comments

Can VLMs recognize famous videos from just their colors?

https://loganbolton.github.io/blog/videocolorbench/
1•septisum•41m ago•0 comments

OpenWAM: An Open Framework for Composable World-Action Models

https://openwam.stanford.edu/
2•ilreb•47m ago•0 comments

La Cueva BBS in Mexico in 1993 (session replay)

https://nanochess.org/la_cueva_bbs.html
9•nanochess•48m ago•2 comments

How do you handle client feedback and revisions on website projects?

1•tweakpage•52m ago•0 comments

Agentty is still the best way to manage agentic sessions

https://github.com/agentty-xyz/agentty
2•minev-dev•53m ago•0 comments

A missing filter looks exactly like a correct query

https://read.mir0n.pro/missing-filter/
1•mir0n•55m ago•0 comments

Building Rome from a Single Image

https://build-rome.github.io/
2•ilreb•59m ago•0 comments

Olmo-core 3: Open, scalable training infrastructure for large MoEs

https://allenai.org/blog/olmocore3
2•gmays•1h ago•0 comments

WSJ: OAI Just Cracked Hundreds More Math Problems

https://www.wsj.com/tech/ai/openai-ai-math-problems-millennium-prize-23d14511
2•talon8635•1h ago•2 comments

Ask HN: Why does steam mobile app guard itself

3•petermcneeley•1h ago•4 comments

AI Usage Typology – Fisherman

1•wwolfson97•1h ago•0 comments

Local subagent orchestration for Codex and Claude

https://github.com/ringlochid/oh-my-subagents
1•ringlochid•1h ago•0 comments

Pretrained Classifiers. CPU Only

https://jeffyclassify.com/
1•anonzzzies•1h ago•0 comments

2D Vehicles (Driving Physics in GTA1)

https://patkerr.co.uk/2d-vehicles/
3•forgingahead•1h ago•0 comments