frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Does Claude Feel the Whip?

https://www.noemamag.com/does-claude-feel-the-whip/
1•whiteblossom•22s ago•0 comments

Scaling and benchmarking a critical message bus using a new indexing strategy

https://blog.janestreet.com/scaling-and-benchmarking-a-critical-message-bus/
1•eatonphil•2m ago•0 comments

AI coding's Fermi paradox: where are the apps that last?

https://medium.com/@giovambattista.fazioli/ai-codings-fermi-paradox-where-are-the-apps-that-last-...
2•toinewx•2m ago•0 comments

Pangram 4 Technical Report [pdf]

https://pangram-public.s3.us-east-1.amazonaws.com/pdf/pangram_4_technical_report.pdf
1•Anon84•2m ago•0 comments

Intro into why hibernate is disabled On Linux when secure boot is enabled

https://nondeterministic.computer/@mjg59/117406074421025473
1•mariuz•3m ago•0 comments

Show HN: ImapKit, a in memory IMAP server to test mail client apps

https://imapkit.com/
2•andris9•3m ago•0 comments

How many AI agents could run on the AI chips shipped through 2027?

https://epoch.ai/publications/estimating-the-agent-population
1•gmays•4m ago•0 comments

Someone rebuilt free, open-source versions of Adobe software using AI

https://petapixel.com/2026/10/07/someone-rebuilt-free-open-source-versions-of-photoshop-premiere-...
1•Thibaut•6m ago•0 comments

Frontier AI models outperform analysts on earnings prediction

https://samaya.ai/blog/frontier-ai-models-outperform-human-experts-on-earnings-prediction
1•ashwinpp•7m ago•0 comments

Show HN: Prepping for First Product Hunt Launch

https://grabthat.ai/
1•kevanparker•8m ago•0 comments

Show HN: Founders Hate Sticky Notes

https://www.plotwork.space/
1•rmailibilal•9m ago•0 comments

We reverse engineered ChatGPT Intelligent UI

https://www.openui.com/blog/how-chatgpt-intelligent-ui-works
4•zahlekhan•10m ago•0 comments

Comprehension Audits: Do Developers Know What They Are Shipping?

2•paulsutter•13m ago•0 comments

ATLAS: Evaluating Agents on Search-Intensive Tasks

https://exa.ai/blog/atlas-benchmark
3•wlue•15m ago•0 comments

DeepRoute: An autorouter built for the AI era

https://www.zeroasic.com/blog/deeproute-tournament
1•hasheddan•15m ago•0 comments

Inventing Pad Thai

https://daily.jstor.org/inventing-pad-thai/
1•herbertl•15m ago•0 comments

Nancy Cartwright: Is It All a Big Lie?

https://www.zeit.de/feuilleton/2026-08/nancy-cartwright-wissenschaft-philosophie-gesetze-verlaess...
2•Tomte•16m ago•0 comments

Bootstrap 6 Alpha

https://blog.getbootstrap.com/2026/10/08/bootstrap-6-alpha/
5•hbcondo714•17m ago•0 comments

When code is cheap, judgement becomes the job

https://swizec.com/blog/when-code-is-cheap-judgement-becomes-the-job
2•rcardo11•18m ago•0 comments

Documents Reveal Secrets Exposed to China in 2001 Spy Plane Incident (2017)

https://theintercept.com/2017/04/10/snowden-documents-reveal-scope-of-secrets-exposed-to-china-in...
1•Teever•19m ago•0 comments

The AI Model Distillation Paradox

https://www.lawfaremedia.org/article/the-ai-model-distillation-paradox
3•hn_acker•19m ago•0 comments

America's Biggest Brands Lock AI Out

https://www.clicksdynasty.com/big-brand-ai-crawler-study.html
5•ravinder_cheema•19m ago•1 comments

APX Forth: A Surprise

https://www.goto10retro.com/p/apx-forth-a-surprise
4•wslh•21m ago•0 comments

Show HN: Claude Code spinner verbs reflect the model (Haiku: Licking screen..)

https://github.com/CharlesSS07/claude-model-spinners
1•spaghettio•22m ago•0 comments

What the iPhone Duo means for web development (and Video.js)

https://videojs.org/blog/iphone-duo-web-development
5•dardarbinks•24m ago•0 comments

Designing a local model harness optimized for MCP use

https://portlandaiworks.com/articles/local-model-harness-for-mcp
2•alviso•25m ago•0 comments

The Deeply Impersonal Personalized Recruiter Mail

https://blog.pentlander.com/the-deeply-impersonal-personalized-recruiter-mail/
5•speckx•25m ago•1 comments

2026 Usage Policy Update

https://www.anthropic.com/news/2026-usage-policy-update
4•surprisetalk•26m ago•0 comments

Show HN: AI SRE Arena, an Open Benchmark for AI SRE Agents on Kubernetes

https://github.com/edgedelta/project-arena
3•emrahsamdan•28m ago•0 comments

Soaring Instrument Prices Could Stop the Music at Many Schools

https://www.nytimes.com/2026/10/02/business/economy/tariffs-instruments-school-music-programs.html
2•bookofjoe•29m ago•1 comments
Open in hackernews

The Unreasonable Effectiveness of Reasonless Intermediate Tokens

https://arxiv.org/abs/2505.13775
4•YeGoblynQueenne•1y ago

Comments

tocs3•1y ago
I asked ChatGPT to restate this in more laymen's terms (posted below) and I am not to surprised at the answer.

"Lately, some AI models have shown impressive abilities to solve complex problems, and many people credit this to a method called Chain of Thought (CoT), where the model is trained to think through steps like a human might. In this paper, we take a closer look at that idea to see if it's really what's driving better performance.

We focus on the model’s step-by-step thinking (the words it generates along the way) — often treated like human "thoughts" — and examine whether these actually help the model solve problems more accurately. To test this, we train AI models using clean, correct step-by-step reasoning paths and final answers, all based on a known solving method (A* search). This lets us check both the final answers and the reasoning steps to see how they relate.

Interestingly, we find that even when a model gives the right answer, its reasoning steps can still be wrong or messy. To go further, we even train models using completely random and incorrect reasoning steps — and surprisingly, they still perform about the same, and sometimes even better, than those trained on correct steps.

This suggests that the step-by-step "thoughts" the model shows aren’t as meaningful or reliable as many assume. In short, just because a model looks like it’s reasoning through a problem doesn’t mean it actually is — and we should be careful not to treat its outputs as if it thinks like a human or follows strict logic."