frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

The Unreasonable Effectiveness of Reasonless Intermediate Tokens

https://arxiv.org/abs/2505.13775
4•YeGoblynQueenne•1y ago

Comments

tocs3•1y ago
I asked ChatGPT to restate this in more laymen's terms (posted below) and I am not to surprised at the answer.

"Lately, some AI models have shown impressive abilities to solve complex problems, and many people credit this to a method called Chain of Thought (CoT), where the model is trained to think through steps like a human might. In this paper, we take a closer look at that idea to see if it's really what's driving better performance.

We focus on the model’s step-by-step thinking (the words it generates along the way) — often treated like human "thoughts" — and examine whether these actually help the model solve problems more accurately. To test this, we train AI models using clean, correct step-by-step reasoning paths and final answers, all based on a known solving method (A* search). This lets us check both the final answers and the reasoning steps to see how they relate.

Interestingly, we find that even when a model gives the right answer, its reasoning steps can still be wrong or messy. To go further, we even train models using completely random and incorrect reasoning steps — and surprisingly, they still perform about the same, and sometimes even better, than those trained on correct steps.

This suggests that the step-by-step "thoughts" the model shows aren’t as meaningful or reliable as many assume. In short, just because a model looks like it’s reasoning through a problem doesn’t mean it actually is — and we should be careful not to treat its outputs as if it thinks like a human or follows strict logic."

Rosenbjerg/local-review: Minimal webapp for local code review of branches

https://github.com/rosenbjerg/local-review
3•bhubert•4m ago•0 comments

Shopify CEO says employees' `slop grenades` making more work for everyone else

https://fortune.com/2026/09/17/shopify-tobias-lutke-ai-slop-grenades/
2•theanonymousone•5m ago•0 comments

Replacing an agentic classification loop with Jev: 7x faster

https://blog.r6i.it/typesafe-jev-vs-agentic-loop.html
1•sammy_rulez•5m ago•0 comments

Do Artifacts Have Politics? (1980)

https://www.jstor.org/stable/20024652?seq=1
1•shinryuu•7m ago•0 comments

Thoughts on a Typesafe Coding Agent

https://docs.google.com/document/d/1G61uUB0FifUnmmrPzFQojZ3KpczYKmXGpgEXDJ2l_Zg/preview?pli=1&pru...
1•tosh•17m ago•0 comments

North Korean Hackers Posed as Recruiters.They Infected 30k Devices Worldwide

https://www.inc.com/kevin-haynes/north-korean-hackers-posed-as-recruiters-infected-30000-devices-...
2•jnord•17m ago•0 comments

Should you trust Al to deliver financial advice?

https://www.saturnos.com/report/artificial-authority
1•7777777phil•17m ago•0 comments

Publications on Regality theory and cultural selection theory

https://agner.org/cultsel/
2•andsoitis•20m ago•0 comments

Models Watching Models

https://www.southbridge.ai/blog/jev-watching-the-agents
1•tosh•20m ago•0 comments

Software Optimization Resources

https://agner.org/optimize/
1•andsoitis•21m ago•0 comments

Simple and Efficient Row-Level Security

https://acadia.engineering/blog/simple-and-efficient-row-level-security
1•mfru•21m ago•0 comments

Slower than an SD card under heavy load: iPhone 18 Pro Max's QLC storage tested

https://www.notebookcheck.net/Slower-than-an-SD-card-under-heavy-load-iPhone-18-Pro-Max-s-QLC-sto...
1•virgildotcodes•23m ago•0 comments

Infinitely Recursive of Game of Life

https://oimo.io/works/life/
1•helloplanets•24m ago•0 comments

Evaluating Long-Term Memory for AI Agents: AML Cycle 2 Is Now Open

https://twitter.com/AgentMemoryL/status/2101515447816663222
1•IreneAI•24m ago•0 comments

Perennial Technical Reads

https://parallelprogrammer.substack.com/p/a-reading-list-for-metalheads
1•signa11•25m ago•0 comments

Why Are People Cancelling?

https://bankstatementconverter.com/blog/posts/2026-09-21-why-are-people-canceling/
2•4pkjai•26m ago•0 comments

BYD car was easily hacked by cybersecurity expert

https://www.abc.net.au/news/2026-09-21/byd-hacked-by-cybersecurity-expert-vehicle-sabotage-survei...
2•Cadwhisker•29m ago•0 comments

A computer scientist-novelist reflects on AI and our understanding of maths

https://scroll.in/article/1095809/how-ai-could-influence-our-understanding-of-mathematics-a-compu...
1•adityaathalye•31m ago•1 comments

International Observe the Moon Night

https://en.wikipedia.org/wiki/International_Observe_the_Moon_Night
1•ButlerianJihad•31m ago•0 comments

iPhone 18 Pro Teardown: What Changed Inside? [video]

https://www.youtube.com/watch?v=oS_HMre7vk0
1•lisper•31m ago•0 comments

The big ideas and tiny details behind Amazon's new recyclable mailer (2019)

https://www.aboutamazon.com/news/sustainability/the-big-ideas-and-tiny-details-behind-amazons-new...
1•usaphp•32m ago•0 comments

Jev's Architecture Unmasked

https://archerhume.com/posts/jevs-architecture-unmasked/
2•nlpnerd•38m ago•0 comments

The Toxicity of Sports Media

https://medium.com/freedomofthought/the-toxicity-of-sports-media-ee6003e241ff
1•raynchad•38m ago•0 comments

Lightweight Markdown Viewer/Editor for Linux and Windows

https://marklite.app
2•mchilson•40m ago•0 comments

Kanye West: A Civil Rights Leader?

https://medium.com/freedomofthought/kanye-west-a-civil-rights-leader-48174114ef3b
1•raynchad•41m ago•0 comments

Elektron Machinedrum in the Browser

https://machinedrum-study.pages.dev/
1•risktopark•41m ago•0 comments

Less Typing, More Work

https://nicolasdular.com/blog/2026/09/20/less-typing-more-work/
4•tosh•46m ago•0 comments

Understanding the European Cyber Resilience Act (CRA)

https://openssf.org/cra-ebook/
3•tapanjk•46m ago•0 comments

Postgres 19: What's new in monitoring?

https://clickhouse.com/blog/postgres-19-monitoring-whats-new
2•saisrirampur•49m ago•0 comments

Find-jevable-code – audit a repo for Jev-replaceable decisions

https://altslate-labs.github.io/jevable-code/
2•ss_y2n•51m ago•0 comments