frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

The Unreasonable Effectiveness of Reasonless Intermediate Tokens

https://arxiv.org/abs/2505.13775
4•YeGoblynQueenne•1y ago

Comments

tocs3•1y ago
I asked ChatGPT to restate this in more laymen's terms (posted below) and I am not to surprised at the answer.

"Lately, some AI models have shown impressive abilities to solve complex problems, and many people credit this to a method called Chain of Thought (CoT), where the model is trained to think through steps like a human might. In this paper, we take a closer look at that idea to see if it's really what's driving better performance.

We focus on the model’s step-by-step thinking (the words it generates along the way) — often treated like human "thoughts" — and examine whether these actually help the model solve problems more accurately. To test this, we train AI models using clean, correct step-by-step reasoning paths and final answers, all based on a known solving method (A* search). This lets us check both the final answers and the reasoning steps to see how they relate.

Interestingly, we find that even when a model gives the right answer, its reasoning steps can still be wrong or messy. To go further, we even train models using completely random and incorrect reasoning steps — and surprisingly, they still perform about the same, and sometimes even better, than those trained on correct steps.

This suggests that the step-by-step "thoughts" the model shows aren’t as meaningful or reliable as many assume. In short, just because a model looks like it’s reasoning through a problem doesn’t mean it actually is — and we should be careful not to treat its outputs as if it thinks like a human or follows strict logic."

Google Was a Lifeline for Publishers. Now Some Are Thinking of Cutting It Off

https://www.wsj.com/business/media/google-search-publishers-ai-content-0fb06e41
1•thm•35s ago•0 comments

How I use LLMs as a staff engineer

https://www.seangoedecke.com/how-i-use-llms/
1•kp25•2m ago•0 comments

Show HN: IMAX Analysis throughout the years (1994 – 2026)

https://github.com/omarirfa/IMAX-Analysis
1•fl4res•5m ago•0 comments

Tesla Balance Bike

https://shop.tesla.com/product/balance-bike-for-kids
2•surprisetalk•6m ago•0 comments

Color-space: open color spaces collection with one API

https://color-space.io/
1•made_by_dy•9m ago•0 comments

Watching a language model think before it speaks

https://blog.nathanlangley.dev/posts/subtext.html
1•ninjahawk1•13m ago•0 comments

IGGT4D: Streaming 4D Instance-Grounded Geometry Transformer

https://iggt4d.github.io/
1•ilreb•15m ago•0 comments

A.I. Drones Are Coming. We Are Not Ready

https://www.nytimes.com/2026/07/20/opinion/ai-drones-modern-warfare.html
2•lxm•18m ago•1 comments

Apply Tracker, Free AI CV, Cover Letter and Job Tracker

https://www.apply-tracker.com/en
2•mirnumbing•19m ago•0 comments

Findborg – an ad-free find engine with community, AI, and real web results

https://www.findborg.com/
1•findborg•20m ago•0 comments

FilmWorld: Agentic Novel-to-Film Generation

https://filmworld-ai.github.io/
1•ilreb•20m ago•0 comments

If AI writes everything, why should anyone trust you?

https://www.adgully.com/post/18212/if-ai-writes-everything-why-should-anyone-trust-you
2•sbulaev•25m ago•0 comments

Nvidia increases the price of Jetson modules and devkits by up to 101%

https://www.cnx-software.com/2026/07/22/nvidia-increases-the-price-of-jetson-modules-and-devkits-...
2•mikece•25m ago•0 comments

Why Create Next Palantir?

https://nimish.ch/next-palantir
3•nimish_chauhan•26m ago•0 comments

A simple API for offering your coding agent a smoke break

https://smoke-break.pineapplefreefall.com
2•cr4shed•26m ago•0 comments

Zynthoro – AI-native ERP for SMEs, replace 15 tools, built solo

https://zynthoro.ai/
1•Zynthoro•27m ago•0 comments

Better Than Free: How to Differentiate in the Age of AI

https://tim.blog/2026/07/17/better-than-free/
1•lxm•27m ago•0 comments

Moravec's Paradox

https://en.wikipedia.org/wiki/Moravec%27s_paradox
1•m-hodges•27m ago•0 comments

Human, All Too Human: Martin Heidegger (1999) [video]

https://archive.org/details/youtube-gQ09a5LJFqo
1•petethomas•27m ago•0 comments

Coding agents cannot remove complexity

https://12gramsofcarbon.com/p/agentics-coding-agents-cannot-remove
1•theahura•31m ago•0 comments

Show HN: Roblox Gakuran Guide and Wiki

https://gakuranguide.wiki/
1•cksmct•34m ago•0 comments

Tangleflow: Converts GitHub Actions workflows to tangled workflows and back

https://github.com/43081j/tangleflow
1•wertyk•38m ago•0 comments

Aurora DSQL: Scalable, Multi-Region OLTP

https://brooker.co.za/blog/2026/07/19/dsql-paper.html
2•mariuz•39m ago•0 comments

1 month and 60 models later

https://www.supcpu.com/research/1-month-and-60-models-later/
3•romellogoodman•39m ago•0 comments

Galloping Galaxies (1985-1986)

https://archive.org/details/s-01-e-01-part-one_20260604
2•petethomas•41m ago•0 comments

A Controlled Study of Attention-Only Transformers

https://arxiv.org/abs/2607.18363
2•E-Reverance•41m ago•1 comments

The first six ideas in machine learning

https://stochastic.blog/the-first-six-ideas-in-machine-learning/
1•Anon84•41m ago•0 comments

Never Enough

https://dark.ronacher.eu/2026/7/21/never-enough/
1•reasonableklout•42m ago•0 comments

Show HN: My Attempt of Ricing macOS

https://github.com/kushvinth/dotfiles
1•Kushvinth•42m ago•0 comments

Air Force conducts live-fire test for Collaborative Combat Aircraft program

https://www.acc.af.mil/News/Article-Display/Article/4547699/air-force-conducts-live-fire-test-for...
4•lxm•45m ago•0 comments