frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

The Unreasonable Effectiveness of Reasonless Intermediate Tokens

https://arxiv.org/abs/2505.13775
4•YeGoblynQueenne•1y ago

Comments

tocs3•1y ago
I asked ChatGPT to restate this in more laymen's terms (posted below) and I am not to surprised at the answer.

"Lately, some AI models have shown impressive abilities to solve complex problems, and many people credit this to a method called Chain of Thought (CoT), where the model is trained to think through steps like a human might. In this paper, we take a closer look at that idea to see if it's really what's driving better performance.

We focus on the model’s step-by-step thinking (the words it generates along the way) — often treated like human "thoughts" — and examine whether these actually help the model solve problems more accurately. To test this, we train AI models using clean, correct step-by-step reasoning paths and final answers, all based on a known solving method (A* search). This lets us check both the final answers and the reasoning steps to see how they relate.

Interestingly, we find that even when a model gives the right answer, its reasoning steps can still be wrong or messy. To go further, we even train models using completely random and incorrect reasoning steps — and surprisingly, they still perform about the same, and sometimes even better, than those trained on correct steps.

This suggests that the step-by-step "thoughts" the model shows aren’t as meaningful or reliable as many assume. In short, just because a model looks like it’s reasoning through a problem doesn’t mean it actually is — and we should be careful not to treat its outputs as if it thinks like a human or follows strict logic."

Ask HN: Is there any Hacker News but for investing/trading/economics etc.?

2•fittingopposite•1m ago•0 comments

State-Oriented Consistency: Why We Stopped Looking for One Right Answer

https://keel-iot.eu/blog/state-oriented-consistency.html
1•fcravio•2m ago•0 comments

Doortrackr

https://doortrackr.com
1•samtato•3m ago•0 comments

Why police raided Starbucks Korea's headquarters

https://www.cnn.com/2026/08/05/asia/south-korea-raid-starbucks-tank-ad-intl-hnk
2•firefax•4m ago•0 comments

Onida – Leading Home Appliance Brand in India

https://onida.com/
1•gurjeet•5m ago•0 comments

The Wind at Your Back

https://yusufaytas.com/the-wind-at-your-back
3•yusufaytas•6m ago•0 comments

Unconscious Inference

https://en.wikipedia.org/wiki/Unconscious_inference
2•binyu•7m ago•0 comments

Ask GitHub SRE: How serious is the situation there?

1•laxk•8m ago•0 comments

Depot CI Compatibility with GitHub Actions

https://depot.dev/docs/ci/compatibility
2•da768•9m ago•0 comments

Google is talking about Gemini 4 while 3.5 Pro remains delayed

https://www.reuters.com/business/google-shakes-up-ai-leadership-deepmind-chief-shifts-role-2026-0...
2•BlueBerry2001•10m ago•0 comments

Agent Plugins Specification

https://github.com/agentplugins/agent-plugins-spec
3•Kydlaw•10m ago•1 comments

The Printer

https://alnwlsn.com/z/printer/index.html
1•Cider9986•10m ago•0 comments

LettuceDetect v2 in Semantic Router: Gen. Hallucination Detection vLLM Endpoint

https://vllm-sr.ai/blog/lettucedetect-v2-generative-hallucination-detection/
1•matt_d•10m ago•0 comments

Reducing AppSec false positives with paraconsistent logic instead of thresholds

https://github.com/castrogusttavo/lpa2v-appsec
1•castrogusttavo•11m ago•0 comments

AI can now design functional viruses – not the computer kind, either

https://www.theregister.com/offbeat/2025/09/18/ai-can-now-design-more-deadly-virus-genomes/1489318
2•joebuckwilliams•11m ago•2 comments

Continual Learning via Real-Time RL for Agents

https://rllm-project.com/post.html?post=realtime_rl.md
1•matt_d•12m ago•0 comments

Cynicism and Naïveté in the Summer of a Thousand Fires

https://www.meditationsinanemergency.com/cynicism-and-naivete-in-the-summer-of-a-thousand-fires/
1•cdrnsf•12m ago•0 comments

The DISTINCT in Your COUNT

https://boringsql.com/posts/distinct-in-your-count/
3•gmcabrita•15m ago•0 comments

Large genome models used to design new viruses

https://arstechnica.com/science/2026/08/large-genome-models-used-to-design-new-viruses/
1•pera•16m ago•0 comments

xAI Ignores Laws and Profits: Rules for Thee, Not for Me

https://illegal.solutions/posts/xai_pollution
29•speckx•17m ago•2 comments

Show HN: After 23 years, I finished my indie video game

https://www.missilemassacre.com
3•jevonsparadox•20m ago•1 comments

The Judgment Decorator

https://www.insideainative.com/p/the-judgment-decorator
1•mmayernick•20m ago•0 comments

Landscape and Perspective on Recursive Self-Improvement from a Neolab

https://poetiq.ai/posts/rsi_perspective/
1•gkapur•20m ago•0 comments

Emacs project adds official AGENTS.md

https://lists.gnu.org/archive/html/emacs-devel/2026-08/msg00077.html
4•untilted•20m ago•0 comments

We're Suing the Makers of I-Ready

https://www.educationprogress.org/p/were-suing-the-makers-of-i-ready
1•barry-cotter•22m ago•0 comments

Why ARM Disassembly Looks "Too Simple"

https://comuniq.xyz/post?t=1512
4•01-_-•22m ago•0 comments

AI designs a novel E. coli killer (first publicly announced AI-designed virus)

https://news.stanford.edu/stories/2026/08/evo-2-ai-tool-e-coli-killer-bacteriophages
5•alephnerd•22m ago•1 comments

Anhedonia

https://en.wikipedia.org/wiki/Anhedonia
2•klaussilveira•23m ago•0 comments

EV8: The Post-Ultimate Alpha [pdf]

https://research.ac.upc.es/pact01/keynotes/emer.pdf
1•twoodfin•26m ago•0 comments

Anybody tried Laguna S 2.1 (by Poolside)?

4•spottedmarley•28m ago•0 comments