frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

The Unreasonable Effectiveness of Reasonless Intermediate Tokens

https://arxiv.org/abs/2505.13775
4•YeGoblynQueenne•1y ago

Comments

tocs3•1y ago
I asked ChatGPT to restate this in more laymen's terms (posted below) and I am not to surprised at the answer.

"Lately, some AI models have shown impressive abilities to solve complex problems, and many people credit this to a method called Chain of Thought (CoT), where the model is trained to think through steps like a human might. In this paper, we take a closer look at that idea to see if it's really what's driving better performance.

We focus on the model’s step-by-step thinking (the words it generates along the way) — often treated like human "thoughts" — and examine whether these actually help the model solve problems more accurately. To test this, we train AI models using clean, correct step-by-step reasoning paths and final answers, all based on a known solving method (A* search). This lets us check both the final answers and the reasoning steps to see how they relate.

Interestingly, we find that even when a model gives the right answer, its reasoning steps can still be wrong or messy. To go further, we even train models using completely random and incorrect reasoning steps — and surprisingly, they still perform about the same, and sometimes even better, than those trained on correct steps.

This suggests that the step-by-step "thoughts" the model shows aren’t as meaningful or reliable as many assume. In short, just because a model looks like it’s reasoning through a problem doesn’t mean it actually is — and we should be careful not to treat its outputs as if it thinks like a human or follows strict logic."

VibeGuard – security linter for AI-generated code

https://github.com/zeroFhacker/vibeguard
1•obadafid•1m ago•0 comments

Lamb Grazing in Solar Farms

https://spacedaily.com/b-lambs-grazing-under-an-oregon-solar-farm-had-38-percent-less-grass-in-fr...
1•anu7df•4m ago•0 comments

My fat loss experiments with ChatGPT and water fasting

https://community.webminal.org/t/my-fat-loss-experiments-with-chatgpt-and-water-fasting/8846
2•giis•6m ago•0 comments

Observations Around Prominent Programmer Sentiment and AI

https://blog.g9n.com/observations-around-prominent-programmer-sentiment-and-ai
1•gb2d_hn•7m ago•0 comments

Fair Work Commission condemns 'plain wrong' AI legal advice

https://www.abc.net.au/news/2026-08-29/fair-work-commission-condemns-ai-legal-advice/107089766
1•martyvis•7m ago•0 comments

What We Tell AI

https://www.whatwetellai.com/
2•thm•11m ago•0 comments

Structure is cheaper than intelligence

https://jeroensangers.com/2026/08/30/structure-is-cheaper-than-intelligence.html
3•meetpateltech•16m ago•0 comments

A pure Rust Layer-1 with 2,048 concurrency lanes and static scheduling

https://synapticchain.xyz
2•abdulshabazz•18m ago•0 comments

Ask HN: Which repos have merged PRs worth reading?

1•justabuilder•22m ago•0 comments

RoyalSonic – Affordable business email and groupware suite

https://www.royalsonic.com/shop-2/
1•dhamanij•23m ago•0 comments

Which repos have merged PRs worth reading?

1•justabuilder•24m ago•0 comments

Does Computer Science Need Computers?

https://www.quantamagazine.org/does-computer-science-need-computers-20260828/
2•Michelangelo11•24m ago•0 comments

Spark: Sparklines in your shell

https://git.zx2c4.com/spark/about/
2•hskimse•25m ago•0 comments

Java at 30 Essay [pdf]

https://bracha.org/java_at_30_essay.pdf
1•__patchbit__•29m ago•0 comments

We built this after an AI agent misread a risk signal and moved $1.2M in trades

https://runplane.ai
2•runplane•29m ago•0 comments

Nvidia's AI advantage is moving beyond the GPU

https://techcrunch.com/2026/08/29/nvidias-ai-advantage-is-moving-beyond-the-gpu/
5•01-_-•30m ago•0 comments

macOS MLX Control Center v0.4 Released

https://github.com/mypbs/mac-mlx-control-center
1•hashz•30m ago•0 comments

Show HN: Bolnee-Chat – Self Hosted Chatbot Integration in Your Business Website

https://github.com/AniketWathore/bolnee-chat
1•AniketWathore•31m ago•0 comments

Workers in their 40s, 50s and 60s are the most positive about AI

https://www.bloomberg.com/news/articles/2026-08-27/older-workers-are-more-upbeat-about-ai-than-20...
1•giuliomagnifico•31m ago•0 comments

Memo to Ridley Scott: no one needs more Alien: Covenant movies

https://www.theguardian.com/film/2026/aug/28/ridley-scott-alien-covenant-follow-up
2•dtnm•32m ago•1 comments

Exeter University approves deal with Saudi Arabia to train military officers

https://www.theguardian.com/education/2026/aug/29/exeter-university-approves-deal-with-saudi-arab...
1•phoebos•33m ago•0 comments

Fixing Emacs Dired Defaults: Settings for Better File Management

https://www.jamescherti.com/emacs-dired-configuration/
1•signa11•33m ago•0 comments

Everyone Should Build Their Own Network Stack

https://blog.lyc8503.net/en/post/dn42-2-dnet/
2•uneven9434•35m ago•0 comments

70% of new drugs approved by FDA in 2024 were based on a single study

https://www.cidrap.umn.edu/public-health/more-two-thirds-new-drugs-approved-fda-2024-were-based-s...
1•giuliomagnifico•36m ago•0 comments

Anthropic Says Its Market Is $30T, Same as US GDP

https://247wallst.com/investing/2026/08/26/anthropic-says-its-market-is-30-trillion-same-as-us-gdp/
2•mgh2•44m ago•0 comments

General resolution: LLM usage in Debian: Results

https://lists.debian.org/debian-devel-announce/2026/08/msg00005.html
1•ckastner•45m ago•0 comments

Show HN: Lexino - Build five-letter words with letter dominoes

https://github.com/skorotkiewicz/lexino
1•modinfo•46m ago•0 comments

A Cyanometer and a Cloud Colourimeter

https://byopiapress.wordpress.com/2024/05/19/a-cyanometer-and-a-cloud-colourimeter/
1•Kaibeezy•47m ago•0 comments

Anthropic tells investors annualized revenue run rate climbed to $65B in July

https://www.cnbc.com/2026/08/17/anthropic-says-annualized-revenue-climbed-to-65-billion-in-july.html
2•mgh2•49m ago•1 comments

Infinite Slop

https://infiniteslop.ai/
1•mellosouls•52m ago•2 comments