frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

The Unreasonable Effectiveness of Reasonless Intermediate Tokens

https://arxiv.org/abs/2505.13775
4•YeGoblynQueenne•1y ago

Comments

tocs3•1y ago
I asked ChatGPT to restate this in more laymen's terms (posted below) and I am not to surprised at the answer.

"Lately, some AI models have shown impressive abilities to solve complex problems, and many people credit this to a method called Chain of Thought (CoT), where the model is trained to think through steps like a human might. In this paper, we take a closer look at that idea to see if it's really what's driving better performance.

We focus on the model’s step-by-step thinking (the words it generates along the way) — often treated like human "thoughts" — and examine whether these actually help the model solve problems more accurately. To test this, we train AI models using clean, correct step-by-step reasoning paths and final answers, all based on a known solving method (A* search). This lets us check both the final answers and the reasoning steps to see how they relate.

Interestingly, we find that even when a model gives the right answer, its reasoning steps can still be wrong or messy. To go further, we even train models using completely random and incorrect reasoning steps — and surprisingly, they still perform about the same, and sometimes even better, than those trained on correct steps.

This suggests that the step-by-step "thoughts" the model shows aren’t as meaningful or reliable as many assume. In short, just because a model looks like it’s reasoning through a problem doesn’t mean it actually is — and we should be careful not to treat its outputs as if it thinks like a human or follows strict logic."

Meta debuts first AI coding agent to take on Anthropic and OpenAI

https://www.cnbc.com/2026/08/05/meta-debuts-muse-code-to-take-on-anthropic-and-openai-.html
5•astlouis44•5m ago•0 comments

Price changes in consumer goods and services in the United States

https://ourworldindata.org/grapher/price-changes-consumer-goods-services-united-states
1•skadamat•5m ago•0 comments

Meta Muse Spark 1.2

https://developer.meta.com/ai/models/muse-spark/
4•wojciem•5m ago•0 comments

Unified Representation for Continuous-Latent Diffusion Language Modeling

https://arxiv.org/abs/2608.02602
1•E-Reverance•5m ago•1 comments

How to Survive in a Louisiana Swamp

https://unherd.com/2026/08/how-to-survive-on-a-louisiana-swamp/
1•bookofjoe•6m ago•0 comments

Ban the Throbber

https://banthethrobber.neocities.org/
4•kyledrake•7m ago•1 comments

TikTok Enters the Multichannel Fulfillment Race with FBT MCF

https://www.geekseller.com/blog/tiktok-enters-the-multichannel-fulfillment-race-with-fbt-mcf/
1•kull•8m ago•0 comments

PEP 842 – Module Exports

https://peps.python.org/pep-0842/
1•Ravencentric•9m ago•0 comments

Microsoft makes OpenAI GPT-5.6 Sol default in GitHub Copilot for staff

https://www.cnbc.com/2026/08/05/microsoft-makes-openai-gpt-5point6-sol-default-in-github-copilot-...
1•kjhughes•9m ago•0 comments

My Son's Internet

https://www.gordonmclean.co.uk/2026/08/04/my-sons-internet-2/
1•speckx•10m ago•0 comments

Tesla, Inc. vs. Angstrom Automotive Group, LLC

https://www.courtlistener.com/docket/73664213/1/tesla-inc-v-angstrom-automotive-group-llc/
2•hnburnsy•11m ago•1 comments

Microsoft's BitNet 2B on a 4 GB Raspberry Pi 5, in one 1.2 GB file

https://github.com/geisten/geistlib
1•geisten•11m ago•0 comments

SparklingTree: 30-40% faster specdec than DSpark by combining DDTree and DSpark

https://jwlabs.vercel.app/post/sparklingtree
1•shreybirmiwal•12m ago•0 comments

People prefer stories written by AI when told they're written by a human

https://techxplore.com/news/2026-08-people-stories-written-ai-told.html
1•wjSgoWPm5bWAhXB•12m ago•0 comments

Show HN: ShiftGrid – open-source transparent prompt-engine for pentests

https://github.com/BuFuuu/shiftgrid/blob/main/demo.gif
1•bufuu•13m ago•0 comments

Unpacking ChatGPT Work: The Agent for a Billion Users

https://www.latent.space/p/unpacking-chatgpt-work
1•haritha1313•13m ago•0 comments

Show HN: My receipt printer prints an original artwork every morning

https://github.com/matt-w-horn/morningprint
1•spectraldrift•13m ago•0 comments

Muse Code and Muse Spark 1.2

https://research.meta.ai/blog/introducing-muse-code-and-muse-spark-1-2
5•paulkrush•14m ago•1 comments

Releasing Muse Code in beta today, and Muse Spark 1.2

https://twitter.com/finkd/status/2085080750034940201
3•pmxi•17m ago•0 comments

DeepSeek V4-Flash-0731 is 12 pts more censored than preview (selectively)

https://www.ctgt.ai/research/v4-flash-0731-drift
3•cgorlla•17m ago•0 comments

Why AI agents aren't adopted widely

https://invertedpassion.com/why-ai-agents-arent-adopted-widely/
2•twapi•18m ago•0 comments

A new kind of rental scam strikes SF: He lost $18,000 to a duped Zillow listing

https://sfstandard.com/2026/08/05/san-francisco-rental-listing-scam/
2•randycupertino•20m ago•1 comments

How to Cross-Compile Rust for Windows from Linux

https://www.rustfaq.org/en/how-to-cross-compile-rust-for-windows-from-linux/
2•auraham•21m ago•0 comments

Show HN: DrakeAI – expense tracker you log by voice or text, no bank sync

https://drakeai.app/
2•a_protsyuk•22m ago•0 comments

Linux's Staging Area to Now Reject LLM Patches, Except for Real Security Fixes

https://www.phoronix.com/news/Linux-Staging-Reject-LLMs
2•Bender•25m ago•1 comments

FSF job opportunity: Engineering and Certification Manager

https://www.fsf.org/news/2026-job-opportunity-fsf-engineering-and-certification-manager
2•infognu•25m ago•0 comments

Show HN: Greenlight – Preflight Scanner for App Store and Google Play

https://github.com/RevylAI/greenlight
2•ethanzhoucool•25m ago•0 comments

Nvidia Becomes a Premier Sponsor of LVFS / Fwupd

https://www.phoronix.com/news/NVIDIA-Premier-Sponsor-LVFS
2•Bender•26m ago•0 comments

"AI" will never become conscious

https://mattbee.mataroa.blog/p/no-ai-will-never-become-conscious/
5•speckx•27m ago•0 comments

TikTok A/B testing withheld safety feature from ~10% of US users, lawsuit claims

https://www.bloomberg.com/news/features/2026-08-04/confidential-tiktok-report-shows-algorithm-saf...
2•anigbrowl•27m ago•0 comments