frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

The Unreasonable Effectiveness of Reasonless Intermediate Tokens

https://arxiv.org/abs/2505.13775
4•YeGoblynQueenne•1y ago

Comments

tocs3•1y ago
I asked ChatGPT to restate this in more laymen's terms (posted below) and I am not to surprised at the answer.

"Lately, some AI models have shown impressive abilities to solve complex problems, and many people credit this to a method called Chain of Thought (CoT), where the model is trained to think through steps like a human might. In this paper, we take a closer look at that idea to see if it's really what's driving better performance.

We focus on the model’s step-by-step thinking (the words it generates along the way) — often treated like human "thoughts" — and examine whether these actually help the model solve problems more accurately. To test this, we train AI models using clean, correct step-by-step reasoning paths and final answers, all based on a known solving method (A* search). This lets us check both the final answers and the reasoning steps to see how they relate.

Interestingly, we find that even when a model gives the right answer, its reasoning steps can still be wrong or messy. To go further, we even train models using completely random and incorrect reasoning steps — and surprisingly, they still perform about the same, and sometimes even better, than those trained on correct steps.

This suggests that the step-by-step "thoughts" the model shows aren’t as meaningful or reliable as many assume. In short, just because a model looks like it’s reasoning through a problem doesn’t mean it actually is — and we should be careful not to treat its outputs as if it thinks like a human or follows strict logic."

New Amazon Data Center Is Set to Have the Most Polluting Power Plant in the U.S.

https://www.nytimes.com/2026/08/08/climate/amazon-data-center-texas-pollution.html
1•sbulaev•3m ago•0 comments

US Military's Cyber Command Unit Grapples with Cluster of Deaths by Suicide

https://www.bloomberg.com/news/articles/2026-08-06/us-military-s-cyber-command-unit-grapples-with...
2•rbanffy•6m ago•0 comments

Concurrency vs. Throughput: why more parallelism can make databases slower

https://planetscale.com/blog/concurrency-vs-throughput-vitess-mysql
1•Shaanveer•6m ago•0 comments

AgentBlog – open-source, AI-native SEO blog

https://agentblog.dev
2•goldkey•8m ago•0 comments

Coca-Cola's personalized cans inconsistently enforcing controversial phrases

https://www.foxbusiness.com/media/coca-cola-personalized-cans-show-inconsistent-enforcement-again...
2•mikhael•9m ago•0 comments

Finance Jobs Just Hit a 4-Year Low

https://www.inc.com/georgia-fearn/finance-jobs-hit-four-year-low-banks-posted-nearly-49000-roles/...
2•01-_-•11m ago•0 comments

Europe's free satellite service just made it easier to track wildfires

https://arstechnica.com/gadgets/2026/08/europes-free-satellite-service-just-made-it-easier-to-tra...
1•01-_-•11m ago•0 comments

FlowChartCharter – A Zero-Hallucination, Fear-Driven GraphRAG Alternative

https://github.com/CharleSpectre13/flowchartcharter
1•charlespectre•11m ago•0 comments

Context Engineering Is a Data Problem

https://davidgasquez.com/context-engineering-is-a-data-problem
1•kalendos•14m ago•0 comments

Looking for Someone in SF?

https://meandi-sf-people.aris-han.chatgpt.site/
1•asd000hh•14m ago•1 comments

Massachusetts prosecutors run their offices on software from the dial-up era

https://nasser.blog/dial-up-justice/
1•Michelangelo11•16m ago•0 comments

Environmentalists Targeted Exxon Mobil. Then Hackers Targeted Them

https://www.nytimes.com/2020/06/09/nyregion/exxon-mobil-hackers-greenpeace.html
1•Djens•17m ago•0 comments

Supply-chain hygiene for Emacs: LLM review of package upgrades

https://blog.fidelramos.net/software/emacs-straight-ai-review
1•fidelramos•21m ago•0 comments

Jason Arday: a question of academic standards (by Richard Dawkins)

https://unherd.com/2026/08/jason-arday-a-question-of-academic-standards/
1•Michelangelo11•23m ago•0 comments

Countersign – one kill switch and audit log for AI agents across wallet vendors

https://countersign.network
1•screan•26m ago•0 comments

Show HN: Tyle – a Kanban board with statistical cycle time forecasting built in

https://tyle-brown.vercel.app/about
1•byronical•27m ago•0 comments

OpenAI provides more details on the Hugging Face incident

https://www.heise.de/en/news/OpenAI-provides-more-details-on-the-Hugging-Face-incident-11403391.html
2•slow_typist•37m ago•1 comments

Just built Dot. Would love feedback

https://www.use-dot.com
1•r_gaur•39m ago•1 comments

What's new in Swift: July 2026 Edition

https://swift.org/blog/whats-new-in-swift-july-2026/
2•frizlab•40m ago•0 comments

How to Notice and Avoid Scams

https://nonconfirmed.com/how-to-notice-and-avoid-scams
1•colenikol2•42m ago•0 comments

You can make any software you want now, then why don't you?

1•sarmadgulzar•43m ago•2 comments

Show HN: Turn an idea into a product you can order

https://www.luphra.com
1•stochtinkerer•48m ago•1 comments

DeepMind's WeatherNext model achieves breakthrough forecasting cyclones

https://deepmind.google/blog/weathernext-ai-model-achieves-breakthrough-in-forecasting-cyclones/
2•bhavansig•51m ago•1 comments

Thimbleweed Park 2 Dev Blog

https://blog2.thimbleweedpark.com/twp2_tech/
1•MrVandemar•56m ago•0 comments

Show HN: Draw Your Room. Share Your Plan

https://freeroomplanner.com/
1•wowinter15•1h ago•0 comments

The phone book that led us to Assad's spy chief in hiding

https://www.bbc.co.uk/news/articles/c4gyrzn8p94o
1•justworks•1h ago•0 comments

Show HN: I mapped business competition density across NYC, block by block

https://worththespot.com/
1•losas28066•1h ago•0 comments

Show HN: Lensa – Open-source MCP connectors and skills for ChatGPT

https://github.com/mitrotasios/lensa-mcp
1•amitrotasios•1h ago•0 comments

Don't Build Mindreading

https://keller.substack.com/p/dont-build-mindreading
1•celer•1h ago•0 comments

Show HN: Hosted LLM Wiki

https://getmana.md/
3•zintus•1h ago•0 comments