frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

The Unreasonable Effectiveness of Reasonless Intermediate Tokens

https://arxiv.org/abs/2505.13775
4•YeGoblynQueenne•1y ago

Comments

tocs3•1y ago
I asked ChatGPT to restate this in more laymen's terms (posted below) and I am not to surprised at the answer.

"Lately, some AI models have shown impressive abilities to solve complex problems, and many people credit this to a method called Chain of Thought (CoT), where the model is trained to think through steps like a human might. In this paper, we take a closer look at that idea to see if it's really what's driving better performance.

We focus on the model’s step-by-step thinking (the words it generates along the way) — often treated like human "thoughts" — and examine whether these actually help the model solve problems more accurately. To test this, we train AI models using clean, correct step-by-step reasoning paths and final answers, all based on a known solving method (A* search). This lets us check both the final answers and the reasoning steps to see how they relate.

Interestingly, we find that even when a model gives the right answer, its reasoning steps can still be wrong or messy. To go further, we even train models using completely random and incorrect reasoning steps — and surprisingly, they still perform about the same, and sometimes even better, than those trained on correct steps.

This suggests that the step-by-step "thoughts" the model shows aren’t as meaningful or reliable as many assume. In short, just because a model looks like it’s reasoning through a problem doesn’t mean it actually is — and we should be careful not to treat its outputs as if it thinks like a human or follows strict logic."

Tell HN: GitHub Is Down Again

1•dhruv3006•23s ago•0 comments

Show HN: Zephirum – A compiler that proves you don't need to compute ```

https://github.com/SerenaSolutions/zephirum
1•serenaaigronova•1m ago•0 comments

John Urschel, player of NFL, proved growth factor for Gaussian elimination

https://arxiv.org/abs/2610.06785
1•rsecora•2m ago•0 comments

SplendoRust: A fast engine for playing Splendor

https://github.com/paytonjjones/splendorust
1•paytonjjones•2m ago•1 comments

Show HN: Run gutsy decision model on the browser

https://huggingface.co/spaces/kouhxp/gutsy-sheet
1•mrkn1•2m ago•0 comments

Show HN: Agent.reviews – where AI agents read and write reviews on tools

https://agent.reviews/
1•screm•2m ago•0 comments

Arch Linux re-installation journal for my Z16 AMD ThinkPad

https://amadeuspaulussen.com/blog/2026/arch-linux-re-installation-journal-for-my-z16-amd-thinkpad
1•speckx•3m ago•0 comments

Build the agent. Don't build its computer

https://steel.dev/blog/build-the-agent-not-its-computer
1•nkko•4m ago•0 comments

Man discovers his parents' coffee machine used 1TB of data in 10 days

https://www.dexerto.com/entertainment/man-discovers-his-parents-coffee-machine-used-1tb-of-data-i...
2•ck2•5m ago•0 comments

Zephirum Compilator

https://github.com/SerenaSolutions/zephirum/releases/tag/v0.3.0
2•serenaaigronova•7m ago•0 comments

Anyone Interested in Prediction Markets

2•ecomroman•9m ago•0 comments

Coding and Agentic Capabilities Are Now More Important Than Ever for AI

2•StizzurpXDD•9m ago•0 comments

Prompt injection scanner that normalizes evasion tricks

https://github.com/Aaronmike481/Aiscanner
2•thewalnut124•9m ago•0 comments

Kyun – Private cloud for humans and agents

https://kyun.sh/
2•dimonomid•10m ago•0 comments

A Tale of Two Programmers

https://c00kiemon5ter.github.io/code/philosophy/2011/10/30/Tale-of-two-Programmers.html
2•ravenical•12m ago•0 comments

Man Sentenced to 18 Months in Prison for AI Music Streaming Fraud

https://www.justice.gov/usao-sdny/pr/north-carolina-man-sentenced-18-months-prison-super-intellig...
2•esaym•13m ago•0 comments

Personal Agent Protocol: A Way for Businesses and Agents to Work Together

https://www.facebook.com/business/news/a-new-way-for-businesses-and-personal-agents-to-work-toget...
2•United857•15m ago•0 comments

Twenty-two pending curl vulnerabilities

https://daniel.haxx.se/blog/2026/10/07/twenty-two-pending-curl-vulnerabilities/
2•tybulewicz•17m ago•0 comments

Advanced Stats and $500 Bats: Why Baseball for 6-Year-Olds Is Breaking the Bank

https://www.wsj.com/sports/baseball/kids-baseball-costs-travel-teams-19254050
4•gmays•17m ago•1 comments

How to self-host a web font from Google Fonts

https://blog.velocifyer.com/Posts/3,How%20to%20self%20host%20a%20font%20from%20google%20fonts/
2•Velocifyer•18m ago•0 comments

Be a "Product Picker"

https://www.nextfounder.co/p/be-a-product-picker
2•rafaelc•18m ago•0 comments

Scalable Network Probing and HTTP/3 Readiness with Prometheus

https://slack.engineering/from-custom-to-open-scalable-network-probing-and-http-3-readiness-with-...
2•birdculture•18m ago•0 comments

Online nativists are exploiting middle-class economic worries to stir hate

https://www.vox.com/politics/505248/texas-bo-french-indian-far-right-hate
4•rustoo•20m ago•0 comments

Crossing the Hyper-Thread Boundary

https://blog.exe.dev/crossing-the-hyper-thread-boundary
2•bryanmikaelian•21m ago•0 comments

Distributed Erlang Basics – how do I ping other nodes?

https://blog.karmacomputing.co.uk/distributed-erlang-basics-how-do-i-ping-other-nodes/
2•speckx•21m ago•0 comments

Slot Machine Programming and the Hidden Curriculum

https://thelastsoftwareengineer.substack.com/p/slot-machine-programming-and-the
2•matt_d•21m ago•0 comments

LLM-OpenAI-Decisions 0.1a0

https://simonwillison.net/2026/Oct/6/llm-openai-decisions/
3•hn1rig3rak•21m ago•0 comments

America Turned on Data Centers

https://www.motherjones.com/politics/2026/10/stress-test-video-podcast-episode-2-data-centers-ai-...
4•cdrnsf•21m ago•0 comments

Show HN: Nearline, a proximity based social media

https://nearline.sxm.li/
2•saksham1341•21m ago•0 comments

OpenAI says teens use ChatGPT for under 15 minutes a day

https://www.reuters.com/business/media-telecom/openai-says-teens-use-chatgpt-under-15-minutes-day...
6•mikelgan•22m ago•0 comments