frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

The Unreasonable Effectiveness of Reasonless Intermediate Tokens

https://arxiv.org/abs/2505.13775
4•YeGoblynQueenne•1y ago

Comments

tocs3•1y ago
I asked ChatGPT to restate this in more laymen's terms (posted below) and I am not to surprised at the answer.

"Lately, some AI models have shown impressive abilities to solve complex problems, and many people credit this to a method called Chain of Thought (CoT), where the model is trained to think through steps like a human might. In this paper, we take a closer look at that idea to see if it's really what's driving better performance.

We focus on the model’s step-by-step thinking (the words it generates along the way) — often treated like human "thoughts" — and examine whether these actually help the model solve problems more accurately. To test this, we train AI models using clean, correct step-by-step reasoning paths and final answers, all based on a known solving method (A* search). This lets us check both the final answers and the reasoning steps to see how they relate.

Interestingly, we find that even when a model gives the right answer, its reasoning steps can still be wrong or messy. To go further, we even train models using completely random and incorrect reasoning steps — and surprisingly, they still perform about the same, and sometimes even better, than those trained on correct steps.

This suggests that the step-by-step "thoughts" the model shows aren’t as meaningful or reliable as many assume. In short, just because a model looks like it’s reasoning through a problem doesn’t mean it actually is — and we should be careful not to treat its outputs as if it thinks like a human or follows strict logic."

Show HN: Termic – desktop app for running CLI coding agents (Claude Code, Codex)

https://termic.dev/
1•SimionBaws•24s ago•0 comments

Show HN: Cfbench – A lightweight native CLI for Cloudflare speed tests

https://github.com/fcendesu/cfbench
1•fcen•1m ago•0 comments

Corner Chess

https://cs.nyu.edu/~shasha/drecco/games/2016/cornerchess/
1•tosh•2m ago•0 comments

Emacs Writing Machine

https://chainsawriot.com/postmannheim/2026/07/25/writeredeck.html
1•birdculture•3m ago•0 comments

Robots programmed by kids to speak their Indigenous languages

https://www.npr.org/2026/07/26/nx-s1-5825798/robot-speaks-endangered-native-american-languages
1•defrost•3m ago•0 comments

Apertus 1.5, Swiss open-weight, open-source, open training data model

https://apertus-ai.org/articles/2026-07-apertus-1-5/
1•flaburgan•4m ago•1 comments

Interactive 3D Dice Roller with Custom Physics and Shaders

https://ttrpg.dies.kmcprizmquest.dev/
1•Radversalo•15m ago•0 comments

How to be useful as a software architect

https://swizec.com/blog/how-to-be-useful-as-a-software-architect
2•meetpateltech•17m ago•0 comments

Elevated Errors for Opus 5

https://status.claude.com/incidents/zftg3gqkmv18
2•TimCTRL•18m ago•0 comments

I built a jav data API, need honest feedback

1•javinfo•19m ago•0 comments

Show HN: OpenStrap – use a WHOOP 4.0 band without a subscription

https://openstrap.github.io/edge/
1•abdulsaheel123•25m ago•0 comments

Show HN: Tawrida – Global Trade Tools – Export-Document Generators

https://tawrida.xyz/tools
2•m-zakarya•28m ago•0 comments

Anthropic should learn from those cotton-picking socialists

https://asteriskmag.com/issues/15/rust-and-boll
2•maxall4•30m ago•0 comments

Practice your YC interview against Garry Tan's gstack specialists, on a meeting [video]

https://www.youtube.com/watch?v=PMiIytmoEuk
4•anand_pattern•31m ago•1 comments

High Volume Transaction Processing Without Concurrency Control (1997)

https://www.researchgate.net/publication/2310700_High_Volume_Transaction_Processing_Without_Concu...
1•tosh•31m ago•0 comments

Ruff v0.16.0 – Significant new updates – 413 default rules up from 59

https://astral.sh/blog/ruff-v0.16.0
5•vismit2000•35m ago•0 comments

Build a self-scaling OCR pipeline with Qwen 3.5 and Kubernetes

https://github.com/neural-maze/production-ocr-course
2•mtrofficus•49m ago•0 comments

Claude Code has a hardcoded instruction telling Opus 5 not to use subagents

https://old.reddit.com/r/ClaudeCode/comments/1v6y5q2/claude_code_has_a_hardcoded_instruction_tell...
4•bredren•52m ago•0 comments

The most capable resume editor I could build, as a solo dev

https://www.roleframe.ai/product/ai-resume-builder
1•larbisahli•54m ago•0 comments

The diminishing returns of productivity culture (2021)

https://annehelen.substack.com/p/the-diminishing-returns-of-productivity
4•downbad_•1h ago•0 comments

Show HN: Reproducibility Benchmark a Risk Quantitative Model

https://github.com/Fluxara-GOD/fluxara1-audit-verification-pack
3•mingshi_tz•1h ago•0 comments

What will more intelligence do for us?

https://www.noahpinion.blog/p/what-will-more-intelligence-actually
3•jinjin2•1h ago•0 comments

HomeLab #2: Immutable GitOps homelab and MikroTik hardening

https://justsomebody.dev/blog/gitops-k3s-flux
1•rafal_opilowski•1h ago•1 comments

If you feel the need to tell me your answer is AI-based, work it more

https://cephalosec.com/blog/the-hidden-cost-of-ai-is-someone-elses-time/
1•Versipelle•1h ago•0 comments

The Context Bootloader

https://adlrocha.substack.com/p/adlrocha-the-context-bootloader
1•adlrocha•1h ago•0 comments

House AI 'kill switch' bill unveiled as OpenAI hack raises alarms

https://www.politico.com/news/2026/07/23/house-ai-kill-switch-bill-unveiled-as-openai-hack-raises...
1•reasonableklout•1h ago•0 comments

Modeling Facts and Reactions with Domain Events

https://deniskyashif.com/2026/07/25/modeling-facts-and-reactions-with-domain-events/
1•deniskyashif•1h ago•1 comments

Anthropic secures its AI-native software development lifecycle

https://claude.com/blog/how-anthropic-secures-its-ai-native-software-development-lifecycle
3•_tk_•1h ago•0 comments

Visaism (Part 1): India as a Waiting Room

https://twitter.com/samyak128/status/2079837839798260154
1•barry-cotter•1h ago•0 comments

Fallback Font

https://en.wikipedia.org/wiki/Fallback_font
1•pndy•1h ago•0 comments