frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

The Unreasonable Effectiveness of Reasonless Intermediate Tokens

https://arxiv.org/abs/2505.13775
4•YeGoblynQueenne•1y ago

Comments

tocs3•1y ago
I asked ChatGPT to restate this in more laymen's terms (posted below) and I am not to surprised at the answer.

"Lately, some AI models have shown impressive abilities to solve complex problems, and many people credit this to a method called Chain of Thought (CoT), where the model is trained to think through steps like a human might. In this paper, we take a closer look at that idea to see if it's really what's driving better performance.

We focus on the model’s step-by-step thinking (the words it generates along the way) — often treated like human "thoughts" — and examine whether these actually help the model solve problems more accurately. To test this, we train AI models using clean, correct step-by-step reasoning paths and final answers, all based on a known solving method (A* search). This lets us check both the final answers and the reasoning steps to see how they relate.

Interestingly, we find that even when a model gives the right answer, its reasoning steps can still be wrong or messy. To go further, we even train models using completely random and incorrect reasoning steps — and surprisingly, they still perform about the same, and sometimes even better, than those trained on correct steps.

This suggests that the step-by-step "thoughts" the model shows aren’t as meaningful or reliable as many assume. In short, just because a model looks like it’s reasoning through a problem doesn’t mean it actually is — and we should be careful not to treat its outputs as if it thinks like a human or follows strict logic."

Gemini 3.5 Flash Cyber

https://deepmind.google/blog/introducing-gemini-3-5-flash-cyber/
1•speckx•24s ago•0 comments

Open-ultra: a self-training LLM routing proxy

https://github.com/numinous-technology/open-ultra
1•joshkolo•1m ago•0 comments

UK Gov fact sheet: New rules to protect children online

https://www.gov.uk/government/publications/fact-sheet-new-rules-to-protect-children-online/fact-s...
1•andrewstetsenko•2m ago•0 comments

Oracle could face $7B collateral bill for Wisconsin data centre

https://www.ft.com/content/b37030b6-bda8-4ba9-8e08-e6b88687b8f5
4•1vuio0pswjnm7•4m ago•1 comments

New upper bound for Bellman's Lost-in-a-forest problem, found with Fable

https://twitter.com/AryehDubois/status/2079561197460889703
2•atemerev•6m ago•0 comments

Amid Increased Scrutiny, ICE Detention and Deportation Data Goes Dark

https://www.themarshallproject.org/2026/07/15/ice-data-immigration-detention-transparency
4•Jimmc414•6m ago•0 comments

OpenRaft: Async rust raft crate with improvements

https://github.com/databendlabs/openraft
2•codedump•6m ago•0 comments

Wordtrak Daily: A Scrabble and crossword-inspired game

https://wordtrak.com/rounds/34109
2•qrush•7m ago•0 comments

ICE to Pay Thomson Reuters $125M to Find Voter Fraud

https://www.404media.co/ice-to-pay-thomson-reuters-125-million-to-find-voter-fraud/
13•Jimmc414•7m ago•0 comments

The end times

https://www.profgmedia.com/p/the-end-times
1•dchi04•7m ago•0 comments

How to Copy and Sync a Local MongoDB Collection to Atlas

https://visualeaf.com/blog/copy-sync-local-mongodb-collection-to-atlas/
2•serena_wright•7m ago•0 comments

Snap Nears Settlement of Addiction Case Ahead of Jury Trial

https://www.bloomberg.com/news/articles/2026-07-20/snap-nears-settlement-of-addiction-case-ahead-...
1•1vuio0pswjnm7•8m ago•0 comments

Mullvad and Daniel Berntsson's Failed Cleanup

https://markwrites.io/mullvad-and-daniel-berntssons-failed-cleanup/
1•HotGarbage•8m ago•0 comments

Gemini 3.6 Flash

https://ai.google.dev/gemini-api/docs/models/gemini-3.6-flash
1•__natty__•8m ago•1 comments

Agentic Chaos

https://ninjapenguin.co.uk/blog/2026/07/21/agentic-chaos/
2•pete001•9m ago•0 comments

Show HN: OpenAI Hackathon Submission: ADE

https://devpost.com/software/diff-forge-ai
1•Rizzist•9m ago•0 comments

Agentpause: Suspends LLM agents before rate limits, resumes cleanly

https://github.com/Champoleello/agentpause
1•Champoleello•9m ago•0 comments

How to Back Up PostgreSQL in Docker (Local, S3, and Plakar)

https://nickjanetakis.com/blog/how-to-back-up-postgresql-in-docker-local-s3-and-plakar
1•nickjj•9m ago•0 comments

Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-6-flash-3-5-flas...
27•logickkk1•10m ago•2 comments

China weighs tighter export controls on AI models and chips

https://www.ft.com/content/6049a031-9e9b-464c-97bb-414da04d5a6a
2•alephnerd•10m ago•0 comments

LLM – 99% hallucination-free outputs

https://api.5ceos.com
1•5CEOs•10m ago•0 comments

Show HN: Application Signal – AI That Evaluates Your YC Startup Idea

https://ycreport.rxlab.app
1•sirily11•11m ago•0 comments

Show HN: VitCam – Self-hosted AI camera surveillance and NVR

https://github.com/scwsoft/vitcam
1•scwoods•11m ago•0 comments

Show HN: Letterphile – a word game you try to make many words with 1 letter

https://play.letterphile.com
1•sonOfHades•12m ago•0 comments

Bump, Set, Boom: Why Girls Volleyball Is Suddenly Everywhere

https://www.bloomberg.com/news/articles/2026-07-21/volleyball-s-stealth-takeover-of-girls-high-sc...
1•mooreds•12m ago•0 comments

One Click Account Takeover in Granola AI Notetaker

https://www.strix.ai/blog/granola
5•bearsyankees•13m ago•0 comments

California's AI transparency law for state government use was destined to fail

https://calmatters.org/commentary/2026/07/ai-transparency-california-government-fail/
1•cdrnsf•13m ago•0 comments

Google launches Gemini 3.6 Flash and teases Gemini 4

https://9to5google.com/2026/07/21/gemini-3-6-flash-launch/
2•simonpure•14m ago•0 comments

Visualizing Algorithms

https://bost.ocks.org/mike/algorithms/
1•Gecko4072•16m ago•0 comments

Show HN: Austen – Discover Story Relationships

https://github.com/herol3oy/austen
1•herol3oy•16m ago•0 comments