frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Ask HN: Who wants to be hired? (August 2026)

141•whoishiring•1d ago•384 comments

Ask HN: Who is hiring? (August 2026)

219•whoishiring•1d ago•228 comments

Ask HN: What is a good format for a tool to report data to a LLM?

4•michaelmure•2h ago•2 comments

Ask HN: Dear Anthropic, can we please have thought traces back?

2•exabrial•4h ago•2 comments

Ask HN: What was your big failure? How did you get around it?

3•jspann•6h ago•1 comments

How to Pivot?

4•noreplydev•7h ago•3 comments

Ask HN: What if there was an 8th day of the week?

6•cyanregiment•4h ago•1 comments

Ask HN: Can we direct AI on our personal projects during work hours?

3•rpnx•8h ago•3 comments

ArXiv

3•fred123123•9h ago•0 comments

Why remote roles are region specific and not 100% remote?

3•moizrocky1•9h ago•8 comments

An agent built for Mobile

2•technicalbot•10h ago•0 comments

Tell HN: Seedance 2.5 API Pricing

3•JimsonYang•10h ago•1 comments

The LLMs Problems

2•noreplydev•11h ago•1 comments

How did Wikipedia layout become so bloated?

21•countWSS•13h ago•27 comments

Ask HN: Does AI research need "world models" more than bigger LLMs?

3•unjuno•1d ago•2 comments

Ask HN: Are you working 996 hours?

3•derwiki•10h ago•15 comments

Local, Open source AI meeting recorder that sees and hears everything

9•puremetrics•1d ago•3 comments

Ask HN: Have LLMs Plateaued?

11•leandrobon•1d ago•26 comments

You've reached the end!

Open in hackernews

Ask HN: Dear Anthropic, can we please have thought traces back?

2•exabrial•4h ago
Dear Anthropic, can we please have thought traces back?

I can't verify whether or not the LLM is arriving at the conclusion from cheating, or if it's fudging or making stuff up.

Opus 4.6 remains the best model because of this.

Comments

bigyabai•3h ago
> Opus 4.6 remains the best model because of this.

Huh? Are you not trying other open-weight models that stream thinking traces?

The new DeepSeek-V4-Flash-0731 should clobber Opus 4.6 in a lot of tasks. Kimi K3 and GLM 5.2 feel like they stand toe-to-toe with Opus 4.8 in my experience. This probably isn't the last stupid decision that Anthropic stands on, you might as well hedge your bet and put some money into another inference provider and see how it goes.

exabrial•3h ago
Yeah, I think we need to. Opus 5 is great, but it's too wordy and it's making mistakes that aren't caught till much later. Usually you can see if the model is going off-course through the thought trace, Opus5 you're flying blind.

Honestly Anthropic keeps clubbing themselves. They're so worried about their competition they're no longer innovating.