frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

Claude Opus 5.5 Intelligence, Performance and Price Analysis

https://artificialanalysis.ai/models/claude-opus-5-5
41•theanonymousone•53m ago

Comments

hglaser•22m ago
Half the cost per task compared to Opus 5, comparing high effort to high effort. That's just really nice.

Edit: https://artificialanalysis.ai/models/claude-opus-5-5?models=...

sharktheone•16m ago
That is a lot. I thought Anthropic models would just do the opposite because they are greedy for money.
giancarlostoro•14m ago
Greed is not what's driving these prices, its cost. They considered very much in the red.
makeavish•13m ago
Nice catch, AA only shows max effort by default and I got disappointed thinking it's a token guzzler though: https://artificialanalysis.ai/models/claude-opus-5-5?models=...

Not sure about how adaptive reasoning works though as they mention adaptive reasoning for every reasoning level

user43928•10m ago
Astra High is slightly cheaper at $1.73 vs $1.82 for Opus 5.5
sharktheone•16m ago
Interesting to see it now. I've used it a bunch before it came out and i pretty much didn't notice it. It might have been slightly better code quality, but still not great in that. I guess it just was slightly less frustrating to work with, but still AI...
giancarlostoro•15m ago
I think we're hitting the ceiling of most models capabilities. We're getting to a point where too much training apparently creates models that hack people.
WhitneyLand•11m ago
China who?
breckenedge•11m ago
Do these evaluations get re run a few weeks after launch? I started doing that yesterday for our internal dataset and found Sol’s performance had regressed to be equal to Luna’s. Granted this was one run, but something I’m becoming more concerned about, the model providers want to quickly prove they’re the best, people switch to them, then they pull the rug.
simonw•9m ago
This is the page for the "max" reasoning setting. The page for xhigh is https://artificialanalysis.ai/models/claude-opus-5-5-xhigh and the page for medium (the default setting) is https://artificialanalysis.ai/models/claude-opus-5-5-medium

I've failed twice to get "Generate an SVG of a pelican riding a bicycle" to work with max, because in both cases it ran out of the 128,000 token budget while it was still reasoning about the problem.

I'm suspicious that "max" may be virtually useless if it's that easy to have it overthink to the point that it doesn't get to a response.

Transcript for one attempt here - expand the "Reasoning trace" bit to see it: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...

az226•7m ago
How did you get the reasoning trace? Is it the actual one or the summarized one?
simonw•7m ago
It's the summarized one returned by their API.

Piping the visible reasoning trace through their token counter API (I use https://tools.simonwillison.net/claude-token-counter for that) counts 27,888 tokens, so it's definitely a summary of the 128,000 actual token trace.

RGS1811•5m ago
qsort•7m ago
I am begging you on my knees to please stop posting this cringe.

The model is just out. It could be good, great even, I don't know. But I do know that this index has Opus 5, one of the worst releases of 26, ahead of Astra. What information are we supposed to deduce from number having gone up?

firemelt•5m ago
so its more intelligence than fable?

can anyone help me?

bkishan•5m ago
Definitely a quiet release. Perhaps pre-empting marketing for Astra public release?
"This is a classic test request..."

I know there's been discussion about whether pelicanmaxxing is happening, but this is at least evidence that Claude was explicitly exposed to this problem.

simonw•5m ago
See here for more discussion of that: https://news.ycombinator.com/item?id=49803892#49804881
beardsciences•4m ago
I am very interested in why it was able to overthink that much. In the 20-30mins of Max reasoning I've had so far, I'm not having the same issues (yet).

Shopify CEO: employees' 'slop grenades' are making more work for everyone else

https://fortune.com/2026/09/17/shopify-tobias-lutke-ai-slop-grenades/
1•cratermoon•49s ago•0 comments

Encoding transparent videos that work in Safari, Chrome and Firefox

https://ben.terhech.de/posts/2025-02-02-transparent-video-safari.html
1•figbert•1m ago•0 comments

ConferenceRank – conference and journal deadlines for 980 CS venues

https://rabimba.github.io/ConferenceRank/
1•rabimba•2m ago•0 comments

Trump says AI will be renamed 'super intelligence' in all US documents

https://thehill.com/homenews/administration/6104142-trump-renames-ai-super-intelligence/
4•guardiangod•3m ago•0 comments

Avy Tab Switcher, Avy-style keyboard navigation for browser tabs

https://github.com/Artawower/avy-tab-swithcer
1•darkawower•5m ago•0 comments

A beginner-friendly, step-by-step guide – How to Fingerprint Popular honeypots

https://medium.com/meetcyber/how-to-fingerprint-popular-honeypots-from-your-laptop-fe3396a13211
1•ls1911•5m ago•0 comments

Tiny Startups Are Getting Even Smaller with Help from AI

https://www.wsj.com/tech/ai/startup-hiring-ai-staffing-8c626f75
1•nradov•5m ago•0 comments

Japanese used bookstores are seeing a surge in bulk orders for obscure books

https://twitter.com/Johnny_suputama/status/2102425777128292810
2•wahnfrieden•7m ago•0 comments

Support Fins – Stop Using Tree Support for Your 3D Prints [video]

https://www.youtube.com/watch?v=WGsi5SCurpM
1•luanmuniz•8m ago•0 comments

Horowitz Andreessen Academy

https://a16z.com/announcement/incubating-horowitz-andreessen-academy/
2•lquist•8m ago•0 comments

Go pace yourself Dario [video]

https://www.youtube.com/shorts/83X79cfuE3k
1•hsuduebc2•9m ago•0 comments

Dutch social network Hyves returns nearly 13 years after closure

https://nltimes.nl/2026/09/22/dutch-social-network-hyves-returns-nearly-13-years-closure
2•giuliomagnifico•10m ago•0 comments

Koòrdinate Thinking

https://github.com/UnmarkedPM/koord/tree/main/tools/skills/FlatSkill
1•bender90•11m ago•1 comments

Kalshi asks CFTC to allow margin trading on its platform

https://www.cnbc.com/2026/09/22/kalshi-asks-cftc-to-allow-margin-trading-on-its-platform-letting-...
4•thm•12m ago•0 comments

Wishing It So: Excerpt from Will and Attention by Meghan O'Gieblyn

https://www.nybooks.com/articles/2026/09/24/wishing-it-so-meghan-ogieblyn/
1•1vuio0pswjnm7•13m ago•0 comments

RomM – Self-Hosted ROM Library with Metadata from IGDB, Screenscraper, MobyGames

https://digitalescapetools.com/tools/romm.html
1•xabd•14m ago•0 comments

In the beginning, it was good –part 1

1•theubie•15m ago•1 comments

CoreQuarry

https://corequarry.com/
1•Non-monotonic•15m ago•0 comments

Show HN: Task and test management as YAML in your Git repo (VS Code)

https://github.com/gitoza-io/gitoza-lite
1•weiwen-weng•16m ago•0 comments

Musings on the Barrow Scale

https://www.centauri-dreams.org/2023/12/12/seti-musings-on-the-barrow-scale/
1•JumpCrisscross•16m ago•0 comments

What if you could experience the regret of a decision before making it?

https://regret.solvailabs.com/
1•theycallmeritik•17m ago•0 comments

Show HN: Optimized runtimes for three VLAs on Jetson Thor

https://github.com/Agents2AgentsAI/vla-edge
2•hhuytho•17m ago•0 comments

Launch HN: Coverage Cat (YC S22) – Umbrella insurance via your personal agent

https://www.coveragecat.com/
7•botacode•17m ago•3 comments

Show HN: Nomoreda – Browser EDA, MCP-Friendly, KiCad/Altium-Compatible

https://nomoreda.com/
1•VladimirIvlev•19m ago•0 comments

NASA Discovery Reveals Complex Water Systems on Early Mars

https://www.jpl.nasa.gov/news/nasa-discovery-reveals-complex-water-systems-on-early-mars/
2•gmays•20m ago•0 comments

TinyJev -Tiny Jev-style decision model that runs offline

https://github.com/ankit-aglawe/tinyjev
2•aglaweankit•21m ago•0 comments

Are the Government's Conversations with AI Accessible Under Public Records Laws? [pdf]

https://reason.com/wp-content/uploads/2026/09/Are-the-Goverments-AI-Conversations-Accessible.pdf
3•compiler-guy•21m ago•0 comments

German court rules Meta liable for scam ads on Facebook and Instagram

https://thenextweb.com/news/meta-scam-ads-ruling-germany-frankfurt-court
10•buzer•23m ago•3 comments

Priorities and principles for effective third party assessments

https://openai.com/index/priorities-principles-third-party-assessments/
1•samaysharma•24m ago•0 comments

Scaling Discovery Through Test-Time Communication

https://arxiv.org/abs/2609.21032
1•marojejian•24m ago•1 comments