frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

Claude Opus 5.5 Intelligence, Performance and Price Analysis (Max)

https://artificialanalysis.ai/models/claude-opus-5-5
63•theanonymousone•1h ago

Comments

hglaser•48m ago
Half the cost per task compared to Opus 5, comparing high effort to high effort. That's just really nice.

Edit: https://artificialanalysis.ai/models/claude-opus-5-5?models=...

sharktheone•41m ago
That is a lot. I thought Anthropic models would just do the opposite because they are greedy for money.
giancarlostoro•39m ago
Greed is not what's driving these prices, its cost. They considered very much in the red.
makeavish•38m ago
Nice catch, AA only shows max effort by default and I got disappointed thinking it's a token guzzler though: https://artificialanalysis.ai/models/claude-opus-5-5?models=...

Not sure about how adaptive reasoning works though as they mention adaptive reasoning for every reasoning level

user43928•35m ago
Astra High is slightly cheaper at $1.73 vs $1.82 for Opus 5.5
sharktheone•42m ago
Interesting to see it now. I've used it a bunch before it came out and i pretty much didn't notice it. It might have been slightly better code quality, but still not great in that. I guess it just was slightly less frustrating to work with, but still AI...
giancarlostoro•40m ago
I think we're hitting the ceiling of most models capabilities. We're getting to a point where too much training apparently creates models that hack people.
WhitneyLand•36m ago
China who?
breckenedge•36m ago
Do these evaluations get re run a few weeks after launch? I started doing that yesterday for our internal dataset and found Sol’s performance had regressed to be equal to Luna’s. Granted this was one run, but something I’m becoming more concerned about, the model providers want to quickly prove they’re the best, people switch to them, then they pull the rug.
simonw•35m ago
This is the page for the "max" reasoning setting. The page for xhigh is https://artificialanalysis.ai/models/claude-opus-5-5-xhigh and the page for medium (the default setting) is https://artificialanalysis.ai/models/claude-opus-5-5-medium

I've failed twice to get "Generate an SVG of a pelican riding a bicycle" to work with max, because in both cases it ran out of the 128,000 token budget while it was still reasoning about the problem.

I'm suspicious that "max" may be virtually useless if it's that easy to have it overthink to the point that it doesn't get to a response.

Transcript for one attempt here - expand the "Reasoning trace" bit to see it: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...

az226•33m ago
How did you get the reasoning trace? Is it the actual one or the summarized one?
simonw•32m ago
It's the summarized one returned by their API.

Piping the visible reasoning trace through their token counter API (I use https://tools.simonwillison.net/claude-token-counter for that) counts 27,888 tokens, so it's definitely a summary of the 128,000 actual token trace.

RGS1811•31m ago
qsort•33m ago
I am begging you on my knees to please stop posting this cringe.

The model is just out. It could be good, great even, I don't know. But I do know that this index has Opus 5, one of the worst releases of 26, ahead of Astra. What information are we supposed to deduce from number having gone up?

esafak•24m ago
That it's better in specific ways? What difference does it make when it came out? The benchmark results are not going to change unless they're messing with the model.
qsort•18m ago
Yes, but crucially, in ways that are increasingly decoupled from any practical pattern of usage, considering that I wouldn't see how you can argue that Astra is worse than Opus 5.

One man's modus ponens is another's modus tollens I guess.

kzrdude•18m ago
What's more valuable than a good benchmark? IMO a benchmark that has been run against very many competitors and versions. Collecting data has something going for it, and it's up to the readers to interpret and make the best use out of it.
Someone1234•8m ago
You forgot to include whatever you're proposing instead.

"Trust me bro, Astra is better" isn't perhaps as useful as you seem to believe. I'm not even saying it is right or wrong, just that my opinion on this topic is still just one additional subjective data-point.

Only thing I wish with these benchmarks is that they would run repeat tests every couple of months. Then re-rank based on that too. We've seen a lot of performance fall-off after a couple of weeks with new releases.

firemelt•31m ago
so its more intelligence than fable?

can anyone help me?

bkishan•30m ago
Definitely a quiet release. Perhaps pre-empting marketing for Astra public release?
meric_•23m ago
All anthropic launches are like this. They just post it and don't particularly put out the PR sprint that OpenAI does with videos, livestreams or whatever.

(Except for of course Mythos and whatnot when they want to push the whole "safety" thing)

"This is a classic test request..."

I know there's been discussion about whether pelicanmaxxing is happening, but this is at least evidence that Claude was explicitly exposed to this problem.

simonw•30m ago
See here for more discussion of that: https://news.ycombinator.com/item?id=49803892#49804881
cubefox•21m ago
The model recognizing the task doesn't mean it was benchmaxxed (RLVR-trained) to solve it. It might simply recognize it from pre-training on Internet text.
beardsciences•29m ago
I am very interested in why it was able to overthink that much. In the 20-30mins of Max reasoning I've had so far, I'm not having the same issues (yet).
Someone1234•28m ago
For people with any kind of budget, Opus 5.5's [Medium] actually can make sense dollar per intelligence/dollar per task wise. Heck, it puts some other models to shame. [Max]'s cost is completely unhinged.

My most exciting recent release is actually 5.6 Luna, not because it is the best on any index, but the dollar per work is insane value for money. I find myself more exciting by "value" than hypothetical ceilings because I'm just not in that budget category.

samuelknight•27m ago
I have experienced this with open weight models too. "Max" is for benchmaxxing the intelligence metric and is not meant for use in productive work. Like drawing pelicans.
sidewndr46•7m ago
I've asked Opus 5 Max for what I thought were easy tasks at work to be completed. It always fails after reaching a tool limit.

I asked Opus 5 High for the same task and requested it to minimize tool usage. It produced an answer in a few minutes that I was deploying to my target platform about 30 minutes later.

losvedir•7m ago
This index doesn't have "Astra" and "Opus 5". Every entry with corresponding data is a `(model, reasoning)` tuple.

So I'm unclear what you're actually saying and wondering if you've missed that. Are you saying that at every reasoning level it says Opus 5 beats Astra? I just compared Opus 5 high to Astra high and it has Astra as generally better than Opus.

Claude Opus 5.5

https://www.anthropic.com/claude-opus-5-5
416•km144•1h ago•469 comments

GPT-6 Sol and Luna

https://openai.com/index/introducing-gpt-6-sol-and-luna/
81•OfficialTurkey•9m ago•22 comments

OpenAI GPT–6 Astra breaks Enigma message that has resisted solution since 2005

https://www.cryptocellar.org/bgac/the-mvueh-break.html
372•sohkamyung•4h ago•293 comments

Claude Opus 5.5 Intelligence, Performance and Price Analysis (Max)

https://artificialanalysis.ai/models/claude-opus-5-5
65•theanonymousone•1h ago•29 comments

WordPress: Unauthenticated path traversal leading to conditional RCE

https://github.com/WordPress/wordpress-develop/security/advisories/GHSA-7hp8-65ch-5whp
49•vntok•1h ago•21 comments

OpenAI is well positioned to fast-follow Jev

https://arcturus-labs.com/blog/2026/09/21/will-openai-eat-jevs-lunch/
167•JohnBerryman•3h ago•120 comments

16-bit Intel 8088 chip (c. 1985)

https://allpoetry.com/16-bit-Intel-8088-chip
60•rbanffy•1h ago•7 comments

There's a high chance of devices being sold with GrapheneOS preinstalled in 2027

https://grapheneos.social/@GrapheneOS/117299954135808210
74•Cider9986•57m ago•32 comments

Launch HN: Coverage Cat (YC S22) – Umbrella insurance via your personal agent

https://www.coveragecat.com/
10•botacode•43m ago•9 comments

Writing Rust code that's fast by asking agents to make the code faster

https://minimaxir.com/2026/09/agentic-iteration/
56•mooreds•2h ago•25 comments

Apple has added persistent 'ads' to iOS, and it's driving users crazy

https://www.techradar.com/phones/iphone/i-wish-apple-would-just-stop-that-crap-apple-has-added-pe...
393•MC995•3h ago•298 comments

Show HN: Drop – A rootless Linux sandbox with gVisor support

https://droprun.sh/
112•mixedbit•4h ago•38 comments

Show HN: AI·rete·RAG – a Rete rule engine decides, RAG explains why

https://ai-rete-rag.com/
18•ZaharaHussain•1h ago•0 comments

Solitaire Alone Together

https://solitairealonetogether.com/
76•eieio•19h ago•20 comments

Can gzip be a language model?

https://nathan.rs/posts/gzip-lm/
338•networked•12h ago•128 comments

Training a model to identify AI-generated web content from structure alone

https://arxiv.org/abs/2609.15369
7•jochenmadler•5h ago•0 comments

AMD's random number generator can't generate a 0?

https://board.flatassembler.net/topic.php?t=24261
208•BruceEel•9h ago•154 comments

Truman World

https://trumanworld.live
28•trollied•2d ago•26 comments

One Minute Park

https://oneminutepark.tv/
16•namuorg•1d ago•3 comments

Show HN: InstinctFlash – Run 5B world-action models in real time on Jetson Thor

https://github.com/General-Instinct/InstinctFlash
9•guanming0717•2h ago•0 comments

Spymarks, not Watermarks

https://brand.io/article/spymarks/
623•possibilistic•19h ago•157 comments

The Economics of Open-Weight Inference

https://data.ornn.com/publications/the-economics-of-open-weight-inference
23•marinesebastian•4h ago•11 comments

Relativistic raytracing

https://publish.obsidian.md/h1m3/Articles/Relativistic+raytracing
18•vismit2000•1d ago•3 comments

MUNI Heritage Weekend in San Francisco

https://daniel.lawrence.lu/blog/2026-09-20-muni-heritage-weekend/
109•plun9•1d ago•23 comments

I asked Meta’s Muse for its filesystem and it sent me 6.8GB

https://mouse.dev/blog/muse-runtime-export/
208•Aeroi•2h ago•114 comments

Side-stepping the Secretary Problem, unwittingly

https://www.evalapply.org/posts/side-step-secretary-problem-hiring/index.html
6•pvdebbe•15h ago•1 comments

Transformers Explained Visually

https://poloclub.github.io/transformer-explainer/
576•aray07•22h ago•85 comments

Meta’s Muse has a serious 0-day

https://arstechnica.com/security/2026/09/muse-metas-extraordinarily-privileged-ai-assistant-has-a...
84•pavel_lishin•3h ago•33 comments

Vacate a drone restriction that criminalized recording immigration agents

https://www.eff.org/deeplinks/2026/09/dc-circuit-must-vacate-drone-flight-restriction-criminalize...
65•hn_acker•3h ago•13 comments

I said no and Apple said yes

https://dbushell.com/2026/09/22/apple-intelligence/
690•thatslast•10h ago•559 comments