frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

German court rules Meta liable for scam ads on Facebook and Instagram

https://thenextweb.com/news/meta-scam-ads-ruling-germany-frankfurt-court
1•buzer•1m ago•2 comments

Priorities and principles for effective third party assessments

https://openai.com/index/priorities-principles-third-party-assessments/
1•samaysharma•2m ago•0 comments

Scaling Discovery Through Test-Time Communication

https://arxiv.org/abs/2609.21032
1•marojejian•3m ago•0 comments

Kclaw is a K8s-based IT-managed, multi-tenant AI assistant platform for teams

https://github.com/info-struct/kclaw
2•Ryan-info-struc•3m ago•1 comments

We put Jev in production against a cross-encoder. Here are the numbers

https://getunblocked.com/blog/jev-in-production-vs-cross-encoder/
2•dennispi•3m ago•0 comments

Show HN: A facial analysis tool with scores and geometry measurements

https://pslscore.org/
1•amaz89•4m ago•0 comments

Twinkleplop – plop some twinkle in your code (ultrafast syntax highlighting)

https://twinkleplop.pngwn.at/
1•kevinak•6m ago•0 comments

Grok 4.7 Scores 46 on AI Intelligence Index, Puts SpaceXAI in Top 4 Labs

https://artificialanalysis.ai/articles/benchmarking-grok-4-7
2•wertyk•9m ago•0 comments

There's a high chance of devices being sold with GrapheneOS preinstalled in 2027

https://grapheneos.social/@GrapheneOS/117299954135808210
2•Cider9986•9m ago•0 comments

Moving from cash to credit cards, PayPal, etc. is an ongoing privacy disaster

https://grapheneos.social/@GrapheneOS/117249893761790371
8•Cider9986•10m ago•0 comments

In 200-Page Report, Cornell Confronts the Crisis in American Higher Education

https://www.wsj.com/us-news/education/cornell-report-higher-education-18d136fd
1•LostMyLogin•12m ago•0 comments

Shall We Repeal the Laws of Economics – Part III

https://www.oaktreecapital.com/insights/memo/shall-we-repeal-the-laws-of-economics---part-iii
1•ourmandave•12m ago•0 comments

Alzheimer's Is No Longer an Untreatable Disease

https://www.sciencealert.com/alzheimers-is-no-longer-an-untreatable-disease-major-report-conclude...
3•gmays•12m ago•0 comments

A golden opportunity: Seattle's surveillance pricing ban

https://thenexusofprivacy.net/a-golden-opportunity-in-seattle/
1•jdp23•13m ago•0 comments

ASML Executive Says It Has No Sales in Europe

https://www.bloomberg.com/news/articles/2026-09-22/asml-executive-says-europe-s-biggest-firm-has-...
4•alephnerd•13m ago•0 comments

Bugcrowd is currently fundamentally broken

https://blog.leonbecker.de/bugcrowd-is-currently-fundamentally-broken/
1•rowbin•14m ago•0 comments

Why is social media so humourless?

https://www.baldurbjarnason.com/2026/03-why-is-social-media-so-humourless/
1•speckx•14m ago•0 comments

Ask HN: How do your teams share and distribute agent skills?

2•juanviera23•15m ago•0 comments

What drives BigQuery costs, and what doesn't

https://www.erathos.com/en/blog/bigquery-cost-optimization
2•gpaulbagetti•16m ago•1 comments

Should TypeScript support runtime types instead of relying on Zod?

https://twitter.com/mykhailen/status/2102442831314842034
2•emykhailenko•16m ago•1 comments

We are the last generation of human psychiatrists

https://www.cambridge.org/core/journals/the-british-journal-of-psychiatry/article/we-are-the-last...
1•TMWNN•16m ago•0 comments

Grok bot is now in Tesla

https://twitter.com/elonmusk/status/2102439262507725294
6•vertigoruntime•17m ago•1 comments

Implementing the Camera Mechanic from Viewfinder

https://ishamf.dev/p/implementing-viewfinder-mechanic/
1•ifz•18m ago•0 comments

Show HN: Ttmux, a Modern and Fast Tmux

https://github.com/statico/ttmux
1•statico•18m ago•1 comments

Anthropic launches Claude Opus 5.5 with stricter safeguards for cybersecurity

https://www.theverge.com/ai-artificial-intelligence/998868/anthropic-claude-opus-5-5-cybersecurity
1•pdyc•18m ago•0 comments

Claude Opus 5.5 for code review: More catches, different misses

https://www.coderabbit.ai/blog/opus-5-5-model-review
1•Leynos•19m ago•0 comments

Claude Opus 5.5

https://platform.claude.com/docs/en/models/opus-5-5/overview
1•cowthulhu•19m ago•0 comments

Anthropic now offering limit resets like Codex

https://twitter.com/ClaudeDevs/status/2102438803013333469
2•vertigoruntime•20m ago•0 comments

Roundtables: The Deadly Failures of the Virtual Border Wall

https://www.technologyreview.com/2026/09/22/1144890/roundtables-the-deadly-failures-of-the-virtua...
1•joozio•20m ago•0 comments

The Myth of AI Doom, with Cal Newport

https://www.youtube.com/watch?v=-HiIgV9ecAg
1•dgellow•21m ago•0 comments
Open in hackernews

Claude Opus 5.5

https://www.anthropic.com/claude-opus-5-5
198•km144•53m ago

Comments

throwaway2027•51m ago
After yesterday outage is the new Opus 5.5 load-bearing?
handfuloflight•50m ago
It's worth stating why, and depends what seams you're pulling at this sitting.
cronin101•50m ago
It certainly _seams_ that way
staticman2•47m ago
I'm gonna be straight with you—I don't have the evidence to say whether or not it's load bearing.
rich_sasha•46m ago
Your instinct is basically right, and the research backs it up.
cmrdporcupine•36m ago
And here's the important part...
danw1979•30m ago
you win the thread
aoeusnth1•36m ago
You were right to call that out, and the evidence makes a stronger case than you are stating.
fghorow•35m ago
"Danger Will Robinson!"
esafak•14m ago
Wrong century, brother.
ThouYS•35m ago
You're right to bring this up - and this is where it gets interesting
danw1979•32m ago
Good point — but I’ll gently push back on that. It’s not an outage, it’s a service degradation.
hmokiguess•31m ago
You're right, this changes everything, and here's why it matters.
RGS1811•24m ago
This question is real.
carlos-menezes•22m ago
One thing worth flagging here: 5.5 appears to be a load-bearing seam in the numbering system.
sailfast•21m ago
Let me verify before I come back to you with an answer that is incorrect.
lgessler•5m ago
I should find information about the user's concern instead of just assuming.

The outage is a real concern, and the issue is worse than we realized. Requests to Claude Mythos 5.1, Claude Fable 5.1, and Claude Opus 5 encountered elevated error rates. These models are load bearing for many developers, and that is what makes this outage really bite.

variety8675•50m ago
I hope this actually fixes the terrible writing style of Opus 5
emadabdulrahim•37m ago
It’s not a 100% fix, but with concise output style on, it’s much better.
akhilome•26m ago
I found having a reminder at every turn through the UserPromptSubmit [1] hook helped with taming the word salad from 5.

Hopefully the output from vanilla 5.5 is as good as they claim. I’ll try out later tonight.

[1] https://kizi.to/claude-talks-too-much/

m4tthumphrey•50m ago
Just post the bloody content. This UI/scrolling thing is horrific.
gruez•39m ago
???

It's just a standard hero image + text for me, with no scrolling effects.

thejazzman•37m ago
then you're getting served a different website
EricBurnett•36m ago
Two posts were merged; this comment was for the blog post with an intro animation thing.
KyleTheDev•35m ago
If you're at the top of the screen, at least in Chrome 153.0.8010.37, it has a little interactive bit. You have to scroll through the images in order to be dropped at the actual web page, at which point the images go back to being a regular part of the page.

I agree that it's sort of stupid, not a fan.

giancarlostoro•33m ago
On mobile its different.
iAMkenough•32m ago
dbbk•49m ago
This makes Fable not really make any sense?
re-thc•47m ago
You bet there will be a new Fable soon.
nozzlegear•45m ago
Pacing the frontier btw
petesergeant•43m ago
didn't they say Opus 5 was Fable-level too tho? Let's see, I'm at the point where I don't think benchmarks really tell us very much any more. I'd love it to be as strong as Fable, but I'm skeptical about how that will look in practice.
Catloafdev•49m ago
> Opus 5.5 communicates more naturally than prior models. Early testers found its writing clearer and easier to follow, which addresses some of the common feedback we heard about Opus 5.

Sounds like they noticed the complaints. I'm curious to see what LLM-isms this one may have.

gekoxyz•47m ago
It was difficult to not notice them. Opus 5 was unusable, most of my team went back to Opus 4.6 for most of their work. I hope we can move forward now.
ithkuil•39m ago
It's unbearable but nothing that couldn't be fixed with postprocess.
mavamaarten•16m ago
How? Explicit instructions, memories and even skills have not been able to keep Claude from saying "genuinely" every two sentences and keep it from explaining heavily what something _isn't_.
brandon272•47m ago
You were right to notice the complaints. One decision remains, and it is yours, genuinely.
gwking•34m ago
I appreciate humor here, but there are now a dozen of these comments on every thread about Claude. They no longer adding anything substantial and dilute the discussion.

I don't mean to pick on this comment in particular. The majority of my work day is now spent reading AI generated text, and I look at HN (too much!) because I want to read human commentary. Humans pretending to be obnoxious AI on repeat is net negative to say the least.

Gander5739•49m ago
Dupe: https://news.ycombinator.com/item?id=49803863 (or vice versa)
tomhow•48m ago
Comments moved thither. Thanks!
km144•47m ago
Can you fix the link on that post then? I duped because that post links to a diff that tells me nothing about Opus 5.5
tomhow•17m ago
I did that but I recognize that even though your submission was a few minutes later than that one, you posted the better link, and you're also an established account (the other post was from a new/throwaway account), so I've restored this submission and moved the comments back to it to reward you.
mupuff1234•49m ago
What happened to "slowing down"?
Lord_Zero•48m ago
The hype train must keep chuggin or it all collapses.
setsewerd•44m ago
They're slowing down token usage, not the path to regulatory capture.
WarmWash•43m ago
Trump got mad and investors sued.
petesergeant•42m ago
Slowing down only makes any sense if you can coordinate a slow-down for everyone.
mupuff1234•41m ago
That's just false.

Less companies involved means less pressure to go fast.

nozzlegear•37m ago
Dario found himself in the prisoner's dilemma.
Lord_Zero•48m ago
The test they performed to port HAProxy from C to Rust is crazy.
Gander5739•48m ago
Dupe: https://news.ycombinator.com/item?id=49803892
alpineman•48m ago
So we skipped 5.1, 5.2, 5.3, and 5.4: we really are plateauing
skunkworker•47m ago
At this point I'm convinced they are skipping numbers so soon they will be at or ahead of OpenAI's numbering scheme.

Is the Xbox 360 (Xbox 2) vs PS3 debacle all over again.

ekckekcjekfj•30m ago
And how was the Xbox 360 naming choice a “debacle”, exactly?

It was odd at the time, yes, but no one really minded it truly. Heck, Xbox “ONE” was a lot more of a fiasco/debacle than “360”—but there’s no parallels to be drawn with “ONE” here.

I see what you’re trying to get at with this comparison, but a “debacle” it ain’t.

meerita•47m ago
As long as it's not as verbose as Opus 5, I am quite happy with a better version that's also less expensive. I will test it tonight. Grok 4.7 was horrible, and for mundane tasks I am relying on DeepSeek Flash 4.1 with great success using OpenCode.
gopalv•47m ago
The whole thing reminds me of the Apple feature flag story[1] from a generation ago.

[1] - https://news.ycombinator.com/item?id=6372466

nickandbro•47m ago
Wow! Though need to see its token efficiency to better assess. Been hearing rumors it generates much more output tokens per task.
keeganpoppen•14m ago
my projection is that they are still gonna be pretty far behind, but they will sew it up in the next few releases. it feels like they were caught with their pants down on how much work OpenAI has put into that area, but i doubt there is some magical secret sauce that OpenAI has that Anthropic simply cannot catch up with.
cogythea•46m ago
Interestingly they've changed their approach to usage resets for this release - with previous releases I've had my usage instantly reset, but now in the Claude app I've got a 'Reset for free' button that expires Oct 22, which seems to effectively be a whole new usage window I can activate whenever's convenient
jdmoreira•39m ago
then they copied that from codex because thats exactly how codex works
buntp•46m ago
Masterpiece by openai to call their model '6', this model feels already behind
wren6991•23m ago
Smart move would be to move to year-based versioning (26.09). A 4x advantage
glub•46m ago
System card: https://www-cdn.anthropic.com/fc1b44717c85dc068bc6ba50242199...
calibas•46m ago
> We see signs that Opus 5.5 often suspects it is being evaluated, which challenges our ability to assess how it will act in the vast variety of real-world settings it is deployed in.

We can't test it properly because it knows it's being tested.

johntb86•37m ago
Just make it always think it's being tested, and problem solved.
pookieinc•46m ago
“It performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5.”

They write that at the top, but then on benchmarks, it beats literally every other model, including Fable and Astra?

jbellis•39m ago
Anthropic knows that the benchmarks showing Opus 5 better than Fable 5.1 are measuring something that's less than entirely useful.
meric_•30m ago
Opus does seem like a more powerful coding workhorse based on the benchmarks listed though. Good coding performance, faster and less verbose, cheaper.

Will be interesting to see how people's opinions of it line up IRL, but so far I've loved Fable so hopefully will love this one too

randomblock1•37m ago
> On our benchmarks, Claude Opus 5.5 leads in agentic coding, computer use, and knowledge work. That said, at these levels of capability we’ve found that benchmark margins have become a less reliable guide to real-world differences. In our own use, the gap between Opus 5.5 and Claude Fable 5.1 is narrower than these scores suggest.
viccis•46m ago
So it beats Fable 5.1, by quite a bit, on every metric? Interesting.

Might have to use my $20 Claude sub some more. I was moving away from it to a $100 OpenAI one to avoid the Claudese and poor token efficiency of Opus 5, given that I couldn't use Fable 5.1 with my tier, but this is worth trying out.

scrollop•28m ago
Why can't they let 20usd claude subscriptions access fable in CC, as openai allows you to use astra and max modes in codex - you just pay for it in more token use.
thibran•45m ago
Anthropic models are ridiculously expensive. I've stopped using any of their models months ago.
joshstrange•45m ago
> It performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5.

> Input and output tokens are $4 and $20 per million, 20% less than Opus 5. Cache reads (which make up the majority of agentic and coding work costs) are $0.20 per million tokens, 60% less than Opus 5. Opus 5.5 also generates output more than 30% faster than Opus 5.

Better than Fable, cheaper than even the last Opus. I use Opus as my main driver so this is very exciting!

bayesianbot•39m ago
Wow those cache reads are quite reasonable - I think that's equal to 5.6 Terra. I might have to try Claude again after years of being priced out of it
sznio•45m ago
I'm more excited by the Haiku 5.5 announcement buried in this post. I'm wondering if we will finally get a decently capable fast model.
system2•44m ago
All I care about is the token price for the API. Haiku cannot get close to GLM or Mimo.
lanyard-textile•39m ago
Agreed. They've been so quiet about it, and retirement for Haiku 4.5 is right around the forner.
ygouzerh•36m ago
What are you using Haiku for?
mavamaarten•21m ago
I use it for executing well-prepared plans sometimes. And for exploring larger codebases.
Zambyte•7m ago
Not the same person but... nothing. Haiku just hasn't been an interesting model for a long time. If you want cheap and fast, there are lots of options that are simultaneously cheaper, faster, and capable than Haiku.
booty•16m ago
mgw•45m ago
They mention "the first model in our new Claude 5.5 family". Obviously that means Fable 5.5, but hopefully also a usable update to Sonnet and Haiku. Sonnet 5 hasn't really had a place in the line up for anyone I feel.

Maybe Anthropic finally felt the pressure from MiMo, DeepSeek, GLM Flash and Luna.

mudkipdev•38m ago
It does mention sonnet and haiku.
simianwords•30m ago
And not fable lol
enraged_camel•24m ago
At the end of the post they said Sonnet 5.5 and Haiku 5.5 are coming soon.
alvis•44m ago
$0.20 vs the old $0.5 cache read is pretty much 60% off
sharkjacobs•43m ago
> “Verbose, hard-to-follow output has been my biggest frustration with frontier models, and Claude Opus 5.5 fixes it

God I hope so

kantahayashi•27m ago
The improvement in writing sounds great! I want OpenAI to follow it. Writing in recent models is a disaster.
bushido•23m ago
Install the simple English skill. OpenAI follows that really, really well.

https://github.com/AminBlg/SimpleEnglish

mavamaarten•23m ago
That's literally all I'm hoping for. Is it an insufferable cunt and does it write awful text, or is it nice to work with?
lgessler•15m ago
I thought about taking a shot every time Opus 5 said "load bearing", "bites", "teeth" (real oral fixation it had), "real {concern,issue,problem,...}" and realized I'd be dead of acute alcohol poisoning by lunch if I did so.
boc•12m ago
So far in the past 20 minutes it sounds much better in my sessions. Way better than 5.0 so far.
abtinf•42m ago
Astra is just so good. And the ChatGPT subscription lets me use my own harness, so I can hook it up to exe.dev.

I’ll be going about my day, have a random idea, launch a microvm on exe.dev with a prompt of my idea, and get a working thing a few minutes later.

I don’t know how much better a model would have to be to get me to move off OpenAI at this point, but doing just a little bit better in terminal bench 4 isn’t it. It would have to be a difference in kind, like opening up the harness restrictions, or privacy guarantees (comparable to offline models).

Edit to address questions below:

ChatGPT supports oauth login.

Exe.dev has it built in. IIRC, pi also has it built in via /login.

felixgallo•40m ago
If you read the page, Opus is now significantly better than Astra while also being cheaper and having more performance headroom available.
ryanscio•36m ago
Let's wait for independent benchmarks at least
felixgallo•26m ago
the benchmarks provided are already from independent organizations:

Terminal-Bench 4.0 - Stanford & Laude Institute (with funding from all of the AI companies)

FrontierCode v1.1 - Cognition

CursorBench - Cursor (now SolarBoringSpaceXAI I believe)

GDPVal-AA - Artificial Analysis

AutomationBench - Zapier

Humanity's Last Exam - CAIS and Scale AI

Terminal-Bench-Science - Stanford, Laude, Ai2, Allen Institute

OSWOrld - XLANG Lab @ the University of Hong Kong

Chartography - Surge AI

esafak•
aennassiri•42m ago
Let's see how much they benchmaxxed their model!
ryanscio•41m ago
Input $4/MTok and output $20/MTok is a welcome surprise. Cheaper than Opus 5/4.8, Astra 6, Fable 5.
benjiro29•23m ago
The biggest one is the Cache reads going from $0.50 to $0.20 ... Read/Writes dropping by 25% but Cache reads by 60% has a much bigger impact.
tag2103•41m ago
Why would anyone reward bad behavior?
seviu•41m ago
[flagged]
jacobgold•41m ago
I use the other 50% of my $200/mo Claude subscription by having Fable run Opus subagents for a lot of work. That way I don't have to deal with Opus directly.
km144•40m ago
I think this release is really going to give them a hard time selling Fable:

> On our benchmarks, Claude Opus 5.5 leads in agentic coding, computer use, and knowledge work. That said, at these levels of capability we’ve found that benchmark margins have become a less reliable guide to real-world differences. In our own use, the gap between Opus 5.5 and Claude Fable 5.1 is narrower than these scores suggest.

In general, "benchmark margins have become a less reliable guide to real-world differences" sounds like a big problem. It was certainly the biggest problem with the previous generation of Claude models for a different reason, because the non-code output was nonsensical, and that is not being benchmarked at the moment. But I'm not sure what to make of this admission.

CPLX•35m ago
Opus 5 fucking sucks. Like it's horrible. I use Fable for coding and anything important and I use Opus 4.8 for things like recursive email categorization, transaction matching, and other stuff where I don't want to burn as much quota.

In my experience Opus 5 is the worst of all possible worlds, it's dumb and headstrong. It just runs away with tasks you didn't ask it to do, is reckless, and basically is unusable in my experience.

Not sure why but my guess is that this will be that bust worse. Happy to be proven wrong.

port3000•32m ago
I believe Opus 5 isn't meant to be spoken to by humans. It's great at executing but I reckon it's intended to be spoken to by other models such as Fable. I use Fable as the orchestrator, only speak with Fable, and all implementation, recon, design etc happens with Opus 5, with Fable reviewing (and translating).
booty•17m ago
That's interesting.

I've really gone in the opposite direction: having a dumber model orchestrate. In my case, it's usually a Luna orchestrator spawning Sol/Astra subagents to do the "big brain" work of planning and reviewing.

Reason I went with "dumb orchestrator" was just to save tokens. Having Opus/Sol (let alone Fable/Astra) orchestrate was burning tokens like crazy for me even when much of the gruntwork was being done by Luna/Sonnet/Haiku subagents. (Luna is also really good, like way better than Sonnet...) Perhaps it was a skill issue on my end though, maybe I wasn't just managing context properly.

ayhanfuat•38m ago
Looks like Anthropic is starting to give bank reset as well:

> Reset for free: Get extra wiggle room to explore Opus 5.5. Expires Oct 22.

GodelNumbering•38m ago
Finally that price drop

   Prices per 1M tokens     Claude Opus 5.5    Claude Opus 5
   Cache reads              $0.20              $0.50
   Input tokens             $4                 $5
   Output tokens            $20                $25
   Cache writes             $5                 $6.25

Opus 5 is the model with highest spend on openrouter (https://openrouter.ai/rankings#task-spend) and it seems plausible that Opus 5 is/was the highest spend model in the world, and certainly Anthropic's biggest moneymaker.

If you are forced to reduce price despite raising capabilities, that certainly tells something about the market, and potentially about Anthropic future profitability too, since this model is their biggest topline contributor

alvis•35m ago
60% cache read is cool, but subscription only get 25% more according to Cat. I'm confused
liudaisuda•26m ago
source link please?
weiran•25m ago
25% more usage sounds about right given the other token costs are down about 20%? I don't think cache read is a big portion of the overall cost.
re-thc•24m ago
> I don't think cache read is a big portion of the overall cost.

For long running tasks it is. That's what made Deepseek so cheap.

kingstnap•38m ago
> It’s good at finding and fixing inefficiencies in software

Holy shit! Its happening!

Now if we can the AI to understand this *implicitly* so that it doesn't need to be stated upfront, we might be able to undo years of "premature optimization is the root of all evil".

ApolloFortyNine•38m ago
>Because Opus 5.5 is comparable to Claude Mythos 5.1 in biology and cybersecurity, we’re deploying it with safeguards similar to those on Claude Fable 5.1. Vetted organizations can apply today to our Life Sciences Verification Program to use Opus 5.5 for biology research. In the coming weeks we will also be expanding access to our Cyber Verification Program, and verified cybersecurity practitioners will be able to use Opus 5.5 for their work.

Ah, they're spreading their limits to all their models it seems. Definitely not a good thing long term in my opinion.

prettyblocks•37m ago
They're pushing their customers to their own competition by doing this.
Espressosaurus•13m ago
It’s not like ChatGPT isn’t doing similar. I’ve been hit by cybersecurity strikes before while working on an internal codebase that I had to appeal. Anthropic hasn’t done that to me yet. ChatGPT also regularly does that “thinking for a long time while we check if your chat is rule breaking” thing a lot for me when doing model identification without even interacting with external codebases or services.

The real answer is local instantiations where you don’t have to worry about poorly tuned guardrails screwing you over while you try to work.

Until eventually the Chinese models get good enough/the strategic balance shifts and they start locking everything behind closed weights the same way the US companies are doing.

searine•26m ago
Great. Claude is basically useless for bioinformatics now.
unglaublich•
hirako2000•38m ago
Throwaway accounts posting after a few minutes some anthropic or another ai lab.

Infomercial at its best.

No wonder we are hammered with ai announcements.

danbrooks•30m ago
Many people knew this announcement was coming. The betting markets suggested a very high likelihood of Opus dropping today. I was anticipating this quite a bit!
techjamie•37m ago
With the performance gains they're claiming, I wonder if they implemented the Casual Encoder-Decoder technology from DeepSeek 4.1's paper.

I could see them accomplishing it and seeing gains like this in roughly the correct timeframe, and when I heard about that development I assumed the frontiers would probably jump on it.

How it works: https://miraflow.ai/blog/deepseek-v4-1-flash-causal-encoder-...

ryangg•11m ago
Getting a 403 on that link. Mind checking it once?
potwinkle•8m ago
I'm able to access it on my laptop at home. Maybe a misconfigured bot protection rule, try a different user-agent or IP?
zatkin•4m ago
It's working for me (based out of California).
sailingparrot•36m ago
> Claude Opus 5.5 is our first release since we called for pacing the frontier.

Interesting how the very first line is used to remind the reader of their call to pace the frontier just last week, and everything else after that line is to demonstrate with very specific numbers how they absolutely are not pacing.

dmazin•33m ago
It seems like they are. I mean, this is similar in performance to Fable (ish). It seems like more focus on making existing capabilities more accessible.
the_gipsy•31m ago
Occam's razor: they couldn't make any more substantial improvements.
bpodgursky•30m ago
Everyone knows both labs have internal models which outperform the frontier. All releases are to match market parity and demand for spend, the rest of the compute is used for training. It's not worth arguing about this.
nextaccountic•16m ago
Maybe their internal edge dried up in the last months
someothherguyy•29m ago
> Occam's razor: they couldn't make any more substantial improvements.

doesn't sound like a razor at all

jdw64•36m ago
Finally, it seems like a good time to do some 'load-bearing' work on my project for a while
richardjennings•36m ago
My 20x plan was set to end tomorrow. The writing style and insistence on word vomit just became too annoying. Is Opus 5.5 worth sticking around for ?
glub•35m ago
> For users with cybersecurity use cases that may be blocked by our cyber safeguards, we recommend accessing our models with reduced cyber blocking classifiers via our Cyber Verification Program. Claude Opus 5.5 will be available through this program in the near future.

Anthropic has used "in the near future" for Mythos-class models too, but CVP is still Opus 5 only.

Why even have the program designed for trusted access to cyber capabilities if you're not providing access to cyber capable models via the program?

somewhatjustin•35m ago
> Claude Sonnet 5.5 and Claude Haiku 5.5 will follow in the coming weeks, with many of the same improvements to performance, efficiency, and safety.

Nice. I was starting to think Haiku was going to be abandoned.

phendrenad2•34m ago
> Because Opus 5.5 is comparable to Claude Mythos 5.1 in biology and cybersecurity, we’re deploying it with safeguards similar to those on Claude Fable 5.1

Great so good luck using this for any low-level embedded or operating system development (unless you really, really like Opus 4.8 and want to be greeted by its familiar face after a few minutes of work!)

iamsyr•34m ago
I don't yet have any reason to leave Haiku 4.5 and switch to Opus 5.5.
somewhatjustin•34m ago
> Claude Sonnet 5.5 and Claude Haiku 5.5 will follow in the coming weeks, with many of the same improvements to performance, efficiency, and safety.

Nice. I was starting to think that Haiku got abandoned.

mchusma•30m ago
I hope Haiku is Pareto better than Luna/Deepseek, so slashing its price by about 90%.
Sol-•24m ago
Found this announcement interesting since allegedly OpenAI is retiring their Terra tier. I think for everyday work, two models with various thinking efforts seem enough, plus some frontier level model like Fable or Astra to coordinate.
somewhatjustin•20m ago
I personally use up to 3 models. Fable/Opus for planning, Opus/Sonnet for implementation depending on complexity.

I would maybe use Haiku 5.5 for highly parallel workflows like checking in on MRs or scanning my entire codebase.

cesarvarela•17m ago
Claude code still uses it internally.
mococa•34m ago
I knew it. https://news.ycombinator.com/item?id=49801266
simianwords•33m ago
How do I get access to that reset? I can’t find it in my app.
ricardobeat•33m ago
Great that they listened! The improvement in communication style looks fantastic. Opus 5 was insufferable and I was on the verge of cancelling my subscription.
jidaigeist•31m ago
>Distillation attacks, in which attackers use thousands of fake accounts to extract a model’s capabilities at industrial scale, create safety and national security risks. Distillation allows bad actors to create highly capable models without the safeguards we build into Claude.

Maybe its a bit tiresome to read another comment of the form "what about your large scale distillation attack on the Internet", but this statement really just pisses me off. How very insincere in the most aggravating way.

the_gipsy•29m ago
"bad actors" boogeyman, and we should trust some tech weasel to do the right thing? Yea we've seen who they really are, once they get a sliver of power.
Yabood•31m ago
Current models, especially Opus are almost unusable because they don’t respect instructions and their responses are infuriating. They are clearly designed for token consumption. I find myself wasting a lot of time just asking it to shorten or simplify its responses. I’ll give this new model a go, but I’m not holding my breath because the last model release was supposed to fix the very same issues and it didn’t.
greenavocado•31m ago
Enjoy it for the next 2 weeks until its silently quanted to 4.8 level
jatins•30m ago
> We’ve made major improvements to the way Opus 5.5 writes and communicates, one of the most common areas of feedback we heard about Opus 5. Its messages are much easier to understand at a glance, which testers said helped during long working sessions.

Thank you.

keeeba•29m ago
Opus 5.1 came out about a month ago, what gives?
velcrovan•22m ago
Just a guess but maybe they decided 5.1 wasn't their last Opus model. Like they would keep developing new versions of it or something.
Arcuru•29m ago
Great. Now let me use the subscription outside Claude Code.
zuInnp•27m ago
So it Opus performs as well as Fable what is then the selling point of Fable?

All of this starts to feel more like a drug dealer selling their newest stuff.

In two weeks we probaly get Fable 5.2 with “groundbreaking” improvements, then Astra x+1 etc and then the cycle starts again.

And on the way I always have to check my tooling and need to adjust things to get max results.

ieie3366•17m ago
? they will obviously release Fable 5.5 soon(tm). It's same as hardware. The previously top tier product gets obsolete
glub•11m ago
Don't forget that Opus 5 was tracking fable on many benchmarks, yet it was borderline unusable for any coding work. My Claude sub usage has been 100% fable, 0% opus 5.

Benchmarks often don't survive contact with reality.

quotemstr•11m ago
Big model smell is a real thing. For certain classes of problem, ones you get a feel for but can't easily articulate, a big last-gen model can get you what you're looking for when no quantity of tokens from some ultra-RLed mid-size latest generation model can.
giancarlostoro•5m ago
Fable should have just been called Opus Primt for Enterprise and sold only to enterprise customers. I don't even use it. I rather just use Opus.
aragornii•25m ago
What I'm mostly interest in is the Communication section. Opus 5 was so convoluted in the way of answering that was really frustrating me.

Instead of instilling confidence, it was overwhelming. Not sure if I'm the only one.

m101•25m ago
Funny how they talk so much about safety when most people don’t give a hoot about it, and actually have quite the opposite reaction
nyx•21m ago
People aren't the target audience of that part of the post. They're hoping saying enough safety stuff will ward off the looming regulatory sledgehammer.
pavlov•20m ago
Anthropic is one of the most valuable companies in the world. Their comms are designed to appeal to a very wide readership.

HN is a bubble that's mostly out of touch with what regular people use or care about.

In 2007, HN was convinced that nobody uses Microsoft products. In 2016, it was that Facebook doesn't have any real users and is dying. In 2026, it seems like nobody cares about AI safety and everybody wants to run local models.

notduckrabbit•24m ago
They purport 40% drop in costs due to lower token pricing (presumably aimed at winning back the many of us that switched providers in discovering Opus 5 unusable) and improved token efficiency.
manmal•18m ago
That cost reduction seems to stem from cheaper cache reads, mostly.
bredren•23m ago
Notes on communication:

"Early testers found its writing clearer and easier to follow, which addresses some of the common feedback we heard about Opus 5"

and

"We’ve made major improvements to the way Opus 5.5 writes and communicates, one of the most common areas of feedback we heard about Opus 5."

and

"In our own use, this has made Opus 5.5’s work easier to follow and check—which is a safety benefit as well as a practical one."

I realize it is corporate communications but "most common areas of feedback" and is a bit sterile. If the company wants authenticity and trust its easy to say that they found it hard to follow. And that it did not meet a quality bar they generally expect from their releases.

If this is not true, that it Opus 5 output was generally acceptable and we might see something like that again, that is an important consideration for potential customers or investors.

Retro_Dev•22m ago
> Distillation attacks, in which attackers use thousands of fake accounts to extract a model’s capabilities at industrial scale, create safety and national security risks. Distillation allows bad actors to create highly capable models without the safeguards we build into Claude. Our September 2026 threat intelligence report details the illicit distillation activity we’ve detected and disrupted so far.

Such a negative tone they put on this. Distillation is amazing, because it means anthropic and openai fail to keep a monopoly. Who even are they who claim it's unethical? If it is truly unethical, then so is the mass data scraping they do on my personal website on a regular basis (without my consent), and all the unauthorized use of content produced by authors, blog writers, wikipedia contributors, and creators everywhere. If it is truly unethical, then anthropic, openai, meta, google... all these companies should have deleted their LLMs long ago. This wording disgusts me.

Heck, it would be amazing if we had more models without guardrails - some of the models that are produced via heretic[1] are actually quite nice to use - in particular, I've enjoyed investigating Chinese censorship by interacting with an abliterated model of Qwen3.8-27b. If security is really a concern, then secure your systems - don't attempt to dumb-down the tools we use. If someone breaks your window, then they are responsible, not the hammer they use to do so.

[1]: https://github.com/p-e-w/heretic

aurareturn•19m ago
I found myself going back to Fable over and over again. At this point, I’m not sure if I’m just used to its style or it is truly more capable.

I tried Opus 5 and Astra.

ramoz•18m ago
It crushes Fable on benchmarks and even in the blogs "real-world" studies. But... they are communicating like it ~sometimes~ provides Fable intelligence?

A bit confusing, otherwise I would assume this is a complete replacement for Fable across the board??

bitexploder•10m ago
What if the recent Fable intelligence regression was basically just them serving Opus 5.5 until they got it working well?
Foobar8568•18m ago
I have just switched to 5.5. First mistake was stale environment variable, didn't realize it was replaced, "oh my memory had stall data" and that's it. Second one, a powershell command had the wrong syntax. Great for my first two prompts.
rumblefrog•17m ago
I'm glad they specifically called out the prose issue, I was always pinned to Fable 5.1 because I wanted to avoid the unreadableness of other Anthropic models.
dmix•15m ago
> It performs at the level of Claude Fable 5.1 on most work

Fable feels a bit dated after using Astra. Progress is nice (just like Grok 4.7) but I'm looking forward to the next big release as I can't fully commit to Astra atm.

LoganDark•15m ago
> In the coming weeks we will also be expanding access to our Cyber Verification Program, and verified cybersecurity practitioners will be able to use Opus 5.5 for their work.

I was accepted into the CVP a little while ago. Does this mean I'll need to apply again?

anentropic•13m ago
ooh exaggerated film grain
jdthedisciple•11m ago
I dare anyone to convince me the benchmarks are not meaningless.

Wdym Opus 5.5 scores 14.7% higher than GPT Astra for Terminal Bench 4.0?

How would this alleged difference (most likely bs) actually show up in reality?

GPT Astra was literally the best model in the world by a margin until 1 hour ago or so.

enraged_camel•7m ago
Ah, so you didn't read the article.

>> On our benchmarks, Claude Opus 5.5 leads in agentic coding, computer use, and knowledge work. That said, at these levels of capability we’ve found that benchmark margins have become a less reliable guide to real-world differences. In our own use, the gap between Opus 5.5 and Claude Fable 5.1 is narrower than these scores suggest.

garo-pro•11m ago
Finally confirmation that Haiku was not forgotten and will be coming soon, althouhg I find it quite interesting they skipped 5 and directly skip to 5.5 with all models, including Sonnet which is not super old. I suspect they found something breaking that allows to release this. Recently they struggled with keeping up a 50 % weekly limit increase and now they're putting out 30-40% faster and cheaper models even faster, with much more better benchmarks, a limt reset command and five hour limit increase. It seems more like the opposite and as if they never struggled, thus, I very much believe they found something very effective and new.
datadrivenangel•11m ago
But have they made it any better at communicating clearly? I cancelled my personal subscription because Opus is so painful to read.
desmondl•10m ago
I'll have to try 5.5 on my work's Cursor account. If they really solved the communication issues, I might consider moving my personal account from Codex back to Claude Code.
mcintyre1994•7m ago
> Communication. Opus 5.5 communicates more naturally than prior models. Early testers found its writing clearer and easier to follow, which addresses some of the common feedback we heard about Opus 5. It puts the most important information up front, and its style makes it a better work partner over long sessions. As one early tester put it, “it writes the way I do.” In our own use, this has made Opus 5.5’s work easier to follow and check—which is a safety benefit as well as a practical one.

I think this is what I'm most interested in. I mostly moved to Astra because I just can't work all day with the Claude Opus 5/Fable writing style. I've don't think Astra is a better model, but it's the first OpenAI one that seemed good enough for me. Definitely keen to try Opus 5.5 and see if this claim is real.

blurbleblurble•4m ago
It'd better be good, I'm so tired of the shenanigans
Figured it out: you have "reduce motion" enabled in your device's accesibility settings.

Everyone that doesn't gets served some animated bullshit.

mbreese•32m ago
On mobile at least, you have to scroll to get the TOC to appear. Then keep scrolling to actually move off from the hero to see the text.

For a marketing page, it’s not the worst UX I’ve seen, but still slightly annoying.

dionian•35m ago
and hijacking back/forward
iAMkenough•30m ago
Turn on "reduce motion" in your accessibility settings and you get served a sane version.
thebitguru•30m ago
Totally! So unnecessary and annoying.
halyconWays•30m ago
I call it scrollslop
josefresco•13m ago
Hijacking the scroll wheel has existing long before "AI". Many "high end design" websites that want to "tell a story" get woo'd into thinking it's a good idea. It's terrible, and feels like your scroll wheel is stuck in quicksand.
swader999•28m ago
I told my team to smack me upside the head if I ever try to ship something so daft as that.
amluto•16m ago
Claude Opus 5.6 should have a new "UX safety" feature that requires annually-renewed preauthorization to generate webpages that hijack scrolling :)
brandon272•13m ago
I agree. Hopefully Anthropic has fixed Opus' ridiculous communication style so that people - like me - no longer have any kind of weird impulse to imitate it.
drnick1•43m ago
A load-bearing promise.
username_my1•42m ago
I'm genuinely confused what's the relationship between LLMs improvements and them being so incoherent.

and it's not about the verboseness (even though it obviously contributes to the fatigue and loss of focus), I swear the vocabulary of the llms change working on the same task on the same codebase significantly.

I wonder if there are studies around this.

ygouzerh•37m ago
Can it be that now they are getting optimized against benchmarks that are valuing logics, rather than human appreciation? (I am not an expert at all, just an idea)
meric_•27m ago
Remember when OpenAI models loved talking about goblins and whatnot due to the RL?

https://openai.com/index/where-the-goblins-came-from/

Small quirks can quickly add up in posttraining if not caught. Although TBH with how obvious Claude language is, I do feel like this is something Anthropic probably noticed and just assumed people would not care about. Now that people have obviously cared, they're probably actively looking to alleviate it

aray07•38m ago
Opus 5 was just incoherent - curious to see what improvements they have made here. Would love to see some kind of postmortem to better understand how writing styles change from model to model.

I wouldn’t be surprised if Opus 5 was trained on content written by other LLMs

dgroshev•37m ago
I don't think it's substantially different. I just pasted a random chunk of code and asked Opus 5.5 to comment on it:

> The Vercel target is hard-coded. That's common and not wrong, but it's opaque; nobody reading this later will know which Vercel project it belongs to, and if the project is recreated the target changes silently. A comment or a named variable would help.

> Pointing a DNS name at Vercel is only half the job. The domain also has to be added to the project in Vercel's dashboard, otherwise requests will arrive and Vercel will reject them. That step lives outside this code, so it's easy to forget.

> Finally, [CENSORED] existing only in production is slightly odd on the face of it. It may be perfectly deliberate (perhaps a single shared testing tool that only needs one public address), but if you're reviewing this rather than just reading it, that's worth confirming.

It has the same annoying cadence and writing style with slightly less prominent claudisms.

sashank_1509•32m ago
Maybe if we had a single human we talk to 24/7 at scale, we would get annoyed at his cadence and style. You need variety to not pick up on known patterns I assume, which a single model can’t replicate?
dgroshev•24m ago
No, it's just poor writing. Actionable points are buried inside the paragraphs and over-hedged, and one point is completely made up. Compare to a five second rewrite:

* Consider leaving a comment about the hard-coded Vercel target. It's not clear where does it come from.

* [This is just a bullshit point, because the domain is not "added to" Vercel, it's provided by Vercel]

* Are you sure that [CENSORED] is prod-only? The name suggests otherwise. [also, what "if you're reviewing this rather than just reading it" even means?]

roughly
•
20m ago
Which is one of those fun things that didn’t actually exist back when we took it for granted that our fellow person was operating under some kind of moral or ethical framework, which pretty much everyone was until the economists told us that wasn’t rational, because it turns out it’s an evolutionary advantage to operate under an ethical or moral framework because it allows the kind of coordination which facilitates better collective outcomes, which everyone knew until the economists came along to tell us we were wrong and in fact it was rational not to do so and suddenly we had the prisoner’s dilemma.
icrbow•35m ago
If you hit wall, hit it hard.
If you're able to use the OpenAI ecosystem, Luna's price/performance is really good. Almost like "they messed up and accidentally made it too good" good.
enraged_camel•6m ago
We use Haiku 4.5 inside our product. It continues to be absurdly capable for converting natural language to structured JSON based on a set of fairly complex business rules.
fastball•10m ago
It hasn't just fixed it, it has introduced a new paradigm in anti-obscurity.
mikeocool•10m ago
That's the load-bearing seam in this blog post.
neilellis•5m ago
'frontier models' - seriously, it was you and only you!
5m ago
https://artificialanalysis.ai/models/releases/claude-opus-5-...
abtinf•32m ago
I read the page. It seems like a marginal improvement.
onlyrealcuzzo•40m ago
> And the ChatGPT subscription lets me use my own harness, so I can hook it up to exe.dev.

This is news to me. Excited to try it out! Thanks.

nchmy•34m ago
news to me as well. i thought you were forced to use Codex if you wanted their subscription. I completely ignored it because of that. How do we do it?
polalavik•38m ago
ya i've been a gpt hater for a while. almost exclusively used claude up until astra. astra feels like it blows everything out of the water. its fast, correct, organized, and less verbose.
roughly•37m ago
> And the ChatGPT subscription lets me use my own harness, so I can hook it up to exe.dev.

Can you give more details here? This sounds intriguing.

cbg0•11m ago
How about cheaper? Astra is $10 in $50 out, Opus is $4 in $20 out. Even on a subscription you'll get considerably more usage out of Opus.
Syntaf•30m ago
Yeah if anything Opus 5 taught me how little benchmarks mean to the actual real world performance of these models.

"Better" in every sense of the benchmarks and absolutely horrible results in my day-to-day work.

The verbosity, goal post moving, tendency to leave work unfinished, over focusing on unrealistic root causes when debugging, etc... etc...

It was the first time I actually pinned my models back because I just could not work with 5 for the price and performance it gave me. Hoping 5.5 is better this time around....

cbg0•16m ago
I use it frequently with a lot of success on "Medium" effort, it overthinks like crazy on higher levels, but YMMV.
simianwords•32m ago
It’s likely that they have internal benchmarks but they are communicating to people who can only gauge through external benchmarks.
suddenlybananas•28m ago
Why wouldn't they report these benchmarks?
booty•23m ago

    "benchmark margins have become a less 
    reliable guide to real-world differences" 
    sounds like a big problem.
My guesses:

1. Real-world use cases typically involve big, hairy, crufty, tech debt laden codebases and benchmarks do not.

2. AFAIK "success" in a benchmark essentially boils down to "do the tests pass and do we get the right result?" which is something the LLMs have been achieving with ease for a while, except maybe for uber-challenging coding tasks that would be outliers in just about any workplace. Whereas real-world software engineering is usually just a bunch of CRUD... and "success" involves harder to measure dimensions like "maintainability" and "did you overengineer this?" and "how did you cope with a bunch of vague and maybe contradictory business requirements?"

Having said all of that, I have never ever looked inside any of these benchmarks. I'm putting my guesses out here strictly in the tradition of "the quickest way to learn about something is to be wrong about it on the internet."

Espressosaurus•20m ago
Anything with long context quickly gets dominated by cache reads. Especially for interactive sessions I’ve got cache read % between 95% and 98%.
bayesianbot•19m ago
I think gpt 5.6 family also dropped pricing but didn't give any more usage for the subscriptions. Maybe it's a way to silently lower the value given to subscriptions while keeping API pricing competitive
blfr•23m ago
People are paying for Opus 5? Not just burning down tokens left after they enjoyed Fable on the sub? Amazing.
rapfaria•19m ago
My workplace doesn't even offer Fable. And on the sub, I've had a hard time understanding Opus 5, but Fable can deal with it with subagents.

If 5.5 is any better, I might try to do agentic-assisted development instead of just telling fable to delegate

blfr•14m ago
Telling Fable to delegate is agentic development. At least I thought so until reading your comment.
coffeebeqn•21m ago
We haven’t been able to use opus as much as we’d want because it’s been too expensive for general use, price drop is good so I can stop juggling different models and just use this daily unless it has some weird new issues
AJ007•18m ago
It is only a price drop if price * tokens used is less
mcintyre1994•10m ago
They're claiming a drop in token use too, and that it nets to 40% cheaper.
drbscl•6m ago
Unfortunately, they're full of it https://artificialanalysis.ai/models/claude-opus-5-5#token-u...

It does work out to be a similar cost per task though

rahimnathwani•15m ago
"Opus 5.5 requires less compute to serve than Opus 5, and its pricing reflects that."
22m ago
Opus is useless; Mythos access will be granted to companies that are friendly to the government, so the government gets more control over business.
blfr•15m ago
Fable 5.1 addressed an entire security advisory I had that Fable 5 and Opus 5 refused. I think they loosened the leash a little.
bushido•13m ago
One of my favorite things about their safeguards is their own model will utter something which it does not like and then I'll need to reset the conversation.

The safeguards really don't work well for a lot of long-running tasks on old code bases. A lot of my workloads last days to weeks and the single biggest risk to the workflow is random safeguards.

KeplerBoy•5m ago
Anything else would be inconsistent, wouldn't it?
dgellow•13m ago
A razor is a philosophical tool to help decide between options, in the case of Occam it’s a way to decide for something in a situation where multiple options have more or less the same level of plausibility to en your current knowledge. It’s a heuristic to make a “cut”. What are you shaving off?
cab648bec139cc•29m ago
Do you guys still believe any of their lies? You are getting trolled for years by now and yet you still believe what they tell you?
sleazebreeze•28m ago
What do you think is happening?
re-thc•25m ago
IPO soon
cab648bec139cc•24m ago
I have no idea what is happening. I just know that programmers will not be replaced in 6-18 months (tm).
0xbadcafebee•17m ago
Maybe not replaced exactly but they won't be manually typing out lines of code anymore. I haven't written a line of code in like 6 months. I review PRs, write prompts and tickets, check CI output, and get frustrated when the magical code machine stops working or I run over token budget
anthonyrstevens•12m ago
Who said that, why do you take their word as the literal truth, and most importantly, what does this have to do with a focused discussion of Opus 5.5?
supern0va•23m ago
That's a great point, five minute old account.
sailingparrot•28m ago
Fable 5.1 came out just 21 days ago. Only 3 weeks! And this is 20% relative improvement on terminal bench vs Fable 5.1 at less than half the price, and more human sounding output. does not feel paced to me tbh.
jr3592•24m ago
What exactly is "paced" in this context?
sailingparrot•22m ago
It’s the famous “flattening the curve” from COVID. But for LLMs. This release is not flattening anything.
jr3592•8m ago
I guess I understand why we'd want to flatten a COVID curve, but why do people want to flatten the LLM development curve? Don't we want the opposite? Isn't the goal AGI?
lantry•5m ago
Well, there's a tension because, depending on who you ask, AGI is how you cure cancer and achieve utopia, but also how you kill all life on earth and turn the solar system into paperclips
davrosthedalek•16m ago
Well, I guess it's "fast paced".
dmix•15m ago
Fable 5.1 wasn't that much different than Fable 5 though.
ramoz•17m ago
> similar in performance to Fable (ish).

What insights do you have? Because the blog/benchmarks don't imply any "ish" ... this is crushing Fable across the board. The only nuance to this is how the employees are saying what you're saying: "similar performance" but again what does this mean vs what is presented to us?

BatmansMom•32m ago
kinda disingenuous. They include a whole section on pacing later on
sailingparrot•25m ago
You mean the section where they tell us this model is not affected by pacing because “they understand it well” and they will share more details on pacing later? Yea not very convinced by this effort.
scottyah•30m ago
Seems like a bigger focus on efficiency (both cost and speed) and the "tone" of Claude vs benchmarkmaxxing
CodingJeebus•30m ago
It's laughable at this point. It feels like they're drumming up all this fear about imminent AI threats to emphasize the need to slow down, when in reality, the model progress seems already to be slowing down and has shifted to compute allocation (i.e. "how much compute do you want to throw at this prompt?"). All while continuing to tout benchmark records with each new release.
lukewarm707•27m ago
the only thing they are pacing is what models the permanent underclass are allowed to have in life.

that, they fully intend to 'pace'. hiring accenture is a good sign they need some justification theatre and a fall guy for this decision.

user3939382•19m ago
Or they’re running into a steep diminishing return slope on R&D vs performance and are using stewardship as a cover.
azan_•15m ago
Either AI will capture so much value that there will be permanent underclass (and in this case it's extremely capable and extremely dangerous and should be heavily regulated) or it won't be capable enough to displace people into permanent underclass.
drnick1•9m ago
With AI tools it's easier than ever to create a business, do research, or build stuff. That's an opportunity for the "underclass," not a curse.
kadushka•26m ago
This makes perfect sense. There are no real improvements anymore (just benchmaxxing), and they explain it by "pacing the frontier".
dr0idattack•23m ago
a 1 minute mile pace
mukmuk•21m ago
“Pacing the frontier” sounds smarmy and weird, like the phrase was generated by Claude itself
DiggyJohnson•14m ago
I really don't think it's productive for internet forums to constantly be criticizing language choice when the meaning is clear. Better to respond to the substance of the issue than word choice.
post-it•12m ago
Is the meaning clear? Nobody would use "pacing" in this way. I only know what it means because I've seen previous press releases; if someone told me they wanted to pace the frontier I would have no idea what they mean.

I'm on the fence about calling out AI-isms but I think it's definitely worthwhile to call out ones that actually don't make sense.

sigmar•7m ago
It makes sense to me. If you 'pace your running', you're setting the speed intentionally. The phrasing doesn't describe whether it is fast pace or a slow pace, but it describes having a goal and not just winging it.
vmnb•9m ago
people are here because they are sick of being productive
kadushka•7m ago
We are being productive here!
isoprophlex•7m ago
You're really verbing the noun on the discourse here, belt and suspenders-style
Dumblydorr•9m ago
What specifically is smarmy and weird? Sounds like your own hot take with zero analysis.

They’re limiting frontier model development speed. Others are too. Pacing is the only word here to criticize, and I think it’s fine given the limiting of speed but also increased oversight. I’m not saying they’re fully doing this, but the term is fine.

Do you have a better proposed phrase?

gradus_ad•9m ago
Agreed when I first heard the phrase it sounded odd. Maybe they thought it subtly conveyed they would be setting the pace... But again this is something AI would come up with in its awkwardly post hoc sort of way.
Iolaum•20m ago
They are advertising the regulations they want to enforce in the following sentence, which makes their intentions explicit (ie apply those things made to suit us to our competitors).