frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Livenerf: Has Opus 5.5 been nerfed yet?

https://github.com/ninjahawk/livenerf
94•bryan0•1h ago•56 comments

U.S. postal inspectors shut down website selling counterfeit postage labels

https://postalemployeenetwork.com/news/2026/09/26/u-s-postal-inspectors-shut-down-website-selling...
152•ilamont•4h ago•79 comments

GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price

https://openai.com/index/introducing-gpt-6-1-sol/
747•crorella•6h ago•685 comments

Show HN: A working 3D model of an Enigma machine

https://enigma.design
17•primitivesuave•6h ago•1 comments

Vermont replacing power plants with home batteries

https://www.bbc.com/future/article/20260928-a-virtual-power-plant-hidden-in-vermont-homes-is-keep...
36•devonnull•5h ago•21 comments

Show HN: Real-time Solar System with 526k asteroids and all tracked satellites

https://space.bl2.net/
82•wanick•4h ago•23 comments

Needed 1+1, built a functional programming language

https://hereticpleb.vercel.app/blog/needed-one-plus-one/
10•birdculture•7h ago•0 comments

How our vibe coded website looks like a designer made it

https://railcode.dev/blog/vibe-coded-website
32•yakkomajuri•1h ago•21 comments

We’re forgetting what darkness feels like

https://www.theguardian.com/environment/2026/sep/29/night-sky-darkness-city-regulation
11•pseudolus•5h ago•6 comments

NAND-16: a computer built from 277,248 NAND gates

https://somethingbig.ai/computer
105•rossant•2d ago•54 comments

America.gov

https://america.gov/
293•plesiv•9h ago•237 comments

UnoDOS

https://github.com/hmofet/unodos
11•nicoloren•1h ago•6 comments

How Delhi cut electricity loss from 50 to 5 percent

https://spectrum.ieee.org/delhi-electricity-loss
421•rbanffy•11h ago•247 comments

PS5 Relapse Exploit

https://github.com/ntfargo/Relapse-Exploit
220•therepanic•8h ago•119 comments

NASA asked several former SR-71A staffers to help secret restart

https://aviationweek.com/defense/aircraft-propulsion/nasa-asked-several-former-sr-71a-staffers-he...
7•ilamont•13h ago•3 comments

When oil prices spike, where does the money go?

https://theconversation.com/when-oil-prices-spike-where-does-the-money-go-280763
12•thelastgallon•19h ago•2 comments

A Privacy Analysis of Web and Mobile Conversational AI Agents [pdf]

https://jorgegarciaherrero.com/wp-content/interactivos/20260916-Prompt-like-a-butterfly-sting-lik...
406•damaru2•14h ago•128 comments

Language models for text classification: From bag-of-words to Jev

https://magazine.sebastianraschka.com/p/classifier-history-and-jev
13•Anon84•12h ago•0 comments

Backblaze drive stats for Q2 2026

https://www.backblaze.com/blog/backblaze-drive-stats-for-q2-2026/
49•HieronymusBosch•10h ago•3 comments

Phyllotaxis: An audio-reactive LED display

https://jagi.studio/posts/phyllotaxis/
257•evakhoury•1d ago•43 comments

Tcl/Tk 9.1

https://www.tcl-lang.org/software/tcltk/9.1.html
229•dmux•6h ago•78 comments

Show HN: Corral – Kill every command your agent starts

https://github.com/Cardinal44/corral
10•CG144•23h ago•1 comments

Virus Stole a Human Gene and Won't Let Go of It

https://www.nytimes.com/2026/09/28/science/virus-molluscum-human-gene.html
79•gumby•1d ago•28 comments

Ask HN: What are you reading?

119•dan-bailey•10h ago•302 comments

Commodore 64: Mercenary

https://gamesexplained.com/c64/mercenary/
19•a1r•3d ago•4 comments

Show HN: NSL – WSL for Linux

https://frostyard.github.io/nsl/
74•bketelsen•9h ago•58 comments

Dots: Always-on agents

https://openai.com/index/introducing-dots/
444•alvis•6h ago•337 comments

Stuck in the Suez Canal – the short version (2021)

https://cathsenker.co.uk/stuck-in-the-suez-canal-the-short-version/
33•joebig•2d ago•2 comments

Deser: Rethinking Rust Serialization

https://lucumr.pocoo.org/2026/9/29/deser/
12•tosh•2h ago•3 comments

A Staff Engineer's Guide to Inventing Work

https://sujithjay.com/inventing-work
175•amortize•1d ago•35 comments
Open in hackernews

Livenerf: Has Opus 5.5 been nerfed yet?

https://github.com/ninjahawk/livenerf
92•bryan0•1h ago

Comments

gigatexal•58m ago
This is genius. I’m so worried opus 5.5 will get nerfed cuz sonnet 5 was such trash I can’t go back.
aabhay•54m ago
Only ten day interval? I felt Astra got nerfed within a week
refulgentis•47m ago
It's always around a week

(which has led me to believe that's a good approximation for hedonic adaptation, I've seen tons of attempts at demonstrating nerfing via benches, none persist)

aloukissas•47m ago
already getting poorer results today
judge2020•45m ago
I wonder if more organizations approving the model on a fast-tracked basis means Anthropic is straining for more compute and thus sheds a tiny bit to handle the increased demand, especially at peak times.
solenoid0937•41m ago
Hot take, none of the models are getting "nerfed", people are just getting used to the new level of intelligence.
solfox•40m ago
No, whether or not it's intentional, maybe can be debated. But there's definitely an experience of a model losing horsepower quickly after launch.
nba456_•39m ago
No there isn't.
omani•37m ago
who is paying you to say that?
nba456_•31m ago
Mr. Dario himself
AnimalMuppet•35m ago
The claim was that there's an experience of a model losing power. Your claim amounts to "No, you are not experiencing what you say". That's quite a claim for you to make with no data and no argument.
voiceeh•35m ago
You mean to write YOU haven't experienced this. Many have indeed experienced this.
solfox•34m ago
It seems as if this is based on demand. Whenever a new model is released, I'm guessing tens of thousands of us switch over to try the latest and greatest, which overloads the servers, leading to nerfing. It's 100% dishonest, but they realized they would lose users a lot quicker if they were honest and just said "our models are overloaded, come back later".

After Fable launch I switched over to Codex and it was simply amazing, with frequent usage resets that seemed never ending. They clearly had more compute than they knew what to do with. Post Astra, Codex has gotten dumb again across all models, increased usage for no real reason, and no resets.

I'm guessing Opus 5.5 will take the heat off Codex for a bit, leading to better performance. So I guess I stick around here instead of switching again?

whs•33m ago
I wonder if API is affected by this issue, especially Claude on public clouds? Would that means the subsidized rate just means they use cheaper quantized models and it's not comparable to API spending.
madeofpalk•27m ago
I've always used Enterprise per-token billing for Claude Code and I've never understood these nerf complaints. I've never noticed any slow downs at certain times of day, or a gradual decline in quality.
ENGNR•21m ago
There’s probably contractual guarantees in the enterprise plans. My understanding of the subscriptions is they can swap the models out if any of them is getting too heavily loaded for a period of time
Computer0•15m ago
Open AI aims to have a stable API and admits to meddling with effort levels and such for subscriptions here: https://news.ycombinator.com/item?id=49804316#49809266
gr_norm•32m ago
All this dishonesty and shadiness is part of why open models feel inevitable. Even if the total cost of ownership is higher (debatable; seems that way at small scales, but likely not as you grow), I'd rather have intelligence controlled by me that works for me.

The current period is as pro-customer as we're ever going to get, with cash still flying around and neither OpenAI nor Anthropic on the public market, and people are already forced into this sort of business to keep them true to their word. The point isn't even whether they're nerfing the models (I don't think they are), but that people can't seem to trust them to do right.

nico•31m ago
Anecdata: I've been running a long-lived claude code session with Opus 4.6 for the last few days. Yesterday, almost right after the Sonnet 5.5 announcement, codex starting asking for permission to run things a lot more often

The quality of the output/work seems the same, but the speed at which it gets stuff done is a lot slower, because it's asking for permission so much more

I don't have any numbers/stats, just my impression. However, I imagine that if Anthropic could make the models ask for permission more often, it could be an interesting way to throttle access, without degrading quality of the output

Computer0•18m ago
I think most people are on 'auto' mode nowadays.
pkaye•5m ago
I just use auto mode but there is also some config settings for more fine grain control. The model could even help you customize them.
LeoPanthera•28m ago
n=1 is useless. The output is not deterministic.
colordrops•28m ago
This repo already has too much visibility now. Anthropic will soon benchmaxx it.
Razengan•27m ago
Theory (Conjecture? Hypothesis?): What we notice as "model nerfing" is the company diverting compute to training/running new unreleased models..

Remember that some people get access to the next flagship version long before us peasants do. I recall seeing the mention of "Astra" more than a month before it was officially announced

jug•26m ago
We also have Nerf Bench:

https://www.bridgebench.ai/nerf-bench

They test it on launch day, then benchmark it against that. A deviation of above 10% is considered a change. They're currently tracking Opus 5.5 and GPT-6 Astra.

This bench famously detected a degradation of Opus 4.6 which Anthropic later blogged about. I personally think people sense nerfs more often than they happen and that it's often about honeymoon effects.

Grimblewald•7m ago
I dunno, I never sense nerfs for local models, but consistently a few months after launch for corpo hosted models, seems odd my internal model for the capacity of a model drifts for anthropic models but not local ones. I've been using LLMs heavily even before ada/babbage/davinci days, and trust my internal calibration over baseless handwavey explanations for why im imagining things, especially when I have data that shows capacity regression on frontier models for tasks, e.g. one shot success at loss, 0 success in 15 attempts once nerf is sensed. Others publish their quantified capability regressions which are also more trust worthy than this kind of handwaving.
octoberfranklin•26m ago
OpenAI will simply set up a classifier to detect if the client is livenerf, and selectively not nerf those requests.

Open models are the endgame.

paradox460•22m ago
So make it so all clients pretend to be livenerf at first. Same as the ol' pretending to be Google UA for free articles
octoberfranklin•7m ago
I'm not talking about HTTP headers.

The classifier is a model; it examines the actual prompt.

They already do this for the safety "guardrails".

johnfn•24m ago
"Nerf"ing models isn't real. Benchmarks like this or the 100 other "let's see if nerfing is real" copies would have shown it by now if it was.

I made a graphic to explain why people feel like the models get nerfed:

https://x.com/thesilenceturns/status/2103551351825543610

The idea is that new models can handle up to a certain level of complexity, at which point they fall apart. Every new model can handle more complexity, so there's a wonderful time upon release when you feel like you can do anything, only for you to hit the complexity ceiling a few days later when you saturate it. Rinse and repeat for the next model.

hbn•21m ago
I have not been doing increasingly complex things since Opus 4.6 when models got really good.

My work at my job has stayed the same. But the model quality has varied.

They definitely tune the models in production after launch, if not only to share load during high traffic times. It’s not a crazy conspiracy that the same model can be stupider at different times.

Computer0•15m ago
Open AI admits to such here: Open AI aims to have a stable API and admits to meddling with effort levels and such for subscriptions -https://news.ycombinator.com/item?id=49804316#49809266
johnfn•14m ago
It's not about doing more complex things - complexity is more dictated by how large your codebase is, etc.

> It’s not a crazy conspiracy that the same model can be stupider

Sorry, I really do think it's a conspiracy. If nerfing were real, it would be trivial to prove. DeepSWE, SWEBench, and other benchmarks are all available for anyone to run. A "nerfing" hypothesis has to survive the fact that a statistically significant dip in benchmarks has never been observed.

xlayn•17m ago
The only reason why claude fable is better than opus in my opinion is that it has more "criteria"... if you present a problem and then ask for his recommendation you can get an opinion on why and reasoning on why that one... Opus is going to vomit 10k lines of extremely dense prose in nerdify++ level.

Yesterday I fought claude fable to not just jump to make changes like a dog following a treat, that we were researching... at some point I introduced the word HAWAI... and only if I say HAWAI the thing can start making changes..

I was going to post here in HN just to have a "I knew this was the reason" when they release fable > 5.1

I had the exact same feeling every time they have a new big release

bethekidyouwant•7m ago
People just tend towards conspiracies you have to actively fight it.
nba456_•31m ago
No you didn't.
doginasuit•28m ago
I'm not sure "experienced" is carrying much weight here. This is why there are benchmarks, anecdote doesn't mean very much.
Craighead•35m ago
prove it
bradfa•25m ago
Literally the point of the linked GitHub repo.
jyoung8607•25m ago
Has there ever been any measurement of this, of any sort? Honest question. I frequently see a plural of anecdotes to that effect, but I've not seen a concrete statement of fact or measurement that could be scrutinized or tested in any way.

If so, please share. This should be measurable, and I'm glad this project is measuring it.

Answers in the form of additional anecdotes, stated with even greater passion but still lacking a statement that could be tested and falsified, would validate my exact concern.

dude250711•38m ago
Suuuuure...
empath75•35m ago
Yeah people push the models to the limits of what they are capable of almost instantly.
wccrawford•35m ago
I was just wondering if, like certain processors, bugs get fixed and the speed goes down. Like, they find it's doing things it shouldn't, restrict it, and harm the throughput.
raincole•34m ago
It's not a hot take at all. Every benchmark shows that.
jascha_eng•23m ago
Yeh it's absurd that people claim this all the time. It's some crazy conspiracy theory and when you ask for examples nothing ever shows up.

It would be economical suicide from anthropic and OpenAI to actually need models intentionally.

But hey I guess it's hard with technology that truly seems like magic. People say if you'd bring electricity to the middle ages you'd be called a witch and burned. The same is happening to the model labs here because they are bringing tech that the world isn't ready for yet.

frde_me•13m ago
> I have not been doing increasingly complex things since Opus 4.6 when models got really good.

This is a more a statement on the work you do and how you work versus the models. I'm doing more complex work since Fable (and now for way cheaper thanks to Opus 5.5)

With 4.6 I would still babysit a lot more code quality and so on. With the newer model I see myself talking about features at a higher level, and then not having to nitpick PRs to death. Which means most of my time is now spent talking to the model about the product instead of the implementation of the product.

Grimblewald•11m ago
Nerf is real, i think we initially get full precision models and later quants. My own logs show it clearly for opus 4.5 to 5, consistently a few months post launch, models start making quant based mistakes, like slipping in inappropriate tokens (e.g. chinese ones in english text) which doesnt happen at all in the first few months and regularly later. Additionally frontier problems previously done well start being done poorly, until later model variants where performance mostly holds, likely due to them training on your data reguardless of what boxes you tick.

My local models don't display that degradation, sensed or measured. They consistently perform equally to what I expect of them, precisely because they don't change.

How does twitter explain that? Is my internal model for expectation of capacity magically not drifting for local models but somehow is for anthropic api call based models?

gobdovan•8m ago
There are recorded cases of real regressions, but they're better characterised as incidents, not nerfs, e.g.: https://www.anthropic.com/engineering/april-23-postmortem

Btw, you have a typo in the twitter handle on your profile (not in your comment), 'thesilencesturns'.

prodigycorp•7m ago
Incorrect.

Anthropic has admitted to nerfing in the past. There have also been inference bugs. On top of that, model performance changes as they move compute to schwaggier providers as well.

Your chart is wrong.

simonw•5m ago
> Anthropic has admitted to nerfing in the past

Where?

swader999•4m ago
Right, and it would be simple to un-nerf or shadow nerf by any kind of angle they want.