frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Get your Windows license refund

https://en.refund4freedom.org/
301•smartmic•3h ago•103 comments

Just the rumour of a bug is enough to find an exploit these days

https://anil.recoil.org/notes/rumour-is-the-exploit
35•avsm•53m ago•10 comments

Inception-style curved map for turn-by-turn directions

https://www.orbify.eu/demo/
212•smoser•4h ago•75 comments

U.S. sanctions against the A/I Collective

https://www.inventati.org/
186•exiguus•3h ago•158 comments

GLM-5.3 is now open-weight

https://twitter.com/Zai_org/status/2093354097122455713
188•jeudesprits•1h ago•60 comments

GUIs should be fully keyboard-driven

https://ckardaris.com/blog/2026/08/28/keyboard-driven-guis.html
39•ckardaris•1h ago•30 comments

Run Qwen3.8 27B locally: real numbers from my Mac Studio

https://terminalbytes.com/run-qwen-3-8-27b-locally/
37•speckx•1h ago•23 comments

Htmx 4.0.0

https://four.htmx.org/announcements/2026-08-28-htmx-4.0.0-is-released
96•rmsaksida•3h ago•20 comments

State of the Map 2026

https://2026.stateofthemap.org/
69•lode•3h ago•20 comments

The Twelve-Factor App

https://12factor.net/
75•jxmorris12•18h ago•39 comments

OpenAI: Migrating to HTTPX2

https://github.com/openai/openai-python/blob/main/httpx2.md
142•tosh•5h ago•58 comments

How Dactyl Works

https://dactyl.dev/blog/how-dactyl-works/
28•anorak27•1h ago•1 comments

The conservationists helping to restore Africa’s wild dog populations

https://www.smithsonianmag.com/science-nature/africa-wild-dogs-most-hated-carnivores-continent-he...
22•speckx•2h ago•7 comments

Hilariously Fast Volume Computation with the Divergence Theorem

https://alyssarosenzweig.ca/blog/hilariously-fast-volume-computation-with-the-divergence-theorem....
192•luu•7h ago•47 comments

“It works better in the app”

https://shkspr.mobi/blog/2026/08/it-works-better-in-the-app/
520•blenderob•4h ago•329 comments

Verschlimmbesserung: The Word Your Software Updates Need

https://geekyschmidt.com/post/2026-08-25-verschlimmbesserung/
18•speckx•2h ago•2 comments

Are Crows Our Friends?

https://www.audubon.org/magazine/are-crows-really-our-friends
47•speckx•3h ago•33 comments

Don't use musl if you care about performance

https://blog.brokk.ai/dont-use-musl-if-you-care-about-performance/
18•jbellis•1h ago•6 comments

Lake formed after ice-rock avalanche remains at a high level and is overflowing

https://kathmandupost.com/national/2026/08/28/barrier-lake-continues-to-pose-flood-risk-china-warns
9•r721•1h ago•1 comments

Luanti removed from Google Play due to baseless AI copyright notice

https://blog.luanti.org/2026/08/27/luanti-dmca-tracer-ai/
255•miniBill•10h ago•76 comments

Smaller reactors bring nuclear power closer to fulfilling its promise

https://www.nature.com/articles/d41586-026-02506-4
46•sohkamyung•4h ago•68 comments

httpx2

https://github.com/pydantic/httpx2
68•tosh•5h ago•21 comments

Judge rules Trump administration’s blacklisting of Anthropic was illegal

https://www.nytimes.com/2026/08/27/technology/anthropic-government-blacklisting-ruling.html
277•jbegley•14h ago•248 comments

Debugging my new network, when 10 Gigabit Ethernet Runs at 300 Megabits

https://www.hanselman.com/blog/debugging-my-new-network-when-10-gigabit-ethernet-runs-at-300-mega...
12•speckx•2h ago•2 comments

“Weird” is a weird word

https://www.deadlanguagesociety.com/p/weird-is-a-weird-word
29•pseudolus•3h ago•8 comments

I used AWS cognito for a startup. I wouldn't do it again

https://joshkaramuth.com/blog/aws-cognito-authentication-startup-nightmare/
123•speckx•3h ago•97 comments

Interactive pattern discovery in binaries (FF-16-TUI)

https://github.com/HexLasso/FF-16-TUI
25•HexLasso•3h ago•2 comments

Interactive Warhammer 40k Galaxy Map

https://cartographia40k.com/
103•gbxyz•8h ago•30 comments

Bhartrhari's Paradox

https://www.futilitycloset.com/2026/08/18/bhartrharis-paradox/
30•surprisetalk•4h ago•31 comments

'Recent revelations' prompt Pflugerville to abruptly kill Flock camera access

https://www.mysanantonio.com/news/austin/article/pflugerville-flock-cameras-22404257.php
18•jkestner•3h ago•1 comments
Open in hackernews

GLM-5.3 is now open-weight

https://twitter.com/Zai_org/status/2093354097122455713
184•jeudesprits•1h ago

Comments

scosman•1h ago
I've been using it more and more. Feels like Opus 4.8, in the best possible way.
jonplackett•1h ago
Can you give ant more details how you are you using it? Which harness / service / what you’re building with it etc?
scosman•1h ago
z.ai coder plan, both in opencode and direct API access. I use it for my side projects like https://github.com/scosman/Biscotti (on-device meeting transcription and summaries).
MaxikCZ•1h ago
> in the best possible way

You implying its better than opus 5?

johnnyApplePRNG•1h ago
I'm starting to think Opus 4.8 is significantly smaller than most people assume.

If it's significantly larger than GLM 5.3 (I've heard some insane guesstimates out there like upwards of 5T params or more), that would prove rather embarrassing for Anthropic.

scosman•1h ago
You can't compare models released 6+ months apart. GLM 5.2 was same architecture as 5.3 and not nearly as good. Takes time to build frontier intelligence and distill down to smaller sizes.
walrus01•49m ago
It's not that GLM5.3 in full precision unquantized is any smaller, it's 141 * 5.4GB files at approx 770GB which is about the same size as 5.2.
jasonjmcghee•59m ago
I hear the argument here, but isn't it possible it has dramatically more knowledge and when you get outside the common cases many of us use it for, it'll have completely different capabilities?

I feel like most benchmarks cluster on a reasonably limited area of human knowledge

everforward•35m ago
Sort of depends on how well the core reasoning works. It’s not a big effort to connect an LLM to a search provider.

You do pay for the tokens, but in theory on a smaller model each token is cheaper.

re-thc•
mlnj•58m ago
I have been only using GLM models since last December and have had the best experience without any drama about tokens and geopolitical restrictions. The quality has been great and I am doing more and more with the latest 5.3 and am really excited that consumer hardware will develop in the next few years where I can run these at home.
amelius•30m ago
Do you use it to write HTML/CSS? Javascript? C++? There's a huge difference in ways people use models and if you are not specific about it then your comment means nothing, unfortunately.
InsideOutSanta•20m ago
I really like how it doesn't have that Claude talk. It just does the thing without Claude's "load-bearing honesty." It's probably my favorite model to interact with, even if it isn't the best or most reliable.
a012•15m ago
My second favorite model by now is GLM 5.3 flash which is very capable of day to day task. I use it as the main model and GLM 5.3 for task that is more complex
m00dy•1h ago
GLM-5.3-Flash is actually cheaper than deepseek and better than deepseek but no one is talking about yet :)
scosman•1h ago
It's actually slightly more expensive ($0.50 vs $0.48), but there's a temporary 50% discount.

I've seen dozens of conversations about it in last 24 hours, and every major inference provided added in first 24 hours. I think it's gaining plenty of traction.

swiftcoder•45m ago
It's interesting that OpenCode Go is treating it as 2x more expensive than DeepSeek Flash, even factoring in the 50% discount
re-thc•39m ago
Go has API pricing + this weird scaling of how much is it worth. Some models get $60 of usage, some $30 and some $15 etc.
chillfox•9m ago
OpenCode Go is becoming less of a good deal by the month. I pretty much only use it for mimo 2.5 pro now, and everything else is either ollama or openrouter.
esafak•1h ago
It is a slow for me through z.ai; it does not feel 'flash' at all. But then neither did the new DS Flash. I think they were getting hammered.
fra•1h ago
h/t to DeepInfra for being the first 3rd party provider for it on OpenRouter (https://openrouter.ai/z-ai/glm-5.3?endpoint=b711bea7-3994-49...).
andrewmunsell•46m ago
It's also now live on Ollama Cloud as of a couple minutes ago
ljlolel•37m ago
on my TrustedRouter:

z-ai/glm-5.3: also Z.ai, Novita, Atlas Cloud, IO.NET

stavros•4m ago
Have you guys been having a good experience with OpenRouter? I tried it out recently with Claude, and it cached no tokens, charging me $200 for one conversation of 11 messages.
revolvingthrow•58m ago
GLM 5.3 is probably the sweet spot open weights model if you want to go beyond deepseek flash or the new glm flash. I used it with pi and had a fairly good time, especially since it’s less touchy about cyber and whatnot than the US guys. It’s slightly behind Kimi in ability but it’s a lot easier to run it, I’d expect prices (and speed!) from third parties to be noticeably better.

Assuming you’re willing to drop a fat stack of cash on the upcoming Mac m5 ultra with 512 gb unified memory, you can even run it locally, quantized to 4 bit. Whether it’s even slightly reasonable, well, my wife would probably skin me alive but maybe yours is more understanding.

walrus01•47m ago
One could also run it locally on a used dual xeon (or amd-equivalent) server with 512GB RAM, albeit slower, if you have a useful workflow for it that's like "take this day's efforts and run it through various analysis agents", combined with giving it one-shot tasks/modules to build overnight. You would want a place like a garage or basement to put the server because it'll be loud.
dataplumb3r•17m ago
You'd also likely spend far more in electricity than the API cost of processing the prompt(s)
walrus01•9m ago
yes, though for some uses, not sending data anywhere to third parties has its own value which is harder to measure.
0xdeadbeefbabe•33m ago
Well if you did get the m5 ultra could you obliterate the guardrails and then your wife can ask it pertinent but unsafe questions about how to punish you. Seems doable.
pal9000i•55m ago
how feasible is it build a SOTA specialized model for some use case e.g. deal sourcing by using this as pre-trained model or a LORA or similar pattern on top? Gonna shoot my shot at a billion dollar business
ChildOfChaos•43m ago
How much usage do you find you get on these kinda models (I know the pricing changes a bit) compared to a $20 sub say for Google AI Pro in anti gravity?

I hate how difficult it is to compare prices when looking at subscriptions.

Would $20 in open router, using models like GLM get me more or less?

nozzlegear•13m ago
I think it'd get you less than a $20 sub to any of the big three. I've used it on OpenRouter and found it kind of expensive for the results, but that might change now that it's open weight and other providers can host it/compete with Z.ai. For the work I did with it, I would've rather used DeepSeek V4 Flash just because it's more economical and still gives good results IMO.

Z.ai does have their own subscription, but I haven't used it because their privacy policy was pretty buns last time I checked.

nkmnz•42m ago
I'd like to ask Sam Altman if he still thinks that it's too dangerous to publish GPT-3. I mean, no one would use it, but what is his reasoning for not publishing it now, in 2026?
gruez•30m ago
>but what is his reasoning for not publishing it now, in 2026?

What's the point of publishing it when it'll likely be outclassed by gpt-oss?

paxys•30m ago
They already publish gpt-oss which is several generations better than gpt-3
bigyabai•20m ago
At release, GPT-OSS was arguably a few generations behind the open frontier.
Philpax•11m ago
No? They were the frontier, or near it, at the time of release: https://artificialanalysis.ai/models/releases/gpt-oss-120b
cptcobalt•11m ago
GPT-3 is a different model than gpt-oss and is therefore not an answer to the question.

I cannot stand using gpt-oss, but I miss some of the creative spark of GPT-3 davinci dearly.

chillfox•41m ago
Seeing the price, I am probably just going to keep using GLM-5.2 until 5.3 gets cheaper.
bel8•18m ago
isn't GLM 5.3 flash better than GLM 5.2 overall?
mmastrac•34m ago
I previously posted that DS4Flash was _good_ but not _great_ on two DGX Sparks, but I have to say that GLM-5.3 is pretty amazing. It's been able to tackle all the random hard problems I've thrown at it and it has the intuition that DS4Flash seems to lack.

We're nowhere near a Fable-class model IMO, but things are going to get interesting in this next year.

villish•5m ago
What quant are you running and tps?
hkalbasi•28m ago
Is it possible to fine tune this model and unlock / extend its cybersecurity capabilities? I'm scared that maybe we are not ready for an open-weight model with high cybersecurity skills.
milkshakes•17m ago
brace yourself
40m ago
> that would prove rather embarrassing for Anthropic

Not really, in that you just work with different constraints.

Anthropic and US labs in general has maybe 100s to 1000s of GPUs per person to experiment. Zai and Chinese labs in general have 1-10.

The priorities are different.

nkmnz•34m ago
It seems like there is tradeoff between model size and the need for tool use, which - in my mind - is quite costly in terms of time and tokens. More detailed world knowledge requires an exponential increase in model size, but most knowledge can be acquired ad hoc using search or database queries. This will fail for questions where the model lacks the knowledge to ask the right questions, but maybe this could be solved by a handful small inquiry models with different knowledge encoded in their weights?
JoeLee1991•1h ago
I've been using it quite a bit too. My main complaint is that it can be really slow sometimes — like, really slow — and the speed feels pretty inconsistent.
scosman•59m ago
z.ai is using all Chinese hardware for flash: https://thenewstack.io/glm-5-3-flash-chinese-chips/

There are other providers with much faster inference, like BaseTen at >100t/s: https://openrouter.ai/z-ai/glm-5.3-flash#performance

malshe•45m ago
How do I find out where the openrouter model providers' servers are located?
RussianCow•32m ago
If you click on the provider name, the panel that pops up shows a "Region" value. Not every provider lists their region, however.
malshe•18m ago
I think the region is just the HQ of the provider. So z.ai's region is Singapore but it's quite likely that their servers are actually in China
_aavaa_•1h ago
It’s cheaper sure, but it’s very slow. It’s not a drop in replacement
natrys•29m ago
I think we don't have a good draft model for better speculative decoding yet (e.g. DFlash 2). Once we do, it will be faster.
_aavaa_•12m ago
It very well could be faster, but right now it isn’t.
lnenad•22m ago
I have just built an Epyc with 512gb DDR4 3200 RAM for a "reasonable" price and I'm hoping to have a setup with GLM as the architect and Qwen 27b/Next Flash as the implementer. This is 1/5 of the price of the Mac, but also probably 1/5 of the speed lol.
jchw•16m ago
Honestly I suspect neither of them will be performing terribly well but with DDR4 3200 RAM I wonder if you'll be counting tokens per second or seconds per token. I mean, you do at least get a lot of memory channels at least, compared to consumer PCs. I am curious to hear what performance you get, I feel there is not enough information out there on what different setups manage to eek out.
Philpax•14m ago
The fastest I was able to get my Threadripper 3960X + 2x 3090s + 256GB DDR4-3200 to run a 2-bit quant of GLM-5.2 was 8 TPS. I would expect to be in seconds-per-token territory for a pure-CPU 4-bit quant.
0x457•9m ago
Depending on which Epyc you got it might be slower than 1/5 of the speed.
futureshock•29m ago
I think it would be an important historical document as well. We are potentially looking at the dawn of AGI and one of the most important models ever created. Each model is also a kind of ultimate time capsule, containing a snapshot of the entire human collective mind. If you wanted to ask a 2002 person what they thought about future historical events you can just ask them directly.
Philpax•5m ago
There are risks associated with releasing historical proprietary models that were not designed for open release:

- It is trivial to extract samples of the training data that was used, which can bolster existing lawsuits/foster new ones.

- Older models are not as safety-hardened, so it is easier to coax unsafe behaviour out of them, which is a PR risk.

- It may be possible to divulge proprietary secrets from the model (e.g. architectural details that may still be relevant still).

For these reasons, and more, it's unlikely that GPT-3/similar models will be released until these concerns are no longer relevant (e.g. when they become a purely historic concern, similar to the open-sourcing of other proprietary software from decades ago).