frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

AWS Acquires DuckDB

https://ducklabs.com/news/2026/08/26/ducklabs-to-join-aws
474•onderkalaci•2h ago•108 comments

GLM-5.3-Flash

https://z.ai/blog/glm-5.3-flash
210•Philpax•1h ago•69 comments

Qwen3.8-Flash-Next: A New Architecture, Towards Ultimate Cost-Efficiency

https://qwen.ai/blog?id=qwen3.8-flash-next
269•tosh•2h ago•84 comments

WebMCP: Teaching Your Website to Talk to AI Agents

https://sreenathmenon.com/blog/2026-08-04-webmcp-teaching-websites-to-talk-to-ai-agents/
7•sreenathmenon•7m ago•0 comments

RAG Is Simpler Than You Think

https://www.lighthousenewsletter.com/p/rag-is-simpler-than-you-think
276•j0selit0•6h ago•117 comments

AI Is a Harsh Mistress

https://cacm.acm.org/opinion/ai-is-a-harsh-mistress-on-anima-machina-herd-acceptance-and-the-poli...
24•fscaramuzza•1h ago•3 comments

Taylor Farms: How One Company's Reach Became a National Risk

https://farmaction.us/taylorfarmsreport/
13•speckx•47m ago•0 comments

A curmudgeon tries a language server

https://entropicthoughts.com/curmudgeon-tries-language-server
45•crescit_eundo•2h ago•22 comments

Oldinsurancemaps.net is now a Charter Project

https://openstreetmap.us/news/2026/08/oim-charter-project/
128•altilunium•6h ago•24 comments

Twitter Viewer – View Twitter Without Account

https://twitterwebviewer.com/
50•motownphilly•57m ago•14 comments

Meta reaches $16.68B settlement over social media harms to children

https://www.reuters.com/world/us/meta-settles-with-us-states-over-social-media-harms-2026-08-26/
228•bhouston•1h ago•178 comments

Proliferate (YC S25) Is Hiring

https://www.ycombinator.com/companies/proliferate/jobs/OgpCKYJ-founding-product-engineer
1•pablo24602•3h ago

Nebula Sans

https://www.nebulasans.com
4•GavinAnderegg•5m ago•0 comments

AurionMail: E2EE suite (CryptPad/Stalwart) with single-password UX

https://github.com/AurionMail/docs
4•polo46•8m ago•0 comments

Z.ai confirms Ox Alpha is a new GLM-series model and will release its weights

https://www.bloomberg.com/news/articles/2026-08-26/china-s-z-ai-made-ox-alpha-stealth-model-that-...
337•garo-pro•5h ago•123 comments

Apple introduces M6 and M5 Ultra

https://www.apple.com/newsroom/2026/08/apple-introduces-m6-and-m5-ultra-for-a-big-leap-in-perform...
1260•interpol_p•1d ago•1197 comments

Beyond Recall and the Illusion of Competence

https://var0.xyz/posts/beyond-recall-and-the-illusion-of-competence.html
59•tuxie_•5h ago•19 comments

Stalking the Wily Hacker: 40 years later – Cliff Stoll [video]

https://www.youtube.com/watch?v=656058JxTM0
190•zoenolan•4d ago•59 comments

Value Classes Still Need Compiler Sympathy

https://johan-sjolen.github.io/post/compiler-sympathy/compiler-sympathy/
66•lichtenberger•6h ago•29 comments

FDA authorizes first wearable device that monitors ketone and blood sugar levels

https://www.fda.gov/news-events/press-announcements/fda-authorizes-first-wearable-device-continuo...
468•sunnynagra•20h ago•216 comments

Queryable Executables

https://fzakaria.com/2026/08/24/actually-queryable-executables
275•rguiscard•14h ago•74 comments

OpenAI Jalapeño: Better than Nvidia Blackwell

https://newsletter.semianalysis.com/p/openai-jalapeno-better-than-nvidia
555•bmulholland•1d ago•351 comments

Radiation link in flight attendant's breast cancer, French court finds

https://www.bbc.com/news/articles/cn0j3z6147jo
15•dazhbog•3h ago•1 comments

New Mac Studio with M5 Max and M5 Ultra

https://www.apple.com/newsroom/2026/08/apple-introduces-new-mac-studio-with-m5-max-and-m5-ultra/
809•interpol_p•1d ago•535 comments

Who Makes Your Supplements?

https://www.worseonpurpose.com/p/who-actually-makes-your-supplements
14•speckx•27m ago•4 comments

LibreOffice 26.8 Released with Many Nice Improvements

https://wiki.documentfoundation.org/ReleaseNotes/26.8
42•DemiGuru•2h ago•3 comments

Harvest (IBM 7950): Supercomputer for cryptanalysis at the NSA in the Cold War

https://spectrum.ieee.org/cold-war-codebreaker-nsa-ibm
71•jnord•10h ago•20 comments

Wiped out: US faces surging toilet paper prices amid trade war with Canada

https://www.theguardian.com/us-news/2026/aug/26/paper-product-toilet-paper-tariffs-us-canada
66•mikhael•1h ago•56 comments

Black hole singularity is a surface not a point

https://arxiv.org/abs/2608.21590
289•raattgift•22h ago•201 comments

New Mac mini, featuring M6 and M5 Pro

https://www.apple.com/newsroom/2026/08/apple-unveils-a-more-powerful-mac-mini-featuring-the-all-n...
529•runako•1d ago•327 comments
Open in hackernews

GLM-5.3-Flash

https://z.ai/blog/glm-5.3-flash
203•Philpax•1h ago

Comments

rahimnathwani•51m ago
Related: https://news.ycombinator.com/item?id=49446422

(281 points, 118 comments)

iamsyr•51m ago
Standard API Pricing for GLM-5.3-Flash (per 1M tokens)

- Input: $0.15 - Output: $0.50 - Cached input: $0.03

Xunjin•48m ago
Is that cheaper than DS4 flash?
javier123454321•43m ago
All I can say is that even if it is, I was almost glad to go back to using DS4 Flash. Because 0XAlpha was just so friggin slow to complete a task because of the level of circular reasoning that it would go over and over into, sometimes even returning no output. If I just wanted something done I would switch from a free model to a paid one which is crazy.
denysvitali•41m ago
Tbh it was also slow because it was being hammered by everyone making use of the free tokens
javier123454321•28m ago
Possibly influenced by that, but I believe that is a different issue. I meant the way it processed a request. It went into so many more loops of thinking.
swiftcoder•21m ago
It's even cheaper than DS4's off-peak pricing. Seems like DeepSeek have some stiff competition now
arizen•2m ago
Few weeks ago, I wouldn't expect this statement to be true. Accelerate!
nateb2022•3m ago
https://openrouter.ai/compare/deepseek/deepseek-v4-flash/z-a...
epolanski•50m ago
I'm starting to think that this whole sanctioning China may motivate and prompt them to do more and better in every field.

It's too big, bright and resourceful of a country to choose confrontation instead of collaboration.

esperent•49m ago
This has been clearly stated as what would happen going back several decades at least.
ricardobeat•39m ago
Starting? This was obvious way back in 2019, when the US decided to give China a little push developing their own silicon industry.
himata4113•29m ago
Well the big problem with china is that they do not respect international law when it comes to technology theft. But that argument is very weak when it appears that a lot of what they do is out in the open for anyone to replicate.
nananana9•11m ago
That's how you catch up when you're behind.

Now the US is behind in EVs can you guess what they're doing? [1]

[1] https://evwire.com/p/video-ford-ceo-jim-farley-says-they-fly...

himata4113•
sunbum•50m ago
> with all of this traffic served on Chinese AI chips

RIP Nivida shareholders

ChoosesBarbecue•47m ago
God I wish I could’ve shorted NVIDIA right now
browningstreet•46m ago
It's earnings day for them...
re-thc•40m ago
Which 9/10 times hasn't been great anyway (stock reaction).
Bluestein•8m ago
Of course the release was not coincidental - with the earnings days - I am sure.-
Bluestein•46m ago
This is the takeaway here: That's how they have been serving it at scale as Ox-Alpha. This is a definitional moment.-

Further quote:

"Compared with our initial baseline on the same hardware, we achieved a 3× improvement in end-to-end serving performance, reaching hardware efficiency and per-token cost comparable to mainstream NVIDIA GPUs. This demonstrates that Chinese chips can support frontier-model inference efficiently and economically at scale."

https://z.ai/blog/glm-5.3-flash

mmastrac•48m ago
Weights on HF here: https://huggingface.co/zai-org/GLM-5.3-Flash

I decided to take the plunge and get myself four sparks at a decent price (and bought the QSFP cables from AliExpress because they are literally 1/2 the price of Amazon), even knowing Apple was going to release new hardware and there's probably a spark 2 on the horizon. It looks like this is going to be a decent fit for what I need. I've been experimenting with a two-node DS4 and it's _good_ at some tasks, but it really just spins its wheels when it hits the limit of what it can reason through.

I can offload mundane/basic tasks to DS4 on two sparks, but I've been pushing it harder on some novel work and it just can't run on its own at all beyond a certain complexity level.

I would love to see an Opus-4.8-level local model but TBH I just haven't got there yet. The models I've tried so far _are_ good but they aren't able to solve tough technical challenges, regardless of harness/prompting/etc.

kilroy123•47m ago
> get myself four sparks at a decent price

Wow, if you don't mind me asking. How and where?

mmastrac•44m ago
I bought 4x Asus GX10 with the 1TB option. I don't understand why, but it's the only model in the whole lineup that isn't priced insanely.

They were briefly on sale with a $200-off coupon, but they show up on warehouse deals from time-to-time as well.

swiftcoder•25m ago
> it's the only model in the whole lineup that isn't priced insanely

$4,000 isn't priced insanely? ye gads

packetlost•48m ago
For those who didn't read, this is the identity of the mysterious "Ox Alpha" model
Bluestein•36m ago
They even give this over the API now:

│ https://openrouter.ai/api/v1/chat/completions model: stealth/ox-alpha auth: OPENROUTER_API_KEY status: 404 Not Found response: {"error":{"message":"Thank you for participating in the Stealth Ox Alpha testing period. This model was ZAI's GLM-5.3 Flash.

│ Use it now: https://openrouter.ai/z-ai/glm-5.3-flash","code":404},"user_...":"}

AbsurdCensor•26m ago
Yeah, made me suspicious of how well the Ox Alpha was performing that it wasn't some 'new group' making the model.
revolvingthrow•47m ago
> 320B total parameters and just 18B active parameters

This is pretty hefty for a "flash" model, even a 256 GB setup is insufficient at q4 - and q4 is already the worst-but-still-acceptable quant in my experience. The benchmarks look great, especially since GLM tends to be more honest than the average Chinese lab, but you’ll need to splurge to run it at home.

@edit: so many releases that I forgot to math. This fits just fine in q4, realistically the minimal hardware would be 192gb - so blazing fast on double rtx 6000 pro and usable on 256gb unified memory. You could even go with 5bit quant on 256gb.

… you’ll still need to splurge, though.

colingauvin•44m ago
That's 160GB-ish for Q4...how is 256 insufficient?
dannyw•34m ago
Looks like the M5 Ultra Studio wait times are going to increase again. Already at 10-12 weeks, I wonder how long it'll go?
speedgoose•24m ago
I guess like the M3 Ultra, at some point normal customers won’t be able to buy it.
Destiner•47m ago
from the article, pareto frontier for open source models is completely dominated by GLM now.
montroser•23m ago
Well, it will be interesting to see where Qwen3.8-Flash-Next ends up landing, also released today. These are exciting times!
TaLiTr•45m ago
> it outperforms GLM-5.2 across benchmarks and real-world workloads at one-tenth the price, while approaching Claude Opus 4.8 on coding and agentic benchmarks.

From a biased source, but would be big if true. I've had great results with GLM 5.2.

From their subscription page, the smallest plan gives you about 97M tokens weekly for 5.3 but 292M for 5.3 Flash. Not exactly 10x the limit.

re-thc•40m ago
> From a biased source, but would be big if true. I've had great results with GLM 5.2.

It's at least close (even if not better) from the Ox Alpha runs. For the price it's definitely great.

wolttam•31m ago
The recent and slightly smaller DSv4 Flash is also GLM 5.2 equivalent (or close enough)
mariopt•43m ago
It's only 320B, local frontier AI is getting closer, sooner than expected.
oceansky•31m ago
Can't come soon enough!
garo-pro•37m ago
> Combined with our latest 30T-token multimodal pre-training corpus [...]

Is the optimal formula still 20x the amount of model params in tokens for training? Could this mean we're getting a GLM with 1.5t params?

swingboy•36m ago
How much is the “discounted” pricing they mention?
Imustaskforhelp•34m ago
> To overcome the relatively limited compute and memory capacity of individual chips, we built a dedicated inference engine for this architecture on top of SGLang. Notably, this effort was accelerated by our GLM-5.3-powered infrastructure agent, which assisted engineers in developing and optimizing kernels, diagnosing performance bottlenecks, and improving the serving stack — creating a feedback loop in which the model helped optimize the system serving the model itself.

> (...) Compared with our initial baseline on the same hardware, we achieved a 3× improvement in end-to-end serving performance, reaching hardware efficiency and per-token cost comparable to mainstream NVIDIA GPUs. This demonstrates that Chinese chips can support frontier-model inference efficiently and economically at scale.

It might be one of the most actually practical tasks that AI might've done because the compounding effects of it and also its implications are/feels so immense. It feels as if Nvidia might be in a slight turbulence from it.

toppy•34m ago
By clicking this link you download some PDF in the background
krystofee•33m ago
Its displayed in the html...
kayleykiwi•25m ago
This looks like it goes hard, can't wait to try it
yipinwong•22m ago
When reading this type of announcements, always have keen eyes on graphs.

e.g. "Agent Coding Performance by Effort Level" cuts Y-axis from 0~20.

- This makes it as if GLM-5.3-Flash made a bigger jump than it claimed as the Y-axis does not increase much (stupid trick used in biz reports)

I did mention that ox was working ok for me, and having an open-weight comparable to close to SOTA makes it very compelling for me to try it out locally (well, only if I got more VRAM)

nchmy•14m ago
they also conspicuously omitted GPT 5.6 Luna from comparison. It scores lower, but is also cheaper. MiMo 2.5 is not a valid comp at this point

edit: nevermind. it is there in the artifical analysis scatter plot, but is greyed-out.

MUCH more interesting is that in that chart, their cost is WAY off. The actual chart shows GLM 5.3 Flash at $0.09, but their chart shows $0.045...

cootsnuck•19m ago
If we fast forward say 5 years, I don't see how we don't end up in world where people (and enterprises) are more savvy with how they use LLMs. Meaning, more models, smaller models, weirder models, more specialized models, etc. And all of it running on a variety of hardware (edge devices, personal computers, on-demand cloud compute).

I don't see how NVIDIA can keep their spot as belle of the ball. If LLMs and friends are truly to become as useful and ubiquitous as everyone thinks they will, then commoditization is the only option.

4m ago
"argument is very weak" regardless as I said.
rvz•39m ago
This is no surprise [0] [1].

>> "They are already there on open weight models and Jensen knows that it is only a matter of time until China catches up with GPUs or other AI accelerators."

It is also why Nvidia becoming a bank for other AI companies who are unable to find VCs to fund them isn't really a good thing and that is bearish.

[0] https://news.ycombinator.com/item?id=49397204

[1] https://news.ycombinator.com/item?id=49431231

ThouYS•32m ago
yay, I called it! :) (in the other thread)
dannyw•30m ago
Another self-inflicted own courtesy of US government policy.

While I think China would always get to hardware self-sufficiency eventually, all export controls have done is (1) accelerate China's development, and (2) divert revenue that would've otherwise gone to NVIDIA/AMD/etc instead.

ignoramous•21m ago
The export controls were revoked before it triggered Chinese protectionism: https://www.silicon.co.uk/e-innovation/artificial-intelligen... / https://archive.vn/B2pah
mlinsey•9m ago
Revoked or not, just ever having those controls signals to the Chinese ecosystem that you're not necessarily a reliable supplier (Would you trust US export policy to remain stable for the next ~decade given the state of US politic?) and to the Chinese government just how strategically important you see these components.

This isn't the kind of thing you can hash out in public and go back and forth on. Once you put it out there, the other party will take steps to make sure they don't have to rely on us in the long run.

re-thc•8m ago
> The export controls were revoked before

Zai is on another "export control" list outside the broader 1. Doesn't help.

redox99•10m ago
Not really a brag: it ran like shit. Very slow (~20tps, VERY high latency) and it would timeout all the time.

I'm sure the chips are fine, but they clearly didn't have enough capacity for the demand they had (that 100T/day claim was asbolute bs)

nchmy•8m ago
seems unlikely that they'll get nearly as much demand now that it isnt free
Aurornis•6m ago
Ox Alpha is a smaller model and it was running very slowly. Chinese AI accelerators are coming along, but nVidia’s lead is huge.
a3w•22m ago
I thought 4000 in sum. No wait, 4000 per, plus tax. Or EUR pricing to similar accord. Ouch.
swiftcoder•18m ago
Yeah, that little cluster costs about the same as a brand-new Dacia Sandero.
esafak•10m ago
Yes, but it was $200 off!
swatcoder•3m ago
[delayed]
bmurphy1976•22m ago
~$4000 USD each on Amazon, $175 for the cable.
cmrdporcupine•14m ago
I mean, I have the same machine and the pricing is only what it is because it has that 1TB nVME in it instead of larger. nVME prices are insane and have been for months.

Reality is on a single spark I'm constantly running out of room and it being an odd size M.2 slot it's a pain to upgrade. I'm setting up a NAS over RDMA via ConnectX though, that's fun.

0xbadcafebee•13m ago
If you used the bare API pricing, 1M tokens @ 30% input/70% output/50% cached, you'd pay $0.05805. Even with four discounted sparks, how much are you paying for the same tokens/distribution?
Aurornis•9m ago
> I would love to see an Opus-4.8-level local model but TBH I just haven't got there yet. The models I've tried so far _are_ good but they aren't able to solve tough technical challenges, regardless of harness/prompting/etc.

Agree. It doesn’t even have to be local, using models in this size class through OpenRouter will reveal their limits if you work side by side with Opus level models regularly.

There are a lot of social media posts about people cancelling their Anthropic or ChatGPT subscriptions after installing a local LLM. I’ve used local LLMs a lot and I spend a lot of time with frontier models and the difference is still huge. As far as I can tell, the social media posts about local LLMs replacing frontier models are either wishful thinking, engagement bait, or people who must be working on much simpler projects with a much higher tolerance for slop than I have.

disiplus•4m ago
To be fair, there is no 3 turns that I don't have to jump in into what Opus 5 is doing. There is either some regression or my prompting skills are so much worse now. Flash is not perfect and honestly some things depend on how big context do you keep. So I'm keeping like a really short context with my flash, but it works okay, even though it has a tendency to overthink, and yeah, I run it always in max effort mode.
disiplus•8m ago
I will give it a try, but from the benchmarks it never exceeds the DS4 flash benchmarks by significant margin and And I feel that the throughput that you will get on those machines or what I'm getting with my local hosted flash will be so much worse that it's not worth it.