I've seen dozens of conversations about it in last 24 hours, and every major inference provided added in first 24 hours. I think it's gaining plenty of traction.
z-ai/glm-5.3: also Z.ai, Novita, Atlas Cloud, IO.NET
Assuming you’re willing to drop a fat stack of cash on the upcoming Mac m5 ultra with 512 gb unified memory, you can even run it locally, quantized to 4 bit. Whether it’s even slightly reasonable, well, my wife would probably skin me alive but maybe yours is more understanding.
I hate how difficult it is to compare prices when looking at subscriptions.
Would $20 in open router, using models like GLM get me more or less?
Z.ai does have their own subscription, but I haven't used it because their privacy policy was pretty buns last time I checked.
What's the point of publishing it when it'll likely be outclassed by gpt-oss?
I cannot stand using gpt-oss, but I miss some of the creative spark of GPT-3 davinci dearly.
We're nowhere near a Fable-class model IMO, but things are going to get interesting in this next year.
Not really, in that you just work with different constraints.
Anthropic and US labs in general has maybe 100s to 1000s of GPUs per person to experiment. Zai and Chinese labs in general have 1-10.
The priorities are different.
There are other providers with much faster inference, like BaseTen at >100t/s: https://openrouter.ai/z-ai/glm-5.3-flash#performance
- It is trivial to extract samples of the training data that was used, which can bolster existing lawsuits/foster new ones.
- Older models are not as safety-hardened, so it is easier to coax unsafe behaviour out of them, which is a PR risk.
- It may be possible to divulge proprietary secrets from the model (e.g. architectural details that may still be relevant still).
For these reasons, and more, it's unlikely that GPT-3/similar models will be released until these concerns are no longer relevant (e.g. when they become a purely historic concern, similar to the open-sourcing of other proprietary software from decades ago).
scosman•1h ago
jonplackett•1h ago
scosman•1h ago
MaxikCZ•1h ago
You implying its better than opus 5?
johnnyApplePRNG•1h ago
If it's significantly larger than GLM 5.3 (I've heard some insane guesstimates out there like upwards of 5T params or more), that would prove rather embarrassing for Anthropic.
scosman•1h ago
walrus01•49m ago
jasonjmcghee•59m ago
I feel like most benchmarks cluster on a reasonably limited area of human knowledge
everforward•35m ago
You do pay for the tokens, but in theory on a smaller model each token is cheaper.
re-thc•
mlnj•58m ago
amelius•30m ago
InsideOutSanta•20m ago
a012•15m ago