I've seen dozens of conversations about it in last 24 hours, and every major inference provided added in first 24 hours. I think it's gaining plenty of traction.
z-ai/glm-5.3: Z.ai, Novita, Atlas Cloud, IO.NET
Assuming you’re willing to drop a fat stack of cash on the upcoming Mac m5 ultra with 512 gb unified memory, you can even run it locally, quantized to 4 bit. Whether it’s even slightly reasonable, well, my wife would probably skin me alive but maybe yours is more understanding.
I hate how difficult it is to compare prices when looking at subscriptions.
Would $20 in open router, using models like GLM get me more or less?
There are other providers with much faster inference, like BaseTen at >100t/s: https://openrouter.ai/z-ai/glm-5.3-flash#performance
scosman•39m ago
jonplackett•35m ago
scosman•33m ago
MaxikCZ•34m ago
You implying its better than opus 5?
johnnyApplePRNG•33m ago
If it's significantly larger than GLM 5.3 (I've heard some insane guesstimates out there like upwards of 5T params or more), that would prove rather embarrassing for Anthropic.
scosman•29m ago
walrus01•17m ago
jasonjmcghee•27m ago
I feel like most benchmarks cluster on a reasonably limited area of human knowledge
re-thc•8m ago
Not really, in that you just work with different constraints.
Anthropic and US labs in general has maybe 100s to 1000s of GPUs per person to experiment. Zai and Chinese labs in general have 1-10.
The priorities are different.
mlnj•26m ago