This is napkin math since I'm mostly just extrapolating from glm 5.2 by assuming it's twice as heavy to serve in every single measurement, but I believe you can easily achieve 2500tok/s aggregate compared to 4500tok/s and up to 8000tok/s for glm5.2.
with nvidia r100 you are likely going to be able to push that number even higher while the cost of hardware appears to be relatively the same, so far I am seeing 21% premium from supermicro which is twice as fast and has nearly twice the vram.
The model weights are supposed to release tomorrow.
Over the next several weeks, I would expect competition among open weight providers to drive down the cost, as I've seen happen with other open weight model releases.
I don't mean to imply that Kimi is not at all cheaper than U.S frontier models. I more wrote this because I believe - since Chinese LLMs entered the public consciousness via DeepSeek R1, which was genuinely ~20x cheaper than o1 - there's a bit of a halo effect around Chinese models which causes people to overestimate the scale of the discount. And relative to that price anchor, Kimi is less extraordinarily cheap.
At the moment Kimi is ~10% cheaper than GPT-5.6 on the AA benchmark, and as you say that could go down to 20-30% cheaper (although I don't know how inference provider discounts play out on real world usage once you account for quantisation etc...). I'm not trying to suggest that that's nothing, but I do think some of the people driving the Chinese AI discourse would have a harder time pitching their conclusions if they were saying "this new Chinese model is 10% cheaper on some tasks, and it might get another 20% cheaper in the future".
GPT-5.6-Sol is pretty competitively priced, but not all American frontier models are, and even 10% to 30% is still significant for any commodity that's as fungible as frontier models often are.
> as you say that could go down to 20-30% cheaper
I never said anything about 20% to 30%. We don't know how much it actually costs to host this model yet, and that will determine the final price. It could be just a little less, or it could be a lot less.
> once you account for quantisation
There will be no need to account for quantization. Kimi models have been 4-bit only since at least K2.5. They don't release or serve models in higher precision than that. This isn't one of those situations where LLM inference providers are debating between serving 16-bit, 8-bit, or 4-bit, and I have never seen a publicly hosted, paid model that was hosted in less than 4-bit, even if hobbyists will use sub-4-bit quantizations sometimes locally.
If you switch the view to "coding tasks" on this website:
Kimi K3: $3.18 per task
GLM 5.2: $6.51 per task
GPT 5.6 Sol: $7.02 per task
Opus 5: 8.23 per task
Fable: 11.70 per task
So it's pretty dang cheap lol. Nobody is using frontier inference for "office tasks".Additionally, I fully expect the frontier labs to continue increasing prices to meet the profit margins they need to to continue existing.
General office work is one of the big frontiers the labs are pushing on, and it's part of how they're justifying the value proposition to enterprise customers. It's also accounts for a big portion of the spend on RL; tasks/environments designed to train agents to navigate Slack or Salesforce. If you're Anthropic pitching Claude to a bank (taking an example I'm familiar with), coding probably accounts for ~20% tops of the workforce, and it doesn't drive direct revenues. The 'agentic coding bump', but for all your analysts, traders, and wealth managers, would be a much more attractive prospect.
I don't disagree that coding is the most successful use case so far (and probably more relevant to a HN audience). But I think the future of the labs is also contingent on them making progress on more general white collar work. I suspect that's why the Opus 5 release blog lists 3 coding benchmarks (FrontierBench, DeepSWE and FrontierCode) to 3 or 4 more general ones applicable to office work - depending on how you slice it (GDPVal, AutomationBench, Legal Agent Benchmark, BrowseComp).
ronsor•1h ago