frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

Kimi K3 is not cheap

https://www.alexinch.com/blog/kimi-k3
18•ainch•59m ago

Comments

ronsor•45m ago
It's cheap because it won't refuse random tasks. You can't get rid of nannying at any price beyond training your own model, and relative to that, K3 is cheap.
cloudie78•44m ago
It’s cheap.
SwellJoe•43m ago
In my testing, I'm finding it more expensive than Opus 4.8/5 and GPT 5.6 Sol at API rates, because it chews so much. And, their plan (at least the $19 tier) is much less generous than the ChatGPT $20 plan, like an order of magnitude less, it's basically a demo not a useful amount of usage.
Stagnant•18m ago
Yeah can't recommend their $19 plan, only took me a day and a half to hit the weekly usage cap. The $39 plan has 5x the limits so I recommend getting that instead. Having now tested it for a few days, it is the first of the chinese models that actually feels comparable to Opus-tier models
himata4113•42m ago
it IS cheap (once the weights are released) and it will only become CHEAPER. For around $3700 a month (via loan purchased hardware + energy cost) you can run around 32 concurrent instances of kimi k3 which can generate nearly a 6.9 billion tokens a day.

This is napkin math since I'm mostly just extrapolating from glm 5.2 by assuming it's twice as heavy to serve in every single measurement, but I believe you can easily achieve 2500tok/s aggregate compared to 4500tok/s and up to 8000tok/s for glm5.2.

with nvidia r100 you are likely going to be able to push that number even higher while the cost of hardware appears to be relatively the same, so far I am seeing 21% premium from supermicro which is twice as fast and has nearly twice the vram.

CamperBob2•41m ago
A lot hinges on what happens tomorrow. I'll believe they'll open the weights when I see the files appear on HF (and when somebody with 24 RTX6000s or whatever reports that they are indeed as good as the closed version.)
coder543•41m ago
This article seems premature to post. Right now, the price is arbitrarily set by a single provider. Why wouldn't Moonshot collect extra revenue during this exclusivity period when they knew there would be hype?

The model weights are supposed to release tomorrow.

Over the next several weeks, I would expect competition among open weight providers to drive down the cost, as I've seen happen with other open weight model releases.

ainch•9m ago
That's a very fair critique.

I don't mean to imply that Kimi is not at all cheaper than U.S frontier models. I more wrote this because I believe - since Chinese LLMs entered the public consciousness via DeepSeek R1, which was genuinely ~20x cheaper than o1 - there's a bit of a halo effect around Chinese models which causes people to overestimate the scale of the discount. And relative to that price anchor, Kimi is less extraordinarily cheap.

At the moment Kimi is ~10% cheaper than GPT-5.6 on the AA benchmark, and as you say that could go down to 20-30% cheaper (although I don't know how inference provider discounts play out on real world usage once you account for quantisation etc...). I'm not trying to suggest that that's nothing, but I do think some of the people driving the Chinese AI discourse would have a harder time pitching their conclusions if they were saying "this new Chinese model is 10% cheaper on some tasks, and it might get another 20% cheaper in the future".

sroerick•40m ago
It feels like this "Kimi is a token hog" meme is 100% astroturfed by Anthropic. It's cheap. Believe your own eyes.
ofjcihen•37m ago
I mean at this point their very existence depends on it so I’m not sure if I’d be surprised
sroerick•29m ago
Not to mention -

If you switch the view to "coding tasks" on this website:

  Kimi K3: $3.18 per task
  GLM 5.2: $6.51 per task
  GPT 5.6 Sol: $7.02 per task
  Opus 5: 8.23 per task
  Fable: 11.70 per task
So it's pretty dang cheap lol. Nobody is using frontier inference for "office tasks".
ofjcihen•26m ago
Right? The availability of this being in the article that’s pushing the opposite narrative is like…what?
sroerick•18m ago
I was actually shocked to see that much improvement on GLM 5.2. I am getting pretty great rates in GLM5.2 right now and I'm extremely happy with the output. I have found Kimi to be generally a little slower but noticeably better at architecture and structuring things. I would have thought is for sure currently more expensive than GLM5.2, particularly with subscriptions etc, but I'm excited for this to decrease further.
ofjcihen•40m ago
I mean you have a chart showing that it’s cheaper than the other models and it also does what I want without argument.

Additionally, I fully expect the frontier labs to continue increasing prices to meet the profit margins they need to to continue existing.

jszymborski•37m ago
K2.6 is cheaper than GLM5.2 (at least on DeepInfra) and I've found it works as good as Sonnet for my purposes. Both tend to think themselves into circles a bit and aren't super token efficient, but I've found GLM5.2 much worse on this count making K2.6 even cheaper than the per token price would make seem.
sroerick•16m ago
I find them to be about comparable, but I use them both for coding tasks and I'm happy with each. I like K3

How Web Browsers Work

https://arnauc.me/blog/how-browsers-work/
2•ErenayDev•4m ago•0 comments

Plasma Tunnels Reveal How Dying Satellites Fall to Earth

https://spectrum.ieee.org/space-debris-atmosphere-burn-up
1•marc__1•4m ago•0 comments

The fifth dimension could be possibility

https://gingerjuice.club/article/the-dimensional-stacking-principle-a-geometric-framework-for-the...
1•streetai•5m ago•0 comments

Ask HN: What's the best hands-on path to learn ML inference infrastructure?

1•censor5•8m ago•0 comments

AI is set to drive surging electricity demand from data centres (2025)

https://www.iea.org/news/ai-is-set-to-drive-surging-electricity-demand-from-data-centres-while-of...
1•dredmorbius•10m ago•0 comments

Making a cache cluster more effective

https://basta.substack.com/p/making-a-cache-cluster-more-effective
1•tyre•13m ago•0 comments

Gary Stevenson to quit YouTube channel, citing health concerns

https://www.theguardian.com/business/2026/jul/26/gary-stevenson-to-quit-youtube-channel-garys-eco...
1•jimnotgym•16m ago•0 comments

CuNi v01

https://cuni-studio.fly.dev/
1•agentrider•17m ago•0 comments

Why Everyone Got Spain's Blackout Wrong [video]

https://www.youtube.com/watch?v=Rb9oWsQuENE
1•pepperoni_pizza•21m ago•0 comments

An Open-Source Static AI Capability and Risk Analyzer – First Week Report

https://github.com/ikaruscareer/SafeAI
1•ikaruscareer•23m ago•0 comments

1877–1878 El Niño event

https://en.wikipedia.org/wiki/1877%E2%80%931878_El_Ni%C3%B1o_event
2•simonebrunozzi•23m ago•0 comments

Doom for PC-FX [video]

https://www.youtube.com/watch?v=wUd5IGMbe48
1•cedel2k1•24m ago•0 comments

Buying a Home Has Gotten Harder for Young Adults in Most U.S. Metro Areas

https://www.pewresearch.org/short-reads/2026/06/24/buying-a-home-has-gotten-harder-for-young-adul...
1•karakoram•25m ago•0 comments

PGSimCity – an explorable 3D model that shows how Postgres works

https://github.com/NikolayS/pgsimcity
1•samokhvalov•25m ago•0 comments

Big Tech accused of stonewalling European social media researchers

https://www.wired.com/story/european-researchers-want-to-study-social-medias-harms-but-cant-get-t...
2•logickkk1•29m ago•0 comments

Wright's Law Edges Out Moore's Law in Predicting Technology Development (2012)

https://spectrum.ieee.org/wrights-law-edges-out-moores-law-in-predicting-technology-development
1•simonpure•34m ago•0 comments

Coding Has Agents. Trading Has One

https://henryzhang.substack.com/p/coding-has-agents-trading-finally
1•henryzhangpku•34m ago•0 comments

Simulate cassette tape audio profiles using FFmpeg

https://github.com/AARomanov1985/Audio-Cassette-Simulation
2•xterminal•34m ago•1 comments

Show HN: Infinite Jigsaw Game

https://infinitejigsaw.com
2•impostervt•35m ago•0 comments

The Sustained Performance Gap: Why Laptop Specs Don't Tell the Whole Story

https://psyll.com/articles/technology/tech-gadgets/thermal-throttling-why-laptops-lie-about-speed
1•lucasfletcher•35m ago•0 comments

The Lego Problem, Revisited

https://seths.blog/2026/07/the-lego-problem/
1•herbertl•37m ago•0 comments

Most Americans Say Financial Milestones Are Harder for Today's Young Adults

https://www.pewresearch.org/short-reads/2026/07/17/majorities-of-americans-say-key-financial-mile...
4•karakoram•40m ago•0 comments

The American E.V. Has Been Crushed. Will It Take the U.S. Auto Industry with It?

https://www.nytimes.com/2026/07/15/magazine/electric-cars-american-evs.html
4•bookofjoe•41m ago•1 comments

Vulnerabilities in CJSON

https://joshua.hu/cjson-json-parser-cve-vulnerabilities
2•ingve•44m ago•0 comments

Automate 7 Projects Marketing with AI for $40

https://apsquared.co/posts/marketing-automation-claude-codex
1•apsquared•44m ago•0 comments

What does GitHub's security team even do?

https://orchidfiles.com/github-security-team/
41•theorchid•45m ago•6 comments

We built phone-only laser tag with computer vision and sensor fusion

https://lightwarsar.com/engineering/targeting-system/
2•mrr7337•47m ago•0 comments

When Compilers Disagree About UTF‑8

https://nemanjatrifunovic.substack.com/p/when-compilers-disagree-about-utf8
2•ingve•50m ago•0 comments

Two AI Futures to Choose From

https://www.rameznaam.com/p/two-ai-futures-to-choose-from
2•simonpure•51m ago•0 comments

against “it’s not that deep”. actually, it is

https://velvetnoise.substack.com/p/against-its-not-that-deep-actually
2•jger15•53m ago•0 comments