frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

Moonshot built on 20k Nvidia chip cluster from Alibaba

https://www.bloomberg.com/news/articles/2026-07-31/moonshot-s-kimi-built-on-20-000-nvidia-chip-cluster-from-alibaba
33•gk1•1h ago

Comments

HarHarVeryFunny•44m ago
Interesting if true - that Moonshot can train a ~3T SOTA model on only 20K NVIDIA GPUs, while others like Musk (who freely admits to distilling OpenAI's models) struggle to build a competitive 1T model (Grok 4.5) with massively more compute (Colossus-1 100-250K GPUs, Colossus-2 500K+ GPUs).

I guess at least partly a reflection of all the optimizations in the Kimi 3 architecture.

In the recent leaked DeepSeek investor meeting, they also mentioned only having a 20K GPU cluster (unclear if NVIDIA, or Huawei).

cubefox•20m ago
> Musk (who freely admits to distilling OpenAI's models)

Source?

HarHarVeryFunny•19m ago
https://www.forbes.com/sites/antoniopequenoiv/2026/04/30/elo...
gyanchawdhary•8m ago
Sounds fair, since Sam can run a for profit NGO
a-priori•16m ago
It kind of confirms a hypothesis I have that the next phase of AI development will be about getting smaller (in terms of model size and compute), because smaller is more capital efficient for training (allowing faster iteration and more iteration cycles for a given amount of capital), allows for denser inference (more inference for a given amount of compute hardware), and allows for more edge inference applications.

The goal will be to develop smaller models with more efficient architectures, that have similar or even better performance than larger models.

trollbridge•10m ago
Getting smaller is how the last revolution in computing happened.

A VAX 11/780 was good, but an 80386 was a lot better, since the latter could run on 3 AA batteries and the former needed 6,000 watts of 3 phase.

infecto•14m ago
Genuine question. Is there a time factor part of that equation?
HarHarVeryFunny•9m ago
It's obviously a question - did they just train for 10x as long due to having 10x fewer GPUs, but then that spoils the claim that they distilled Fable which was only recently introduced. Now doubt they did use some training data generated from older US models though.
Zigurd•41m ago
I would not minimize the extent to which AI progress depends on things like the effort that goes into training material acquisition and data labeling. In those areas, improvement is a direct function of how much you invest.

I wonder how much spending is motivated by the phenomenon of sudden emergent performance in LLMs. Clearly some people who are smarter than me expect something like emergent AGI, or at least they think the odds justify spending whatever it takes to see if that would happen.

That leaves a lot of room for efficient aggressive followers.

idoxer•17m ago
https://archive.is/vgmMB
infecto•13m ago
My thesis is still that model building has no moat. Folks continue to migrate around between the big labs. There is a lot of value in having good taste around the harness and how the models are used. The medium to long term winners will be the folks that control the compute.
ycui7•10m ago
K3 is natively trained to mxfp4, if they cannot get a hold of Blackwell chip, it is meaningless. Hopper does not do native 4-bit floating math.

Either they have Blackwell with native 4-bit floating math, or they use have Chinese domestic NPU that support mxfp4 natively.

The article’s statement does not make sense.

avipars•5m ago
https://archive.md/vgmMB

DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis

https://artificialanalysis.ai/models/deepseek-v4-flash-ga
306•theanonymousone•7h ago•150 comments

The Maxwell Conjecture Is False (GPT 5.6 Sol)

https://arxiv.org/abs/2607.27197
71•rahen•3h ago•39 comments

Google fixed more Chrome bugs in June than over the past two years, thanks to AI

https://blog.google/security/chrome-stronger-with-every-update/
325•Garbage•7h ago•285 comments

The session you cannot take with you

https://earendil.com/posts/session-portability/
613•apitman•11h ago•167 comments

Winding Down Artichoke Ruby

https://hyperbo.la/w/winding-down-artichoke-ruby/
18•ksec•5d ago•1 comments

Show HN: Gander, an Android file viewer that asks for no permissions at all

https://github.com/mokshablr/gander
157•mokshablr•9h ago•59 comments

JEP 401: Value Objects (Preview) merged to OpenJDK master

https://github.com/openjdk/jdk/pull/31120
198•mfiguiere•10h ago•116 comments

DeepSeek-V4-Flash Update

https://api-docs.deepseek.com/updates/
476•dnhkng•9h ago•240 comments

Stacked PRs are now live on GitHub

https://github.blog/changelog/2026-07-30-stacked-pull-requests-are-now-in-public-preview/
738•tomzorz•22h ago•253 comments

The End of an Era

https://hughhowey.com/the-end-of-an-era/
238•harscoat•3h ago•258 comments

Tasklet (YC P26) Is Hiring a Customer Success Engineer

https://tasklet.ai/careers/customer-success-engineer
1•mayop100•3h ago

Gemini Robotics 2 brings whole body intelligence to robots

https://deepmind.google/blog/gemini-robotics-2-brings-whole-body-intelligence-to-robots/
599•ai2027•1d ago•487 comments

Detect Dark Matter's Mark from Your Backyard

https://spectrum.ieee.org/dark-matter
18•Brajeshwar•51m ago•1 comments

Moonshot built on 20k Nvidia chip cluster from Alibaba

https://www.bloomberg.com/news/articles/2026-07-31/moonshot-s-kimi-built-on-20-000-nvidia-chip-cl...
33•gk1•1h ago•14 comments

I flagged two research papers for fake authors and both were accepted as orals

https://geospatialml.com/posts/reviewing-ai-slop/
255•volumes94•16h ago•128 comments

Better to Beg Forgiveness

https://pluralistic.net/2026/07/31/just-do-it/
14•hn_acker•1h ago•4 comments

The mean means nothing: data visualization to debug a latency problem

https://fzakaria.com/2026/07/27/the-mean-means-nothing
90•fanf2•2d ago•17 comments

Premier league bans gambling sponsors

https://www.footyheadlines.com/2646571793/betting-ban-takes-effect-no-more-gambling-sponsors-in-t...
171•paoliniluis•15h ago•50 comments

We shall dwell amidst wonder and glory for ever: On weird fiction

https://clereviewofbooks.com/we-shall-dwell-amidst-wonder-and-glory-for-ever-on-weird-fiction/
38•plimp•1d ago•1 comments

The fragile foundations of CoT monitoring

https://web.stanford.edu/~cgpotts/blog/cot/
3•jxmorris12•3d ago•0 comments

The Religion of Speed

https://graybeard.ing/the-religion-of-speed/
236•MobiusHorizons•15h ago•119 comments

Read this before you buy that TV streaming stick

https://krebsonsecurity.com/2026/07/read-this-before-you-buy-that-tv-streaming-stick/
771•speckx•22h ago•481 comments

Situational Awareness Down 67% in July in AI Stock Rout

https://www.wsj.com/finance/investing/situational-awareness-down-67-in-july-in-ai-stock-rout-cd19...
83•pondsider•1h ago•85 comments

Ruby Central's Destructive Legacy

https://andre.arko.net/2026/07/30/ruby-centrals-destructive-legacy/
59•4d66ba06•3h ago•31 comments

Where USB Memory Sticks are Born (2013)

https://www.bunniestudios.com/blog/2013/where-usb-memory-sticks-are-born/
81•jacquesm•3d ago•5 comments

Simulating TCP loss and congestion in browser using Go/WASM

https://ccsim.fly.dev
56•dilyevsky•2d ago•1 comments

GCC steering committee announces AI policy

https://lwn.net/Articles/1086041/
328•arto•1d ago•379 comments

The Economic Benefit of Refactoring

https://martinfowler.com/articles/exploring-gen-ai/refactoring-economic-benefit.html
267•javaeeeee•1d ago•115 comments

Memo-1: A 6502 computer built from scratch, using a Minitel as its terminal

https://github.com/MemoireMorte/Memo-1
96•sciences44•3d ago•13 comments

Bad Apple but It's Traceroute

https://jssfr.de/2026-07-27-bad-apple-but-traceroute.html
147•jssfr•3d ago•48 comments