I wonder how much spending is motivated by the phenomenon of sudden emergent performance in LLMs. Clearly some people who are smarter than me expect something like emergent AGI, or at least they think the odds justify spending whatever it takes to see if that would happen.
That leaves a lot of room for efficient aggressive followers.
Either they have Blackwell with native 4-bit floating math, or they use have Chinese domestic NPU that support mxfp4 natively.
The article’s statement does not make sense.
HarHarVeryFunny•44m ago
I guess at least partly a reflection of all the optimizations in the Kimi 3 architecture.
In the recent leaked DeepSeek investor meeting, they also mentioned only having a 20K GPU cluster (unclear if NVIDIA, or Huawei).
cubefox•20m ago
Source?
HarHarVeryFunny•19m ago
gyanchawdhary•8m ago
a-priori•16m ago
The goal will be to develop smaller models with more efficient architectures, that have similar or even better performance than larger models.
trollbridge•10m ago
A VAX 11/780 was good, but an 80386 was a lot better, since the latter could run on 3 AA batteries and the former needed 6,000 watts of 3 phase.
infecto•14m ago
HarHarVeryFunny•9m ago