> We match DeepSeek V4 Pro Base using ~50x fewer FLOPs – that’s around half of GPT3’s pretraining compute, or ~$0.5M on GB200.
If this holds up that's a really big deal.
wayfwdmachine•21m ago
Huge if true. As it were.
ismael_rr•12m ago
Super awesome. Wish they would release the paper about what they did to achieve this. I remember nous released the token superposition paper which improved pretraining FLOPs some, but not 50x: https://nousresearch.com/token-superposition. Wondering if they also found some cool tokenization strategiesa
monneyboi•7m ago
Imagine the sheer amount of power you could save by releasing the paper.
simonw•37m ago
If this holds up that's a really big deal.
wayfwdmachine•21m ago