Unbelievably slow, but we managed to fit a 27b model on the kv260 board, using an FPGA to hardcode the model architecture. Part of prototyping at Lamb Labs!
This was a PoC, now we're post-training the models to run faster, so hopefully this can be useful to someone. Bonus points if you can guess the quantization of the model.
thomaslanning•1h ago
This was a PoC, now we're post-training the models to run faster, so hopefully this can be useful to someone. Bonus points if you can guess the quantization of the model.