+
https://poolside.ai/assets/laguna/laguna-m1-xs2-technical-re...
This is pretty impressive.
Even more important, subjectively, is that this model will run very well on Strix Halo (e.g. Framework Desktop), DGX Spark kinds of devices. Looking forward to Unsloth dynamic mtp quants.
P.S. Looking at the HF release they already offer Q4_K_M and DFlash drafter for speculative decoding!
For a while there's been nothing to run on my Strix Halo that's notably better than what I can run on my dual 32GB GPU desktop (Gemma 4 or Qwen 3.6 dense models), but this seems likely to be the step up in size that actually works better than those.
Anyways, keep 'em coming.
The pricing here is incredible. This is the first US release that's competitive with DeepSeek V4 Flash. Very excited about this.
That said, if someone would kindly quantise this down for the 64GB paupers, that would be appreciated. (I know there’s likely degradation, but some people reported good results with a 2 bit version of Qwen 3.5 122B, and this is starting from a higher point. Would be interesting to try, at least.)
Edit: someone in the process of doing so: https://huggingface.co/vcruz305/Laguna-S-2.1-GGUF
tosh•1h ago
similar performance to deepseek v4, inkling at size of nemotron 3 super (!)