I didn’t see a gguf yet, going to be most interesting if there’s a quant that fits nicely into about 90FN so it can run in 128GB unified memory.
It’s 12B active so should hopefully be pretty fast at 2 bit quant if it fits, as in the ratio of total to active params is big compared to e.g. Qwen 3.5 122A10 which seems more commom
andy99•44m ago
It’s 12B active so should hopefully be pretty fast at 2 bit quant if it fits, as in the ratio of total to active params is big compared to e.g. Qwen 3.5 122A10 which seems more commom
Edit: and it’s out - excited to try https://huggingface.co/unsloth/Inkling-Small-GGUF