it also has too many false positives
> We achieved a total 78.2% model size reduction compared with GLM 5.3 at full precision, and a 32.8% reduction compared with GLM 5.3 with an AWQ INT4 quantization. Altar retains 168 of the original 256 routed experts in each backbone expert layer, removing 88, or 34.4% of them. The router still selects eight per token, now from the smaller bank.
This might explain it
zshn•57m ago