> Coder: a coding version with half of the experts removed. It reaches 91% of the full model's SWE-bench Verified score (measured by its authors) and fits 32 GB of RAM.
https://github.com/Niko1221/Strata#which-model-should-i-pick
We observe that such degradation varies substantially across these factors: 4-bit quantization usually preserves performance, 2-bit often causes broad degradation
This repo uses 2-bit quantization and removes some of the experts for its smallest fastest model. Make of that what you will.ISTA IQ3_XXS does ~21 tok/s decode and ~240t/s prompt processing
And I thought piping to bash was bad
The Readme doesn't say, but it's all AI generated, so..
snehesht•1h ago
https://huggingface.co/Qwen/Qwen3.8-Flash-Next
proc0•1h ago
incognito124•43m ago
snehesht•40m ago
nicce•23m ago
mickeyp•37m ago
It is also a competent tool caller when quantised to NVFP4 for use with ninfer; my own harness only reports the occasional hiccup and it is only because the model will sometimes emit tool calling tokens in its reasoning loop.
snehesht•24m ago
thatsabadlook•19m ago
geye1234•13m ago
thatsabadlook•21m ago