The fix here works around a problem where the VM was causing llama.cpp to select the wrong kernels.
Nobody tells you how, they just act like you're an idiot. So I always say fuck this and install Ollama.
Then everyone goes BLASPHEMY "just use llama.cpp"
Yeah I did, and it's slow as hell. It doesn't work well. Idk why.
"You aren't doing it right"
Okay tell me how
"NO"
-----
For this reason, Ollama is the superior solution. I know, downvote, everyone hates Ollama here but until llama.cpp gets their shit together on developer experience it doesn't exist as far as I'm concerned
For reference, I get ~26 tok/sec with the new Muse 30B model.
I wonder if their work is related?
thehamkercat•31m ago
So this was the comparison, for me the title was a bit confusing
frabonacci•16m ago