frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

Apple Silicon and macOS VMs: 11–16× Faster LLM Inference with Llama.cpp

https://github.com/trycua/cua/blob/main/blog/gpu-passthrough-macos-vms.md
45•frabonacci•52m ago

Comments

thehamkercat•31m ago
> 11.08× faster and generated tokens 16.36× faster than the same workload in the same stock VM.

So this was the comparison, for me the title was a bit confusing

frabonacci•16m ago
yeah fair point. it's always tricky to get the whole idea across within HN's title limit. tldr: we ran the same workload in the same Lume macOS VM on the same Apple Silicon host, first with stock Metal capability reporting and then with our process-scoped dynamic library. The 11.08x figure is prompt processing, while 16.36x is token generation. the mechanism technically extends to graphics workloads too but these figures are specifically from llama.cpp
azinman2•26m ago
I don’t understand what Apple 1-9 are. At first I thought it was M series chips but there is no M9 (yet)
niklasbuschmann•22m ago
https://developer.apple.com/documentation/metal/mtlgpufamily
wtallis•8m ago
So those generation numbers aren't really anchored to Apple's hardware designs. It's just counting from when Apple introduced the Metal API, and the first several generations were when the GPU cores Apple was using were still nominally PowerVR designs.
simonw•24m ago
It looks to me like this won't speed up llama.cpp for everyone, just for users running it in this particular kind of Virtualization.framework VM.

The fix here works around a problem where the VM was causing llama.cpp to select the wrong kernels.

engzaanin•14m ago
That makes sense. The title initially sounded like a general llama.cpp speedup on Apple Silicon, but if the improvement comes from fixing kernel selection inside Virtualization.framework VMs, that distinction is pretty important.
purplemoonx•12m ago
Still no easy way to get up and running performantly with Llama.cpp

Nobody tells you how, they just act like you're an idiot. So I always say fuck this and install Ollama.

Then everyone goes BLASPHEMY "just use llama.cpp"

Yeah I did, and it's slow as hell. It doesn't work well. Idk why.

"You aren't doing it right"

Okay tell me how

"NO"

-----

For this reason, Ollama is the superior solution. I know, downvote, everyone hates Ollama here but until llama.cpp gets their shit together on developer experience it doesn't exist as far as I'm concerned

unglaublich•11m ago
I think people generally throw Claude or Codex at the configuration challenge, so they don't know either.
purplemoonx•4m ago
Maybe the llama.cpp dev loved webpack as a child, or just loves making the most simple thing complicated as hell for no reason lol
drittich•6m ago
There are certainly challenges. When setting up a new model, I get AI to walk me through the commands using llama-benchmark that determine the best parameters for my particular configuration and needs. Once you've got that it's pretty easy to port those parameters to llama-server. It takes me about an hour to run through this process. It would be great if there was a registry of hardware, models, configuration parameters, and resulting tokens per second. Maybe one day we'll get there.
kevin42•5m ago
What hardware do you run? I have a first-gen mac studio, and I just run cmake and build with no special options. Same thing with llama-server, I just specify the model and use the built-in web UI.

For reference, I get ~26 tok/sec with the new Muse 30B model.

shay_ker•6m ago
I recall there was another YC startup that was working on Mac-specific ML optimizations for local inference (and perhaps fine-tuning).

I wonder if their work is related?

Mother of all VPNs v2 release: Same 16 protocols, rebuilt underneath

https://twitter.com/MotherofallVPNs/status/2087169933604008016
1•shayanbahal•1m ago•1 comments

We built our own status page with AI, replacing a $65k SaaS product

https://www.fivetran.com/blog/we-built-our-own-status-page-with-ai-replacing-a-65k-saas-product
1•georgewfraser•1m ago•0 comments

Jackie the bald eagle dead at 14

https://www.popsci.com/environment/jackie-bald-eahttps://www.popsci.com/environment/jackhttps://w...
1•cyndunlop•2m ago•1 comments

Nvidia Nemotron 3.5 Lightning

https://twitter.com/NVIDIAAI/status/2087162151995629926
1•tosh•2m ago•0 comments

What does International Vlogging Day mean in 2026?

https://sijobling.com/micro/what-does-international-vlogging-day-mean-in-2026/
1•mooreds•3m ago•0 comments

PDCA

https://en.wikipedia.org/wiki/PDCA
1•tosh•3m ago•0 comments

AI Is Accelerating Vulnerability Discovery

https://ai-updates.net/ai-accelerating-vulnerability-discovery-inventory/
1•ashurandi•4m ago•1 comments

Scrunch: Rewriting the Web for AI

https://www.decibel.vc/articles/scrunch-rewriting-the-web-for-ai
1•mooreds•5m ago•0 comments

Pure-Rust, Sandboxed, Browser for Claude Code / Codex

https://www.reddit.com/r/ClaudeAI/comments/1vknrjx/building_a_purerust_sandboxed_localfirst_browser/
1•syumei•5m ago•0 comments

The Vertical AI Bubble: We Keep Forgetting That LLMs Roll Dice

https://medium.com/@MirArshadTalpur/the-vertical-ai-bubble-we-keep-forgetting-that-llms-roll-dice...
1•Arshad-Talpur•6m ago•0 comments

Share What You Learned, Not Just Your AI Conversation

https://vcfg.me/writing/share-what-you-learned-not-just-your-ai-conversation/
1•vcfgdev•7m ago•0 comments

Low-Tech Ceramic Water Filter

https://wiki.lowtechlab.org/wiki/Filtre_%C3%A0_eau_c%C3%A9ramique/en
1•Bluestein•7m ago•0 comments

Learning from Historical Mistakes

https://www.johndcook.com/blog/2026/08/10/learning-from-historical-mistakes/
1•ibobev•7m ago•0 comments

Manually Unbreakable Cryptography

https://www.johndcook.com/blog/2026/08/11/manually-unbreakable-cryptography/
1•ibobev•7m ago•0 comments

Dogs and Fat Tails

https://www.johndcook.com/blog/2026/08/11/dogs-and-fat-tails/
1•ibobev•7m ago•0 comments

Why did I build a product no one needs?

https://metedata.substack.com/p/017-why-did-i-build-a-product-no
1•young_mete•7m ago•0 comments

Free tokens for sale: How fake signups drive AI fraud

https://www.okta.com/blog/threat-intelligence/free_tokens_for_sale/
1•mooreds•8m ago•0 comments

A quick look at zero-knowledge proofs

https://bernsteinbear.com/blog/zkp/
1•tekknolagi•8m ago•0 comments

How the Heck Does JPEG Work?

https://perthirtysix.com/how-the-heck-does-jpeg-work
2•shriracha•10m ago•0 comments

Show HN: A football SIM that has run 8 seasons – every match deterministic

https://danfootball.com
1•declaj•10m ago•0 comments

Summer 2026 on course to be UK's hottest on record, says Met Office

https://www.theguardian.com/environment/live/2026/aug/11/uk-europe-france-heatwave-climate-crisis...
1•ljf•11m ago•0 comments

Spotify Xirp

https://xirp.spotify.com/
1•tosh•11m ago•0 comments

Stealing Reasoning Traces from Proprietary LLM APIs [pdf]

https://stolen-thoughts.com/paper.pdf
1•amrrs•13m ago•0 comments

Program with Paint Brushes, Not Pencils

https://blog.pickcode.io/program-with-paint-brushes-not-pencils/
1•skadamat•13m ago•0 comments

Show HN: Prompting Refinement Tool [requesting testing]

https://www.promptme.host/
1•markquis91•14m ago•0 comments

Microsoft responds to unwanted OneDrive feature installs

https://www.neowin.net/news/following-backlash-microsoft-rushes-to-add-uninstall-option-for-unwan...
1•stalkHN24x7•15m ago•1 comments

Changes to the Open Science Framework

https://www.cos.io/blog/osf-changes-a-note-to-users
1•EndXA•17m ago•0 comments

Harvesting Ethereum Traces Without an Archive Node

https://fables-for-robots.ch/blog/harvesting-ethereum-traces-without-an-archive-node/
1•draganm•19m ago•0 comments

Mozilla revokes Firefox signing key after unencrypted copy lands in GitHub

https://www.theregister.com/security/2026/08/11/mozilla-revokes-firefox-signing-key-after-unencry...
1•connorboyle•19m ago•0 comments

Leanroute Is Live One AI Gateway for Models and Tools

1•leanroute_ai•19m ago•0 comments