frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

Apple Silicon and macOS VMs: 11–16× Faster LLM Inference with Llama.cpp

https://github.com/trycua/cua/blob/main/blog/gpu-passthrough-macos-vms.md
77•frabonacci•1h ago

Comments

thehamkercat•1h ago
> 11.08× faster and generated tokens 16.36× faster than the same workload in the same stock VM.

So this was the comparison, for me the title was a bit confusing

frabonacci•47m ago
yeah fair point. it's always tricky to get the whole idea across within HN's title limit. tldr: we ran the same workload in the same Lume macOS VM on the same Apple Silicon host, first with stock Metal capability reporting and then with our process-scoped dynamic library. The 11.08x figure is prompt processing, while 16.36x is token generation. the mechanism technically extends to graphics workloads too but these figures are specifically from llama.cpp
azinman2•57m ago
I don’t understand what Apple 1-9 are. At first I thought it was M series chips but there is no M9 (yet)
niklasbuschmann•53m ago
https://developer.apple.com/documentation/metal/mtlgpufamily
wtallis•39m ago
So those generation numbers aren't really anchored to Apple's hardware designs. It's just counting from when Apple introduced the Metal API, and the first several generations were when the GPU cores Apple was using were still nominally PowerVR designs.
simonw•55m ago
It looks to me like this won't speed up llama.cpp for everyone, just for users running it in this particular kind of Virtualization.framework VM.

The fix here works around a problem where the VM was causing llama.cpp to select the wrong kernels.

frabonacci•31m ago
> this won't speed up llama.cpp for everyone, just for users running it in this particular kind of Virtualization.framework VM.

correct. these figures apply to llama.cpp inside the macOS guest configuration we tested. Lume is the VM frontend we used, while Apple's Virtualization.framework provides the virtual GPU. bare-metal llama.cpp is unaffected.

> The fix here works around a problem where the VM was causing llama.cpp to select the wrong kernels.

mostly, with one nuance: llama.cpp is selecting the correct kernels for the capability answers it receives. the stock guest reports an older Apple GPU family and a 32 KB threadgroup memory limit, so llama.cpp chooses slower kernels. Our process-scoped layer reports the tested Apple 9 and 64 KB values while allowing llama.cpp to select newer paths that the paravirtual GPU successfully execute

the layer itself though works at the Metal API boundary, independently of llama.cpp. other Metal compute and graphics apps now may select newer paths from the same capability answers, although this is still preliminary and each app needs separate testing. for example, MLX-LM stayed flat in our tests

historically related limitations have been coming up across Apple Silicon VM frontends for a while e.g. Tart tracked MPS/GPU support back in 2023: - https://github.com/openai/tart/issues/501 - https://github.com/openai/tart/issues/1032

UTM also has related cases where apps detect the Apple paravirtual Metal device but falls back to software rendering: https://github.com/utmapp/UTM/issues/7671

engzaanin•44m ago
That makes sense. The title initially sounded like a general llama.cpp speedup on Apple Silicon, but if the improvement comes from fixing kernel selection inside Virtualization.framework VMs, that distinction is pretty important.
shay_ker•36m ago
I recall there was another YC startup that was working on Mac-specific ML optimizations for local inference (and perhaps fine-tuning).

I wonder if their work is related?

frabonacci•20m ago
RunAnywhere or Conifer?
myshapeprotocol•24m ago
Massive speedups for LLM inference on Apple Silicon VMs. Unlocking local hardware performance like this is huge for local-first workflows.
woadwarrior01•24m ago
The Claudish in the blogpost makes it really hard to ready. Also, TinyLlama 1.1B lol.

England set to be one of the first countries to eliminate hepatitis C

https://www.bbc.com/news/articles/c75gk620r22o
268•stevekemp•3h ago•177 comments

Apple Silicon and macOS VMs: 11–16× Faster LLM Inference with Llama.cpp

https://github.com/trycua/cua/blob/main/blog/gpu-passthrough-macos-vms.md
79•frabonacci•1h ago•19 comments

Show HN: Git-knife – edit commit messages, authors, and dates like a spreadsheet

https://github.com/TheRealYT/git-knife
44•YonathanTesfaye•1h ago•19 comments

Stealing Reasoning Traces from Proprietary LLM APIs

https://stolen-thoughts.com/
114•quantumgarbage•2h ago•43 comments

Manus will return to operating as an independent company

https://manus.im/blog/a-note-to-our-users
31•thm•1h ago•12 comments

Launch HN: Keet (YC S24) – An app to create video courses on anything

https://www.trykeet.com/
10•zackashen•1h ago•10 comments

Nvidia's Risky Business

https://stratechery.com/2026/nvidias-risky-business/
128•jonbaer•6h ago•35 comments

France to ban unsolicited telemarketing calls

https://www.lemonde.fr/en/france/article/2026/08/06/france-to-ban-unsolicited-telemarketing-calls...
806•aziaziazi•7h ago•392 comments

As AI eats the web, the internet’s collective memory is disappearing

https://thewalrus.ca/google-search-is-dying/
643•awnird•17h ago•700 comments

H3-metal – Native MiniMax-H3 inference for Apple Silicon

https://github.com/antirez/h3.c
385•swyx•14h ago•87 comments

University of Michigan Drops First-Semester Grades To'Curb Mental Health Crisis'

https://www.wsj.com/us-news/education/university-of-michigan-grades-mental-health-1a5701d4
14•cwwc•19m ago•1 comments

What I learned by putting GitHub Copilot behind a MitM proxy

https://www.lighthousenewsletter.com/p/i-put-github-copilot-behind-a-mitm
73•j0selit0•5h ago•6 comments

Halcyon Video – a 3D video store for your media server

https://github.com/halcyon-video/halcyon-video
22•Gander5739•4d ago•7 comments

Flock wanted to tap dashcams in rideshare vechicles to add to surveillance data

https://flowingdata.com/2026/08/10/flock-wanted-to-tap-dashcams-in-rideshare-vechicles-to-add-to-...
45•skadamat•2h ago•2 comments

Beltrunner: Game Design Postmortem

https://blog.gingerbeardman.com/2026/07/30/beltrunner-game-design-postmortem/
6•surprisetalk•1d ago•2 comments

Why Did OpenAI's Head of Ethics Chloé Bakalar Leave?

https://aimagazine.com/news/why-did-openai-head-of-ethics-chloe-bakalar-leave
55•ashurandi•2h ago•51 comments

The US tried to stop cartel money-laundering; devastated mom-and-pop businesses

https://www.theguardian.com/us-news/2026/aug/11/us-mexico-border-area-money-transfer-rule-change-...
49•hedora•1h ago•27 comments

Federal vendor with $50M in contracts leaves portal broken for a month

https://www.propublica.org/article/foia-requests-responses
26•ams1•1h ago•1 comments

Show HN: Needle2: 14MB agentic LLM for phones, wearables, smart home and robots

https://cactuscompute.com/needle
464•HenryNdubuaku•22h ago•159 comments

Jolt: Clojure compiler implemented with Chez Scheme

https://jolt-lang.github.io
7•mark_l_watson•2d ago•2 comments

Mark Zuckerberg attacks 'closed' AI rivals as Meta returns to open models

https://www.ft.com/content/4e3957f8-ea7c-4c46-a3de-cdce8e526878
583•root-parent•1d ago•557 comments

Chicken Scheme 6.0

https://code.call-cc.org/releases/6.0.0/NEWS
258•eatonphil•15h ago•40 comments

LFM2.5 2.6B model competitive with 4x larger models

https://huggingface.co/LiquidAI/LFM2.5-2.6B
127•nateb2022•6d ago•35 comments

Show HN: Scroll through all 43252003274489856000 Rubik's Cube states

https://everycube.alen.is/
256•Alen123•16h ago•91 comments

$580M undersea cable rerouted to avoid the grave of Dobby the House Elf

https://www.tomshardware.com/networking/usd580-million-undersea-cable-rerouted-to-avoid-the-grave...
48•rbanffy•2h ago•48 comments

How Claude marks AI-generated content

https://support.claude.com/en/articles/16266773-how-claude-marks-ai-generated-content
334•mfiguiere•18h ago•303 comments

The “mechanical miracle” that ruined Mark Twain’s life

https://resobscura.substack.com/p/the-mechanical-miracle-that-ruined
184•benbreen•6d ago•100 comments

Nvidia Nemotron 3.5 Lightning

https://huggingface.co/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4
59•beklein•2h ago•15 comments

Stowaway – Take the window seat on any plane or satellite overhead

https://stowaway.live/
364•thunderbong•4d ago•48 comments

Faster floating point math with Rust's new API

https://pythonspeed.com/articles/faster-float-math-rust/
83•subset•5d ago•28 comments