Skip to content

Comment on Apple Silicon and macOS VMs: Faster LLM Inference with llama.cpp

Comments

11.08× faster and generated tokens 16.36× faster than the same workload in the same stock VM.

So this was the comparison, for me the title was a bit confusing

yeah fair point. it's always tricky to get the whole idea across within HN's title limit. tldr: we ran the same workload in the same Lume macOS VM on the same Apple Silicon host, first with stock Metal capability reporting and then with our process-scoped dynamic library. The 11.08x figure is prompt processing, while 16.36x is token generation. the mechanism technically extends to graphics workloads too but these figures are specifically from llama.cpp

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.