Skip to content

Show HN: Shoehorn, a library to quantize an LLM to fit your Mac's VRAM

github.com/notactuallytreyanastasio
6 pointsrhgraysoniidiscuss
On HN

I made this after seeing someone posit the idea online yesterday over lunch then spent some time refining it. So far it's pretty impressive IMO! Right now I am running Qwen3-30B-A3B on my 24gb unified memory m4 MacBook Pro at 50 tok/sec and this should definitely not be working for such a large model on my middling hardware.

Things are detailed in the README to get up and running and DESIGN.md has details on all the choices and such made along the way.

Comments

No comments yet.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.