Skip to content

Show HN: Shoehorn – Quantize any model down to run on your machine

notactuallytreyanastasio.github.io
99 pointsrhgraysonii16 comments
On HN

Working on Mac, Linux, and Windows now. I include a simple GUI to find new models and get things built and set up. It is working quite well across a few models for me. The GitHub README and DESIGN.md files go into detail of the how/why and it's working remarkably well so far. https://github.com/notactuallytreyanastasio/shoehorn

Comments

This is interesting. I wonder how it could work with something like https://github.com/JustVugg/colibri.

LLMFit tells you what can run on something. I built something quite similar to their search into Shoehorn now.

I gotta laugh at some of the models it suggests, for example:

AnkitAI/Parable-Qwen3-4B-Claude-Fable-5-GGUF

you’re telling me you managed to fit Fable 5 into just 4B?

Fyi that model name to me reads

Qwen3 4b params distilled/trained with fable 5

This is really impressive. Can you say a bit about the underlying process? I'm guessing this is post-training qantization? Isn't PTQ also resource-intensive? (Ie might not work on any machine)

The project name is perfect!

does this work similar to airllm? i am wondering how it would handle something like quantizing kimi k3 on a budget of 8 gbs, or is that something you are not attempting to solve yet?

Yes that is exactly what this does.

Could you explain what happens when you try to shoehorn a 2.4T parameter model into a 24gb m4 mac?

extreme divergence would be my guess

Wondering the same thing but for 48gb M5 Max.

tried it out but based on the model sizing result i got i got an insufficient memory error when the server started running

If you could post an issue if you still have the error around that would be awesome.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.