Skip to content

Show HN: Slotstream, run Qwen3.8-Flash-Next 4-bit on a low-memory Mac

github.com/carloslfu
3 pointscarloslfu1 comment
On HN

I built slotstream, a way to run Qwen3.8-Flash-Next 4-bit on a low-memory mac starting from 16GB, a 125B parameter model that would need 100GB+ memory/RAM, thanks to expert-offloading/ssd-streaming. Easy to install/update, and mac-native using MLX and Swift.

It ships with auto-mode, which makes a good tradeoff between memory usage and speed.

I'll be implementing and porting the MTP module for speculative decoding next

Local models really are the future of computing!

Comments

crazy

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.