Skip to content

Comment on Show HN: Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/sparent

Comments

thanks!

ditched oLlama"

yeah! this is interesting.

8.1GB per slotserve process is a lot! Is that in your control?

yes, it is hard, but I agree the smaller the better. I'll work on that

If it's local and open-weight, this could be marketed this way I think.

I like this!

what's the actual, real use case for slotserve?

I'm working rn on an app on top of it that closes the loop and is a fully local AI app, an experiment. I'll publish it as soon as it is usable!

built a small html hello world served via Python

What did you use as a harness here?

For the harness... just the shell. No client library. Does that answer your question?

thanks! in part, I was wondering how you got the code into files. I guess you copy pasted it inside a file, am I right?

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.