Comment on Show HN: Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/sparentComments−carloslfuOP10dthanks!ditched oLlama"yeah! this is interesting.8.1GB per slotserve process is a lot! Is that in your control?yes, it is hard, but I agree the smaller the better. I'll work on thatIf it's local and open-weight, this could be marketed this way I think.I like this!what's the actual, real use case for slotserve?I'm working rn on an app on top of it that closes the loop and is a fully local AI app, an experiment. I'll publish it as soon as it is usable!built a small html hello world served via PythonWhat did you use as a harness here?−baristaGeek10dFor the harness... just the shell. No client library. Does that answer your question?−carloslfuOP10dthanks! in part, I was wondering how you got the code into files. I guess you copy pasted it inside a file, am I right?
Comments
thanks!
yeah! this is interesting.
yes, it is hard, but I agree the smaller the better. I'll work on that
I like this!
I'm working rn on an app on top of it that closes the loop and is a fully local AI app, an experiment. I'll publish it as soon as it is usable!
What did you use as a harness here?
For the harness... just the shell. No client library. Does that answer your question?
thanks! in part, I was wondering how you got the code into files. I guess you copy pasted it inside a file, am I right?